6DOF rendering of multi-point higher order ambisonics
The apparatus and method address discontinuities in 6DoF audio rendering by determining listener position and switching rendering modes, ensuring continuous audio quality when HOA source data is unavailable.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2025-10-30
- Publication Date
- 2026-06-04
AI Technical Summary
Current 6DoF audio rendering systems face challenges when they do not have immediate access to audio signals for higher-order Ambisonics (HOA) sources at the listener's location, leading to discontinuities and reduced rendering quality.
An apparatus and method for generating binaural audio that determines the listener's position relative to available HOA sources, switching between interior and exterior rendering modes, and using neighboring sources when necessary to maintain plausible audio rendering quality.
Enables seamless 6DoF audio rendering by smoothly transitioning between rendering modes, ensuring continuous and responsive audio experiences even when HOA source data is unavailable.
Smart Images

Figure EP2025081374_04062026_PF_FP_ABST
Abstract
Description
[0001] 6DOF RENDERING OF MULTI-POINT HIGHER ORDER AMBISONICS
[0002] Field
[0003] The present application relates to apparatus and methods for audio rendering with 6 degree of freedom systems of multi-point higher order ambisonics audio.
[0004] Background
[0005] The current MPEG-I Immersive audio standard (ISO / IEC 23090-4 WD3) renderer supports 6 degrees-of-freedom (6DoF) rendering of audio scenes comprising multiple first order or higher-order Ambisonics (FOA, HOA) microphone recordings or synthesized signals. These renderers are able to provide a binaural signal at the listener position (and orientation) based on the recorded FOA / HOA signals and positions associated with the signals and the listener. That is, the renderer is able to provide a binaural signal at a non-sampled position in the scene, thus providing a 6DoF experience for the listener.
[0006] For example Fig.1 shows an example scene with microphone or FOA / HOA source positions (mi 111, m2 113, m3 115,
[0007]
[0008] 117, and ms 119) and a listener position p / 101. The microphone or FOA / HOA source positions (mi 111, m2 113, m3 115, n 117, and ms 119) furthermore can be used to define a series of constellations which can be used to locate a listener. For example a first triangle or constellation can be defined by the positions mi 111, m2 113, m3 115, a second triangle or constellation defined by the positions mi 111, m3115, ms 119, a third triangle or constellation defined by the positions m2113, m3115, n 117 and a fourth triangle or constellation defined by the positions m3115, n 117, ms 119. In this example the listener position is within the first triangle or constellation.
[0009] An example of such a scene is a recording of a couple of musicians playing on a busy street. There are people around the musicians as well as people walking on the street. Some people on the street are listening and contribution to the audio scene by applauding or talking. The song is recorded using multiple Ambisonics mics, placed some distance apart. The positions of at least some of the musicians may be known. The scene is captured and then subsequently turned into a VR scene by the content creator. The content creator creates a scene description with the positions of the recording microphones and any of the known sources. The scene description is then encoded into a bitstream along with the audio and is provided to the listener for consumption for example via a VR headset.
[0010] Multi-point HOA rendering uses multiple HOA / FOA source positions and the listener position to be able to render audio scenes comprising HOA / FOA content which can be recorded or synthesized. Current rendering approaches typically employ a selection of the HOA / FOA sources immediately adjacent the current listener position, for example selection of HOA / FOA sources which define a constellation of locations within which the listener or user is currently located.
[0011] There can be situations, for example during streaming, where the renderer does not currently have access to the audio signals for HOA / FOA sources associated with the current listener or user location.
[0012] Summary
[0013] There is provided according to a first aspect an apparatus for generating binaural audio based on a listener position within an audio scene, the apparatus comprising means configured to: obtain, for at least one audio source at least one audio source position from the audio scene; obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; determine whether the listener position is located within the at least one internal region or at least one external region; determine whether the at least one audio source is available for rendering based on the determined listener position; determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and render the binaural audio signal based on the determined rendering mode.
[0014] The means configured to determine whether the at least one audio source is available for rendering based on the determined listener position may be further configured to determine an estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
[0015] The means configured to determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position may be configured to determine the rendering mode based on the estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
[0016] The means configured to determine whether the listener position is located within the at least one internal region or at least one external region may be configured to: determine whether the listener position is within a perimeter defined by the at least one audio source position; or determine whether the listener position is outside the perimeter defined by the at least one audio source position.
[0017] The means configured to determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region may be configured to determine at least one of: an interior rendering mode when all of the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one internal region; a delayed interior rendering mode when one of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one internal region; an external rendering mode when the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one external region; an external rendering mode when one or more of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region; a no rendering mode when all of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region.
[0018] The means configured to render the binaural audio signal based on the determined rendering mode may be configured to: identify, a neighbouring region to the at least one internal region, and at least one audio source defining the neighbouring region is available; determine at least two audio signals based on the at least one audio source.
[0019] The means configured to render the binaural audio signal based on the determined rendering mode may be further configured to determine at least two audio signals based on respective metadata associated with the at least one audio source.
[0020] The means may be further configured to obtain the at least one audio source from microphone arrangements, wherein each microphone arrangement may be at a respective position and comprises one or more microphones.
[0021] The means configured to obtain a listener position may be configured to obtain the listener position from a further apparatus.
[0022] According to a second aspect there is provided a method for an apparatus for generating binaural audio based on a listener position within an audio scene, the method comprising: obtaining, for at least one audio source at least one audio source position from the audio scene; obtaining a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; determining whether the listener position is located within the at least one internal region or at least one external region; determining whether the at least one audio source is available for rendering based on the determined listener position; determining a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and rendering the binaural audio signal based on the determined rendering mode. Determining whether the at least one audio source is available for rendering based on the determined listener position may further comprise determining an estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
[0023] Determining a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position may comprise determining the rendering mode based on the estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
[0024] Determining whether the listener position is located within the at least one internal region or at least one external region may further comprise: determining whether the listener position is within a perimeter defined by the at least one audio source position; or determining whether the listener position is outside the perimeter defined by the at least one audio source position.
[0025] Determining the rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region may comprise determining at least one of: an interior rendering mode when all of the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one internal region; a delayed interior rendering mode when one of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one internal region; an external rendering mode when the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one external region; an external rendering mode when one or more of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region; a no rendering mode when all of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region.
[0026] Rendering the binaural audio signal based on the determined rendering mode may comprise: identifying, a neighbouring region to the at least one internal region, and at least one audio source defining the neighbouring region is available; determining at least two audio signals based on the at least one audio source.
[0027] Rendering the binaural audio signal based on the determined rendering mode may comprise determining at least two audio signals based on respective metadata associated with the at least one audio source. The method may further comprise obtaining the at least one audio source from microphone arrangements, wherein each microphone arrangement is at a respective position and comprises one or more microphones.
[0028] Obtaining the listener position may comprise obtaining the listener position from a further apparatus. According to a third aspect there is provided an apparatus for generating binaural audio based on a listener position within an audio scene, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain, for at least one audio source at least one audio source position from the audio scene; obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; determine whether the listener position is located within the at least one internal region or at least one external region; determine whether the at least one audio source is available for rendering based on the determined listener position; determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and render the binaural audio signal based on the determined rendering mode.
[0029] The apparatus caused to determine whether the at least one audio source is available for rendering based on the determined listener position may be further caused to determine an estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
[0030] The apparatus caused to determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position may be caused to determine the rendering mode based on the estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
[0031] The apparatus caused to determine whether the listener position is located within the at least one internal region or at least one external region may be caused to: determine whether the listener position is within a perimeter defined by the at least one audio source position; or determine whether the listener position is outside the perimeter defined by the at least one audio source position.
[0032] The apparatus caused to determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region may be caused to determine at least one of: an interior rendering mode when all of the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one internal region; a delayed interior rendering mode when one of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one internal region; an external rendering mode when the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one external region; an external rendering mode when one or more of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region; a no rendering mode when all of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region.
[0033] The apparatus caused to render the binaural audio signal based on the determined rendering mode may be caused to: identify, a neighbouring region to the at least one internal region, and at least one audio source defining the neighbouring region is available; determine at least two audio signals based on the at least one audio source.
[0034] The apparatus caused to render the binaural audio signal based on the determined rendering mode may be further caused to determine at least two audio signals based on respective metadata associated with the at least one audio source.
[0035] The apparatus may be further caused to obtain the at least one audio source from microphone arrangements, wherein each microphone arrangement may be at a respective position and comprises one or more microphones.
[0036] The apparatus caused to obtain a listener position may be caused to obtain the listener position from a further apparatus.
[0037] According to a fourth aspect there is provided an apparatus for generating binaural audio based on a listener position within an audio scene, the apparatus comprising: means for obtaining, for at least one audio source at least one audio source position from the audio scene; means for obtaining a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; means for determining whether the listener position is located within the at least one internal region or at least one external region; means for determining whether the at least one audio source is available for rendering based on the determined listener position; means for determining a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and means for rendering the binaural audio signal based on the determined rendering mode.
[0038] According to a fifth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus for generating binaural audio based on a listener position within an audio scene to perform at least the following: obtain, for at least one audio source at least one audio source position from the audio scene; obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; determine whether the listener position is located within the at least one internal region or at least one external region; determine whether the at least one audio source is available for rendering based on the determined listener position; determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and render the binaural audio signal based on the determined rendering mode.
[0039] According to a sixth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for generating binaural audio based on a listener position within an audio scene to perform at least the following: obtain, for at least one audio source at least one audio source position from the audio scene; obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; determine whether the listener position is located within the at least one internal region or at least one external region; determine whether the at least one audio source is available for rendering based on the determined listener position; determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and render the binaural audio signal based on the determined rendering mode.
[0040] According to a seventh aspect there is provided an apparatus for generating binaural audio based on a listener position within an audio scene, the apparatus comprising: obtaining circuitry configured to obtain, for at least one audio source at least one audio source position from the audio scene; obtaining circuitry configured to obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; determining circuitry configured to determine whether the listener position is located within the at least one internal region or at least one external region; determining circuitry configured to determine whether the at least one audio source is available for rendering based on the determined listener position; determining circuitry configured to determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and rendering circuitry configured to render the binaural audio signal based on the determined rendering mode. According to an eighth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus for generating binaural audio based on a listener position within an audio scene to perform at least the following: obtain, for at least one audio source at least one audio source position from the audio scene; obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position; determine whether the listener position is located within the at least one internal region or at least one external region; determine whether the at least one audio source is available for rendering based on the determined listener position; determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; and render the binaural audio signal based on the determined rendering mode.
[0041] An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above.
[0042] A computer program comprising program instructions for causing a computer to perform the method as described above.
[0043] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.
[0044] An electronic device may comprise apparatus as described herein.
[0045] A chipset may comprise apparatus as described herein.
[0046] Embodiments of the present application aim to address problems associated with the state of the art.
[0047] Summary of the Figures
[0048] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:
[0049] Fig.1 shows schematically an example scene comprising microphone positions (mi to ms) arranged into microphone constellations and listener position p / ;
[0050] Figs.2a to 2c show schematically an example scene comprising microphone positions (mi to ms and a listener position p / outside or external to a microphone constellation;
[0051] Fig.3 shows schematically an example scene comprising microphone positions (mi to ms) arranged into microphone constellations and time varying listener positions p / which progress from one microphone constellation to a different microphone constellation;;
[0052] Fig.4 shows schematically apparatus suitable for performing conventional multi-point HOA / FOA rendering; Fig.5 shows an example direction parameter adjustment based on an exterior projection radius; Fig.6 shows an example apparatus suitable for implementing some embodiments;
[0053] Fig.7 shows a flow diagram of the operation of the example apparatus shown in Fig.6 with respect to rendering according to some embodiments;
[0054] Fig.8 shows a further flow diagram showing the operation of external rendering switching as shown in Fig.7 in further detail according to some embodiments; and
[0055] Fig.9 shows schematically an example device suitable for implementing the apparatus shown.
[0056] Embodiments of the Application
[0057] The concept as discussed herein in further detail with respect to the following embodiments is related to the rendering of audio scenes of 6DoF audio scenes captured with multiple higher-order Ambisonics microphones (HOA) where there is provided a method for rendering when audio data is not immediately available for at least one of the HOA sources required for rendering at a listener position to enable plausible rendering quality by performing exterior rendering while the HOA source audio data is not available and switching smoothly to regular rendering when the HOA audio data becomes available.
[0058] The improved plausibility of rendering by the proposed embodiments is enabled by restoring responsiveness to change in listener position, since without the application of the embodiments described herein, the listener position remains fixed to the previous position in the previous triangle formed by the HOA sources.
[0059] In the following examples the audio signal sets are generated by microphones (or microphonearrays). For example a microphone arrangement may comprise one or more microphones and generate for the audio signal set one or more audio signals. In some embodiments the audio signal set comprises audio signals which are virtual or generated audio signals (for example a virtual speaker audio signal with an associated virtual speaker location). In some embodiments the microphone-arrays are furthermore separate from or physically located away from any processing apparatus, however this does not preclude examples where the microphones are located on the processing apparatus or are physically connected to the processing apparatus.
[0060] Before discussing the concept in further detail we will initially describe in further detail some aspects of spatial reproduction. For example with respect to Figs.2a to 2c is shown an example of spatial playback where the listener or user position is outside of a constellation of FOA / HOA position or locations.
[0061] Fig.2a, for example differs from the example shown in Fig.1 wherein the listener position is located outside or external to any of the constellations defined by the location of available FOA / HOA positions. For example, as shown in Fig.2a, the listener position p / 201 is located outside of the first triangle or constellation defined by the positions mi 111, m2 113, m3 115. The rendering of audio signals when the listener position is external to the constellations has been described previously, for example in EP application EP21201766.9.
[0062] In summary, in normal operation (where the listener positon is located internally to the FOA / HOA source defined area, or in other words in an interior region), spatial parameters at the listener position are calculated based on weighted interpolation of spatial parameters calculated from the microphone signals of the microphones defining the triangle or constellation shape that the listener is located in. The closer a listener is to a FOA / HOA source or microphone, the more weight is given to the spatial parameters calculated at that source or microphone position.
[0063] However, when the listener moves outside of the source or microphone area (or in other words the listener is located in an external region with respect to at least one audio source position), and thus is not inside any of the constellation shapes or triangles, the weighted interpolation methods as such cannot be used. In the following examples the regions, interior and exterior are described with respect to constellation shapes or triangles which are defined with respect to the audio source positions. However in some embodiments the regions can be defined with respect to one or more audio source positions, for example by defining a source position radius or distance or shape relative to the audio source position or positions.
[0064] As described in EP21201766.9, the solution is to use the spatial parameters interpolated from the spatial parameters calculated at the microphones defining the outside edge that the listener is closest to and modifying the parameters based on the distance of the listener to the edge and an exterior rendering radius. Such a situation is shown in Fig.2b where the listener position p / 201 is located outside of the first triangle or constellation defined by the positions mi 111, m2 113, m3115 and projected to a position 211 on the outside edge between the positions mi 111, m2 113 (which is the closest edge with respect to the listener position p / 201).
[0065] Additionally as shown in Fig.2c the direction of arrival and diffuseness parameters (which form a subset of spatial parameters) are modified based on an assumption that audio sources 251 outside of the capture area are approximately at a distance from the microphone array area equal to the exterior rendering radius of the circle 221. Thus, the spatial parameters calculated at the projected position 211 are modified based on the distance of the listener from the projected position and a predetermined projection radius around the projected listener position. The audio rendered to the user is projected on the circle defined by the radius and thus the direction of arrival of audio changes (213, 223, 243) as the listener moves further from / towards the edge (shown by the position change from 211, 231, 241) giving an impression of 6D0F audio rendering.
[0066] To obtain the binaural signal at the listener position, the renderer, among other things, performs STFT processing on the FOA / HOA source or microphone audio signals. To avoid unnecessary processing, the STFT is only performed for the sources that comprise the constellation shape or triangle that the listenr is located in (or projected on). Since a STFT takes a few frames of audio to provide meaningful output, special care needs to be taken when the listener moves between constellation shapes or to new triangle.
[0067] Fig.3 shows an example of a situation where the listener moves to a new triangle or region. In this example, at time t-1, the listener is located at Pz (t - 1) 303 and within the triangle or constellation defined by the positions mi 111, ms 119, m3 115. STFT processing is therefore only performed for the microphone or source signals of the microphones or sources at positions mi 111, ms 119, m3115 comprising the triangle.
[0068] As the listener moves at time t, to the location p{(t) 301 which is within the triangle (or region) or constellation defined by the positions mi 111, m2113, m3115, STFT processing is started for the microphone or source signal of microphone or source located at m2. It takes, however, several frames of audio for the processing to provide meaningful output. For this reason, a delayed switching between constellations or triangles is performed where the listener position is kept at pi(t-1) until the STFT processing provides meaningful output with respect to the new triangle or constellation selection. A cross-fade can then be performed when the STFT for the audio signals associated with the microphone or audio source m2 becomes available to avoid a discontinuity or snap in the output audio.
[0069] Fig.4 shows an example apparatus for implementing current MPEG-I Audio system for 6D0F MultiPoint HOA rendering of audio scenes comprising HOA / FOA content (recorded or synthesized). This is an example implementation of rendering such a scene, and modifications to this implementation and other implementations are possible. This approach is described in further detail in GB application GB2007710.8 and GB2315366.1.
[0070] The apparatus is configured to receive as inputs positions of HOA / FOA sources p1 Vsor pt(where i=1 to Ns) 402, the listener position 404, the head related impulse responses (HRIRs) 400, orientation of the head orientation of the listener (yaw, pitch and roll) Oi406 and orientations of HOA / FOA sources o1 Vsor Ot (where i=1 to Ns) 408.
[0071] In some embodiments the renderer comprises a pre-precessor 401. The pre-processor 401 is configured to initialize the rendering after receiving information of the scene such as positions of HOA / FOA sources p1 Vs402 and also receiving head-related impulse responses (HRIRs) 400.
[0072] The pre-processor 401 is configured to obtain the positions 402 and performs Delaunay triangulation to provide a set of triangles T1 MT412 that partition the scene into triangle sections. This triangulation is for example shown in the Figures 1, 2a to 2c. The pre-processor 401 is also configured to convert the head-related impulse responses (HRIRs) 400 which are converted to frequency domain head-related transfer functions (HRTFs) 420. In some embodiments this can be implemented by a short-term Fourier transform. In a MPEG-I Immersive audio case, the alias-free STFT algorithm is used such as described in Pulkki, V., S. Del ikaris-Mani as, and A. Politis, Parametric Time-frequency Domain Spatial Audio. 2018: John Wiley & Sons, Incorporated. In some embodiments the HRTFs 420 can be used to calculate an Ambisonics-to-binaural transform matrix MH0A2binb'), for each frequency band b.
[0073] In some embodiments the renderer further comprises a position pre-processor 403. The position preprocessor 403 is configured to determine HOA / FOA sources that are close to the listener for later processing purposes.
[0074] During rendering, for each input frame j, the position pre-processor 403 is configured to take as an input the listener position pt404, the HOA / FOA source positions p1 Vs402 and the set of triangles T1 MT412 created in the pre-processor 401. Based on the input, the position pre-processor 403 determines interpolation weights wc(i,j) 416 for the HOA / FOA sources pt402. This is done by determining in which triangle Ththe listener is in and calculating the barycentric coordinates for the triangle Ttand the listener position. The barycentric coordinates are used as the weights. The weights sum to one and the closer the listener is to an HOA / FOA source, the higher the weight will be. At an edge of a triangle, the weight for the HOA / FOA source that is not part of that edge is 0. The weights wc(i,j) 416 and the triangle that the listener is in (active triangle) 7^ (J) 414 are provided as outputs.
[0075] In some embodiments the renderer comprises a spatial analyzer 409. The spatial analyzer 409 provides spatial metadata parameters for frames of HOA / FOA signals that positioned near the listener. These are later used, in the spatial metadata interpolatior 411, to estimate spatial metadata parameters at the listener position.
[0076] The spatial analyzer 409 is configured to receive as an input a frame (for example, 256 samples) of HOA / FOA source signals sESD(i,f) 410 for each source i and frame j and the current active triangle 7^ (J) 214. For each HOA / FOA source i, belonging to the triangle TA(j) 414, the spatial analyzer 409 performs STFT processing to obtain time-frequency domain signals S(i,j, k, h) 424, where k refers to a sub-frame (for example, 128 samples), and b the frequency bin. The spatial analyzer 409 then calculates spatial metadata from the time-frequency domain signals. The spatial metadata comprises the energy, direction information (azimuth and elevation) and diffuseness information (direct-to-total energy ratio) and is obtained as follows:
[0077] First, a signal analysis vector s i, j, k, h) is calculated:
[0078] * 1.0
[0079] M_ Sb'2(i,j, k) * 0.5774
[0080] * o.5774
[0081]
[0082] fc) * 0.5774
[0083] where Sb ci,j, k~) is the value in matrix fc) corresponding to channel c and frequency bin b. From the signal analysis vector, a signal intensity vector is calculated: 's-tCij, k, b) * s4(i,j, fc, h)T
[0084] i(i, j, k, b) = Re Si tiJ, k, b) * s2(i,j, k, b) •
[0085]
[0086] k, b) * s3(i,j,k, b)L where sc(i, j, k, h) denotes the complex conjugate of sci,j, k, b).
[0087] Signal energy is then calculated as follows:
[0088] 14
[0089] e(i,j, k, b) = fc, b) * sc(i,j, k, b)
[0090]
[0091] C = 1
[0092] An average intensity vector and average energy is calculated as follows:
[0093] Nsf
[0094] e(i,j,b) = —^ e(i,j,k,b)
[0095] sfk=l
[0096] Nsf
[0097] Ki,j, b) = — ) i(i,j, k, b)
[0098] NsfZ— i
[0099]
[0100] 1k=l
[0101] The direction data (azimuth and elevation) are calculated as follows:
[0102] §(i,j, b) = atan2(i2(i,j, b))
[0103] (p(i,j, b) = atan2 h) £)1 \
[0104]
[0105] where ini> j, b) is the nth element of the average intensity vector i(i,j, h), 6 i,j, b) is the azimuth and p i,j, b) is the elevation.
[0106] The direct-to-total energy ratio is calculated as follows:
[0107] M W’J’ W
[0108]
[0109] e(i,j, b)
[0110] The energy for the subframes k are then obtained as follows:
[0111] e(i,j, k, ) = e(i,j, b), k E l.. NSf
[0112] Thus the output also comprises q(i,j,b) 428, j(i,j,b) 430 r(i,j,b) 432 and e(i,j,k,b) 434.
[0113] The renderer furthermore comprises a spatial metadata interpolator 411 which is configured to provide an estimate of the spatial metadata at the listener position based on the spatial metadata for the HOA / FOA sources belonging to the triangle that the listener is in, and the interpolation weights calculated in the position pre-processing block.
[0114] First, the spatial metadata is converted into vector form:
[0115] — sin(0(i, j, h)) cos(<jo(i,j, h))
[0116] v(i,j, b) = sin(<p(i,j, h)) r(i,j, b)
[0117]
[0118] —cos((p(i>j> &)) cos(<jo(i, j, h))
[0119] The vectors are rotated according to the listener’s head and source orientations: v
[0120]
[0121] (i,j, b) = R head (j)R source (i,j)v(i,j,b)
[0122] The rotation matrix Rhead(j) is calculated based on the head orientation and the rotation matrix RSOurce(i>iscalculated based on the source orientation.
[0123] An interpolated spatial metadata vector is then calculated by a weighted average of spatial metadata vectors:
[0124] Ns
[0125] v(j> k,b) = ^ wc.(j, k)v(i,j, b)
[0126]
[0127] i=l
[0128] where the interpolation weights are the barycentric coordinates calculated in the position preprocessing block.
[0129] And finally, the interpolated spatial metadata vector is converted to spatial metadata parameters as follows:
[0130] Azimuth:
[0131] 0
[0132]
[0133] (j, k, b) = atan2(— v^j, k, b), k, h))
[0134] Elevation:
[0135] <p(j, k, b) = atan2 v2(J, k, b), y / ^—v-^fJ. k. b))2+ (— v3(j, fc, h))2)
[0136]
[0137] Direct-to-total energy ratio:
[0138] >
[0139] f
[0140]
[0141] (j, k, b) = ^(v^j. k. b))2+ (v2( / , fc, h))2+ (v3(j, fc, 6))2Energy:
[0142] Ns
[0143] e(j, k, b) = ^ wc.(j, k)e(i,j, k, b)
[0144]
[0145] i=l
[0146] These interpolated values 0(J, k, h) 436, <p(j, k, h) 438, r(j, k, h) 440 and e(j, k, h) 442 can be output.
[0147] The signal interpolator 405 is configured to provide an interpolated signal at the listener position, that is, an estimate of the signal at the listener position. This is used later (in the mixer 407) in conjunction with the interpolated metadata at the listener position to provide the final binaural output.
[0148] The interpolated signal Sb c(j, k~) 418 is obtained by taking the HOA / FOA signal that is closest to the listener and applying an EQ on the signal:
[0149] S
[0150]
[0151] b,c(j>k) = Geq(J,k,b)SbiC(,mc(J),j,k)
[0152] where mc(j) is index of the HOA source chosen for interpolation (closest to listener in most cases) and: / e(j, k, b) \
[0153] Geq(j, k, b) = min I+^Max •
[0154]
[0155] The value S (J, k, h) 418 can be output to the mixer 407.
[0156] The renderer further can comprise a mixer 407, which takes as an input the interpolated spatial metadata 436, 438, 440, 442 at the listener position the interpolated signal 418 at the listener position, the orientation of the head of the listener (yaw, pitch and roll) 0 / 406 and orientations of HOA / FOA sources o1 Nsor Ot (where i=1 to Ns) 408 and provides the binaural time-frequency domain output signal O(j,k,b) 422. The mixer 407 can create a signal covariance matrix from the interpolated spatial metadata which describes the desired (or target) spatial characteristics of the signal at the listener position. An optimal mixing algorithm is then used to obtain a mixing matrix that when multiplied with the interpolated signal, the resulting signal is inline with the desired spatial characteristics.
[0157] First, a binaural prototype signal is created from the interpolated signal:
[0158] B(j, b) — MH0A2bin(b') * RSh(D *
[0159]
[0160] Sb, Ncfl(j> 1) ■■■ Sb,Nch(j, Nsf) Where Rsh(j) is a rotation matrix taking into account the listener and HOA / FOA source orientations or rotation and MH0A2bin(b) is the Ambisonics to binaural matrix. In the mixer, the rotation matrix Rsh(j) is calculated based on the head orientation.
[0161] Then, a signal covariance matrix Cxis calculated for the prototype signal:
[0162] c
[0163]
[0164] rw(j^) = S(j, b)B(j, b)H
[0165] Recursive averaging is applied to get the signal covariance matrix for frame j:
[0166] Cxj,b) = (1 - d)Cxew(j, b) + dCxj - l,b)
[0167] Where d = 0.9.
[0168] Next, a signal covariance matrix Cyis calculated from the interpolated spatial metadata at the listener position. First the direct portion of Cyis calculated:
[0169] Nsf
[0170] Cdir^tj b) = ^ eQ, k, b}fQ, k, b}H(b, d}H^(b, (T)
[0171]
[0172] k=l
[0173] where W(h, d) refers to the HRTF value at frequency bin b, in direction.
[0174] Second, the diffuse portion of Cyis calculated.
[0175] Nsf
[0176]
[0177] k=l
[0178] where:rlNdiiH(b’d^b’ ^
[0179] C
[0180]
[0181] ^ = - Nd~ - Cyis then obtained as follows:
[0182] C
[0183]
[0184] yew(j, b) = tfirect(j,b) + Cyiffuse(j, b)
[0185] And recursive averaging:
[0186] Cy(j, b) = (1 - d C"ew(j, b) + dCy(j - 1, 6)
[0187] where d = 0.9.
[0188] The signal covariance matrices are then used to obtain mixing matrices which are used to obtain the binaural output like so:
[0189] O
[0190]
[0191] (j, k, 6) = M(j, k, 6) * B(j — 1, k, 6) * D(j, 6)
[0192] Where D(j, b) is a decorrelated time-frequency domain signal obtained from a buffer of previous binaural signals B.
[0193] The matrices M(j, k, 6) and Mr(j, k, 6) are obtained from the optimal mixing procedure outlined in Vilkamo, J., Backstrbm, T., and Kuntz, A. (2013). Optimized covariance domain framework for timefrequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. After applying the mixing matrices on the binaural prototype signal B, the result output 0 422 has the spatial characteristics of the spatial metadata at the listener position.
[0194] The renderer can furthermore comprise an output processor 413 which is configured to receive the obtained binaural output signal O(j, k, 6) and perform an inverse STFT on it to produce the final time domain binaural output signal sout(i,j) 446.
[0195] A summary of the example operations of the renderer according to some embodiments is as follows: Obtaining HRIR and positions p,.;
[0196] Pre-processing to generate TI.. NT and HRTFs;
[0197] Obtaining p / in Figure 3 by 305;
[0198] Position-preprocessing to generate TA(J) and wc(i,j);
[0199] Obtaining SESD(i,j)
[0200] Spatial analyzing to generate S(i,j,k,b), e(i,j,k,b), q(i,j,b), j(i,j,b), r(i,j,b);
[0201] Spatial metadata interpolating to generate e(j,k,b) q(j,k,b) j(j,k,b), r(j,k,b);
[0202] Signal interpolating to generate S(j,k,b);
[0203] Generation by mixing O(j,k,b);
[0204] Outputting of Spatialized audio Sout(ij) (e.g., binaural, surround loudspeakers, Ambisonics).
[0205] When the listener moves outside of the microphone area, the interpolated spatial metadata
[0206]
[0207] (0(j, k, b), k, b), r(j, k, 6)) is adjusted based on a listener distance to the microphone area and the exterior rendering radius to give the listener a 6DoF listening experience outside of the microphone area as described above. Thus when implementing exterior rendering there are modifications to the position preprocessing and spatial metadata interpolation blocks, as described in the following.
[0208] In the position pre-processor 403, the triangle in which the listener is in is determined. This is done by calculating barycentric coordinates for the listener position for all triangles TI--TNTfound by the Delaunay triangulation. The listener is inside a triangle if all barycentric coordinates are positive. In the case there are no triangles for which all barycentric coordinates are positive, it is determined that the listener is outside of the microphone area. In such cases, the following is performed.
[0209] Firstly the closest HOA source (mk) and edge (eki) to the listener position (pL) is calculated. Secondly a distance to the closest HOA sources to the listener is obtained by calculating the squared horizontal distance from the listener to the HOA sources as follows:
[0210]
[0211] PL, Z) " T (Pmk,y Pt,y)
[0212] The closest edge is calculated by finding the distance Dklfrom the listener position pLto all edges
[0213] Pt.m / j ’ Yl-ki, if dotv n— 0
[0214] -> / PL, mi, ’ ^kl
[0215] Dij = < PL,mk■ nkt~dotv n
[0216] 1, else
[0217]
[0218] dotv’n~ dot^f
[0219] where pL mfcis the position of the listener in HOA source mkspace:
[0220] PL,mkPL Pmk
[0221] and where dotv nis the dot product between the normalized edge vector vk(and the normal of the edge n^-.
[0222] dotv nvkl• Tlki.
[0223] If the closest HOA source to the listener is closer than any of the outside edges, the listener position is projected to the HOA source position:
[0224] P.proj Pmciosest
[0225] When this is not the case, the listener is projected on to the closest exterior edge (
[0226]
[0227] e0):
[0228] P
[0229]
[0230] L,pro j WklPmk" T (1 Wk()pm /
[0231] The weights wiare obtained by:
[0232] . PL,mk’ ^kl
[0233] , if dotv n— 0
[0234] |Vfcj|
[0235] ~ I PL,mk’ Tlfti \
[0236] wki= PL,mk■ Vkl ~ Idotv nI
[0237] 1 - — - 11 vkl|, else
[0238]
[0239] dotv’n~ dotv n In the spatial metadata interpolator 411, the interpolation weights are calculated differently when doing exterior rendering and an adjustment to the obtained spatial metadata is performed as described below:
[0240] If the listener was projected on a HOA source, the interpolation weight for that source is set to 1 and the interpolation weight for all other sources is set to 0.
[0241] If the listener was projected to an outside edge, the interpolation weights are calculated as follows (for frame j):
[0242] w(k,j) = wM
[0243]
[0244] = 1 - wkl
[0245] Where k and I are the indices of the HOA sources defining the edge that the listener is closest to. Spatial metadata is interpolated as described above, for the non-exterior rendering case (except with the interpolation weights calculated as directly above). After interpolation the spatial metadata can then be adjusted based on the geometry shown in Fig.5, where the projected listener position pL>proj501 from the pL505 is shown with a Rprojthe exterior projection radius 509 and DOA vector uD(M509 and modified DOA vector uDOA mod507. The operations for adjusting the spatial metadata can be as follows:
[0246] First a DOA vector is calculated (for each frequency band) from the interpolated spatial metadata:
[0247] — sin($) cos(<p)
[0248] UDCM sin(<p)
[0249]
[0250] — cos($) cos(<p)
[0251] where 0 and <p are the interpolated azimuth and elevation parameters. Note that the audio frame ( / )and frequency band indices ( / ) are left out here for clarity. A weighted DOA vector is also calculated:
[0252] U
[0253]
[0254] DCM —UDO A Rproj
[0255] where Rprojis the exterior projection radius. A modified DOA vector is calculated:
[0256] ^DOA,mod PL,proj " T UQO 1 PL
[0257] A unit length version is also calculated:
[0258] > ^DOA,mod
[0259] ^DOA,mod ~ 1 F
[0260]
[0261] I^DOA,mod |
[0262] A modified direct-to-total energy ratio is also calculated:
[0263] rm0d = min (r, — - - )
[0264]
[0265] Kpro j
[0266] A directional weighting is applied to the modified parameters to modify parameters more heavily from sound sources outside the capturing region, while applying less modifications to sources inside the capturing region: W
[0267]
[0268] dir 22(1 + ^L,proj ’ ^DO / 1)
[0269] Where nL projis a “listener normal” which is a vector pointing away from the HOA source area.
[0270] ^DOA, mod, weighted ^dir^DOA,mod " F (1 ^dir)^DOA
[0271]
[0272] ^mod, weighted mtn (1, WdirTmod”F (1
[0273] Modified spatial metadata values are then obtained from viDOAimodiWeightedand processing is continued as in the non-exterior rendering case:
[0274] Azimuth:
[0275] 6 atan2 ^DOA, mod, weighted^’ ^DOA, mod, weighted3Elevation:
[0276] (p atan2 ^DOA,mod,weighted2' -J ^DO A, mod, weighted " F ^DOA, mod, weighted
[0277]
[0278] Direct-to-total energy ratio:
[0279]
[0280] T-mod, weighted
[0281] An example system employing some embodiments is described in Fig.6. The figure illustrates an end to end system overview for an audio scene comprising multiple HOA sources, which is rendered according to the above examples and suitable for implementing the embodiments which are described hereafter. The renderer receives the scene description and audio bitstreams and performs rendering accordingly. The system can comprise a content creator 600 which can be implemented on any suitable computer or processing device. The content creator 600 comprises an (MPEG-I Immersive audio) encoder 601 which is configured to receive the audio scene description 602 and the audio signals or data 604. The audio scene description 602 can be provided in the MPEG-I Immersive audio Encoder Input Format (EIF) or in other suitable format. Generally, the audio scene description contains an acoustically relevant description of the contents of the audio scene, and contains, for example, the scene geometry as a mesh or voxel, acoustic materials, acoustic environments with reverberation parameters, positions of sound sources, and other audio element related parameters such as whether reverberation is to be rendered for an audio element or not. The MPEG-I Immersive audio encoder 601 is configured to output encoded data 606.
[0282] The content creator 600 furthermore in some embodiments comprises a MPEG-H encoder 603 and generates MPEG-H audio bitstream 605.
[0283] The MPEG-H audio bitstream 605 and 6DoF audio bitstream 606 in some embodiments can be streamed to end-user devices or made available for download or stored. Additionally the system comprises a server 610 configured to obtain the bitstreams, and store them (in bitstream storage 611) and supply them (for example as a six degrees of freedom (6-DoF) audio bitstream 612) to the player 620.
[0284] The relevant (6DoF audio) bitstream 612 is retrieved by the player 620. In some embodiments other implementation options are feasible such as broadcast, multicast.
[0285] The player 620 in some embodiments comprises a playback device 621 configured to obtain or receive the 6DoF audio bitstream 612, and furthermore can be configured to receive or otherwise obtain the 6 DoF tracking information (listener orientation or position information) 628 from a suitable listener user interface, for example from the head mounted device (HMD) 629. These can for example be generated by sensors within the HMD 629 or from sensors in the environment sensing the orientation or position of the listener.
[0286] In some embodiments the playback device 621 comprises a bitstream parser 623 configured to obtain the encoded bitstream 612 and decode these in an opposite or inverse operation to the encoders 601, 603 to generate audio and metadata 622 which can be passed to a MPEG-I Immersive audio renderer 625.
[0287] In some embodiments the playback device 621 comprises the MPEG-I Immersive audio renderer 625 configured to implement the rendering operations as described above and hereafter and generate audio output signals 630 which can be output to the head mounted device 629.
[0288] The playback device 621 can be implemented in different form factors depending on the application. In some embodiments the playback device is equipped with its own listener position tracking apparatus or receives the listener position information from an external apparatus. The playback device can in some embodiments be also equipped with headphone connector to deliver output of the rendered binaural audio to the headphones.
[0289] Additionally the player 620 comprises a data handler 627. The data handler 627 communicates with the MPEG-I audio renderer 625 via an availability API 624 and utilization API 626. The Availability API 624 can be configured to indicate what audio is available to the MPEG-I audio renderer 625 while the utilization API 626 configured to indicate which audio is to be requested.
[0290] The data handler 627 is configured to communicate with the server using request audio 632 channel to request audio signals which can then be provided to the player 620 from the server 610.
[0291] In the multi-point ambisonics 6DoF rendering operations described above during “normal operation”, only the HOA sources which comprise the triangle that the user is located inside are used for processing. A scene might comprise, for example, 12 HOA sources, but only a subset of those (e.g., typically three) are utilized for rendering.
[0292] Thus, time-frequency domain conversion (STFT) is required to be performed only for these three sources out of the total number of sources (12). This enables computational complexity savings to be achieved, with the potential issue that when the listener moves from one triangle to another then there is noticeable delay before any relevant or meaningful time-frequency domain conversion for an HOA source needing to employed in the new triangle can be generated. Thus when the listener moves to a new triangle, there would be a wait for the timefrequency domain conversion to start providing meaningful output.
[0293] For example in MPEG-I immersive audio specification applications, the initialization time for the timefrequency domain conversion is 1792 samples, or 7 blocks of 256 samples in the MPEG-I immersive audio renderer. This is around 37ms of delay. For this duration, as discussed above, the listener position is kept at the previous position before switching triangles. In other words, the switch to a new triangle of HOA sources is delayed. This is not usually noticeable to the listener. Furthermore, the duration of the initialization is known or deterministic, which makes the handling of the delay in switching triangles simpler.
[0294] The above assumes that the audio which is input to the time-frequency domain conversion is immediately available when the listener moves and changes triangles.
[0295] This, however, may not always be true. In some cases, an implementation of the player 620 housing the renderer 625 may, for example, limit the bitrate requirements by only retrieving / downloading / streaming audio for the HOA sources comprising the triangle the listener is currently located in (or a subset of the HOA sources in the audio scene).
[0296] The player 620 receives information about which HOA sources are being used at any given time from the MPEG-I immersive audio renderer 625 through the audio utilization interface 626 and can be configured to determine which audio data to obtain.
[0297] Thus in examples where the audio data for an HOA source is not immediately available for the timefrequency domain conversion an extra delay of unknown duration is added to the switching of the triangle.
[0298] If this delay becomes significant then the listener will hear it as a discontinuity in the audio responsiveness to listener movement and as a mismatch between the audio the listener hears and any VR visual elements.
[0299] Thus, instead of the known delay of 1792 samples, there can be an extra delay caused by the player first requesting audio for an HOA source and the server starting to provide the HOA source audio data. This extra delay can depend on network conditions and rather than the delay of around 37ms, the delay may be more than 1 second for example.
[0300] More generally a renderer for performing 6DoF HOA rendering as described in the following embodiments is one which is configured to render when audio data from all the relevant HOA sources is available as well as when the audio data for a subset of audio sources is not available.
[0301] The concept as discussed in the following embodiments relates to (binaural) rendering of 6DoF audio scenes captured with multiple higher-order Ambisonics microphones (HOA) where there is provided apparatus and methods for rendering even when audio data is not immediately available for at least one of the HOA sources required for rendering at a listener position. The scene, as shown in the examples presented herein can be captured with higher-order ambisonics (HOA), but could also be captured by first-order ambisonics (FOA) microphones or other types of microphone arrays (such as mobile phones with multiple microphones). A conversion to ambisonics from such microphone arrays could then be implemented and the converted signals could be processed in a manner as described herein.
[0302] These embodiments aim to produce plausible rendering quality audio signals by performing exterior rendering while the HOA source audio data is not available and furthermore to switch smoothly to regular rendering when the HOA audio data becomes available.
[0303] The embodiments thus aim to improve plausibility of rendering by restoring responsiveness to change in listener position, since without this the listener position will remain fixed to the previous position in the previous triangle formed by the HOA sources.
[0304] Furthermore the embodiments enable a smooth switching to regular rendering methods by waiting for the required processing delay with the interior or regular rendering as well handling any issues with respect to the availability of audio data for the required HOA sources during interior or regular rendering. The switching from the exterior rendering back to regular rendering can be implemented when sufficient audio data is available for regular rendering and the listener position is updated without any drastic changes. This differs from conventional rendering where the listener position would jump from the earlier triangle to the new triangle with a sudden discontinuity.
[0305] In the present disclosure the term regular rendering refers to the rendering approach which utilizes 3 (HOA) sources for 6D0F rendering, thus providing superior localization. The terms regular rendering and interior rendering have been used interchangeably herein.
[0306] In some embodiments, the regular or interior rendering mode can also utilize audio source position information in the audio scene to leverage informed source rendering.
[0307] In some embodiments, the exterior rendering method or mode can employ audio source position information in the audio scene to leverage informed source rendering.
[0308] There are several implementation alternatives, such as, but not exclusively when a listener moves to a new constellation or triangle and audio data (for example at least one audio signal from the audio source) is not available at the renderer:
[0309] Employ or perform exterior rendering such as described herein as if the scene only comprises the (HOA) source(s) that are available for which the audio data is available for the duration until the audio data is not available (usually the HOA sources defining previous triangle based on the listener position); and then Resume or perform interior rendering (the 6D0F HOA rendering approach such as detailed herein) when audio data for all the HOA sources based on the listener position.
[0310] A further implementation can be also when the audio data is not available and the same exterior rendering / resume interior rendering modes but decide based on information from the player on an estimated time of availability for the missing audio whether to switch to exterior rendering (or exterior rendering mode) or use the delayed switching mode (with an extra delay for the purpose of waiting for the audio to be available).
[0311] Additional implementations can be i n scene previews where the player may offer to the listener scene previews, i.e. scenes with reduced number of HOA sources (only 1, for example) for the purpose of quickly browsing the scenes. In such implementations:
[0312] If the user decides to continue with a scene past the preview, switching to using more HOA sources is performed as described above. For this case, the scene metadata may include instructions for the renderer that which HOA sources are rendered initially (preview / default source indicator bit).
[0313] In the first part of this section, the use case for partial audio data availability for the HOA sources that are currently required for interior or regular 6D0F rendering is described.
[0314] With respect to Fig.7 is shown an example flow chart showing the implementation of some embodiments with respect to the operation of the renderer.
[0315] For example Fig.7 shows by 701 the operation of obtaining the source position information. The positions of the (HOA) sources comprising the scene can in some embodiments be obtained from the metadata bitstream describing the scene. For example as shown in Fig.6 the bitstream parser 623 can be configured to generate the audio and metadata 622 which is passed to the MPEG-I audio renderer 625. In some embodiments, as described above, during the initialization of the renderer, the (HOA) source positions are triangulated (for example based on a Delauney triangulation) to obtain a triangulation of the (HOA) source positions. The source position information comprises the triangulation information and the position information of the HOA sources. This facilitates the renderer to determine based on the listener position, which HOA sources form a triangle for a given listener position. (As indicated above the regions can be defined with respect to the source position or source positions in any suitable manner according to the implementation embodiment).
[0316] Furthermore as shown in Fig.7 by 703 is the operation of obtaining listener position information. The playback device 621 such as shown in Fig.6 can in some embodiments be configured to generate or determine or otherwise obtain the listener position. For example the HMD 629 can be equipped with location sensors which determine the position of the listener within the room and map this position within the audio scene being presented to the listener. However the listener position determination can be implemented based on any suitable means, for example a user input device operated by the listener or the content creator determining the position within the audio scene.
[0317] Additionally the operation of obtaining the listener position information further comprises determining which triangle that the listener is located. Consequently, the renderer is able to determine which of the (HOA) sources form the triangle that the listener is located.
[0318] As further shown in Fig.7 by 705 is the operation of determining (HOA) sources required for interior or regular 6DoF rendering information, based on the (HOA) source position information and the listener position information. The renderer can determine which (HOA) sources are required for “normal” interior or regular rendering operation. This for example, in most cases, the renderer is configured to determine that the set of (HOA) sources that are required are the (HOA) sources comprising the triangle that the listener is located. In some embodiments there could be implemented other schemes for interpolation of spatial metadata at the listener position. For example in the example described herein, only the sources comprising the triangle have a greater than 0 interpolation weight and are thus used for interpolation. However in some embodiments the scene could be partitioned in some other manner or shape other than triangles and have a weight calculation system for this partition method.
[0319] Furthermore as shown in Fig.7 by 707 is the operation of obtaining or otherwise determining (HOA) source audio data availability information. For the set of (HOA) sources for which audio data is required for, availability information can be obtained via the audio Availability API or interface 624 such as shown in Fig.6. The Availability API or interface 624 is able to inform the renderer of which audio is immediately available for processing. In some embodiments, the availability information contains information (for example a time estimate or estimated delay) about when the audio for an (HOA) source will become available. The playback device 621 or player 620 may be configured to obtain or determine this information based on network status information such as network bitrate, network congestion (data that is still in the network), round-trip-time (RTT), etc.
[0320] Fig.7 furthermore by 709 shows the operation of obtaining the (HOA) source audio spatial processing availability information. In some embodiments the spatial processing availability information is obtained for the set of (HOA) sources for which audio data is required. In other words the renderer 626 can be configured to check whether the STFT processing for a source has been performed long enough to provide meaningful output. For implementation embodiments that do not perform STFT processing, for example, with zero latency, this delay can be assumed to be zero.
[0321] Finally Fig.7 by 711 shows the operation of determining the 6DoF rendering mode to be performed based on the determined (HOA) source required for interior or regular 6DoF rendering, (HOA) source audio data availability and (HOA) source audio processing availability information. In other words based on the information obtained or determined in the earlier operations, the renderer makes a decision on a rendering mode to be used for providing the spatial audio (for example binaural) output at the listener position. In some embodiments the renderer is configured to select or choose when there is a determined lack of availability (either audio data availability or processing availability) between performing either a “normal” or interior rendering with a delayed switching operation such as described above or an exterior rendering with the rendering performed as if the scene comprises only the HOA sources for which the audio data are currently available.
[0322] For example in some embodiments the selection of the rendering mode can be implemented according to a table, for example as presented below:
[0323] Availability for all required (HOA) STFT status for all required (HOA) Rendering sources sources mode
[0324] All available Ready Interior or Regular All available Not ready for at least one (HOA) source Delayed switch
[0325] Not available for at least one (HOA) - Exterior source rendering No (HOA) sources available - No
[0326] rendering
[0327]
[0328] In some embodiments when switching back to interior or regular rendering from delayed switch or exterior rendering, a crossfade of the interpolation weights can be performed.
[0329] For the duration of the crossfade, interpolation weights can be determined or calculated in the interior or regular rendering mode as well as either the delayed switch or exterior rendering mode (For example such as described above with respect to the application of a Cross-fade after a delayed switch). After the cross-fade only the interior or regular rendering mode is employed. If the renderer switches to exterior rendering mode, the renderer waits for both the audio to become available and for the STFT to provide meaningful output before switching back to the normal interior mode (which can also be considered to be a regular or common mode) rendering.
[0330] As mentioned above, in some embodiments, the audio availability information may include an estimate of delay or when an audio source will become available. For example, the audio availability information may indicate that the audio for HOA source 1 will be available in 1 second. The renderer may then, based on this information, and a time threshold decide to use the delayed switch mode even though the audio is not immediately available. The threshold may be set by the playback device, it can be a hard-coded value in the renderer or in some cases it could be set in the scene bitstream (set by the content creator).
[0331] The operation as shown in Fig.7 by 711 is shown in Fig.8 in further detail. The flow diagram described in Fig.8 therefore describes one example implementation method for the decision process for switching to the exterior rendering mode. In Fig.8 the solid lines describe the forward path and the dashed lines describe the reverse or feedback path.
[0332] With respect to Fig.8 by 801 is the first operation of, receiving or obtaining information. The information can be about HOA source positions, listener position, HOA source audio data availability and current rendering mode information.
[0333] Then as shown in Fig.8 by 803 is the operation of determining mode selection information. The determining of the mode selection information can comprise, determining, based on the current rendering mode, the processing delay. For example, the use of time-frequency domain transform has a higher delay compared to the time domain processing method. Subsequently, the delay for availability of audio data for the HOA sources currently required for processing is determined. This delay can be considered zero in case the data is already available with the player.
[0334] Additionally as shown in Fig.8 by 805, there is a determined a cumulative delay. The cumulative delay can be a sum of the processing method dependent delay and audio data availability related delay.
[0335] Furthermore, as shown in Fig.8 by 807 is the operation of obtaining or determining a delay tolerance threshold value. The delay tolerance threshold value can be obtained by the renderer implementations with preset values or obtained via the player in situations where the threshold is set or defined based on the player preferences. In some embodiments of the implementation, the threshold value can also be obtained from the audio rendering bitstream. For example the threshold value can be defined or set by the content creator based on the sensitivity of the content.
[0336] Then as shown in Fig.8 by 809 the operation of comparing the delay tolerance for switching modes with the total expected delay.
[0337] If the delay tolerance is less than the switching delay threshold then, as shown in Fig.8 by 811, maintain the interior or regular rendering mode as described herein. Then a feedback path can return the operation to 801.
[0338] If the delay tolerance is more than the switching delay threshold, then, as shown in Fig.8 by 813, switch to the exterior rendering mode.
[0339] Following the switch to the exterior rendering mode then, as shown in Fig.8 by 815 is the operation of determining whether the ‘new’ source delayed or is not available. When the comparison determines only one HOA source audio data is available then when switching to the exterior rendering mode a one source exterior rendering mode is used and the feedback path can return the operation to 801.
[0340] But when the comparison determines two HOA source audio data is available then when switching to the exterior rendering mode a two source exterior rendering mode is used and the feedback path can return the operation to 801.
[0341] In some embodiments as discussed previously in addition to embodiment being implemented in normal rendering applications, a further application can be in scene previewing. In this use case, the user or listener is able to sample VR scenes and decide which of them they want to start consuming. The playback device, in such embodiments, for reduced bandwidth, is configured to provide the user with reduced versions of the scenes while the listener is previewing them. The reduced version could, for example, be a version of the scene where only three (HOA) sources are present (a single triangle) or perhaps only a single (HOA) source. When previewing the scene the listener may move freely in the scene, but such that only the (HOA) sources belonging to the reduced set are indicated as available for example through the use of the availability API or a suitable interface. If the user or listener decides that a scene is to their liking, the renderer can then be configured to switch to a normal rendering as described above. The (HOA) sources used for the reduced set may be indicated in the bitstream (for example set by the content creator).
[0342] The following tables show example bitstream signaling which could be employed to control switching between rendering modes (for example switching to an exterior rendering mode)
[0343] Syntax No. of bits Mnemonic hoaGroups()
[0344] {
[0345] hoaGroupsCount = GetCountOrIndex();
[0346] for (int i = 0; i < hoaGroupsCount; i++) {
[0347] hoaGroupId = GetID();
[0348] hoaGroupHasRegion; 1 bslbf
[0349] if (hoaGroupHasRegion) {
[0350] hoaGroupRegionId = GetID();
[0351] }
[0352] coSourceCount = GetCountOrIndex();
[0353] for (int j = 0; j < coSourceCount; j++) {
[0354]
[0355] coSourceId = GetID();
[0356] }
[0357] FreqBandConfig();
[0358] enableNoAudioExteriorSwitch; 1 bslbf hoaGroupHasInformedSources; 1 bslbf if (hoaGroupHasInformedSources == True) {
[0359] informedSourceCount; 16 uimsbf maxSimulinformedSources; 8 uimsbf for G = 0; j < informedSourceCount; j++) {
[0360] InformedSourceInfoStruct()
[0361] }
[0362] }
[0363] }
[0364]
[0365] In some embodiments, an additional parameter is signaled to indicate if the adaptive switch to exterior rendering mode in case the audio data is not available to one or more HOA sources that are required for 6D0F HOA rendering.
[0366] enableNoAudioExteriorSwitch equal to 1 indicates the renderer can switch to exterior rendering mode in case audio data is not available for at least one HOA source. A value equal to 0 indicates, the renderer will always continue in interior or regular rendering mode if the listener is within a triangle formed by the HOA sources.
[0367] In another embodiment of the implementation, additional value of exterior switch tolerance can be provided to the renderer by the content creator via the bitstream.
[0368] Syntax No. of bits Mnemonic hoaGroups()
[0369] {
[0370] hoaGroupsCount = GetCountOrIndex();
[0371] for (int i = 0; i < hoaGroupsCount; i++) {
[0372] hoaGroupId = GetID();
[0373] hoaGroupHasRegion; 1 bslbf
[0374] if (hoaGroupHasRegion) {
[0375]
[0376] hoaGroupRegionId = GetID();
[0377] }
[0378] coSourceCount = GetCountOrIndex();
[0379] for (int j = 0; j < coSourceCount; j++) {
[0380] coSourceId = GetID();
[0381] }
[0382] FreqBandConfig();
[0383] enableNoAudioExteriorSwitch; 1 bslbf if (enableNoAudioExteriorSwitch) {
[0384] exteriorSwitchThreshold = GetCountOrIndex
[0385] 0;
[0386] }
[0387] hoaGroupHasInformedSources; 1 bslbf if (hoaGroupHasInformedSources == True) {
[0388] informedSourceCount; 16 uimsbf maxSimulInformedSources; 8 uimsbf for G = 0; j < informedSourceCount; j++) {
[0389] InformedSourceInfoStruct()
[0390] }
[0391] }
[0392] }
[0393]
[0394] exteriorSwitchThreshold indicates the delay in milliseconds cumulative delay as a guide for the renderer to switch to exterior rendering from interior or regular rendering mode. In some implementation embodiments, the threshold can be in terms of number of audio segments processed by the renderer at a given time. For example, the audio segment can be 256 samples for 48KHz sampled audio stream.
[0395] In another implementation embodiment, the default mode is the single source rendering and the switch to the regular or interior rendering is performed when the listener prefers to have a higher quality localization and greater freedom of movement in terms of listener position translation extent and listener position translation speed.
[0396] With respect to Fig.9 an example electronic device which may be used as the computer, encoder processor, decoder processor or any of the functional blocks described herein is shown. The device may be any suitable electronics device or apparatus. For example in some embodiments the device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.
[0397] In some embodiments the device 1600 comprises at least one processor or central processing unit 1607. The processor 1607 can be configured to execute various program codes such as the methods such as described herein.
[0398] In some embodiments the device 1600 comprises a memory 1611. In some embodiments the at least one processor 1607 is coupled to the memory 1611. The memory 1611 can be any suitable storage means. In some embodiments the memory 1611 comprises a program code section for storing program codes implementable upon the processor 1607. Furthermore in some embodiments the memory 1611 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1607 whenever needed via the memory-processor coupling.
[0399] In some embodiments the device 1600 comprises a user interface 1605. The user interface 1605 can be coupled in some embodiments to the processor 1607. In some embodiments the processor 1607 can control the operation of the user interface 1605 and receive inputs from the user interface 1605. In some embodiments the user interface 1605 can enable a user to input commands to the device 1600, for example via a keypad. In some embodiments the user interface 1605 can enable the user to obtain information from the device 1600. For example the user interface 1605 may comprise a display configured to display information from the device 1600 to the user. The user interface 1605 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1600 and further displaying information to the user of the device 1600.
[0400] In some embodiments the device 1600 comprises an input / output port 1609. The input / output port 1609 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1607 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
[0401] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802. X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The transceiver input / output port 1609 may be configured to transmit / receive the audio signals, the bitstream and in some embodiments perform the operations and methods as described above by using the processor 1607 executing suitable code.
[0402] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0403] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media, and optical media.
[0404] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
[0405] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0406] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0407] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.
Claims
CLAIMS:
1. An apparatus for generating binaural audio based on a listener position within an audio scene, the apparatus comprising means configured to:obtain, for at least one audio source at least one audio source position from the audio scene; obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position;determine whether the listener position is located within the at least one internal region or at least one external region;determine whether the at least one audio source is available for rendering based on the determined listener position;determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; andrender the binaural audio signal based on the determined rendering mode.
2. The apparatus as claimed in claim 1, wherein the means configured to determine at least one of:whether the at least one audio source is available for rendering is further configured to determine an estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable; andthe rendering mode based on the estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
3. The apparatus as claimed in any of claim 1 or 2, wherein the means configured to determine whether the listener position is located within the at least one internal region or at least one external region is configured to:determine whether the listener position is within a perimeter defined by the at least one audio source position; ordetermine whether the listener position is outside the perimeter defined by the at least one audio source position.
4. The apparatus as claimed in any of claims 1 to 3, wherein the means configured to determine the rendering mode is configured to determine at least one of:an interior rendering mode when all of the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one internal region;a delayed interior rendering mode when one of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one internal region;an external rendering mode when the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one external region;an external rendering mode when one or more of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region; anda no rendering mode when all of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region.
5. The apparatus as claimed in claims 1 to 4, wherein the means configured to render the binaural audio signal is configured to at least one of:identify, a neighbouring region to the at least one internal region, and at least one audio source defining the neighbouring region is available;determine at least two audio signals based on the at least one audio source; anddetermine at least two audio signals based on respective metadata associated with the at least one audio source.
6. The apparatus as claimed in any of claims 1 to 5, wherein the means is further configured to obtain the at least one audio source from microphone arrangements, wherein each microphone arrangement is at a respective position and comprises one or more microphones.
7. The apparatus as claimed in any of claims 1 to 6, wherein the means configured to obtain a listener position is configured to obtain the listener position from a further apparatus.
8. A method for an apparatus for generating binaural audio based on a listener position within an audio scene, the method comprising:obtaining, for at least one audio source at least one audio source position from the audio scene; obtaining a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position;determining whether the listener position is located within the at least one internal region or at least one external region;determining whether the at least one audio source is available for rendering based on the determined listener position;determining a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; andrendering the binaural audio signal based on the determined rendering mode.
9. The method as claimed in claim 8, wherein determining whether the at least one audio source is available for rendering further comprises determining an estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
10. The method as claimed in claim 9, wherein determining the rendering mode comprises determining the rendering mode based on the estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
11. The method as claimed in any of claims 8 to 10, wherein determining whether the listener position is located within the at least one internal region or at least one external region further comprises:determining whether the listener position is within a perimeter defined by the at least one audio source position; ordetermining whether the listener position is outside the perimeter defined by the at least one audio source position.
12. The method as claimed in any of claims 8 to 11, wherein determining the rendering mode comprises determining at least one of:an interior rendering mode when all of the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one internal region;a delayed interior rendering mode when one of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one internal region;an external rendering mode when the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one external region;an external rendering mode when one or more of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region;a no rendering mode when all of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region.
13. The method as claimed in claims 8 to 12, wherein rendering the binaural audio signal based on the determined rendering mode comprises at least one of:identifying, a neighbouring region to the at least one internal region, and at least one audio source defining the neighbouring region is available;determining at least two audio signals based on the at least one audio source; and determining the at least two audio signals based on respective metadata associated with the at least one audio source.
14. The method as claimed in any of claims 8 to 13, further comprising obtaining the at least one audio source from microphone arrangements, wherein each microphone arrangement is at a respective position and comprises one or more microphones.
15. The method as claimed in any of claims 8 to 14, wherein obtaining a listener position comprises obtaining the listener position from a further apparatus.
16. An apparatus for generating binaural audio based on a listener position within an audio scene, the apparatus comprising at least one processor and at least one memory including a computer program code,the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to:obtain, for at least one audio source at least one audio source position from the audio scene; obtain a listener position within the audio scene, wherein the audio scene comprises: at least one internal region; and at least one external region, the regions defined with respect to the at least one audio source position;determine whether the listener position is located within the at least one internal region or at least one external region;determine whether the at least one audio source is available for rendering based on the determined listener position;determine a rendering mode based on the determination of the at least one audio source is available for rendering based on the determined listener position and whether the listener position is located within the at least one internal region or at least one external region; andrender the binaural audio signal based on the determined rendering mode.
17. The apparatus as claimed in claim 16, wherein the apparatus is caused to determine whether the at least one audio source is available for rendering based on the determined listener position is further configured to determine an estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
18. The apparatus as claimed in claim 16, wherein the apparatus is caused to determine the rendering mode is configured to determine the rendering mode based on the estimated time until the at least one audio source is available for rendering when the at least one audio source is unavailable.
19. The apparatus as claimed in any of claims 16 to 18, wherein the apparatus is caused to determine whether the listener position is located within the at least one internal region or at least one external region is configured to:determine whether the listener position is within a perimeter defined by the at least one audio source position; ordetermine whether the listener position is outside the perimeter defined by the at least one audio source position.
20. The apparatus as claimed in any of claims 16 to 19, wherein the apparatus is caused to determine the rendering mode is configured to determine at least one of:an interior rendering mode when all of the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one internal region;a delayed interior rendering mode when one of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one internal region;an external rendering mode when the at least one audio source defining the internal region is available for rendering based on the determined listener position and the listener position is located within the at least one external region;an external rendering mode when one or more of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region; anda no rendering mode when all of the at least one audio source defining the internal region is not available for rendering based on the determined listener position and the listener position is located within the at least one interior region.
21. The apparatus as claimed in claims 16 to 20, wherein the apparatus is caused to render the binaural audio signal is configured to at least one of:identify, a neighbouring region to the at least one internal region, and at least one audio source defining the neighbouring region is available;determine at least two audio signals based on the at least one audio source; anddetermine the at least two audio signals based on respective metadata associated with the at least one audio source.
22. The apparatus as claimed in any of claims 16 to 21, wherein the apparatus is further configured to obtain the at least one audio source from microphone arrangements, wherein each microphone arrangement is at a respective position and comprises one or more microphones.
23. The apparatus as claimed in any of claims 16 to 22, wherein the apparatus is configured to obtain the listener position from a further apparatus.