Spatial Audio Representation and Rendering

The binaural renderer addresses low directional resolution issues by combining user-provided and predefined datasets, ensuring accurate spatial audio rendering with enhanced timbre and directionality.

JP7818660B2Active Publication Date: 2026-02-20NOKIA TECHNOLOGIES OY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024131830
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-11
Filing Date
2024-08-08
Publication Date
2026-02-20
Estimated Expiration
2040-09-29

AI Technical Summary

Technical Problem

Existing binaural rendering technologies face challenges in accurately rendering spatial audio with low directional resolution and measurement quality issues, leading to artifacts such as timbre coloration and inaccurate localization, especially when using user-provided datasets.

Method used

A binaural renderer that combines a loaded binaural dataset with a predefined, high-quality dataset to enhance directional accuracy and timbre quality, using perceptual matching procedures to adjust spectral and interaural properties, and supports rendering with and without room effects.

Benefits of technology

Enables accurate directional perception and uncolored timbre in binaural audio reproduction, even with low directional resolution datasets, by integrating a predefined dataset to improve spatial audio rendering quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007818660000018
    Figure 0007818660000018
  • Figure 0007818660000019
    Figure 0007818660000019
  • Figure 0007818660000020
    Figure 0007818660000020
Patent Text Reader

Abstract

To provide a space audio expression and rendering.SOLUTION: The device includes means for acquiring a space audio signal including at least one audio signal and space metadata related to the at least one audio signal, acquiring at least one dataset defined in advance related to binaural rendering, acquiring at least one dataset related to rendering, and generating a binaural audio signal on the basis of a combination of at least a part of the at least one dataset and the at least one dataset defined in advance and the space audio signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to apparatus and methods for spatial audio representation and rendering, although not limited to audio representation for audio decoders. [Background technology]

[0002] Immersive audio codecs have been implemented that support multiple operating points ranging from low bitrate operation to transparency. One example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed for use over communication networks such as 3GPP® 4G / 5G networks, including for immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose audio. Furthermore, it is expected to support channel-based audio and scene-based audio input, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services under various transmission conditions and support high error robustness.

[0003] An input signal can be presented to an IVAS encoder in one of several supported formats (and several allowed combinations of formats). For example, a mono audio signal (without metadata) can be encoded using an EVS (Enhanced Voice Service) encoder. Other input formats can utilize new IVAS encoding tools. One input format proposed for IVAS is the Metadata-Assisted Spatial Audio (MASA) format, which allows encoders to utilize a combination of mono and stereo encoding tools and metadata encoding tools for efficient transmission of the format. MASA is a parametric spatial audio format suitable for spatial audio processing. Parametric spatial audio processing is a branch of audio signal processing in which the spatial aspects of a sound (or sound scene) are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and useful choice to estimate a set of parameters from the microphone array signal, such as the direction of sound in a frequency band or the relative energy of the directional and non-directional parts of the captured sound in a frequency band, expressed as, for example, the direct-to-whole ratio or the ambient-to-whole energy ratio in a frequency band. These parameters are known to well describe the perceived spatial characteristics of the captured sound at the location of the microphone array, and can be used accordingly for spatial sound synthesis, binaurally with headphones, loudspeakers, or other formats such as Ambisonics.

[0004] For example, there may be two channels (stereo) of audio signals and spatial metadata. The spatial metadata may further define parameters such as a direction index describing the direction of sound arrival in the time-frequency parameter interval, a level / phase difference, a direct-to-total energy ratio representing the energy ratio of the directional index, diffuseness, coherence such as diffuse coherence representing the spread of energy representing the directional index, a diffuse-to-total energy ratio representing the energy ratio of omnidirectional sound to the surrounding directions, a surround coherence representing the coherence of omnidirectional sound to the surrounding directions, a reverberant-to-total energy ratio representing the energy ratio of reverberant sound (such as microphone noise) whose energy ratios must sum to 1, a distance representing the distance of the sound originating from the direction of the index in meters on a logarithmic scale, a covariance matrix for multichannel loudspeaker signals, or any data related to these covariance matrices, and other parameters guiding a specific decoder, such as a central prediction coefficient or a 1:2 decoding coefficient (used in MPEG Surround, for example). Any of these parameters may be determined in the frequency domain.

[0005] Hearing a natural audio scene in an everyday environment is not just about sounds from a particular direction. Even without background ambience, it is typical that most of the sound energy reaching the ears is not from direct sound, but from indirect sound (i.e., reflections and reverberation) from the acoustic environment. Based on room effects, including discrete reflections and reverberation, listeners auditorily perceive sound source distance and room characteristics (small, loud, wet, reverberant), among other features, and the room adds to the perceived sensation of audio content. In other words, the acoustic environment is an intrinsically and perceptually relevant feature of spatial sound.

[0006] A listener listens to music in a normal room (as opposed to, for example, an anechoic chamber), and the music (e.g., stereo or 5.1 content) is typically produced in a way that would be expected to be heard in a room with normal reverberation, which creates an envelope and spaciousness to the sound. Listening to normal music in an anechoic chamber is known to be unpleasant due to the lack of room effects. Therefore, normal music will be heard with reverberation in a normal room (as it is essentially always heard). Summary of the Invention

[0007] According to a first aspect, there is provided an apparatus comprising means for obtaining a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal, obtaining at least one data set associated with binaural rendering, obtaining at least one predefined data set associated with binaural rendering, and generating the binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one predefined data set and the spatial audio signal.

[0008] The at least one dataset related to binaural rendering may comprise at least one of a set of binaural room impulse responses or transfer functions, a set of head-related impulse responses or transfer functions, a dataset based on binaural room impulse responses or transfer functions, and a dataset based on head-related impulse responses or transfer functions.

[0009] The at least one predefined data set associated with the binaural rendering may comprise at least one of a set of predefined binaural room impulse responses or transfer functions, a set of predefined head-related impulse responses or transfer functions, a predefined data set based on binaural room impulse responses or transfer functions, and a predefined data set based on captured head-related impulse responses or transfer functions.

[0010] The means may be further configured to divide the at least one dataset into a first portion and a second portion, and the means may be configured to generate a combination of the first portion of the at least one dataset with the at least one predefined dataset.

[0011] The means configured to generate the binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one predefined data set and the spatial audio signal may be configured to generate the first portion binaural audio signal based on a combination of the first portion of the at least one data set, the at least one predefined data set and the spatial audio signal.

[0012] The means configured to generate a combination of at least a portion of the at least one dataset and the at least one predefined dataset may be further configured to generate a second portion combination including one of: a combination of a second portion of the at least one dataset and at least a portion of the at least one predefined dataset, at least a portion of the at least one predefined dataset where the second portion of the at least one dataset is a null set, and at least a portion of the at least one predefined dataset where the second portion of the at least one dataset is determined to have substantial errors, be noisy, or be corrupted.

[0013] The means configured to generate a binaural audio signal based on a combination of at least a portion of the at least one data set and the at least one predefined data set, and the spatial audio signal may be configured to generate a second partial binaural audio signal based on the second partial combination and the spatial audio signal.

[0014] The means configured to generate a binaural audio signal based on a combination of at least a portion of the at least one data set and the at least one predefined data set, and the spatial audio signal may be configured to combine the first portion of the binaural audio signal and the second portion of the binaural audio signal.

[0015] The means configured to divide the at least one data set into a first portion and a second portion may be configured to generate a first window function having a roll-off function based on an offset time from the determined time of maximum energy and a crossover time, the first window function being applied to the at least one data set to generate the first portion, and to generate a second window function having a roll-on function based on the offset time from the determined time of maximum energy and a crossover time, the second window function being applied to the at least one data set to generate the second portion.

[0016] The means may be configured to generate a combination of at least a portion of the at least one data set with at least one predefined data set.

[0017] The means configured to generate a combination of at least a portion of the at least one dataset and the at least one predefined dataset generates an initial combined dataset based on the selection of the at least one dataset, determines at least one gap in the initial combined dataset defined by at least one pair of adjacent elements of the initial combined dataset having a direction difference greater than the determined threshold, and for each gap: The method may be configured to identify elements of the at least one predefined dataset having an orientation that is located within a gap in the at least one predefined dataset, and combine the identified elements of the at least one predefined dataset with the initial combined dataset.

[0018] The determined thresholds may include an azimuth threshold and an elevation threshold.

[0019] The combination of at least a portion of the at least one data set with the at least one predefined data set may be defined over a range of directions, and over the range of directions, the combination does not include a direction gap exceeding a defined threshold.

[0020] At least one portion of at least one data set is free from substantial error; The elements of at least one data set may be at least one of substantially free of noise and substantially free of corruption.

[0021] The means configured to obtain a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal may be configured to receive the spatial audio signal from a further device.

[0022] The means configured to obtain at least one data set related to binaural rendering may be configured to receive the at least one data set from a further device.

[0023] According to a second aspect, a method is provided for obtaining a spatial audio signal comprising: obtaining a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; A method is provided that includes the steps of obtaining at least one data set related to binaural rendering, obtaining at least one predefined data set related to binaural rendering, and generating a binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one predefined data set and a spatial audio signal.

[0024] The at least one dataset related to binaural rendering may comprise at least one of a set of binaural room impulse responses or transfer functions, a set of head-related impulse responses or transfer functions, a dataset based on binaural room impulse responses or transfer functions, and a dataset based on head-related impulse responses or transfer functions.

[0025] The at least one predefined data set associated with the binaural rendering may comprise at least one of a set of predefined binaural room impulse responses or transfer functions, a set of predefined head-related impulse responses or transfer functions, a predefined data set based on binaural room impulse responses or transfer functions, and a predefined data set based on captured head-related impulse responses or transfer functions.

[0026] The method may further include dividing the at least one dataset into a first portion and a second portion, and generating a combination of the first portion of the at least one dataset with the at least one predefined dataset.

[0027] Generating the binaural audio signal based on a combination of the at least one data set, the at least one pre-defined data set, and at least a portion of the spatial audio signal may include generating a first portion binaural audio signal based on a combination of a first portion of the at least one data set, the at least one pre-defined data set, and the spatial audio signal.

[0028] Generating a combination of at least a portion of the at least one dataset with at least a portion of the at least one predefined dataset may further comprise generating a second portion combination comprising one of: a combination of a second portion of the at least one dataset with at least a portion of the at least one predefined dataset; at least a portion of the at least one predefined dataset, where the second portion of the at least one dataset is a null set; and at least a portion of the at least one predefined dataset, where the second portion of the at least one dataset is determined to be substantially erroneous, noisy, or corrupted.

[0029] The method may include generating a binaural audio signal based on at least a portion of a combination of the at least one data set and the at least one predefined data set, and generating a second portion of the binaural audio signal based on the combination of the second portion and the spatial audio signal.

[0030] Generating a binaural signal based at least in part on a combination of the at least one data set and the at least one predefined data set, and the spatial audio signal may include combining a first partial binaural audio signal and a second partial binaural audio signal.

[0031] Dividing the at least one data set into a first portion and a second portion can comprise generating a first window function having a roll-off function based on an offset time from the determined time of maximum energy and a crossover time, the first window function being applied to the at least one data set to generate the first portion; and generating a second window function having a roll-on function based on the offset time from the determined time of maximum energy and a crossover time, the second window function being applied to the at least one data set to generate the second portion.

[0032] The method includes generating a combination of at least a portion of at least one data set with at least one predefined data set.

[0033] The step of generating a combination of at least a portion of the at least one dataset and the at least one predefined dataset may include the steps of generating an initial combined dataset based on the selection of the at least one dataset; determining at least one gap in the initial combined dataset defined by at least one pair of adjacent elements of the initial combined dataset with an orientation difference greater than a determined threshold; identifying, for each gap, an element of the at least one predefined dataset having an orientation located within the gap in the at least one predefined dataset; and combining the identified element of the at least one predefined dataset with the initial combined dataset.

[0034] The determined thresholds may include an azimuth threshold and an elevation threshold.

[0035] The combination of at least a portion of the at least one data set with the at least one predefined data set may be defined over a range of directions, over which the combination does not include a directional gap exceeding a defined threshold.

[0036] At least one portion of at least one data set is free from substantial error; The elements of at least one data set may be at least one of substantially free of noise and substantially free of corruption.

[0037] Obtaining the spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal may comprise receiving the spatial audio signal from a further device.

[0038] Obtaining at least one data set related to binaural rendering may comprise receiving at least one data set from a further device.

[0039] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one storage device containing computer program code, the at least one storage device and the computer program code being configured, using the at least one processor, to cause the apparatus to obtain a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal, obtain at least one dataset related to binaural rendering, obtain at least one predefined dataset related to binaural rendering, and generate the binaural audio signal based on a combination of the at least one dataset and at least a portion of the at least one predefined dataset and the spatial audio signal.

[0040] The at least one dataset related to binaural rendering may comprise at least one of a set of binaural room impulse responses or transfer functions, a set of head-related impulse responses or transfer functions, a dataset based on binaural room impulse responses or transfer functions, and a dataset based on head-related impulse responses or transfer functions.

[0041] The at least one predefined data set associated with the binaural rendering may comprise at least one of a set of predefined binaural room impulse responses or transfer functions, a set of predefined head-related impulse responses or transfer functions, a predefined data set based on binaural room impulse responses or transfer functions, and a predefined data set based on captured head-related impulse responses or transfer functions.

[0042] The apparatus may further be adapted to divide the at least one dataset into a first portion and a second portion and generate a combination of the first portion of the at least one dataset with the at least one predefined dataset.

[0043] The apparatus for generating a binaural audio signal based on a combination of at least one data set and at least a portion of at least one predefined data set with a spatial audio signal is configured to generate a first portion binaural audio signal based on a combination of a first portion of the at least one data set, the at least one predefined data set and the spatial audio signal.

[0044] The apparatus for generating a combination of at least a portion of the at least one dataset and at least one predefined dataset may generate a second portion combination including one of: a combination of a second portion of the at least one dataset and at least a portion of the at least one predefined dataset; at least a portion of the at least one predefined dataset, wherein the second portion of the at least one dataset is a null set; and at least a portion of the at least one predefined dataset, wherein the second portion of the at least one dataset is determined to be substantially erroneous, noisy, or corrupted.

[0045] The device for generating a binaural audio signal based on a combination of at least a portion of at least one data set with at least one predefined data set and the spatial audio signal may be configured to generate the second portion of the binaural audio signal based on the combination of the second portion and the spatial audio signal.

[0046] The apparatus for generating a binaural audio signal based on a combination of at least a portion of at least one data set with at least one predefined data set and the spatial audio signal may be configured to combine a first portion of the binaural audio signal with a second portion of the binaural audio signal.

[0047] The apparatus for dividing at least one data set into a first portion and a second portion may be configured to generate a first window function having a roll-off function based on an offset time and a crossover time from the determined time of maximum energy, the first window function being applied to the at least one data set to generate the first portion, and to generate a second window function having a roll-on function based on the offset time and the crossover time from the determined time of maximum energy, the second window function being applied to the at least one data set to generate the second portion.

[0048] The apparatus may be adapted to generate a combination of at least a portion of the at least one data set with at least one predefined data set.

[0049] An apparatus for generating at least a partial combination of at least one dataset and at least one predefined dataset may generate an initial combined dataset based on the selection of the at least one dataset, determine at least one gap in the initial combined dataset defined by at least one pair of adjacent elements of the initial combined dataset having an orientation difference greater than a determined threshold, identify, for each gap, an element of the at least one predefined dataset having an orientation located within the gap in the at least one predefined dataset, and combine the identified element of the at least one predefined dataset with the initial combined dataset.

[0050] The determined thresholds may include an azimuth threshold and an elevation threshold.

[0051] The combination of at least a portion of the at least one data set with the at least one predefined data set may be defined over a range of directions, and over the range of directions, the combination does not include a direction gap exceeding a defined threshold.

[0052] At least one portion of the at least one data set may be at least one of substantially error-free, substantially noise-free, and substantially corruption-free elements of the at least one data set.

[0053] An apparatus adapted to obtain a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal may be adapted to receive the spatial audio signal from a further apparatus.

[0054] The device adapted to obtain at least one data set related to binaural rendering may be adapted to receive at least one data set from a further device.

[0055] According to a fourth aspect, there is provided an apparatus comprising: obtaining a circuit configured to obtain a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtaining a circuit configured to obtain at least one data set related to binaural rendering; obtaining a circuit configured to obtain at least one pre-defined data set related to binaural rendering; and generating a circuit configured to generate the binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one pre-defined data set, the at least one pre-defined data set, and the spatial audio signal.

[0056] According to a fifth aspect, there is provided a computer program comprising instructions (or a computer-readable medium comprising program instructions) for causing an apparatus to: obtain a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtain at least one data set associated with binaural rendering; obtain at least one pre-defined data set associated with binaural rendering; and generate the binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one pre-defined data set and the spatial audio signal.

[0057] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing an apparatus to: acquire a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; acquire at least one data set associated with binaural rendering; acquire at least one pre-defined data set associated with binaural rendering; and generate the binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one pre-defined data set and the spatial audio signal.

[0058] According to a seventh aspect, there is provided an apparatus comprising: means for obtaining a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; means for obtaining at least one data set related to binaural rendering; means for obtaining at least one predefined data set related to binaural rendering; and means for generating the binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one predefined data set and the spatial audio signal.

[0059] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions for causing an apparatus to obtain a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal, obtain at least one data set associated with binaural rendering, obtain at least one predefined data set associated with binaural rendering, and generate the binaural audio signal based on a combination of the at least one data set and at least a portion of the at least one predefined data set and the spatial audio signal.

[0060] An apparatus configured to perform the operations of the above-described method.

[0061] A computer program comprising program instructions for causing a computer to carry out the method described above.

[0062] A computer program product stored on the medium can cause an apparatus to perform the methods described herein.

[0063] The electronic device may comprise an apparatus as described herein.

[0064] The chipset may comprise an apparatus as described herein.

[0065] Embodiments of the present application aim to address problems associated with the state of the art. [Brief explanation of the drawings]

[0066] For a better understanding of the present application, reference is made, by way of example, to the accompanying drawings, in which: [Figure 1] FIG. 1 shows a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 2] FIG. 2 illustrates a flow diagram of the operation of an exemplary device according to some embodiments. [Figure 3] FIG. 3 illustrates a schematic diagram of a synthesis processor such as that shown in FIG. 1, according to some embodiments. [Figure 4] FIG. 4 illustrates a flow diagram of the operation of an exemplary device such as that shown in FIG. 3, according to some embodiments. [Figure 5] FIG. 5 illustrates an example of an early / late portion divider according to some embodiments. [Figure 6] FIG. 6 illustrates a flow diagram of an exemplary method for generating combined pre-part rendering data, according to some embodiments. [Figure 7] FIG. 7 illustrates an example interpolation or curve fitting of rendering data according to some embodiments. [Figure 8] FIG. 8 illustrates in more detail an example of early and late renderers as shown in FIG. 3, according to some embodiments. [Figure 9] FIG. 9 shows an example of a device suitable for implementing the apparatus shown in the previous figure. DETAILED DESCRIPTION OF THE INVENTION

[0067] Below we describe in more detail suitable devices and possible mechanisms for using the loaded binaural dataset to render a spatial audio stream (or spatial audio signal) containing (carrying) audio signal(s) and spatial metadata associated with the audio signal(s). The goal is to allow the binaural renderer to load HRTFs and BRIRs with suboptimal directional resolution, while still providing optimal reproduced sound quality (accurate directional perception and low-frequency timbre). This is important when listeners load individual HRTFs / BRIRs, which typically cannot be measured at high directional resolution.

[0068] Using individually measured HRTFs / BRIRs has been shown to improve localization and enhance timbre. Therefore, listeners may be interested in loading individual responses to a binaural renderer (and / or a codec that includes a binaural renderer, such as IVAS). However, because obtaining such responses is uncommon (as of the time of writing), there is no regular or standardized way to measure them. As a result, they may be measured in a variety of ways, which can lead to responses with arbitrary directional resolution (i.e., the number of responses and the spacing between data points of available responses can vary significantly between various measurement methods). In practice, fewer HRTFs may be available than expected in known binaural rendering methods that aim to render audio in all directions with high spatial fidelity.

[0069] This diverse effect is even more evident in the context of BRIR databases used to render spatial audio signals. They typically have lower directional resolution than HRTF databases, even for professionally generated datasets (and typically have lower resolution in user-provided datasets). The practical reason for this is that installing a custom binaural measurement system in a typical room is difficult and time-consuming. Therefore, typically, only a few data points are available, corresponding to common multichannel speaker layouts, such as 5.1 and / or 7.1+4. The sparsity of HRTF / BRIR datasets poses challenges for binaural rendering. For example, HRTF / BRIR datasets may only include the horizontal direction, while rendering may need to support rendering height as well. Renderers must accurately render sound in directions where the dataset is sparse (e.g., a 5.1 binaural rendering dataset does not have HRTF / BRIR at 180 degrees). Furthermore, rendering may require head tracking on any axis, thus rendering in any direction with good spatial accuracy becomes relevant. While interpolation between data points is in principle optional when a dataset is sparse, interpolation with sparse data points can introduce serious artifacts such as timbre coloration of the sound, inaccurate and non-point-like localization, etc. Furthermore, a user-provided dataset can also be corrupted; for example, it may have a low SNR or, if not binaural, may have a distorted or corrupted response that affects the quality of the binaural rendering (e.g., timbre, spatial accuracy, externalization).

[0070] Furthermore, if the loaded dataset is an HRTF dataset, by definition, the dataset contains transfer functions only in anechoic space, not reflections or reverberation. However, rendering room effects (including reflections and / or reverberation) is known to be beneficial for certain signal types, such as multi-channel signals (e.g., 5.1). Multi-channel signals are generated to be listened to in a normal room with reverberation. When listened to in an anechoic space (to which the HRTF rendering corresponds), they are perceived as lacking width and envelope, thus degrading the perceived audio quality. Therefore, a binaural renderer should support adding room effects in all cases (even if the loaded dataset is an HRTF dataset).

[0071] Thus, the concept is that a renderer is provided that allows loading HRTF and BRIR sets of any resolution and potentially with measurement quality issues. Furthermore, in some embodiments, the renderer described is configured to render binaural audio from data formats that can have sound sources in any direction (such as MASA format and / or head-tracked binauralization). Furthermore, in some embodiments, the renderer is configured to render binaural audio with and without additive room response from any loaded HRTF and BRIR dataset.

[0072] Furthermore, embodiments can be configured to operate without requiring high directional resolution datasets (which cannot be guaranteed in all cases, especially with listener-loaded datasets), and still perform binaural rendering with good quality for any direction (leading to timbre coloration and suboptimal spatialization).

[0073] The embodiments relate to binaural rendering of a spatial audio stream including carrier audio signal(s) and spatial metadata using a loaded binaural dataset (e.g., based on HRTF and BRIR). Accordingly, the embodiments describe how binaural spatial audio with good directional accuracy and uncolored timbre can be generated even with a binaural dataset having low directional resolution. Furthermore, in some embodiments, this can be achieved by combining the loaded binaural dataset with a predefined binaural dataset (including a perceptual matching procedure) and using the combined binaural dataset to render the spatial audio stream into a binaural output.

[0074] In some embodiments, the binaural renderer may be part of, for example, a decoder (such as an IVAS decoder). It may therefore receive or retrieve spatial audio streams that are rendered into binaural outputs. Additionally, the binaural renderer supports the loading of binaural datasets. These binaural datasets may be loaded by, for example, a listener and may include, for example, individual responses tailored for the listener.

[0075] The binaural renderer further includes, in some embodiments, a predefined binaural dataset. Typically, the predefined binaural rendering dataset is characterized by spatial accuracy, and by this means is based on a spatially dense BRIR / HRTF dataset. Thus, the predefined dataset represents a reliable, high-quality default dataset that is pre-existing in the renderer.

[0076] The loaded binaural rendering data set may consist of responses selected to be used for rendering (e.g., personal responses), but may be suboptimal in some sense. Suboptimal may mean, for example: The dataset is based on a sparse measurement set (e.g., corresponding to 22.2 or 5.1 directions). Some directions (e.g., elevation, side) may have no response. The present invention allows for a load as low as a single (binaural) response, while still providing rendering to any direction. · The data set is subject to noise or corrupted measurement procedures.

[0077] In some embodiments, the loaded binaural dataset is combined with a predefined dataset, for example, by: Adding predefined datasets to the loaded dataset to substantially utilize predefined data in directions where the loaded data is sparse (i.e., large angular gaps in the dataset). Partially or completely replace the loaded binaural rendering data with predefined binaural rendering data.

[0078] Additionally, embodiments describe implementations that perform a perceptual matching procedure on the combined datasets, for example, by: Adjust the spectral characteristics of the combined dataset based on the loaded dataset. Adjust the interaural phase / time properties of the combined dataset based on the loaded dataset.

[0079] The resulting binaural data set is therefore spatially dense and can match the characteristics of the loaded binaural data set. Spatial audio is then rendered using this data set. As a result, listeners get personalized binaural spatial audio reproduction with accurate directional perception and uncolored timbre.

[0080] In some embodiments, if the loaded dataset is an HRTF dataset and binaural reverberation needs to be rendered, predefined binaural reverberation data (or "late part rendering data") is used to render the binaural reverberation.

[0081] Additionally, in some embodiments, when the predefined dataset is a BRIR dataset, the earlier portion of the predefined dataset is extracted for use in processing operations, as described in more detail herein.

[0082] In some embodiments, if the loaded dataset is a BRIR dataset, an earlier portion of the loaded dataset is extracted and used in processing operations as described in detail herein.

[0083] Furthermore, in some embodiments, when binaural reverberation needs to be rendered, a later portion of the loaded dataset is extracted to be used to render the binaural reverberation, which in some embodiments may be used directly, or predefined late reverberant binaural data may be modified to match the characteristics (e.g., reverberation time or spectral characteristics) of the dataset into which it was loaded.

[0084] Referring to FIG. 1, an exemplary apparatus and system for performing audio capture and rendering is shown, according to some embodiments.

[0085] System 199 is shown with an encoder / analyzer 101 portion and a decoder / synthesizer 105 portion.

[0086] The encoder / analyzer 101 portion in some embodiments includes an audio signal input configured to receive an input audio signal 110. The input audio signal may be obtained from any suitable source, such as, for example, two or more microphones on a mobile phone, e.g., a B-format microphone or other microphone array such as an Eigenmike, an Ambisonic signal, e.g., First Order Ambisonic (FOA), Higher Order Ambisonic (HOA), a loudspeaker surround mix, and / or an object. The input audio signal 110 may be provided to an analysis processor 111 and a transport signal generator 113.

[0087] The encoder / analyzer 101 portion may include an analysis processor 111. The analysis processor 111 is configured to perform spatial analysis on the input audio signal, generating appropriate metadata 112. Thus, the purpose of the analysis processor 111 is to estimate spatial metadata in frequency bands. For all of the aforementioned input types, known methods exist for generating appropriate spatial metadata, such as direction and direct-to-total energy ratios in frequency bands (or similar parameters such as diffuseness, i.e., ambient-to-total ratios). These methods are described in detail herein, but some examples may include performing an appropriate time-frequency transform on the input signal, then estimating delay values ​​for microphone pairs that maximize inter-microphone correlation in frequency bands when the input is a mobile phone microphone array, formulating direction values ​​corresponding to the delays (as described in GB Patent Application No. 1619573.7 and PCT Patent Application No. PCT / FI2017 / 050778), and formulating ratio parameters based on the correlation values.

[0088] The metadata can take various forms and can include spatial and other metadata. A typical parameterization of spatial metadata is one directional parameter in each frequency band θ(k,n) and an associated direct-to-total energy ratio in each frequency band r(k,n), where k is the frequency band index and n is the time frame index. Determining or estimating the direction and ratio depends on the device or implementation from which the audio signal is obtained. For example, the metadata can be obtained or estimated using spatial audio capture (SPAC) using the methods described in GB Patent Application No. 1619573.7 and PCT Patent Application No. PCT / FI2017 / 050778. In other words, in this particular context, the spatial audio parameters include parameters intended to characterize the sound field. In some embodiments, the generated parameters may differ for each frequency band. Thus, for example, all parameters are generated and transmitted in band X, but only one of the parameters is generated and transmitted in band Y, and no parameters are generated or transmitted in band Z. A practical example of this may be that for some frequency bands, such as the highest band, some of the parameters are not needed for perceptual reasons.

[0089] If the input is an FOA signal or a B-format microphone, the analysis processor 111 can be configured to determine parameters such as intensity vectors from which directional parameters are created and compare the intensity vector lengths with an overall sound field energy estimate to determine ratio parameters, a method known in the literature as Directional Audio Coding (DirAC).

[0090] If the input is an HOA signal, the analysis processor can either take an FOA subset of the signal and use the above method, or split the HOA signal into multiple sectors, each of which utilizes the above method. This sector-based method is known in the literature as High-Order DirAC (HO-DirAC). In this case, there are two or more simultaneous directional parameters per frequency band.

[0091] If the input is a loudspeaker surround mix and / or objects, the analysis processor 111 may be configured to convert the signal into a FOA signal (through the use of spherical harmonic encoding gain) and analyze the direction and ratio parameters as described above.

[0092] The output of the analysis processor 111 is therefore spatial metadata determined in frequency bands. The spatial metadata can include direction and rate in frequency bands, but can also have any of the metadata types listed above. The spatial metadata can vary over time and frequency.

[0093] In some embodiments, spatial analysis can be performed outside of system 199. For example, in some embodiments, spatial metadata associated with an audio signal may be provided to the encoder as a separate bitstream. In some embodiments, spatial metadata may be provided as a set of spatial (directional) index values.

[0094] The encoder / analyzer 101 portion may comprise a carrier signal generator 113. The carrier signal generator 113 is configured to receive an input signal and generate an appropriate carrier audio signal 114. The carrier audio signal may be a stereo or mono audio signal. The generation of the carrier audio signal 114 may be performed using known methods as summarized below.

[0095] If the input is a mobile phone microphone array audio signal, the carrier signal generator 113 may be configured to select a left and right microphone pair and apply appropriate processing to the signal pair, such as automatic gain control, microphone noise cancellation, wind noise cancellation, and equalization.

[0096] If the input is a FOA / HOA signal or a B-format microphone, the transport signal generator 113 may be configured to formulate a left-right oriented directional beam signal, such as two opposing cardioid signals.

[0097] If the input is a surround mix of loudspeakers and / or objects, the carrier signal generator 113 can be configured to combine the left-hand side channels into a left downmix channel, generate the same downmix signal for the right-hand side, and add a center channel to both carrier channels with an appropriate gain.

[0098] In some embodiments, the transport signal generator 113 is configured to bypass the input. For example, there are situations in which analysis and synthesis are performed in the same device in a single processing step, without intermediate encoding. The number of transport channels can also be any suitable number (preferably one or two channels, as discussed in the examples).

[0099] In some embodiments, the encoder / analyzer unit 101 may comprise an encoder / multiplexer 115. The encoder / multiplexer 115 may be configured to receive the carrier audio signal 114 and the metadata 112. The encoder / multiplexer 115 may be further configured to generate the metadata information and the carrier audio signal in encoded or compressed form. In some embodiments, the encoder / multiplexer 115 may further interleave, multiplex, or embed the metadata within the encoded audio signal into a single data stream 116 prior to transmission or storage. The multiplexing may be performed using any suitable scheme.

[0100] The encoder / multiplexer 115 may be implemented, for example, as an IVAS encoder or any other suitable encoder, and is thus configured to encode the audio signal and metadata to form a bitstream 116 (e.g., an IVAS bitstream).

[0101] This bitstream 116 may then be transmitted / stored 103, as indicated by the dashed line. In some embodiments, the encoder / multiplexer 115 is not present (and therefore the decoder / demultiplexer 121 is not present, as described below).

[0102] System 199 may further include a decoder / synthesizer unit 105 configured to receive, extract, or otherwise obtain bitstream 116 and generate from the bitstream an appropriate audio signal that is presented to a listener / listener playback device.

[0103] The decoder / synthesizer unit 105 may include a decoder / demultiplexer 121 configured to receive the bitstream, demultiplex the encoded stream, and then decode the audio signal to obtain a transport signal 124 and metadata 122.

[0104] Furthermore, in some embodiments, as described above, the demultiplexer / decoder 121 may not be present (e.g., if both the encoder / analyzer unit 101 and the decoder / synthesizer 105 are located in the same device and therefore there is no associated encoder / multiplexer 115).

[0105] The decoder / synthesizer unit 105 may comprise a synthesis processor 123 configured to take the carrier audio signal 124, the spatial metadata 122, and the loaded binaural rendering dataset 126 corresponding to the BRIR or HRTF, and generate a binaural output signal 128 that can be played over headphones.

[0106] The operation of the system is summarized with respect to the flow diagram shown in FIG. 2, which illustrates an example of receiving an input audio signal as shown in step 201.

[0107] Next, the flow diagram shows the analysis (spatial) of the input audio signal to generate spatial metadata as shown in FIG. 2 by step 203 .

[0108] A carrier audio signal is then generated from the input audio signal, as shown in FIG. 2, via step 204 .

[0109] The generated carrier audio signal and metadata may then be multiplexed as shown in Figure 2 by step 205. This is shown in Figure 2 as an arbitrary dashed box.

[0110] The encoded signal can be further demultiplexed and decoded to produce the carrier audio signal and spatial metadata, as shown in Figure 2 by step 207. This is also shown as an optional dashed box.

[0111] Next, as shown in FIG. 2 by step 209, a binaural audio signal can be synthesized based on the carrier audio signal, the spatial metadata, and the binaural rendering dataset corresponding to the BRIR or HRTF.

[0112] The synthesized binaural audio signal can then be output to an appropriate output device, for example a set of headphones, as shown in FIG. 2 by step 211 .

[0113] Referring to FIG. 3, the synthesis processor 123 is shown in more detail.

[0114] In some embodiments, the synthesis processor 123 includes an early / late portion divider 301. The early / late portion divider 301 is configured to receive the binaural rendering dataset 126 (corresponding to a BRIR or HRTF). In some embodiments, the binaural rendering dataset may be in any suitable form. For example, in some embodiments, the dataset is in the form of an HRTF (head-related transfer function), an HRIR (head-related impulse response), a BRIR (binaural room impulse response), or a BRTF (binaural room transfer function) for the set of directions determined. In some embodiments, the dataset is a parameterized dataset based on an HRTF, an HRIR, a BRIR, or a BRTF. The parameterization may be, for example, time difference and spectrum in a frequency band such as the Bark band. Furthermore, in some embodiments, the dataset may be an HRTF, an HRIR, a BRIR, or a BRTF transformed into another domain, for example, spherical harmonics.

[0115] In the following example, the rendering data is a typical format of an HRIR or BRIR (i.e., a set of time-domain impulse response pairs) for a set of determined directions. If the responses are HRTFs or BRTFs, they can be inverse time-frequency transformed into HRIRs or BRIRs for the following processing. Other examples are also described.

[0116] The early / late part divider 301 is configured to divide the loaded binaural rendering data into parts defined as loaded early data 302 which is fed to an early part rendering data combiner 303 and loaded late data 304 which is fed to a late part rendering data combiner 305.

[0117] In some embodiments where the data set includes only HRIR data, this is provided directly as loaded pre-data 302. Loaded pre-data 302 may, in some embodiments, be converted to the frequency domain at this point. Loaded delay data 304 in such an example only indicates the absence of a delay portion.

[0118] In some embodiments where the data set is a BRIR data set, windowing can be applied to split the response to the loaded early data 302 into mostly directional (including the direct portion and potentially first reflections) and the loaded late data 304 into mostly reverberant. The splitting can be performed, for example, by the following steps:

[0119] First, measure the time of maximum energy in the BRIR (which gives an approximation of the time of the first sound arrival).

[0120] Second, a window function is designed. Figure 5 shows an example of a designed window function. Figure 5 shows, for example, a window function with a first window 551 for extracting the early part. This window function is single after the time of maximum energy 501 until a defined offset 503. The function of the first window 551 then decreases throughout the time of crossover 505 until it becomes zero.

[0121] The window function further comprises a second window 553 for extracting the later portion which has a zero value until the beginning of the crossover 505 time. The function value of the second window 553 increases to 1 throughout the time of the crossover 505 and remains 1 thereafter.

[0122] This is just one example of a suitable function, and other functions can be used. In some embodiments, the offset time can be, for example, 5 ms, and the crossover time can be, for example, 2 ms. Third, a window function can be applied to the BRIR to obtain a windowed early portion and a windowed late portion.

[0123] Fourth, the windowed portion is provided to a portion rendering data combiner 303 as loaded portion data 302. In some embodiments, the loaded portion data may be transformed to the frequency domain at this point.

[0124] Fifth, the windowed delayed portion is provided to the delayed portion rendering data combiner 305 as loaded delayed data 304.

[0125] In some embodiments, the synthesis processor also includes predefined early data 300 and predefined late data 392, which may have been generated with steps equivalent to those described above based on predefined HRIR, BRIR, etc. responses. In those embodiments where the data set does not include a slow portion, predefined slow portion 392 indicates only the absence of a slow portion.

[0126] In some embodiments, the compositing processor 123 comprises a pre-part rendering data combiner 303. The pre-part rendering data combiner 303 is configured to receive pre-defined pre-part data 300 and loaded pre-part data 302. The pre-part rendering data combiner 303 is configured to evaluate whether the loaded pre-part data is spatially dense.

[0127] For example, in some embodiments, the partial rendering data combiner 303 is configured to determine whether the data is spatially dense based on a horizontal density criterion. In these embodiments, the partial rendering data combiner can check that the horizontal resolution of the responses is sufficiently dense. For example, the maximum azimuth gap between horizontal responses is not greater than a threshold. This horizontal response distance threshold can be, for example, 10 degrees.

[0128] For example, in some embodiments, the early part rendering data combiner 303 is configured to determine whether the data is spatially dense based on an altitude density criterion. In these embodiments, the early part rendering data combiner may check that there are no directions in elevation where the nearest responses are angularly separated by more than a threshold. This vertical response distance threshold may be, for example, 10 degrees or 20 degrees.

[0129] If these conditions are met, the partial rendering data combiner 303 is configured to provide the loaded initial data 302 without modification to the initial partial renderer 307 as combined initial partial rendering data 306 .

[0130] If the condition is not met, the initial partial rendering data combiner 303 is configured to also use the predefined initial data 300 to form the combined initial partial rendering data.

[0131] In the examples described herein, it is assumed that the predefined data 300 meets the horizontal and elevation density criteria, as described above. Furthermore, although in the embodiments described herein, merging is based on loaded data sets that do not meet the appropriate density criteria, merging may also be performed in situations where the above density criteria are met but the loaded data has separate defects, for example, the data has a poor SNR or is otherwise corrupted.

[0132] The partial rendering data combiner 303 can be configured to combine data in a manner such as that described in Figure 6. In this approach, the loaded partial rendering data 302 is used to render sound in those directions where the loaded data exists, and in other directions predefined partial rendering data 300. This approach is useful when the loaded partial rendering data is known to contain high quality measurements (e.g., good SNR, valid measurement procedure), but is sparse and therefore needs to be added in some directions.

[0133] FIG. 6 shows a flow diagram of combining loaded portion data 302 with predefined portion data 300 according to these embodiments.

[0134] 6 by step 601. In other words, the partial rendering data combiner 303 first generates the preliminarily combined data as a copy of the loaded data by simply copying the loaded data to the combined partial rendering data 306.

[0135] The next operation is to evaluate whether there are any horizontal gaps in the combined data if the gap is greater than the threshold, which is shown in step 603 of Figure 6.

[0136] If such a gap is found, the response from the predefined early data 300 to the combined early partial data 306 is added to the gap, as shown in step 605 of FIG.

[0137] The operation can then loop back to further evaluation checks, as indicated by the arrow returning to step 603. In other words, the evaluation and filling, if necessary, procedure is repeated until there are no horizontal gaps in the combined data that are larger than the threshold.

[0138] If there were no original horizontal gaps in the combined data, or if gaps were filled, the pre-part rendering data combiner 303 can be configured to check all directions of the pre-defined pre-data, in other words, to find the direction from the pre-defined pre-data that has the largest angle difference to the nearest data point in the combined pre-part data, and determine whether this difference is greater than a threshold, as shown in FIG. 6 by step 607.

[0139] If the difference is greater than the threshold, the corresponding response is added to the combined early portion data 306 from the predefined early portion data 300 as shown in FIG. 6 by step 609 .

[0140] Operation then returns to step 607 where the procedure is repeated as long as the maximum angle difference estimate is greater than the threshold.

[0141] If the angle difference is less than the threshold, step 611 outputs the combined partial data as shown in FIG.

[0142] In some embodiments, the pre-part rendering data combiner 603 is configured to directly use the pre-defined pre-part data 600 as the combined pre-part data, without using the loaded pre-part data 602. This approach is useful when the loaded data set may be suboptimal (e.g., insufficient SNR, inadequate measurement procedures).

[0143] The resulting combined pre-phase data 306 therefore has data points (response directions) with a density such that the horizontal and vertical density criteria described above are met.

[0144] In some embodiments, the parts rendering data combiner 303 is configured to apply a perceptual matching procedure to data points in the combined parts data 306 from the predefined parts data 300 .

[0145] Thus, in some embodiments, the part rendering data combiner 303 is configured to perform spectral matching.

[0146] As a preliminary step, the energy of all data points (directions) of the original predefined and loaded data set is measured in frequency bands.

number

[0147] Even if a representative HRTF is used, the response may not be anechoic, but may correspond to the early part of the BRIR response. In some embodiments, the HRTF(b,ch,q) c ) indicates the complex gain of the combined partial data 306 as a corresponding data set index.

[0148] In some embodiments, two angle values ​​are defined. α l,c (q l ,q c ) is the q l th data point and the combined prior data set, q c is the angular difference between the th data point, α p,c (q p ,q c ) is the q in the predefined early data set p th data point and the qth data point in the combined prior data set c is the angular difference between the second data point.

[0149] Next, in some embodiments, the following operations are performed for each data point in the combined partial data originating from the predefined partial data 300:

[0150] First, find the weighted average energy value of the initial data set loaded.

number

number

[0151] Second, find the weighted energy value of the predefined initial data set.

number

[0152] Third, we formulate the equalization gain to correct the average energy.

number

[0153] Fourth, for all bins b belonging to band k, the equalization gain g EQ (k) is the q in the combined early data (arising from the predefined early partial data) c Apply to the th response.

number

[0154] The above operations can then be repeated for all indices in the combined partial data resulting from the predefined partial data and for all frequency bands k.

[0155] In some embodiments, the part rendering data combiner is configured to optionally apply phase / time matching that takes into account the difference in maximum interaural time delay between the data sets. For example, for phase / time matching, the following operations can be performed:

[0156] First, the inter-binaural time difference (ITD) in the low frequency range (e.g., up to 1.5 kHz) is estimated from the initial partial response in the horizontal plane. The inter-binaural time difference can be found, for example, by the difference in the median group delay (in this frequency range) of the left and right ear responses. The estimated ITD value is ITD(θ p ), where θ p is the orientation value, p=1...P, where P is the number of responses in the horizontal plane.

[0157] Second, the response index p derived from the predefined early partial data set and the response index p derived from the loaded early partial data set are separately applied to the ITD data using a sinusoidal ITD function. max sinθ, where ITD max are the variables to solve for. The fitting is done for ITDs from 0.7 to 1.0 ms (or some other interval). max This can be easily done by testing a value (eg 100) and seeing which value provides the smallest difference e.

number

[0158] ITD max can be estimated from the index p derived from a predefined dataset, and the result is the ITD max,pre and the index p comes from the loaded dataset, and the result is ITD max,loaded Figure 7 shows two examples of fitting sinusoidal curves (dotted lines) to exemplary ITD data (shown as circles).

[0159] Third, the ITD scaling term is

number

[0160] Fourth, the response in the combined data from the predefined partial data set is calculated at least in the low frequency range (e.g., up to 1.5 kHz).

number

[0161] In the above example, the horizontal response is used to determine the ITD and the ITD max In some embodiments, for example, if the response is not in the horizontal plane (but instead, for example, in a uniform spherical distribution), all responses, or responses in a particular elevation angle range, are found to be ITD max The aforementioned error measure can then be selected for the determination, e.g.

number

[0162] The combined previous part rendering data may then be output to the previous part renderer 307.

[0163] In some embodiments, even if the representation HRTF'(b,ch) is used, the response may not be anechoic but may correspond to the early part of the BRIR response.

[0164] In some embodiments, the compositing processor 123 comprises a late part rendering data combiner 305. The delayed part rendering data combiner 305 may be configured to receive predefined delayed part data 392 and loaded delayed part data 304 and generate combined delayed part rendering data 312 that is output to a delayed part renderer 309.

[0165] In some embodiments, the predefined and loaded late part rendering data, if present, includes a late part windowing response based on the BRIR. The late part rendering data combiner 305 in such embodiments may be configured as follows:

[0166] First, it determines whether there is loaded delayed part data 304. If there is loaded delayed part data 304, it directly uses the loaded delayed part data 304 as the combined delayed part rendering data 312. As an example, all available responses are forwarded to the late part renderer 309, which then determines how to use these responses. In some embodiments, a subset of the responses may be selected (e.g., one response pair to the left and another response pair to the right) to be used as the combined late part rendering data 312 and forwarded to the late part renderer 309.

[0167] If there is no loaded delayed portion data 304 but there is predefined delayed portion data 392, the predefined delayed portion data is used as the combined delayed portion rendering data 312. However, in this case, equalization is applied to the combined delayed portion rendering data 312. The equalization gain can be, for example,

number

[0168] The equalization gain may be applied, for example, by frequency transforming the combined delayed partial rendering data 312, applying the equalization gain in the frequency domain, and transforming the result back to the time domain.

[0169] If there is neither loaded delayed portion data 304 nor predefined delayed portion data 392, the combined delayed portion rendering data 312 simply indicates that there is no delayed reverberation data, which, when delayed portion rendering is performed, triggers a default delayed portion rendering procedure in delayed portion renderer 309, as described below.

[0170] The combined late part rendering data 312 is then provided to the late part renderer 309 .

[0171] In some embodiments, the synthesis processor 123 comprises a renderer that may be split into an early portion renderer 307 and a late portion renderer 309. The early portion renderer 307 is shown in more detail with respect to Figure 8 and is configured to receive the carrier audio signal 122, the spatial metadata 124, the synthesis early portion rendering data 306 and to generate a suitable binaural early portion signal 308 to the synthesizer 311.

[0172] In some embodiments, the partial renderer 307, shown in more detail in Figure 8, comprises a time-frequency transformer 801 configured to receive (time-domain) carrier audio signals 122 and transform them into the time-frequency domain. Suitable transforms include, for example, a short-time Fourier transform (STFT) and a complex-modulated quadrature mirror filter bank (QMF). The resulting signal is i The time-frequency signal may be represented here in vector form (e.g., for two channels), e.g., (b,n), where i is the channel index, b is the frequency bin index of the time-frequency transform, and n is the time index.

number

[0173] The following processing operations can then be performed in the time-frequency domain across frequency bands. A frequency band can be one or more frequency bins (individual frequency components) of an applied time-frequency transformer (filter bank). In some embodiments, the frequency bands can approximate a perceptually relevant resolution such as Bark frequency bands, which are spectrally more selective at low frequencies than at high frequencies. Alternatively, in some implementations, the frequency bands can correspond to frequency bins. The frequency bands are typically frequency bands (or approximate frequency bands) for which spatial metadata has been determined by an analysis processor. Each frequency band k is a subset of the lowest frequency bin b low (k) and the highest frequency bin b high It can be defined in terms of (k).

[0174] The time-frequency carrier signal 802 in some embodiments may be provided to a covariance matrix estimator 807 and a mixer 811 .

[0175] The partial renderer 307, in some embodiments, comprises a covariance matrix estimator 807 configured to receive the time-frequency domain carrier signals 802 and estimate the covariance matrix of the time-frequency carrier signals and their overall energy estimate (within frequency bands). The covariance matrix may, for example, in some embodiments be:

number

[0176] The covariance matrix estimator 807 also calculates the overall energy estimate E(k,n) 808, i.e., C x The sum of the (k,n) diagonal values ​​can be generated and this overall energy estimate can be configured to be provided to the target covariance matrix determiner 805 .

[0177] In some embodiments, the initial part renderer 307 comprises an HRTF determiner 833. The HRTF determiner 833 may receive the combined initial part rendering data 306, which may be a suitably dense set of HRTFs. The HRTF determiner is configured to determine a 2x1 complex-valued head-related transfer function (HRTF) h(θ(k,n),k) for the angle θ(k,n) and frequency band k. In some embodiments, the HRTF determiner 833 is configured to receive the spatial metadata 124 from which the angle θ(k,n) is derived and to determine an HRTF for the output HRTF data 336.

[0178] For example, the HRTF determiner 833 may determine an HRTF at the mid-frequency of band k. If tracking of the listener's head orientation is involved, the directional parameters θ(k,n) may be modified before obtaining the HRTF to take into account the current head orientation. In some embodiments, the HRTF determiner 833 may determine a diffuse field covariance matrix for each band k, which may be, for example, a parameter for the direction θ(k,n), where d=1...D. d , which may be formulated based on the combined initial partial rendering data 306 by taking a uniformly distributed set of , and may also be determined by estimating the diffuse field covariance matrix as follows:

number

[0179] The HRTF determiner 833 may apply HRTF interpolation by using any suitable method (when the HRTF for the direction θ(k,n) is determined). For example, in some embodiments, a set of HRTFs is decomposed into interaural level difference and left-ear and right-ear energy as a function of frequency. Then, when the HRTF at a given angle is needed, the closest existing data point in the HRTF set is found, and the delay and energy at the given angle are interpolated. These energies and delays can then be transformed as complex multipliers are used.

[0180] In some embodiments, HRTFs are interpolated by converting the HRTF dataset into a set of spherical harmonic beamforming matrices in the frequency band. The HRTF for any angle at a frequency can then be determined by formulating a spherical harmonic weight vector for that angle and multiplying that vector by the beamforming matrix for that frequency. The result is again a 2x1 HRTF vector.

[0181] In some embodiments, the HRTF determiner 833 simply selects the closest HRTF from the available HRTF data points.

[0182] In some embodiments, the partial renderer 307 comprises a target covariance matrix determiner 805, which in this example determines at least one directional parameter θ(k,n), at least one direct-to-total energy ratio parameter r(k,n), a total energy estimate E(k,n) 808, and the HRTF h(θ(k,n),k) and the diffuse field covariance matrix C D(k) and HRTF data 336 consisting of (k). The covariance matrix determiner 805 is then configured to determine a target covariance matrix 806 based on the spatial metadata 124, the data 306, and the overall energy estimate 808. For example, the target covariance matrix determiner 805 may formulate the target covariance matrix according to the following equation:

number

[0183] Next, the target covariance matrix C y The (k,n) 806 can be fed to a mixing rule determiner 809 .

[0184] In some embodiments, the partial renderer 307 comprises a blending rule determiner 809. The blending rule determiner 809 is configured to receive the target covariance matrix 806 and the estimated covariance matrix 810. The blending rule determiner 809 determines the target covariance matrix C y (k,n)806 and the measured covariance matrix C x (k,n) 810 to generate a mixing matrix M(k,n) 812 based on the (k,n)

[0185] In some embodiments, the mixing matrix is ​​generated based on the method described in "Optimized covariance domain framework for time-frequency processing of spatial audio," J Vilkamo, T Backstrom, A Kuntz - Journal of Audio Engineering Society 61, no. 6 (2013): 403-411.

[0186] In some embodiments, the mixing rule determiner 809 determines the prototype matrices that guide the generation of the mixing matrix.

number

[0187] In summary, the covariance matrix C x (k,n), it can provide the mixing matrix M(k,n), which is obtained by the least squares optimization of the covariance matrix C y (k,n). The matrix Q guides the signal content in such mixing. In this example, the matrix is ​​simply the identity matrix, since the left and right processed signals should resemble the original left and right signals as closely as possible. In other words, the design requires that for the processed output, C y The goal is to minimally modify the signal while obtaining (k,n). A mixing matrix M(k,n) is formulated for each frequency band k and provided to the mixer 811. In some embodiments involving head tracking, the matrix Q can be adapted based on the head orientation. For example, if the user rotates 180 degrees, the matrix Q will have zeros on the diagonal and ones off the diagonal. This means in practice that the left output channel should resemble as similar as possible to the original right channel (for a 180 degree head rotation), and vice versa.

[0188] The partial renderer 307 in some embodiments comprises a mixer 811. The mixer 811 receives the audio signal 802 and a mixing matrix 812. The mixer 811 is configured to process the time-frequency audio signal (input signal) in each frequency bin b to generate two processed (initial partial) time-frequency signals 814, which can be formed, for example, based on the following equations:

number

[0189] The above procedure assumes that the input signals x(b,n) have adequate incoherence between them to render the output signal y(b,n) with the desired target covariance matrix characteristics. In some situations, the input signals do not have adequate inter-channel incoherence, for example, when only a single channel carrier signal is present, or when the signals are otherwise highly correlated. Therefore, in some embodiments, a decorrelation operation is implemented to generate a decorrelated signal based on x(b,n) and mix the decorrelated signal into a specific residual signal that is added to the signal y(b,n) in the above equation. Procedures for obtaining such residual signals are known and are described, for example, in the above standards.

[0190] The processed binaural (early part) time-frequency signal y(b,n) 814 is fed to an inverse T / F transformer 813 .

[0191] In some embodiments, the earlier portion renderer 307 includes an inverse T / F transformer 813 configured to receive the binaural (early portion) time-frequency signal y(b,n) 814 and apply an inverse time-frequency transformation corresponding to the applied time-frequency transformation applied by the T / F transformer 801. The output of the inverse T / F transformer 813 is the binaural (early portion) signal 308, which is passed to the combiner 311 (as shown in FIG. 3).

[0192] If the combined late part rendering data 312 indicates only that a late part response is absent, the late part renderer 309 is configured to generate the binaural late part signal 310 using a default binaural late part response. For example, the late part renderer 309 can generate a pair of white noise responses processed to have decay times and spectra according to predefined sets corresponding to binaural diffusion technology field inter-binaural correlations and typical listening rooms. Each of the aforementioned parameters may be defined as a function of frequency. In some embodiments, these sets may be user-definable.

[0193] In some embodiments, the delayed part renderer 309 may also receive an indication determining whether a delayed part renderer should be rendered. If a delayed part renderer is not required, the delayed part renderer 309 provides no output. If a delayed part renderer is required, the delayed part renderer 309 is configured to generate and add reverberation according to an appropriate method.

[0194] For example, in some embodiments, a convolver is applied to generate a late partial binaural output. Several signal processing structures are known for performing convolution. Convolution can be applied efficiently using FFT convolution or partial FFT convolution, see, e.g., Gardner, William G. "Efficient convolution without input / output delay," Audio Engineering Society Convention 97. Audio Engineering Society, 1994.

[0195] In some embodiments, the late part renderer 309 can receive late part BRIR responses from many directions (from the late part rendering data combiner 305). Selecting a BRIR pair for rendering involves at least the following steps: For example, in one embodiment, the carrier audio signal is summed into a single channel that is processed with a pair of reverberation responses. As in a typical set of BRIRs, there are responses from several directions, and a response may be selected as one of the response pairs in the set, such as a center-front BRIR tail. The reverberation response may also be a combined (e.g., averaged) response based on the BRIRs from multiple directions. In some embodiments, the carrier audio channels (e.g., two channels) are processed with different pairs of reverberation responses. The results of the convolution are summed together (separate left and right ear outputs) to obtain a two-channel binaural delayed partial output. In this example of two transport channels, the reverberation characteristics of the left transport signal can be selected, for example, from the left 90-degree BRIR (or the closest available response), and the right-hand side can be selected accordingly. Again, the reverberation responses can be combined (eg, averaged) based on the BRIRs from multiple directions.

[0196] The binaurally delayed partial signals can then be fed to a combiner 311 block.

[0197] In some embodiments, the synthesis processor may comprise a combiner 311 configured to receive a binaural early part signal 308 from an early part renderer 307 and a binaural late part signal 310 from a late part renderer 309 and combine or sum them (separately for the left and right channels). This signal may be played over headphones.

[0198] Referring to FIG. 4, a flow diagram illustrating the operation of the synthesis processor is shown.

[0199] The flow diagram illustrates the operations of receiving inputs such as a carrier audio signal, spatial metadata, and a loaded binaural rendering data set, as illustrated in FIG. 4 by step 401 .

[0200] Further, the method includes determining early / late part rendering data sets from the loaded binaural rendering data sets, as shown in FIG. 4 by step 403 .

[0201] By step 405, FIG. 4 illustrates generating initial part rendering data based on the determined loaded initial part rendering data and the pre-determined initial part rendering data.

[0202] The generation of the deferred partial rendering data based on the determined loaded deferred partial rendering data and the pre-determined deferred partial rendering data is illustrated in FIG.

[0203] Additionally, as indicated in FIG. 4 by step 407, there may be a binaural rendering based on the pre-part rendering data as well as the carrier audio signal and spatial metadata.

[0204] Additionally, as indicated in FIG. 4 by step 408, there may be binaural rendering based on late part rendering data and a carrier audio signal (and optionally a late rendering control signal).

[0205] The early and late rendering signals may then be combined or summed, as shown in FIG. 4, by step 409 .

[0206] The combined binaural audio signal may then be output as shown in FIG. 4, per step 411.

[0207] The above describes an example situation in which a binaural rendering dataset consists of responses from a single set of directions. While this is a typical format, the binaural data may be in other formats. For example, the rendering data (predefined and / or loaded) can be in the spherical harmonic domain. For example, it is known that it is possible to approximate an HRTF dataset as a filter or complex-valued spherical harmonic coefficients. When an Ambisonic signal is processed with such a filter or gain, the result is a binauralized audio signal. In such an embodiment, when the loaded binaural rendering data is in the spherical harmonic domain, it does not correspond to an arbitrary discrete set of directions. In other words, density considerations are no longer important. However, if the loaded rendering dataset has other quality issues (e.g., noise), it can be replaced with predefined rendering data and the perceptual matching procedure described above can be used.

[0208] In some embodiments, the predefined pre-part rendering data is stored in the spherical harmonic domain (e.g., the third or fourth order ambisonic domain) because such a data set can be used both to render ambisonic audio into binaural output and to determine HRTFs for any angle. Then, when a user loads personalized HRIRs or BRIRs (e.g., a sparse set) into the system, the following steps can be performed to determine the combined pre-part rendering data:

[0209] First, a set of HRTFs, eg, a spherically spaced HRTF data set, is determined based on predefined (spherical harmonic domain) rendering data.

[0210] Second, perform the combining and perceptual matching procedures as described above.

[0211] Third, the resulting combined pre-rendered partial data set is returned to the spherical harmonic domain, for example, by finding a spherical harmonic gain that approximates the combined pre-rendered partial data set in a least squares manner.

[0212] The rendering data can be stored in a parameterized form, i.e., not as a response in any domain. For example, this can be stored in a form of left and right ear energy and interaural time difference, stored in a set of directions. In this case, the parameterized form can be directly converted to HRTFs, and all the procedures previously exemplified can be applied. Also, the late part rendering data can be parameterized, for example, as a spectrum as a function of reverberation time and frequency.

[0213] The concepts detailed herein show how to generate a dense dataset even if the loaded dataset is spatially sparse. During the rendering stage, when a sound needs to be rendered at a specific angle, the system: Selecting the closest response from the combined early data set (especially if a dense early data set has been generated); Interpolating between the nearest data points using known methods (e.g., formulating a weighted average of the response (in the time or frequency domain) over the nearest data points, as if performing amplitude panning); Interpolating between data points in a parametric way, for example by interpolating energy and ITD separately, and using pre-rendering data in the spherical harmonic domain (SHD) (which essentially means interpolation in any direction); You can do one of the following:

[0214] In some embodiments, the combined binaural rendering dataset created in the present invention may be stored or used in any domain, such as the spherical harmonic domain (SHD), the time domain, the frequency domain, and / or the parametric domain.

[0215] In the examples described herein, an exemplary situation was described in which the late part rendering was based on a late part response and convolution. However, there are many existing reverberator structures that perform reverberation in a more efficient manner, for example.

[0216] One can implement a feedback delay network (FDN), which is a reverberation signal processing structure that circulates a signal in multiple interconnected feedback loops, outputting delayed reverberations.

[0217] "The reverberator" Vilkamo, J., Neugebauer, Band Plogsties, J., 2012. Sparse frequency-domain reverberator, Journal of Audio Engineering Society, 59(12), pp.936-943, uses a simpler loop structure than FDN, but has a larger number of frequency bands.

[0218] Any reverberator capable of generating two substantially incoherent reverberation responses (e.g., any of those mentioned above) can be used to generate the binaural delayed partial signals. Typically, the reverberator structure generates substantially incoherent signals, which are then mixed in a frequency-dependent manner to obtain inter-binaural correlations that are natural to humans in a reverberant sound field. If the late-part rendering data is a representation of a BRIR late-part response, some reverberators (e.g., those in the publications mentioned above) can be used to adjust the reverberation parameters to approximate the BRIR late-part response. This typically involves setting the reverberation time according to the frequency and spectral gain of the reverberator to match the corresponding characteristics of the BRIR late response.

[0219] In some embodiments, the combined late part rendering data is typically in a format related to the particular signal processing architecture used by the late part renderer, e.g., if convolution is used. If a reverberator such as those described above is used, the late part rendering data is a representation of configuration parameters such as reverberation time as a function of frequency. Such parameters can be estimated from the reverberation response when the user loads the BRIR dataset used for rendering.

[0220] In some embodiments, the perceptual matching procedure can be performed during spatial audio rendering instead of running on a dataset.

[0221] In this example, the mixing matrix is ​​defined based on the input being a two-channel carrier audio signal, however, these methods can be adapted for implementation for any number of carrier audio channels.

[0222] The above describes how predefined binaural rendering datasets can be used in conjunction with loaded binaural rendering datasets, which in some embodiments can improve the playback quality of binaural rendering according to the loaded binaural rendering dataset by using high-quality predefined binaural rendering datasets.

[0223] While the foregoing description may imply a situation in which processing occurs on a single processing entity (e.g., handling the loading of binaural rendering data sets and the rendering of binaural audio output), it should be understood that processing may occur on multiple processing entities, e.g., processing may occur on different software modules and / or devices, such that some of the processing may be offline and some of the processing may be real-time.

[0224] It will therefore be apparent to those skilled in the art that the processing steps may be distributed across two or more different devices or software modules. In one practical example, some of the processing steps may be performed in a first program running on a computer, while other parts of the processing may be performed in another program, for example, an audio processing library running on a separate computer or mobile phone.

[0225] The steps associated with analyzing the binaural rendering dataset may be performed on any suitable platform that is capable of data visualization and thus capable of detecting potential errors in any of the response feature estimations.

[0226] As a practical example, when using a suitable program to carry out part of the processing, the steps involved may include: A set of binaural room impulse responses (BRIRs) is loaded into the program, In the program, the BRIR dataset is split into an early and a late part. The program estimates the spectral information of the early and late parts. In the program, the reverberation time as a function of frequency (e.g., the average of a set of BRIRs) is estimated, Spectral information and reverberation times are exported from the program and incorporated into an audio processing software module, where the software module has a predefined HRTF data set and a configurable reverberator; Audio processing software is enabled to use the spectral information to modify the processing spectrum based on predefined HRTF data sets; Audio processing software is enabled to configure the reverberator using the reverberation time (and spectral information), The software can be compiled and run on a mobile phone, for example, and thus makes it possible to render binaural audio with room effects based on the BRIR dataset loaded with room effects, but also using predefined HRTF datasets.

[0227] In the above, the "combined binaural dataset" consists of a predefined HRTF dataset, spectral information retrieved based on the loaded BRIR dataset, and reverberation parameters retrieved based on the loaded BRIR dataset. As illustrated by this example above, those skilled in the art will understand that processing can be distributed across various platforms in a variety of ways.

[0228] 9 illustrates an exemplary electronic device that may be used as any of the device components of the system, as described above with respect to FIG. 9. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1700 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device can be configured to implement, for example, encoder / analyzer unit 101 or decoder / synthesizer unit 105 as shown in FIG. 1, or any of the functional blocks described above.

[0229] In some embodiments, the device 1700 comprises at least one processor or central processing unit 1707 .

[0230] The processor 1707 may be configured to execute various program codes such as the methods described herein.

[0231] In some embodiments, the device 1700 comprises a memory device 1711 .

[0232] In some embodiments, the at least one processor 1707 is coupled to a storage device 1711. The storage device 1711 may be any suitable storage means.

[0233] In one embodiment, storage device 1711 includes program code sections for storing program code implementable on processor 1707. Additionally, in some embodiments, storage device 1711 may further comprise stored data sections for storing data, e.g., data that has been processed or is to be processed in accordance with embodiments described herein. The implemented program code stored in the program code sections and the data stored in the stored data sections can be retrieved by processor 1707 whenever needed via the memory-processor coupling.

[0234] In some embodiments, device 1700 comprises a user interface 1705. User interface 1705, in some embodiments, can be coupled to a processor 1707. In some embodiments, processor 1707 can control the operation of user interface 1705 and receive input from user interface 1705. In some embodiments, user interface 1705 can allow a user to input commands into device 1700, for example, via a keypad. In some embodiments, user interface 1705 can allow a user to obtain information from device 1700. For example, user interface 1705 can comprise a display configured to display information from device 1700 to the user. User interface 1705, in some embodiments, can comprise a touchscreen or touch interface that can both allow information to be input into device 1700 and further display information to the user of device 1700. In some embodiments, user interface 1705 can be a user interface for communication.

[0235] In some embodiments, apparatus 1700 comprises an input / output port 1709. The input / output port 1709, in some embodiments, comprises a transceiver. The transceiver in such embodiments may be coupled to processor 1707 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means, in some embodiments, may be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.

[0236] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data channel (IRDA).

[0237] The transceiver input / output port 1709 may be configured to receive signals.

[0238] In some embodiments, device 1700 may be used as at least part of a synthesis device. Input / output port 1709 may be coupled to headphones (which may be head-tracked or non-head-tracked headphones) or the like.

[0239] In general, various embodiments of the present invention may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention may be illustrated and depicted as block diagrams, flowcharts, or using some other pictorial representations, it should be appreciated that these blocks, devices, systems, techniques, or methods depicted herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controller, or other computing device, or some combination thereof.

[0240] Embodiments of the present invention may be implemented in computer software executable by a data processor in a mobile device, for example, as a processor entity, or by hardware, or by a combination of software and hardware. It should be further noted in this regard that any block of a logic flow, such as that shown in the figures, may represent program steps, interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. This software may be stored on physical media, such as memory chips, or memory blocks implemented within a processor, magnetic media, such as a hard disk or floppy disk, and optical media, such as a DVD or its data variant, a CD.

[0241] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0242] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is a highly automated process and is large-scale. Complex and powerful software tools are available to convert logic-level designs into complete semiconductor circuit designs ready to be etched and formed on semiconductor substrates.

[0243] Programs such as those offered by Synopsys, Inc., of Mountain View, California, and Cadence Design, of San Jose, California, use well-established rules of design and a library of pre-stored design modules to automatically route conductors and identify the location of components on a semiconductor chip.

[0244] Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) may be sent to a semiconductor manufacturing facility or "fab" for manufacturing.

[0245] The foregoing description provides a complete and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting examples.

[0246] However, various modifications and adaptations will become apparent to those skilled in the art in view of the foregoing description, upon perusal of the accompanying drawings and the appended claims.

[0247] However, all such similar modifications of the teachings of this invention will still fall within the scope of the present invention as defined in the appended claims.

Claims

1. An apparatus, comprising: at least one processor; When executed by the at least one processor, the device includes at least: obtaining a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtaining at least one dataset related to binaural rendering; obtaining at least one predefined data set related to binaural rendering; determining at least one combined dataset comprising at least a portion of the at least one dataset and at least a portion of the at least one predefined dataset; the at least one combined data set; the spatial metadata; and the at least one audio signal; generating a binaural audio signal based at least in part on a combination of dividing the at least one data set into a first portion and a second portion; generating a first combination based on the first portion of the at least one data set and the at least one predefined data set; generating a first window function having a roll-off function based on an offset time from the determined time of maximum energy, the first window function being applied to the at least one data set to generate the first portion; generating a second window function based on the offset time from the determined time of maximum energy, the second window function being applied to the at least one data set to generate the second portion; at least one memory storing instructions for executing the An apparatus comprising:

2. The at least one dataset related to binaural rendering is: A set of binaural room impulse responses or transfer functions, a set of head-related impulse responses or transfer functions; a dataset based on binaural room impulse responses or transfer functions, or datasets based on head-related impulse responses or transfer functions; The apparatus of claim 1 , comprising at least one of:

3. The at least one predefined data set related to binaural rendering is: A set of predefined binaural room impulse responses or transfer functions, a set of predefined head-related impulse responses or transfer functions; a predefined dataset based on binaural room impulse responses or transfer functions, or a predefined data set based on captured head-related impulse responses or transfer functions; The apparatus of claim 1 , comprising at least one of:

4. When performed by the at least one processor, generating the binaural audio signal causes the device to: generating a first partial binaural audio signal based on the first combination and the spatial audio signal; The apparatus of claim 1 comprising the instructions.

5. The instructions, when executed by the at least one processor, cause the device to: the second portion of the at least one data set and at least a portion of the at least one predefined data set; the at least a portion of the at least one predefined data set, wherein the second portion of the at least one data set is a null set; or the at least one portion of the at least one predefined data set, wherein the second portion of the at least one data set is determined to be substantially at least one of erroneous, noisy, or corrupted; generating a second combination based on one of 5. The apparatus of claim 4.

6. When performed by the at least one processor, generating the binaural audio signal causes the device to: generating a second partial binaural audio signal based on the second combination and the spatial audio signal; The apparatus of claim 5 comprising the instructions.

7. When performed by the at least one processor, generating the binaural audio signal causes the device to: generating the binaural audio signal based on a combination of the first partial binaural audio signal and the second partial binaural audio signal; including the instructions, the second partial binaural audio signal is based, at least in part, on the spatial audio signal.

5. The apparatus of claim 4.

8. The determining of the at least one combined data set, when performed by the at least one processor, causes the apparatus to: generating an initial combined dataset based on the selection of the at least one dataset; determining at least one gap in the initial combined data set defined using at least one pair of adjacent elements of the initial combined data set having an orientation difference greater than a determined threshold; for a gap of the at least one gap, identifying, within the at least one predefined data set, an element of the at least one predefined data set having an orientation that is located within the gap; combining the identified elements of the at least one predefined dataset with the initial combined dataset; including instructions to execute 10. The apparatus of claim 1.

9. The apparatus of claim 8 , wherein the determined thresholds include an azimuth threshold and an elevation threshold.

10. 2. The apparatus of claim 1, wherein the at least one combined data set is defined over a range of directions, and over the range of directions, the at least one combined data set does not include a directional gap that exceeds a threshold.

11. The instructions, when executed by the at least one processor, cause the device to: receiving the spatial audio signal from a further device to obtain the spatial audio signal comprising the at least one audio signal and the spatial metadata associated with the at least one audio signal; or receiving from said further device said at least one dataset related to binaural rendering to obtain said at least one dataset; Execute at least one of the following:

10. The apparatus of claim 1.

12. obtaining a spatial audio signal comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtaining at least one dataset related to binaural rendering; obtaining at least one predefined data set related to binaural rendering; determining at least one combined dataset comprising at least a portion of the at least one dataset and at least a portion of the at least one predefined dataset; the at least one combined data set; the spatial metadata; and the at least one audio signal; generating a binaural audio signal based on a combination of dividing the at least one data set into a first portion and a second portion; generating a first combination based on the first portion of the at least one data set and the at least one predefined data set; generating a first window function having a roll-off function based on an offset time from the determined time of maximum energy, the first window function being applied to the at least one data set to generate the first portion; generating a second window function based on the offset time from the determined time of maximum energy, the second window function being applied to the at least one data set to generate the second portion; A method comprising:

13. The method of claim 12 , wherein generating the binaural audio signal comprises generating a first partial binaural audio signal based on the first combination and the spatial audio signal.

14. the second portion of the at least one data set and at least a portion of the at least one predefined data set; the at least a portion of the at least one predefined data set, wherein the second portion of the at least one data set is a null set; or the at least one portion of the at least one predefined data set, wherein the second portion of the at least one data set is determined to be substantially at least one of erroneous, noisy, or corrupted; generating a second combination based on one of The method of claim 12.

15. generating the binaural audio signal generating a second partial binaural audio signal based on the second combination and the spatial audio signal; 15. The method of claim 14, comprising:

16. generating the binaural audio signal generating the binaural audio signal based on a combination of the first partial binaural audio signal and a second partial binaural audio signal; the second partial binaural audio signal is based, at least in part, on the spatial audio signal. The method of claim 13.

17. A non-transitory computer-readable medium containing instructions stored thereon for performing at least the operations comprising the method of claim 12.

Citation Information

Patent Citations

  • SOUND REPRODUCTION SYSTEM, PROGRAM AND DATA CARRIER

    JP2006500818A

  • Coefficient calculation device for head-related transfer function interpolation, sound localizer, coefficient calculation method for head-related transfer function interpolation and program

    JP2010171785A

  • Determination of individual hrtfs

    JP2015019360A

  • Information processing device and information processing method

    JP2017143469A

  • Method for generating customized spatial audio with head tracking

    JP2019146160A