Efficient head-related filter generation
The structured representation of basis functions for HR filter generation addresses inefficiencies in existing methods, enhancing computational efficiency and memory usage while improving audio localization in spatial audio systems.
Patent Information
- Application Number
- JP2025047630
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-07-07
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-23
AI Technical Summary
Existing methods for generating head-related filters for spatial audio rendering are inefficient, particularly in terms of computational complexity and memory usage, especially in limited-capacity systems like mobile devices, leading to inaccuracies in audio localization and increased complexity in HR filter evaluation.
A method and apparatus for generating head-related filters using a structured representation of basis functions, involving sampling and compact storage of basis function shapes, along with metadata, to optimize HR filter evaluation, reducing computational complexity and memory requirements.
The solution enables efficient and accurate HR filter generation with reduced computational complexity and memory usage, allowing for real-time audio rendering with improved localization accuracy in spatial audio systems.
Smart Images

Figure 2025108446000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments related to methods and systems for efficient head-related filter generation are disclosed.
Background Art
[0002] The human auditory system comprises two ears that capture sound (audio) waves propagating towards the listener. In the present disclosure, the words "sound" and "audio" are used interchangeably. FIG. 1 shows sound waves propagating towards a listener from a direction of arrival (DOA) specified by a pair of elevation and azimuth angles in a spherical coordinate system. Along the propagation path towards the listener, each sound wave interacts with the listener's upper torso, head, outer ears, and the surrounding material before reaching the listener's left and right tympanic membranes. This interaction results in temporal and spectral changes in the sound waveforms reaching the left and right tympanic membranes, some of which are DOA-dependent. The human auditory system has learned to interpret these changes in order to infer various spatial characteristics of the sound waves themselves, as well as the acoustic environment in which the listener is located. This ability is called spatial hearing, and spatial hearing relates to how the listener evaluates binaural signals, i.e., the spatial cues embedded in the sound signals in the right and left ear canals, in order to infer the location of auditory events induced by sound events (physical sound sources) and the acoustic characteristics produced by the physical environment in which the listener is located (e.g., a small room, a tiled bathroom, an auditorium, a windowless room (cave)). This human ability, i.e., spatial hearing, can be utilized to create a spatial audio scene by reintroducing the spatial cues, which would result in a spatial perception of sound, into the binaural signals.
[0003] The main spatial cues include (1) angular cues: the binaural cues, i.e., the interaural level difference (ILD) and the interaural time difference (ITD), and the monaural (or spectral) cues, and (2) distance cues: intensity and the direction-to-reverberation (D / R) energy ratio. The mathematical representation of the short-time (e.g., 1 to 5 milliseconds) DOA-dependent or angular cue-related temporal and spectral variations of the waveform is the so-called head-related (HR) filter. The frequency-domain (FD) representation of the HR filter is the so-called head-related transfer function (HRTF), and the time-domain (TD) representation of the HR filter is the so-called head-related impulse response (HRIR). Figure 2 shows the sound wave propagating towards the listener and the difference in the sound paths to the two ears, which gives rise to the ITD. Figure 14 shows an example of the spectral cue (HR filter) of the sound wave shown in Figure 2. The two plots shown in Figure 14 show the magnitude responses of a pair of HR filters obtained at an elevation angle (θ) of 0 degrees and an azimuth angle (φ) of 40 degrees. This data is from the Center for Image Processing and Integrated Computing (CIPIC) database for image processing and integrated computing: subject ID28. This database is publicly available and can be accessed from the link https: / / www.ece.ucdavis.edu / cipic / spatial-sound / hrtf-data / .
[0004] HR-filter-based binaural rendering techniques are being gradually established, where the spatial audio scene is generated by directly filtering the audio source signal using a pair of HR filters at the desired location. This technique is particularly attractive for many emerging applications such as virtual reality (VR), augmented reality (AR), or mixed reality (MR) (sometimes collectively called extended reality (XR)), and mobile communication systems where headsets are commonly used.
[0005] The HR filter is often estimated from measurements of the impulse response of a linear dynamic system that converts an original sound signal (i.e., the input signal), which can be measured within the ear channels of a listening subject (e.g., an artificial head, a mannequin, or a human subject), into left and right ear signals (i.e., the output signals) in a predefined set of elevation and azimuth angles on the surface of a sphere of a certain radius around the listening subject. The estimated HR filter is often provided as a finite impulse response (FIR) filter and can be used directly in that format. To achieve efficient binaural rendering, a pair of HRTFs can be converted into an interaural transfer function (ITF) or a modified ITF to prevent sharp spectral peaks. Alternatively, the HRTF can be described by a parametric representation. Such parameterized HRTFs can be easily integrated with parametric multichannel audio coders (e.g., MPEG Surround and Spatial Audio Object Coding (SAOC)).
[0006] To explain the quality of different spatial audio rendering techniques, the concept of the minimum audible angle (MAA) can be useful. The MAA characterizes the sensitivity of the human auditory system to the angular displacement of a sound event. Regarding localization in azimuth, studies have reported that the MAA is smallest (about 1 degree) at the front and back for broadband noise bursts and much larger (about 10 degrees) for lateral sound sources. The MAA in the median plane increases with elevation angle. An MAA that is on average as small as about 4 degrees at elevation has been reported for broadband noise bursts.
[0007] Audio spatial rendering that leads to a convincing spatial perception of sound at arbitrary locations in space requires a pair of HR filters that represent the location within the MAA of the corresponding location. If the angular mismatch for the HR filters is below a limit (i.e., if the angle for the HR filter is within the MAA), the mismatch is not noticed by the listener. However, if the mismatch is greater than this limit (i.e., if the angle for the HR filter is outside the MAA), such a greater location mismatch can lead to a correspondingly more prominent inaccuracy in the position perceived by the listener. SUMMARY OF THE INVENTION
[0008] HR filter measurements are taken at finite measurement locations, but audio rendering may need to determine HR filters for any possible location on a sphere surrounding the listener (e.g., 150 in FIG. 1). Thus, the mapping method needs to convert from individual measurements taken at finite measurement locations to a continuous spherical angular region. There are several methods for such mapping. This method includes directly using the nearest available measurement, using an interpolation method, and / or using a modeling technique.
[0009] 1. Direct use of the nearest neighbor measurement point
[0010] The simplest technique for mapping is to use the HR filter at the point that is closest (i.e., nearest) among the set of measurement points. Some computational work may be required to determine the nearest neighboring measurement points, and such work can be nontrivial for an irregularly sampled set of measurement points on the sphere around the listener. In the case of general object location, there can be some angular error between the desired filter location (corresponding to the object location) and the nearest available HR filter measurement point. In the case of a sparsely sampled set of HR filter measurements, this can lead to significant error in the object location. The error can be reduced or virtually eliminated when a more densely sampled set of measurement points is used. In the case of a moving object, the HR filter changes in a stepwise fashion that does not correspond to the intended smooth movement.
[0011] Generally, densely sampled measurements of the HR filter are difficult to take for a human subject, as this requires the subject to sit still during data collection and small accidental movements of the subject limit the achievable angular resolution. Also, the measurement process is time-consuming for both the subject and the technician. Instead of taking such densely sampled measurements, inferring the spatial relationship information for the missing HR filters can be more efficient assuming a sparsely sampled HR filter dataset (described below). Densely sampled HR filter measurements are easy to capture for a dummy head, but the resulting HR filter set is not always suitable for all listeners and can lead to inaccurate or ambiguous perception of object location.
[0012] 2. Interpolation between Neighboring Measurement Points
[0013] If the sample measurement points are not sufficiently closely spaced, interpolation between adjacent measurement points can be used to generate an approximate filter for the required DOA. The interpolation filter varies in a continuous fashion between individual sample measurement points, avoiding the abrupt changes that can occur when the above method (i.e., Method 1) is used. This interpolation method introduces additional complexity when generating the interpolated HR filter values, and the resulting HR filter has a DOA that is perceived as being spread out (as if there were fewer points) by mixing filters from different locations. Also, measures need to be taken to prevent the phase alignment problems that arise from directly mixing the filters, which can add complexity.
[0014] 3. Model - based Filter Generation
[0015] More advanced techniques can be used to construct a model for the system that underlies how the HR filter and the HR filter varies with angle. Given a set of HR filter measurements, the model parameters are tuned to reproduce the measurements with minimal error and thereby create a mechanism for generating the HR filter not only at the measurement locations but more globally as a continuous function over the angular space.
[0016] There are other methods for generating the HR filter as a continuous function of DOA that do not require an input set of measurements but instead use high - resolution 3D scans of the listener's head and ears to model the wave propagation around the listener's head in order to predict the behavior of the HR filter.
[0017] A category of HR filter models that utilize weighted basis functions and vectors to represent the HR filter is presented below.
[0018] 3.1. HR Filter Model Using Weighted Basis Vectors - Mathematical Framework
[0019] Consider a model for an HR filter having the following form. TIFF2025108446000002.tif18170
[0020] Here, TIFF2025108446000003.tif9170 is the estimated HR filter, a vector of length K, α, for a particular (θ, φ) angle n,k is a set of scalar weighting values that do not depend on the angle (θ, φ), F k,n (θ, φ) is a set of scalar-valued functions that depend on the angle (θ, φ), e k is TIFF2025108446000004.tif10170 a set of orthogonal basis vectors spanning the K-dimensional space of the filter.
[0021] The model function F k,n (θ, φ) is determined as part of the model design and is typically selected such that the variation in the set of HR filters over the elevation and azimuth dimensions is well captured. For a specified model function, the model parameter α n,k can be estimated using a data fitting method such as the minimized least squares method.
[0022] It is not uncommon to use the same modeling function for all of the HR filter coefficients, which gives rise to a particular subset of this type of model, where the model function F k,n (θ, φ) does not depend on the position k within the filter. F k,n (θ, φ)=F n (θ, φ), ∀k (2)
[0023] Thus, the model can be represented as follows. TIFF2025108446000005.tif19170
[0024] In one embodiment, e kThe basis vectors are the natural basis vectors e1 = [1,0,0,...0], e2 = [0,1,0,...0],... that are aligned with the coordinate system being used. For compactness, when the natural basis vectors are used, the vectors can be rewritten as follows. TIFF2025108446000006.tif19170
[0025] Here, α n is a vector of length K. This leads to the following equivalent equation for the model. TIFF2025108446000007.tif18170
[0026] That is, when the parameter α n,k is estimated, TIFF2025108446000008.tif9170 can be expressed as a linear combination of the fixed basis vectors α n , where the angular variation of the HR filter is captured at the weighting value F n (θ,φ).
[0027] Accordingly, the individual filter coefficients k are obtained as follows. TIFF2025108446000009.tif20170
[0028] This equivalent equation is a compact equation when the unit basis vectors are natural basis vectors. However, the following method can be applied (without this convenient notation) to models that use any selection of basis vectors in any domain (including non-orthogonal as well as orthogonal basis vectors). Other embodiments of the same underlying modeling technique may be different selections of basis vectors in the time domain (e.g., Hermite polynomials, sinusoids, etc.), or in domains other than the time domain such as the frequency domain (e.g., via Fourier transform), or in any other domain where it is natural to represent the HR filter.
[0029] TIFF2025108446000010.tif8170 is the result of the model evaluation specified in equation (5) and should be similar to the measurement of h at the same location. For test points (θ test , φ test ) where the actual measurement of h is known, h(θ test , φ test ) and TIFF2025108446000011.tif9170 can be compared to evaluate the quality of the model. If the model is considered accurate, the model can be used to estimate TIFF2025108446000012.tif9170 for some general point which is not necessarily one of the points where h was measured.
[0030] The equivalent matrix formulation of equation (5) is as follows. TIFF2025108446000013.tif11170
[0031] Here, f(θ, φ) = the row vector of weighted values for one ear, which has length N, i.e., f(θ, φ) = [F1(θ, φ), F2(θ, φ),..., F N (θ, φ)], and α = the basis function for one ear, which is configured as a row in the matrix K rows × N columns, i.e., as follows. TIFF2025108446000014.tif26170
[0032] As described in WO2021 / 074294 (incorporated herein by reference), B-spline functions are suitable basis functions for HR filter modeling for elevation angle θ and azimuth angle φ. This shows that the function F n (θ, φ) can be determined as follows. F N (θ, φ) = Θ p (θ) Φ p,q (φ) (8)
[0033] For p = 1, ..., P and q = 1, ..., Qp, n = (p - 1)Q p + q. P is the number of elevation basis functions, and Q p is the number of azimuth basis functions that can vary for different elevation angles p. In the case of elevation angles, standard B-spline functions can be used, and in the case of azimuth angles, periodic B-spline functions can be used.
[0034] As described above, the three types of methods for inferring an HR filter over a continuous region of angles have varying levels of computational complexity and varying levels of perceived location accuracy. The direct use of the nearest neighbor measurement points is the simplest but requires densely sampled measurements of the HR filter, which are not easy to obtain and usually result in large amounts of data. In contrast, the methods using models for the HR filter have the advantage that they can generate an HR filter with location-specific properties such as points that vary smoothly as the DOA changes. These methods also represent a set of HR filters in a more compact form and thus may require fewer resources for transmission and / or storage (including storage in program memory when they are in use). These advantages come at the expense of numerical complexity (the model must be evaluated before the filter can be used to generate the HR filter). Such complexity is a problem for rendering systems with limited computational capacity, as such limited capacity limits the number of audio objects that can be rendered, for example, in a real-time audio scene.
[0035] In a spatial audio renderer, it is desirable to be able to evaluate an HR filter for any elevation-azimuth angle in real time from a model evaluation equation such as equation (5). Therefore, the HR filter evaluation specified in equation (5) needs to be executed extremely efficiently.
[0036] The repeated evaluation of the HR filter model has the drawback of complexity not only when evaluating the model output but also when evaluating the basis functions of the model. Furthermore, the contribution of a certain basis function can be negligible (e.g., 0) for the evaluation in a certain HR filter direction. This means that the filter evaluation becomes unnecessarily complex. On the other hand, it is extremely important that the memory consumption required for HR filter evaluation does not increase significantly, especially for use in mobile devices where both memory availability and computational complexity availability are limited.
[0037] (For example, as described in WO2021 / 074294) From the B-spline basis functions, the filter evaluation described in equation (5) will involve the determination of F n (θ, φ), and it can be understood that In the evaluation of TIFF2025108446000015.tif9170, for each elevation angle p, P·Q p multiplication, and further for each coefficient n, P·Q p multiplication and addition are involved. These operations are later performed for each filter coefficient k, which results in a significant number of operations for the evaluation of TIFF2025108446000016.tif9170. TIFF2025108446000016.tif9170.
[0038] Figures 3(a) and 3(b) show periodic B-spline basis functions.
[0039] Figure 3(a) shows an example of four periodic B-spline basis functions for the [0, 360]-degree modeling range. The knot points are at 0 (= 360) degrees, 90 degrees, 180 degrees, and 270 degrees. In this example, all basis functions within each segment between the knot points are non-zero.
[0040] Figure 3(b) shows an example of eight periodic B-spline basis functions for the [0, 360]-degree modeling range. The knot points are at 0 (= 360) degrees, 45 degrees,..., 315 degrees. In this case, the non-zero part of each basis function covers only 1 / 2 of the modeling range, i.e., only 180 degrees.
[0041] As shown in FIGS. 3(a) and 3(b), for some B-spline settings, only a few B-spline functions are non-zero for a certain direction (θ, φ). For example, the B-spline function starting at 0 degrees in FIG. 3(b) can be zero for any angle between 180 and 360 degrees. This means that the HR filter evaluation of equation (5) can involve a significant number of multiplications and additions with zero components. The result is a complex and inefficient model-based HR filter evaluation.
[0042] According to some embodiments of the present disclosure, the problem of inefficient HR filter evaluation can be solved by a memory-efficient structured representation for complex-efficient HR filter evaluation and / or by avoiding multiplications and additions by zero-value components.
[0043] Thus, in one aspect, a method for generating a head-related (HR) filter for audio rendering is provided. The method includes generating HR filter model data representing an HR filter model. Generating the HR filter model data includes selecting at least one set of one or more basis functions. The method further includes, based on the generated HR filter model data, (i) sampling the one or more basis functions and (ii) generating first basis function shape data and shape metadata. The first basis function shape data identifies one or more compact representations of the one or more basis functions, and the shape metadata includes information regarding the structure of the one or more compact representations of the one or more basis functions. The method further includes providing the first generated basis function shape data and the shape metadata for storage in one or more storage media.
[0044] In some embodiments, the method may further include detecting the occurrence of a triggering event. Such a triggering event may indicate that a head-related (HR) filter should be generated for audio rendering, which may be induced by an audio renderer when an HR filter is required, for example, to render an audio frame or to prepare for rendering by generating an HR filter to be stored in memory for later use. In some embodiments, the triggering event is merely a determination to retrieve basis function shape data and / or shape metadata from one or more storage media. As a result of detecting the occurrence of the triggering event, the method may further include outputting second basis function shape data and shape metadata for audio rendering.
[0045] In another aspect, a method for generating a head-related (HR) filter for audio rendering is provided. The method includes obtaining shape metadata indicating whether to obtain a converted version of one or more compact representations of one or more basis functions. The method further includes obtaining basis function shape data that identifies (i) the one or more compact representations of the one or more basis functions or (ii) a converted version of the one or more compact representations of the one or more basis functions. The method further includes generating an HR filter by using (i) the one or more compact representations of the one or more basis functions or (ii) a converted version of the one or more compact representations of the one or more basis functions based on the obtained shape metadata and the obtained basis function shape data.
[0046] In another aspect, an apparatus is provided for generating a head-related (HR) filter for audio rendering. The apparatus is adapted to generate HR filter model data indicative of an HR filter model. Generating the HR filter model data includes selecting at least one set of one or more basis functions. The apparatus is further adapted to, based on the generated HR filter model data, (i) sample the one or more basis functions and (ii) generate first basis function shape data and shape metadata. The first basis function shape data identifies one or more compact representations of the one or more basis functions, and the shape metadata includes information regarding the structure of the one or more compact representations of the one or more basis functions. The apparatus is further adapted to provide the generated first basis function shape data and shape metadata for storage in one or more storage media.
[0047] The apparatus is further adapted to detect the occurrence of a triggering event and, as a result of detecting the occurrence of the triggering event, output second basis function shape data and shape metadata for audio rendering. Such a triggering event may indicate that a head-related (HR) filter should be generated for audio rendering, which may be induced by an audio renderer when an HR filter is required, for example, to render an audio frame or to prepare for rendering by generating an HR filter to be stored in memory for later use. In some embodiments, the triggering event is merely a determination to retrieve basis function shape data and / or shape metadata from one or more storage media. In one embodiment, the apparatus comprises a processing circuit and a storage unit storing instructions for configuring the apparatus to perform any of the processes disclosed herein.
[0048] In another aspect, an apparatus for generating a head-related (HR) filter for audio rendering is provided. The apparatus is adapted to obtain shape metadata indicating whether to obtain a converted version of one or more compact representations of one or more basis functions. The apparatus is further adapted to obtain basis function shape data that identifies (i) the one or more compact representations of the one or more basis functions or (ii) a converted version of the one or more compact representations of the one or more basis functions. The apparatus is further adapted to generate an HR filter by using (i) the one or more compact representations of the one or more basis functions or (ii) a converted version of the one or more compact representations of the one or more basis functions, based on the obtained shape metadata and the obtained basis function shape data.
[0049] In another aspect, a computer program is provided that, when executed by a processing circuit, comprises instructions to cause the processing circuit to perform the method described above. In one embodiment, a carrier containing the computer program is provided, and the carrier is one of an electronic signal, an optical signal, a wireless signal, and a computer-readable storage medium.
[0050] Embodiments of the present disclosure enable perceptually transparent (inaudible) optimization for a spatial audio renderer that utilizes a model-based HR filter, for example, to render a monaural source at a position (r, θ, φ) relative to a listener, where r is the radius and (θ, φ) are the elevation angle and the azimuth angle, respectively.
[0051] The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate various embodiments.
Brief Description of the Drawings
[0052]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 11
Figure 12
Figure 13
Figure 14
Modes for Carrying Out the Invention
[0053] Some embodiments of the present disclosure are directed to a binaural audio renderer. The renderer may operate stand-alone or in conjunction with an audio codec. Potentially compressed audio signals and their associated metadata (e.g., data specifying the location of the rendered audio source) may be provided to the audio renderer. The renderer may also be provided with head-tracking data obtained from a head-tracking device (e.g., one or more inside-out inertial-based tracking devices such as an accelerometer, gyroscope, compass, etc., or one or more outside-in based tracking devices such as LIDAR). Such head-tracking data may affect the metadata (i.e., rendering metadata) used for rendering (e.g., such that an audio object (source) is perceived at a fixed position in space independent of the listener's head rotation). The renderer also obtains HR filters to be used for binauralization. Embodiments of the present disclosure provide an efficient representation and method for HR filter generation based on WO2021 / 074294 or weighted basis vectors according to equation (1).
[0054] Scalar-valued function F n (θ, φ) is a set Θ of P elevation basis functions p (θ), p = 0, ..., p - 1 and a set Φ of Q azimuth basis functions q (φ) and is assumed to be a function g(·). As described in WO2021 / 074294, the set of azimuth basis functions or elevation basis functions may also vary for different p or q (e.g., the azimuth basis function Φ p,q (θ) varies the number of, which means that the number of azimuth basis functions Q p depends on p). In one embodiment, F n (θ, φ) may be selected as the product of Θ p (θ) and Φ p,q (φ). In other words, F nψ(θ, φ) = g(Θ p (θ), Φ p,q (φ)) = Θ p (θ)Φ p,q (φ) (9) is as follows.
[0055] Some embodiments of the present disclosure are perceptually based on the efficient structure of the (one or more) HR filter models, and the elevation basis function Θ p (θ) and the azimuth basis function Φ q (φ) are based on spatial sampling.
[0056] 1. HR Filter Model Design
[0057] First, the HR filter model (corresponding to Equation (1)) can be designed by the selection of the HR filter length K, the number P of elevation basis functions, the number Q p of azimuth basis functions, and the sets of basis functions Θ p (θ) and Φ p,q (φ). Each basis function can be smooth and impose more weight on some segments (angles) of the elevation modeling range and the azimuth modeling range (for example, some parts of [-90,..., 90] and [0,..., 360], respectively). Thus, for some segments of the modeling range, a basis function can be zero.
[0058] In some embodiments, the elevation basis function and the azimuth basis function are designed / selected using some properties for efficient use in HR filter modeling and efficient structured HR filter generation. The basis functions can be defined over a periodic modeling range (for example, continuous at the 0 / 360 degree azimuth boundary as shown in FIGS. 3(a) and 3(b), or an aperiodic range, for example, defined over [-90, 90] degrees of elevation as shown in FIG. 5).
[0059] Thus, according to some embodiments,
[0060] [Property 1] At least one of the basis functions has a first segment with a non-zero value and another segment with a zero value, and / or
[0061] [Property 2] The non-zero portion of said at least one of the basis functions a. is equal to the non-zero portion of another basis function, or b. has a non-zero portion length that is a unit fraction of the non-zero portion length of another basis function having the same shape, i.e., TIFF2025108446000017.tif10170, where L1 and L2 are the respective lengths and x = 1, 2, 3,..., and / or c. is symmetric, or d. is a mirror (inverse) of the non-zero portion of another basis function.
[0062] Having more basis functions with the same properties can lead to a more efficient implementation. However, there may be other factors such as modeling efficiency and performance that can also affect the selection of the basis functions. For example, depending on the sampling grid of the measured HR filter data, different numbers of basis functions should be selected to avoid obtaining an ill-determined system. The basis functions can generally be analytically described (e.g., as splines by polynomials).
[0063] In some embodiments, cubic B-spline functions (i.e., of degree 4 or degree 3) are used as the basis functions Φ p,q (φ) and Θ p (θ) for azimuth and elevation, respectively.
[0064] Figures 3(a) and 3(b) show the periodic B-spline basis functions for azimuth, and Figure 5 shows the corresponding standard B-spline basis functions for elevation. The points are marked with different symbols for better discrimination in the figures, but the functions are continuous and can be evaluated at any angle.
[0065] 2. HR Filter Modeling
[0066] Model design parameters that define the model (e.g., K, P, Q p , Θ p (θ) and Φ p,q (φ)) can be used later for HR filter modeling, where the model parameter α n,k can be estimated using a data fitting method such as the minimized least squares method (described, for example, in WO2021 / 074294).
[0067] 3. Basis Function Sampling
[0068] One aspect of the embodiments of the present disclosure is the perceptually motivated sampling of the basis functions Φ p,q (θ) and Θ p (θ). As research has shown, there is a minimum audible angle (MAA). Angle changes smaller than the MAA are not perceived. Based on this observation, the azimuth sampling interval ΔΦ and the elevation sampling interval ΔΘ can be selected. Research proposes ΔΦ = 1° and ΔΘ = 4° for transmission quality (i.e., inaudible loss), but larger sampling intervals can be selected as a compromise between the spatial accuracy requirements, memory requirements, and complexity requirements (regarding calculations) for HR filter evaluation.
[0069] If the selected sample spacing values ΔΦ, ΔΘ are larger than the MAA, interpolation can be used to generate a smoothly varying curve and avoid the step-like changes that can occur with a very coarsely spaced set of sample points (this approach further reduces memory usage but increases numerical complexity). Basis function sampling can generally be performed in a preprocessing stage, where the sampled basis functions to be used for HR filter evaluation are generated and stored in memory.
[0070] 3.1. Efficient Representation of Periodic B-Spline Basis Functions
[0071] Figures 3(a) and 3(b) show two examples of periodic B-spline functions for azimuth angles, each showing a set of basis functions that cover 360 degrees. As shown in the figures, in both examples, all equal symmetric non-zero parts of the basis functions (coherent with Properties 2a and 2c described above) are obtained, which always occurs as long as there is a constant spacing between knot points.
[0072] This means that each of the periodic B-spline basis functions can be efficiently represented by 1 / 2 of its non-zero shape (due to its symmetric properties). The B-spline basis functions can be calculated during runtime, but it is more efficient in terms of computational complexity to store in memory the pre-calculated shapes (i.e., numerical sampling) of the B-spline basis functions. On the other hand, it is generally desirable to minimize the memory requirements (i.e., the memory capacity required to store the pre-calculated shapes). The structure of the (one or more) B-spline basis functions according to embodiments of the present disclosure provides a good compromise between the computational complexity requirements and the memory requirements.
[0073] The number of HR filter measurement points is generally highest at 0° elevation angle and decreases towards ±90°, so fewer basis functions can be utilized towards the polar regions of the sampling sphere.
[0074] Using a varying number of azimuth B-spline basis functions for each elevation angle, a compact representation can be obtained for a set of periodic B-spline functions with different knot point intervals I K (p).
[0075] The knot point interval is for an integer decimation factor M When TIFF2025108446000018.tif11170, the non-zero part of the basis function is coherent with Property 2b described in Section 1 of the present disclosure above. Although no separate shape needs to be stored, only the decimation factor M is necessary to restore the shape. In this case, the maximum knot point interval I K (p1) of the shape has every M-th point corresponding to a sample of the shape with knot point interval I K (p2)=I K / M. This is shown in FIGS. 4(a) to 4(c).
[0076] FIGS. 4(a) to 4(c) show a compact representation of the B-spline basis functions of FIGS. 3(a) to 3(b). Since the non-zero part of the periodic basis function is symmetric, only 1 / 2 of the shape is required to represent the complete shape. Further, the B-spline basis functions of the sample points (○ (circle)) in FIG. 3(b) are obtained by subsampling the sample points (+ (plus)) in FIG. 3(a). In FIG. 4(a), + represents 1 / 2 of the sample points of the basis function in FIG. 3(a). In FIG. 4(b), ○ represents 1 / 2 of the sample points of the basis function in FIG. 3(b). FIG. 4(c) shows the overlaid shape function of (a) and (b). + represents the range of [0,..., 180] degrees, and ○ represents the range of [0,..., 90] degrees, but the shape function (b) can be obtained by subsampling the shape function (a).
[0077] As described above, in FIGS. 4(a) to 4(c), the sample points (○) of the shape in FIG. 3(b) can be obtained as every other sample point (+) for the shape in FIG. 3(a).
[0078] 3.2 Efficient Representation of Standard B-Spline Basis Functions
[0079] Regarding the periodic B-spline basis function, a compact representation can be obtained by sampling the standard B-spline basis function.
[0080] Figure 5 shows the standard elevation B-spline basis functions for P = 9. Some of the basis functions shown in Figure 5 are not symmetric as in the case of the periodic B-spline basis functions (e.g., the basis functions shown in FIGS. 3(a) and 3(b)), but it can be seen that the first and last spline functions (from the left) have a mirrored shape with respect to each other for the non-zero portions (coherent with Property 2d described in Section 1 of the present disclosure above). Similarly, the second and second-to-last non-zero spline functions have a mirrored shape with respect to each other, and the third and third-to-last non-zero spline functions have a mirrored shape with respect to each other. These properties of having a mirrored shape enable memory-efficient storage of the basis functions. Thus, in some embodiments, a constant spacing for knot points can be selected and used. For model evaluation, the stored shape can be read forward or backward depending on the segment being evaluated. The fourth to fourth-to-last (fourth, fifth, and sixth) B-spline basis functions shown in Figure 5 retain the same properties as the azimuth B-spline basis functions, i.e., they are symmetric and equal for the non-zero portions.
[0081] FIGS. 6(a) - 6(b) show a compact representation of the standard B-spline basis functions shown in Figure 5.
[0082] FIG. 6(a) shows a compact representation of the first and last basis functions of Figure 5. This corresponds to the mirrored shape of the non-zero portion of the last basis function.
[0083] FIG. 6(b) shows a compact representation of the second and second-to-last basis functions of Figure 5. This corresponds to the mirrored shape of the non-zero portion of the second-to-last basis function.
[0084] FIG. 6(c) shows a compact representation of the third and third-to-last basis functions of Figure 5. This corresponds to the mirrored shape of the non-zero portion of the third-to-last basis function.
[0085] Figure 6(d) shows a compact representation of the fourth, fifth, and sixth basis functions of Figure 5. This corresponds to 1 / 2 of the symmetric non-zero portion of the basis function.
[0086] Regardless of the total number of B-spline basis functions covering the modeling range (in this case, between -90° and 90°), only four independent non-zero B-spline basis function shapes are required. Further, one of these non-zero B-spline function shapes (e.g., the function shown in Figure 6(d)) is symmetric with respect to the periodic spline function, and thus only 1 / 2 of the non-zero portion needs to be stored.
[0087] 3.3 Storage in Memory
[0088] As a result of basis function sampling, a compact representation of the basis function (i.e., the basis function shape) is stored in memory along with shape metadata. The shape metadata can comprise information representing any one or combination of the following. 1. The number of basis functions (the number of azimuth basis functions can be different for different elevation angles), 2. The starting point of each basis function (within the modeling interval), 3. A shape index for each basis function (identifying which of the stored shapes should be used for the basis function), 4. A shape resampling factor M for each basis function, 5. A flip indicator for each basis function (indicating whether the stored shape should be flipped for that particular basis function), 6. The basis function structure such as B-spline, and 7. The width of the non-zero portion of each basis function.
[0089] In some embodiments, if the flip indicator indicates that the stored shape needs to be flipped, the shape stored in the storage medium can be read backwards from the storage medium so that the flipped shape is provided to the renderer.
[0090] In some embodiments, some parameters (e.g., the inversion indicator and basis function structure) may be stored in the renderer and need not be transmitted, especially when the model structure is already known to the renderer. For example, when a standard cubic B-spline is used as in the case of FIG. 5, it is known that if both basis function sampling and structured HR filter generation assume that the first four shapes (the first three shapes and half of the fourth shape) are stored in that order, there is no need to signal that the last three basis functions need to be inverted. It may further be known that all basis functions between the first and last three basis functions can be composed of the fourth stored shape. In the case of B-splines, the shape metadata may instead include information regarding knot points. It may also be known that a periodic B-spline function is used for the azimuth basis function and a standard B-spline function is used for the elevation angle. This is an example of how shape metadata parameters may be stored in different storage media.
[0091] Furthermore, the HR filter model parameter α n,k is stored in memory along with the basis function shapes and corresponding shape metadata. In other embodiments, the HR filter model parameters, basis function shapes, and / or shape metadata may be stored in different storage media.
[0092] 4. HR Filter Generation
[0093] Based on the stored shapes and parameters, structured HR filter generation can be performed by reading the basis function shapes from memory, correctly applying them for each basis function based on the shape metadata, and avoiding unnecessary computational complexity (e.g., unnecessary multiplications and additions), thereby resulting in a very efficient evaluation of the HR filter using the HR filter model parameter α n,k .
[0094] Sampling of the B-spline basis functions can reduce the computational complexity (involved in audio rendering) by means of a structured tabulation of the sampled basis functions, but the HR filter generation (or model evaluation) can also be optimized to further reduce the computational complexity.
[0095] Assuming the structure of the azimuth basis function and elevation basis function (i.e., the cubic B-spline basis function) according to FIGS. 3 and 5 for all directions (θ, φ), there are at most four non-zero B-spline basis functions for all azimuths and elevations to be evaluated. Thus, for the evaluation of F n (θ, φ) in Equation (8), there will be at most 4·4 = 16 non-zero components. Thus, the filter evaluation in Equation (5) can be reduced to the following equation. TIFF2025108446000019.tif22170 Here, TIFF2025108446000020.tif9170 represents all non-zero components of F n (θ, φ).
[0096] Compared to the complete evaluation of N = P·Q (where, assuming a constant azimuth basis function, i.e., Q p = Q for all p), the HR filter generation based on Equation (9) provides a significant reduction in complexity, which increases as more basis functions are used to model the HR filter data.
[0097] At most points, there are four non-zero basis functions, but at the knot points, fewer than four basis functions contribute to the non-zero components.
[0098] The following describes a method for providing an optimized model evaluation for the generation of the HR filter.
[0099] 4.1 Basis Evaluation for Periodic B-Spline Basis Functions (for Azimuth)
[0100] (1) Knot segment index I n Determine (θ, φ). TIFF2025108446000021.tif17170 Here, φ is the azimuth angle to be evaluated, and I m (0) is the azimuth angle at the first knot point, and I K (p) is the knot point interval for the azimuth B-spline function at the elevation angle of index p.
[0101] (2) Determine the closest segment sample point. TIFF2025108446000022.tif18170 Here, round() is the rounding function, and N s (p) is the number of samples per segment (for example, TIFF2025108446000023.tif12170), and M(p) is the decimation factor for the elevation angle of index p. An example of a suitable rounding function is as follows. TIFF2025108446000024.tif18170 Here, TIFF2025108446000025.tif8170 indicates the floor function that outputs the largest integer less than or equal to its input.
[0102] (3) Determine the number of non-zero basis functions for the azimuth angle TIFF2025108446000026.tif9170. TIFF2025108446000027.tif57170
[0103] (4) Calculate the B-spline sample values and shape indices. TIFF2025108446000028.tif66170 Here, S p is the 1 / 2 sampled shape function at the elevation angle p, subsampled by the factor M(p) (described in Section 3.1 above). The stored shape values Index of TIFF2025108446000029.tif at 10170 TIFF2025108446000030.tif at 10170 is also stored. Q p is the total number of azimuth B - spline basis functions for the elevation angle index p. mod(·) is the modulo function used to determine whether the evaluated azimuth φ is on a knot point.
[0104] 4.2 Basis evaluation for the standard B - spline function (for elevation angle)
[0105] (1) Knot segment index I n Determines (θ, p). TIFF2025108446000031.tif at 18170 where θ is the elevation angle to be evaluated, and I m (0) is the elevation angle at the first knot point, and I K is the knot point interval for the elevation B - spline function.
[0106] (2) Determine the closest segment sample point. TIFF2025108446000032.tif at 18170 where round() is the rounding function, and N s is the number of samples per segment (e.g., TIFF2025108446000033.tif at 12170). The rounding function can be the same as that used for the periodic B - spline basis functions.
[0107] (3) Number of non - zero basis functions TIFF2025108446000034.tif at 10170 to determine TIFF2025108446000035.tif at 47170
[0108] At the first and last knot points, TIFF2025108446000036.tif at 10170 can also be used.
[0109] Calculate B-spline sample values and shape indices TIFF2025108446000037.tif131170 Here, I S is the related sampled shape function at the elevation angle p TIFF2025108446000038.tif8170 is the index representing it.
[0110] P is the total number of elevation B-spline basis functions. If the basis function index (i + I n ) is greater than P - 4, the shape is read backwards. Otherwise, if the shape index, which can occur in the case of symmetric shapes, is greater than the length of the shape stored, the shape is also read backwards. The stored shape value TIFF2025108446000039.tif10170's index TIFF2025108446000040.tif10170 is also stored. len(·) determines the length of the input vector, and min(·,·), max(·,·) determine the minimum and maximum values of the input arguments respectively.
[0111] 4.3 HR filter evaluation
[0112] When the azimuth B-spline basis function and the elevation B-spline basis function are evaluated, F n (θ,φ) can be determined as follows. TIFF2025108446000041.tif41170
[0113] Then, each HR filter coefficient TIFF2025108446000042.tif9170 can be determined as follows. TIFF2025108446000043.tif23170 where the HR filter tap index k = 0,..., K - 1.
[0114] 5. Binocular rendering
[0115] In some embodiments, the methods described above may be used for the zero-time delay portion of the HR filter, i.e., to exclude the onset time delay of each filter or the delay difference between the left and right HR filters due to the interaural time difference. The methods described above may be utilized, in an equivalent manner, to evaluate the interaural time difference modeled in a similar fashion by B-spline basis functions (as described, for example, in WO2021 / 074294). In such a case, a single ITD is determined, i.e., K = 1 as opposed to an HR filter where the number of filter taps is K≫1. The resulting interaural time difference may then be taken into account either by modification of the generated HR filter ( TIFF2025108446000044.tif10170) or by applying an offset during the filtering step.
[0116] Separate weight matrices TIFF2025108446000045.tif8170 are used, but the same basis functions, i.e., the same TIFF2025108446000046.tif9170 are used to generate HR filters TIFF2025108446000047.tif8170 for the left and right sides, respectively. Thus, TIFF2025108446000048.tif9170 is only evaluated once for each updated direction (θ,φ).
[0117] Next, a binaural audio signal for the mono source u(n) can be obtained by filtering the audio source signal using a left HR filter and a right HR filter, respectively (e.g., by using well-known techniques). The filtering can be performed using normal convolution techniques in the time domain or in a more optimized manner, e.g., using the overlap-add technique in the discrete Fourier transform (DFT) domain when the filter is long. K = 96 taps corresponds to a 2 ms filter for a 48 kHz sample rate.
[0118] Embodiments of the present disclosure are based on two main categories of optimization, pre-computed sampled basis functions and structured HR filter evaluation. In some embodiments, the sampled basis functions are computed and stored in memory in a preprocessing stage. Also, the structured HR filter evaluation can be performed at runtime within the renderer or pre-computed and stored as a set of sampled HR filters. Since the memory required to store a set of HR filters sampled with high-precision azimuth and elevation resolution is large, in some embodiments, the HR filters are evaluated during runtime.
[0119] FIG. 7 shows an exemplary system 700 according to some embodiments. The system 700 includes a preprocessor 702 and an audio renderer 704. The preprocessor 702 and the audio renderer 704 can be included in the same entity or in different entities. Also, different modules (e.g., 710, 712, 714, and / or 716) included in the preprocessor 702 can be included in the same entity or in different entities, and different modules (718 and / or 720) included in the audio renderer 704 can be included in the same entity or in different entities.
[0120] In one example, the pre-processor 702 is included within any one of an audio encoder, a network entity (e.g., in the cloud), and an audio decoder (i.e., the audio renderer 704). The audio renderer 704 can be included within any electronic device capable of generating an audio signal (e.g., a desktop, laptop computer, tablet, mobile phone, head-mounted display, XR simulation system, etc.).
[0121] The pre-processor 702 includes an HR filter model design module 710, an HR filter modeling module 712, a basis function sampling module 714, and a memory 716. The HR filter model design module 710 is configured to output design data 720 to the HR filter modeling module 712. The HR filter modeling module 712 can receive HR filter data 722 and obtain an HR filter model based on the received design data 720 and the received HR filter data 722. In some embodiments, the HR filter model is designed according to the properties (1) and (2)(a)-(2)(d) described above.
[0122] Obtaining the HR filter model can include selecting a basis function structure, i.e., selecting a set of basis functions for azimuth (the "azimuth basis functions") and / or a set of basis functions for elevation (the "elevation basis functions"). The azimuth basis functions can be selected to be periodic over a modeling range (e.g., between 0° and 360°). The modeling range can be divided into N seg equal-sized segments defined by knot points. The basis functions can be selected such that at least one basis function is 0 in one or more segments. Also, the basis functions can be such that at most N b <{P,Q p} basis functions are non-zero within segment i (i.e., at most less than (P)) TIFF2025108446000049.tif has 9170 non-zero elevation basis functions and / or at most (smaller than Q p than) TIFF2025108446000050.tif can be selected such that 9170 azimuth basis functions are non-zero, where P is the total number of elevation basis functions and Q p is the total number of azimuth basis functions for elevation p. Further, the basis functions (azimuth basis functions and / or elevation basis functions) can be selected such that for the purpose of utilizing the optimization techniques described in this disclosure, the non-zero portions of some of the basis functions are symmetric, mirror, or subsampled versions of the non-zero portions of other basis functions.
[0123] After obtaining the HR filter model, the HR filter modeling module 712 outputs the HR filter model data 724 to the basis function sampling module 714. The HR filter model data 724 can indicate the obtained HR filter model (i.e., the selected basis function structure). Based on the received HR filter model data 724, the basis function sampling module 714 samples the basis functions at intervals ΔΦ (for azimuth basis functions) and ΔΘ (for elevation basis functions) and can obtain a compact representation of the (non-zero portions of) the azimuth basis functions and / or elevation basis functions. Since not all portions of the basis functions are required to represent the basis functions, a compact representation of the basis functions can be obtained. For example, in the case of symmetric non-zero portions of the basis functions, only 1 / 2 of the shape of the basis function is required to represent the shape. In the case of mirror or inverted non-zero portions of the basis functions, only one of the mirror portions is required to represent the shape of the basis function. In the case of subsampled non-zero portions of the basis functions, only the largest shape is required to represent the shape of the basis function.
[0124] After obtaining a compact representation of the basis functions, the basis function sampling module 714 may store the basis function shape data 728 and the shape metadata 730 in the memory 716. The basis function shape data 728 may indicate the shape of the compact representation of the basis functions. The shape metadata 730 may include information regarding the structure of the compact representation with respect to the HR filter model basis functions. For example, the shape metadata 730 may include information regarding the shape, orientation (e.g., whether it is inverted), and subsampling factor M with respect to the model basis functions. Detailed information regarding the shape metadata 730 was provided above in Section 3.3 of the present disclosure.
[0125] In addition to the basis function shape data 728 and the shape metadata 730, the memory 716 may also store additional HR filter model parameters 726 (e.g., the α parameter).
[0126] The audio renderer 704 includes a structured HR filter generator 718 and a binaural renderer 720. The structured HR filter generator 718 reads the basis function shape data 732, the shape metadata 734, and one or more additional HR filter model parameters 736 from the memory 716 and receives the rendering metadata 738. The basis function shape data 732 may be the same as or related to the basis function shape data 728. Similarly, the shape metadata 734 and the one or more model parameters 736 may be the same as or related to the shape metadata 730 and the one or more model parameters 726, respectively.
[0127] Based on (i) the basis function shape data 732, (ii) the shape metadata 734, (iii) one or more additional HR filter model parameters 736, and (iv) the rendering metadata 738, the structured HR filter generator 718 may generate HR filter information 740 indicating an HR filter. The rendering metadata 738 may define the direction (θ, φ) to be evaluated.
[0128] Figure 8 shows an exemplary process 800 according to some embodiments. Process 800 may be implemented by a structured HR filter generator 718 included in audio renderer 704.
[0129] Process 800 may begin at step s802. In step s802, the structured HR filter generator 718 identifies segments within the modeling range based on the received rendering metadata 738. For example, the rendering metadata 738 defines a particular direction (θ, φ) to be evaluated, and the generator 718 identifies the segment to which the defined direction belongs.
[0130] After performing step s802, in step s804, the structured HR filter generator 718 identifies sample points within the segment identified in step s802.
[0131] After performing step s804, in step s806, the generator 718 identifies a compact representation of the basis functions (i.e., azimuth basis functions and elevation basis functions) based on the basis function shape data 732.
[0132] After performing step s806, in step s808, the generator 718 determines whether the identified compact representation should be read normally, inverted, or subsampled according to the subsampling factor M, and performs inversion and / or subsampling if necessary.
[0133] After performing step s808, in step s810, the generator 718 evaluates at most N b basis functions. Such evaluation is for at most N bObtaining the sample values within each of the compact representations of the non-zero basis functions. A detailed description of how the basis functions are evaluated was provided in Sections 4.1 and 4.2 above.
[0134] After performing step s810, in step s812, based on (i) the obtained azimuth basis function values, (ii) the obtained elevation basis function values, and (iii) the additional model parameter(s) 736 (e.g., parameter α), the structured HR filter generator 718 generates an HR filter. The HR filter can be generated separately as the sum of the multiplied values of the azimuth basis function values and the elevation basis function values weighted by the corresponding model weight parameter (α) for each filter tap k. A detailed description of how the HR filter is generated was provided in Section 4.3 above.
[0135] The HR filters (for the left and right sides) generated by the structured HR filter generator 718 are then provided to the binaural renderer 720.
[0136] Using the HR filter generated by the generator 718, the binaural renderer 720 binauralizes the audio signal 742, i.e., generates two audio output signals (for the left and right sides).
[0137] FIG. 9 shows an exemplary system 900 for creating sounds for an XR scene. The system 900 includes a controller 901, a signal modifier 902 for a first audio stream 951, a signal modifier 903 for a second audio stream 952, a speaker 904 for the first audio stream 951, and a speaker 905 for the second audio stream 952. Although two audio streams, two modifiers, and two speakers are shown in FIG. 9, this is for illustrative purposes only and does not limit the embodiments of the present disclosure in any way. For example, in some embodiments, there may be N audio streams corresponding to N audio objects to be rendered, and the audio streams may include a single mono signal corresponding to a single audio object. Further, FIG. 9 shows that the system 900 receives and modifies the first audio stream 951 and the second audio stream 952 separately, but the system 900 may receive a single audio stream representing a plurality of audio streams. The first audio stream 951 and the second audio stream 952 may be the same or different. If the first audio stream 951 and the second audio stream 952 are the same, a single audio stream may be split into two audio streams that are equivalent to a single audio stream, thereby generating the first audio stream 951 and the second audio stream 952.
[0138] The controller 901 may be configured to receive one or more parameters and trigger the modifiers 902 and 903 to perform modifications to the first audio stream 951 and the second audio stream 952 based on the received parameters (e.g., increase or decrease the volume level according to a gain function). The received parameters are (1) information 953 regarding the position of the listener (e.g., distance and direction to the audio source), and (2) metadata 954 regarding the audio source. The information 953 may include the same information as the rendering metadata 738 shown in FIG. 7. Similarly, the metadata 954 may include the same information as the shape metadata 734 shown in FIG. 7.
[0139] In some embodiments of the present disclosure, the information 953 may be provided from one or more sensors included in the XR system 1000 shown in FIG. 10A. As shown in FIG. 10A, the XR system 1000 is configured to be worn by a user. As shown in FIG. 10B, the XR system 1000 may include an orientation sensing unit 1001, a position sensing unit 1002, and a processing unit 1003 coupled to the controller 1001 of the system 1000. The orientation sensing unit 1001 is configured to detect changes in the orientation of the listener and provide information regarding the detected changes to the processing unit 1003. In some embodiments, the processing unit 1003 determines the absolute orientation (with respect to some coordinate system) on the premise of the detected changes in the orientation detected by the orientation sensing unit 1001. There may also be different systems for determining orientation and position, such as the HTC Vive system using a lighthouse tracker (lidar). In one embodiment, the orientation sensing unit 1001 may determine the absolute orientation (with respect to some coordinate system) on the premise of the detected changes in the orientation. In this case, the processing unit 1003 may simply multiplex the absolute orientation data from the orientation sensing unit 1001 and the absolute position data from the position sensing unit 1002. In some embodiments, the orientation sensing unit 1001 may include one or more accelerometers and / or one or more gyroscopes. The type of the XR system 1000 and / or the components of the XR system 1000 shown in FIGS. 10A and 10B are provided for illustrative purposes only and do not limit the embodiments of the present disclosure in any way. For example, although an XR system 1000 including a head-mounted display covering the user's eyes is shown, the system may not be equipped with such a display in, for example, an audio-only implementation.
[0140] FIG. 11 is a flowchart showing a process 1100 for generating an HR filter for audio rendering. The process 1100 may begin at step s1102.
[0141] Step s1102 includes generating HR filter model data indicating an HR filter model. Generating the HR filter model data may include selecting at least one set of one or more basis functions.
[0142] Step s1104 includes sampling the one or more basis functions (s1104) based on the generated HR filter model data.
[0143] Step s1106 includes generating first basis function shape data and shape metadata based on the generated HR filter model data. The first basis function shape data identifies one or more compact representations of the one or more basis functions, and the shape metadata includes information regarding the structure of the one or more compact representations regarding the one or more basis functions.
[0144] Step s1108 includes providing the generated first basis function shape data and shape metadata for storage in one or more storage media.
[0145] Step s1110 includes detecting the occurrence of a triggering event.
[0146] Step s1112 includes outputting second basis function shape data and shape metadata for audio rendering as a result of detecting the occurrence of the triggering event.
[0147] Such a triggering event may indicate that a head-related (HR) filter should be generated for audio rendering, which may be induced from an audio renderer when an HR filter is required, for example, to render an audio frame or to prepare for rendering by generating an HR filter to be stored in memory for later use. In some embodiments, the triggering event is merely a determination to retrieve basis function shape data and / or shape metadata from one or more storage media.
[0148] In some embodiments, the at least one set of one or more basis functions is subject to the following conditions, (i) the at least one set of one or more basis functions is periodic over the modeling range, (ii) at least one basis function included in the at least one set is 0-valued in one or more segments included in the modeling range, (iii) at most N basis functions included in the at least one set are non-zero in segments included in the modeling range, where N is a positive integer and is less than the total number of basis functions included in the at least one set, and (iv) at least one non-zero portion of the one or more basis functions is either (1) symmetric or a mirror of another non-zero portion of the one or more basis functions, or (2) a subsampled version of another non-zero portion of the one or more basis functions is selected such that any one or combination thereof is satisfied.
[0149] In some embodiments, the compact representation of the one or more basis functions shows the shape of the non-zero portions of the one or more basis functions, and the shape of the non-zero portions of the one or more basis functions is symmetric or a mirror of the shape of another non-zero portion of the one or more basis functions.
[0150] In some embodiments, the shape metadata includes the following information: (i) The number of basis functions, and (ii) The starting point of each basis function, and (iii) One or more shape indices that each identify a specific shape to be used for audio rendering, and (iv) A shape resampling factor for one or more of the basis functions, and (v) An inversion indicator for one or more of the basis functions, the inversion indicator indicating whether an inverted version of the one or more compact representations of the one or more basis functions stored in the one or more storage media should be obtained, and (vi) A basis function structure, and (vii) The width of the non-zero portion of each basis function and comprises any one or combination thereof.
[0151] In some embodiments, the method further includes providing additional HR filter model parameters for storage in the one or more storage media.
[0152] In some embodiments, the method is performed by a pre-processor prior to the occurrence of an event that triggers audio rendering.
[0153] In some embodiments, the method is performed by a pre-processor included in a separate and distinct network entity from the audio renderer.
[0154] In some embodiments, the second basis function shape data and the shape metadata are used to generate an HR filter.
[0155] In some embodiments, the first basis function shape data and the second basis function shape data are the same.
[0156] In some embodiments, the second basis function shape data identifies a converted version of the one or more compact representations of the one or more basis functions, and the converted version of the one or more compact representations of the one or more basis functions is a symmetric or mirror version and / or a subsampled version of the one or more compact representations of the one or more basis functions.
[0157] FIG. 12 is a flowchart showing a process 1200 for generating an HR filter for audio rendering. The process 1200 may begin at step s1202.
[0158] Step s1202 includes obtaining shape metadata indicating whether to obtain a converted version of the one or more compact representations of the one or more basis functions.
[0159] Step s1204 includes obtaining basis function shape data that identifies (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions.
[0160] Step s1206 includes generating an HR filter by using (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions based on the obtained shape metadata and the obtained basis function shape data.
[0161] In some embodiments, the method further includes obtaining data corresponding to the one or more compact representations of the one or more basis functions from a storage medium after obtaining shape metadata indicating how to obtain a converted version of the one or more compact representations of the one or more basis functions. The data is obtained in a predefined manner such that a converted version of the one or more compact representations of the one or more basis functions is obtained.
[0162] In some embodiments, the method includes receiving data identifying the one or more compact representations of the one or more basis functions and providing the received data for storage in another storage medium. Obtaining basis function shape data identifying a converted version of the one or more compact representations of the one or more basis functions includes reading the stored received data from the another storage medium in a predefined manner.
[0163] In some embodiments, the converted version of the one or more compact representations of the one or more basis functions is a symmetric or mirror version and / or a subsampled version of the one or more compact representations of the one or more basis functions.
[0164] In some embodiments, obtaining data in a predefined manner includes (i) obtaining data in a predefined sequence and / or (ii) obtaining data partially.
[0165] In some embodiments, the converted version of the compact representation of the one or more basis functions is a symmetric or mirror version and / or a subsampled version of the compact representation of the one or more basis functions.
[0166] In some embodiments, the method further includes obtaining rendering metadata indicating a particular direction or location to be evaluated and identifying sample points related to the particular direction or location to be evaluated based on the obtained rendering metadata.
[0167] In some embodiments, the one or more compact representations of the one or more basis functions indicate the shape of the non-zero portions of the one or more basis functions, and the shape of the non-zero portions of the one or more basis functions is symmetric or mirror with respect to the shape of another non-zero portion of the one or more basis functions.
[0168] In some embodiments, the shape metadata comprises any one or combination of the following information: (i) the number of basis functions, (ii) the start point of each basis function, (iii) one or more shape indices each identifying a particular shape to be used for HR filter generation, (iv) a shape resampling factor for one or more basis functions, (v) an inversion indicator for one or more basis functions, the inversion indicator indicating whether an inverted version of the one or more compact representations of the one or more basis functions stored in a storage medium should be obtained, (vi) a basis function structure, and (vii) the width of the non-zero portion of each basis function.
[0169] In some embodiments, the method further includes obtaining an audio signal and filtering the obtained audio signal to generate a left audio signal for the left side and a right audio signal for the right side using the generated HR filter. The left audio signal and the right audio signal are associated with a particular direction and / or location indicated by the rendering metadata.
[0170] FIG. 13 is a block diagram of an apparatus 1300 according to some embodiments for implementing the preprocessor 702 or the audio renderer 704 shown in FIG. 7. As shown in FIG. 13, the apparatus 1300 may include a processing circuit (PC) 1302 including one or more processors (P) 1355 (e.g., one or more other processors such as a general-purpose microprocessor and / or an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.), where the processors may be co-located in a single housing or in a single data center or may be geographically distributed (i.e., the apparatus 1300 may be a distributed computing device), a processing circuit (PC) 1302, at least one network interface 1348, where each network interface 1348 includes a transmitter (Tx) 1345 and a receiver (Rx) 1347 for enabling the apparatus 1300 to send data to and receive data from other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 1348 is (directly or indirectly) connected (e.g., the network interface 1348 may be wirelessly connected to the network 110, in which case the network interface 1348 is connected to an antenna configuration), at least one network interface 1348, and one or more storage units (also referred to as a “data storage system”) 1308 that may include one or more non-volatile memory devices and / or one or more volatile memory devices. In embodiments where the PC 1302 includes a programmable processor, a computer program product (CPP) 1341 may be provided. The CPP 1341 includes a computer-readable medium (CRM) 1342 that stores a computer program (CP) 1343 comprising computer-readable instructions (CRI) 1344. The CRM 1342 may be a non-transitory computer-readable medium such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., a random access memory, a flash memory), etc.In some embodiments, when the CRI 1344 of the computer program 1343 is executed by the PC 1302, the CRI is set to cause the apparatus 1300 to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, the apparatus 1300 may be set to perform the steps described herein without the need for code. That is, for example, the PC 1302 may simply consist of one or more ASICs. Thus, the features of the embodiments described herein may be implemented in hardware and / or software.
[0171] Although various embodiments have been described herein, it should be understood that those embodiments are presented by way of example and not limitation. Accordingly, the breadth and scope of the present disclosure should not be limited by any of the exemplary embodiments described above. Moreover, unless otherwise indicated herein or clearly contradicted by context, any combination of all possible variations of the elements described above in all conceivable permutations is encompassed by the present disclosure.
[0172] Furthermore, although the processes and message flows described above and shown in the drawings are presented as a sequence of steps, this was done for illustrative purposes only. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be rearranged, and some steps may be performed in parallel.
[0173] 6. Abbreviations TIFF2025108446000051.tif132170
Claims
1. A method (1100) for generating a head-related (HR) filter for audio rendering, the method comprising: generating (s1102) HR filter model data representing an HR filter model, wherein generating the HR filter model data includes selecting at least one set of one or more basis functions, generating (s1102) HR filter model data; based on the generated HR filter model data, (i) sampling (s1104) the one or more basis functions and (ii) generating (s1106) first basis function shape data and shape metadata, wherein the first basis function shape data identifies one or more compact representations of the one or more basis functions, and the shape metadata includes information regarding the structure of the one or more compact representations of the one or more basis functions, generating (s1106) first basis function shape data and shape metadata; providing (s1108) the generated first basis function shape data and the shape metadata for storage in one or more storage media; A method (1100) comprising:
2. The method further comprising: detecting (s1110) the occurrence of a triggering event; in response to detecting the occurrence of the triggering event, outputting (s1112) second basis function shape data and the shape metadata for the audio rendering; The method according to claim 1, further comprising:
3. The at least one set of one or more basis functions satisfies the following conditions: (i) the at least one set of one or more basis functions is periodic over a modeling range; (ii) at least one basis function included in the at least one set has a value of 0 in one or more segments included in the modeling range; (iii) at most N basis functions included in the at least one set are non-zero in segments included in the modeling range, where N is a positive integer and is less than the total number of basis functions included in the at least one set; and (iv) At least one non-zero portion of the one or more basis functions is either (1) symmetric or mirror to another non-zero portion of the one or more basis functions, or (2) a subsampled version of another non-zero portion of the one or more basis functions, or a combination of any one of them. The method according to claim 1 or 2, wherein the selection is made such that any one or combination of the above is satisfied.
4. The compact representation of the one or more basis functions indicates the shape of the non-zero portions of the one or more basis functions. The shape of the non-zero portions of the one or more basis functions is symmetric or mirror to the shape of another non-zero portion of the one or more basis functions. The method according to any one of claims 1 to 3.
5. The shape metadata includes the following information: (i) The number of basis functions, (ii) The starting point of each basis function, (iii) One or more shape indices each identifying a specific shape to be used for audio rendering, (iv) A shape resampling factor for one or more basis functions, (v) An inversion indicator for one or more basis functions, where the inversion indicator indicates whether a reversed version of the one or more compact representations of the one or more basis functions stored in the one or more storage media should be obtained. (vi) The basis function structure, (vii) The width of the non-zero portion of each basis function. The method according to any one of claims 1 to 4, comprising any one or combination of the above.
6. Further comprising providing additional HR filter model parameters for storage in the one or more storage media. The method according to any one of claims 1 to 5.
7. The method according to any one of claims 1 to 6, wherein the method is performed by a pre-processor prior to the occurrence of an event that triggers the audio rendering.
8. The method according to any one of claims 1 to 7, wherein the method is performed by a pre-processor included in a separate network entity distinct from the audio renderer.
9. The method according to any one of claims 1 to 8, wherein the second basis function shape data and the shape metadata are used to generate the HR filter.
10. The method according to any one of claims 1 to 9, wherein the first basis function shape data and the second basis function shape data are the same.
11. The second basis function shape data identifies a converted version of the one or more compact representations of the one or more basis functions, wherein the converted version of the one or more compact representations of the one or more basis functions is a symmetric or mirror version and / or a subsampled version of the one or more compact representations of the one or more basis functions, The method according to any one of claims 1 to 9.
12. A method (1200) for generating a head-related (HR) filter for audio rendering, the method comprising: obtaining (s1202) shape metadata indicating whether a converted version of one or more compact representations of one or more basis functions should be obtained; obtaining (s1204) basis function shape data that identifies (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions; generating (s1206) the HR filter by using (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions, based on the obtained shape metadata and the obtained basis function shape data; The method (1200) comprising.
13. The method comprises: after obtaining the shape metadata indicating how to obtain the converted version of the one or more compact representations of the one or more basis functions, obtaining data corresponding to the one or more compact representations of the one or more basis functions from a storage medium; further comprising The data is obtained in a predefined manner such that a converted version of the one or more compact representations of the one or more basis functions is obtained. The method according to claim 12.
14. The method is receiving data identifying the one or more compact representations of the one or more basis functions, and providing the received data for storage in a storage medium and obtaining basis function shape data identifying a converted version of the one or more compact representations of the one or more basis functions includes reading the stored data from the storage medium in a predefined manner. The method according to claim 12.
15. The converted version of the one or more compact representations of the one or more basis functions is a symmetric or mirror version and / or a subsampled version of the one or more compact representations of the one or more basis functions. The method according to any one of claims 12 to 14.
16. Obtaining the data in the predefined manner includes (i) obtaining the data in a predefined sequence and / or (ii) obtaining the data partially. The method according to any one of claims 13 to 15.
17. The method is obtaining rendering metadata indicating a specific direction or location to be evaluated, and identifying sample points related to the specific direction or location to be evaluated based on the obtained rendering metadata and further includes. The method according to any one of claims 12 to 16.
18. The one or more compact representations of the one or more basis functions indicate the shape of the non-zero portions of the one or more basis functions, and the shape of the non-zero portions of the one or more basis functions is symmetric or mirror to the shape of another non-zero portion of the one or more basis functions. The method according to any one of claims 12 to 17.
19. The shape metadata includes the following information: (i) the number of basis functions, and (ii) the starting point of each basis function, (iii) one or more shape indices each identifying a specific shape to be used for generating the HR filter; (iv) a shape resampling factor for one or more basis functions; (v) an inversion indicator for one or more basis functions, the inversion indicator indicating whether the inverted version of the one or more compact representations of the one or more basis functions stored in the storage medium should be obtained; an inversion indicator for one or more basis functions; (vi) a basis function structure; (vii) the width of the non-zero portion of each basis function The method according to any one of claims 12 to 18, comprising any one or a combination of the above.
20. The method is obtaining an audio signal; filtering the obtained audio signal to generate a left audio signal for the left side and a right audio signal for the right side using the generated HR filter further comprising wherein the left audio signal and the right audio signal are associated with the specific direction and / or location indicated by the rendering metadata; The method according to any one of claims 12 to 19.
21. A computer program (1343) comprising instructions for causing a processing circuit (1302) to perform the method according to any one of claims 1 to 20 when executed by the processing circuit.
22. A carrier comprising the computer program according to claim 21, the carrier being one of an electronic signal, an optical signal, a wireless signal, or a computer-readable storage medium (1342).
23. An apparatus (1300) for generating a head-related (HR) filter for audio rendering, the apparatus comprising: generating (s1102) HR filter model data indicative of an HR filter model, the generating (s1102) HR filter model data including selecting at least one set of one or more basis functions; Based on the generated HR filter model data, (i) sampling the one or more basis functions (s1104); and (ii) generating first basis function shape data and shape metadata (s1106), wherein the first basis function shape data identifies one or more compact representations of the one or more basis functions, and the shape metadata includes information regarding the structure of the one or more compact representations of the one or more basis functions, generating first basis function shape data and shape metadata (s1106); providing the generated first basis function shape data and the shape metadata for storage in one or more storage media (s1108); An apparatus (1300) configured to perform the above. **Claim 24** The apparatus according to claim 23, further configured to perform the method according to any one of claims 2 to 11. **Claim 25** An apparatus (1300) for generating a head-related (HR) filter for audio rendering, the apparatus being configured to: obtaining shape metadata indicating whether to obtain a converted version of one or more compact representations of one or more basis functions (s1202); obtaining basis function shape data that identifies (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions (s1204); generating the HR filter by using (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions based on the obtained shape metadata and the obtained basis function shape data (s1206); An apparatus (1300) configured to perform the above. **Claim 26** The apparatus according to claim 25, further configured to perform the method according to any one of claims 13 to 20. **Claim 27** An apparatus (1300) for representing an audio object in an extended reality scene, the apparatus including: a storage unit (1308); A processing circuit (1302) coupled to the memory unit and comprising, the apparatus being generating HR filter model data representing an HR filter model (s1102), wherein generating the HR filter model data includes selecting at least one set of one or more basis functions, generating the HR filter model data (s1102); and based on the generated HR filter model data, (i) sampling the one or more basis functions (s1104) and (ii) generating first basis function shape data and shape metadata (s1106), wherein the first basis function shape data identifies one or more compact representations of the one or more basis functions, and the shape metadata includes information regarding the structure of the one or more compact representations of the one or more basis functions, generating first basis function shape data and shape metadata (s1106); and providing the generated first basis function shape data and the shape metadata for storage in one or more storage media (s1108) and an apparatus (1300) configured to perform. **Claim 28** The apparatus according to claim 27, wherein the memory unit (1308) comprises a memory (1342) storing instructions for configuring the apparatus to perform the method according to any one of claims 2 to 11. **Claim 29** An apparatus (1300) for representing an audio object in an extended reality scene, the apparatus comprising a memory unit (1308), a processing circuit (1302) coupled to the memory unit, and comprising, the apparatus being acquiring shape metadata indicating whether a converted version of one or more compact representations of one or more basis functions should be acquired (s1202); and acquiring basis function shape data identifying (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions (s1204); Based on the acquired shape metadata and the acquired basis function shape data, generating an HR filter (s1206) by using (i) the one or more compact representations of the one or more basis functions or (ii) the converted version of the one or more compact representations of the one or more basis functions An apparatus (1300) configured to perform the above. **Claim 30** The apparatus according to claim 29, wherein the storage unit (1308) comprises a memory (1342) storing instructions for configuring the apparatus to implement the method according to any one of claims 13 to 20.
Citation Information
Patent Citations
Method and system for adjusting hrtf measured for smooth 3d digital audio
JP2000166000A