Head Related (HR) filter

By segmenting the alpha matrix of the HR filter model and dynamically adjusting the allocation of computing resources, the problem of high computational complexity of the HR filter is solved, achieving efficient real-time rendering under limited resources and ensuring the accuracy of spatial audio scenes.

CN116597847BActive Publication Date: 2026-05-15TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310553744.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-17
Filing Date
2020-11-20
Publication Date
2026-05-15
Estimated Expiration
2040-11-20

AI Technical Summary

Technical Problem

Existing technologies suffer from high computational complexity in generating spatial audio scenes due to the high computational complexity of HR filters, making it difficult to achieve efficient real-time rendering with limited computing resources, especially in mobile devices or complex audio scenes.

Method used

By dividing the alpha matrix of the HR filter model into multiple parts and using only some sub-vectors for computation, the allocation of computational resources is dynamically adjusted to reduce computational complexity and maintain the accuracy of spatial awareness.

Benefits of technology

With limited computing resources, efficient evaluation and real-time rendering of HR filters were achieved, reducing computational complexity while maintaining the quality of spatial audio rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597847B_ABST
    Figure CN116597847B_ABST
Patent Text Reader

Abstract

A method (1000) for generating an estimated HR filter equation (I) consisting of a set of S head-related (HR) filter portion equations (II), where s = 1 to S. The method includes obtaining an alpha matrix (e.g., an NxK matrix), where the alpha matrix consists of S portions, where each portion of the alpha matrix corresponds to a different HR filter portion of the S HR filter portions, each portion of the alpha matrix consists of N sub-vectors (N > 1), and each sub-vector includes a plurality of scalar values. The method also includes separately computing each of the S HR filter portions, where for at least certain HR filter portion equations (II) of the S HR filter portions, the step of computing equation (II) includes computing equation (II) using no more than a predetermined number q s of sub-vectors within the portion of the alpha matrix corresponding to equation (II), where q s is less than N.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis

[0002] This application is a divisional application of the invention patent application filed on November 20, 2020, with application number 202080102069.9 and invention title "Head-Related (HR) Filter". Technical Field

[0003] This disclosure relates to HR filters. Background Technology

[0004] The human auditory system is equipped with two ears to capture sound waves that travel toward the listener. Figure 1 The diagram illustrates sound waves propagating towards the listener from a direction of arrival (DOA) specified by a pair of elevation and azimuth angles in a spherical coordinate system. Along the propagation path towards the listener, each sound wave interacts with the upper trunk, head, outer ear, and surrounding matter before reaching the left and right eardrums. These interactions result in temporal and spectral variations in the waveforms reaching the left and right eardrums, some of which are DOA-dependent.

[0005] Our auditory system has learned to interpret these variations to infer various spatial characteristics of the sound waves themselves and the acoustic environment in which the listener finds themselves. This ability is known as spatial hearing, and it involves how we evaluate spatial cues embedded in the binaural signals (i.e., the sound signals in the left and right ear canals) to infer the location of the auditory event caused by the sound event (the physical sound source) and the acoustic characteristics caused by our physical environment (e.g., a small room, a tiled bathroom, an auditorium, a cave). By reintroducing spatial cues from the binaural signals that lead to spatial perception of sound, this human ability (spatial hearing) can be used in turn to create spatial audio scenes.

[0006] The main spatial cues include 1) angle-dependent cues: binaural cues (i.e., interaural intensity difference (ILD) and interaural time difference (ITD)) and monoaural (or spectral) cues; and 2) distance-dependent cues: intensity and direct reverberation (D / R) energy ratio. The mathematical representation of the short-time DOA-dependent or angle-dependent temporal and spectral variations (1 to 5 milliseconds) of a waveform is the so-called head-related (HR) filter. The frequency domain (FD) representation of the HR filter is the so-called head-related transfer function (HRTF), and the time domain (TD) representation is the head-related impulse response (HRIR). This HR filter can also represent DOA-independent temporal and spectral characteristics corresponding to the interaction between the sound wave and the listener (e.g., related to resonance in the ear canal).

[0007] Figure 2Examples of ITD and spectral cues for sound waves propagating toward the listener are shown. The two graphs illustrate the amplitude response of a pair of HR filters obtained at 0 degrees elevation and 40 degrees azimuth. (Data from the CIPIC database: Topic-ID 28. This database is publicly available at www.ece.ucdavis.edu / cipic / spatial-sound / hrtf-data / ).

[0008] A binaural rendering method based on HR filters has been gradually established, in which spatial audio scenes are generated by directly filtering the audio source signal using a pair of HR filters at the desired location. This method is particularly attractive for many emerging applications such as extended reality (XR) (e.g., virtual reality (VR), augmented reality (AR), or mixed reality (MR)) and mobile communication systems where headphones are often used.

[0009] HR filters are typically estimated from measurements as the impulse response of a linear dynamic system that converts the raw sound signal (input signal) into left and right ear signals (output signals), which can be measured within the ear canal of the listening object (e.g., an artificial head, mannequin, or human subject) at a predefined set of elevation and azimuth angles on a constant radius sphere. The estimated HR filter is typically provided as a Finite Impulse Response (FIR) filter and can be used directly in this format. For efficient binaural rendering, a pair of HRTFs can be converted to interaural transfer functions (ITFs) or modified ITFs to prevent abrupt spectral peaks. Alternatively, the HRTF can be described parametrically. Such parametric HRTFs can be easily integrated with parametric multichannel audio encoders (e.g., MPEG surround sound and Spatial Audio Object Coding (SAOC)).

[0010] To discuss the quality of different spatial audio rendering techniques, it is useful to introduce the concept of minimum audible angle (MAA), which characterizes the sensitivity of the human auditory system to the angular displacement of sound events. Regarding azimuth localization, MAA is smallest in front of the listener (approximately 1 degree), while it is much larger for lateral sound sources (approximately 10 degrees). Regarding elevation localization, MAA in the mid-plane is smallest in front of the listener (approximately 4 degrees), while it increases at elevation angles away from the horizontal plane.

[0011] Audio spatial rendering that results in a convincing spatial perception of a sound source (object) at any location in space requires a pair of head-and-background (HR) filters representing the position within the MAA (Main Aspect) of the corresponding location. If the angular difference between the HR filters is less than the MAA of the object's location, the listener will not notice the difference; beyond this limit, a larger positional difference leads to a more noticeable inaccuracy in the listener's perceived location. Because head-related filtering varies from person to person, personalized HR filters may be necessary to achieve accurate rendering without audible inaccuracies in perceived location.

[0012] HR filter measurements must be performed at finite measurement locations, but rendering may require filters for any possible location on a sphere surrounding the listener. Therefore, a mapping method is needed to convert the discrete measurement set into a continuous spherical angular domain. Several methods exist, including: directly using the nearest available measurement, using interpolation methods, or using modeling techniques. These three rendering techniques are discussed below.

[0013] 1. Directly use the nearest neighbor point

[0014] The simplest technique for rendering is to use a High-Rate Filter (HR) that uses the nearest neighbor from the set of measurement points. Note that some computational work may be required to determine the nearest neighbor measurement points, which can become important for an irregularly sampled set of measurement points on a 2D sphere. For general object positions, there will be some angular error between the desired filter position (corresponding to the object position) and the nearest available HR filter measurement. For sparsely sampled HR filter measurement sets, this will result in a significant error in the object position; this error is reduced or effectively eliminated when using a denser set of measurement points. For moving objects, the HR filter changes in a stepwise manner, which does not correspond to the expected smooth movement.

[0015] Typically, densely sampled HR filter measurements are difficult to perform on human subjects because subjects must remain seated during data collection: small, accidental movements of the subject limit the achievable angular resolution. The process is also time-consuming for both subjects and technicians. Given a sparsely sampled HR filter dataset (as outlined below), inferring spatially relevant information about missing HR filters may be more efficient than performing such detailed measurements. Densely sampled HR filter measurements are easier to capture with artificial heads, but the resulting HR filter set is not always well-suited to all listeners, sometimes leading to inaccurate or blurred perception of object location.

[0016] 2. Interpolation between adjacent points

[0017] If the sample points are not densely spaced enough, interpolation between adjacent points can be used to generate an approximate filter for the desired DOA. The interpolated filter varies continuously between discrete sample points, avoiding the abrupt changes that can occur with the nearest neighbor method. This method introduces additional complexity in generating the interpolated HR filter values, as the resulting HR filter has a broadened (less point-like) perceived DOA due to the mixing of filters from different locations. Furthermore, measures need to be taken to prevent phase problems caused by directly mixing filters, which can increase complexity.

[0018] 3. Model-based filter generation

[0019] More advanced techniques can be used to build a model for the underlying system that generates HR filters and how they vary with angle. Given a set of HR filter measurements, the model parameters are tuned to reproduce the measurements with minimal error, thus creating a mechanism for generating HR filters not only at the measurement locations but also, more generally, as a continuous function of angle space.

[0020] There are other methods for generating HR filters as continuous functions of DOA that do not require an input measurement set, but instead use high-resolution 3D scans of the listener’s head and ears to model wave propagation around the subject’s head to predict HR filter behavior.

[0021] In the next section, we will present a class of HR filter models that use weighted basis vectors to represent HR filters.

[0022] 3.1 HR Filter Model Using Weighted Basis Vectors - Mathematical Framework

[0023] Consider an HR filter model of the following form:

[0024]

[0025] in

[0026] It is the estimated HR filter, a vector of length K for a specific (θ,φ) angle;

[0027] α n,k It is a set of scalar values ​​independent of the angle (θ, φ), where the set of scalar values ​​is organized into N basis vectors (α1, α2, ..., α...). N Each has a length of K (therefore, the scalar value α) n,k It is vector α n (the kth value);

[0028] F k,n (θ,φ) is a set of scalar-valued functions that depend on the angle (θ,φ); and

[0029] e k It is a leap The set of orthogonal basis vectors in the K-dimensional space of the filter.

[0030] Model function F k,n (θ, φ) are determined as part of the model design and are typically chosen to capture well variations in the elevation and azimuth dimensions of the HR filter set. After the model function is specified, the model parameters α... n,k It can be estimated using data fitting methods such as least squares minimization.

[0031] It is not uncommon to use the same modeling function for all HR filter coefficients, which leads to a specific subset of this type of model where the model function F k,n (θ,φ) is independent of the position within the filter:

[0032]

[0033] The model can then be represented as:

[0034]

[0035] In one application, e k The basis vectors are natural basis vectors aligned with the coordinate system being used, e1 = [1, 0, 0... 0], e2 = [0, 1, 0... 0], ... . For compactness, when using natural basis vectors, we can write it as

[0036]

[0037] Where, α n It is a vector of length . This leads to the equivalent expression of the model.

[0038]

[0039] That is, once the parameter α n,k It has already been estimated. Represented as a fixed basis vector α n A linear combination, where the angle change of the HR filter varies with the weighting value F. n Captured in (θ,φ). That is, each basis vector α n With weighted value F n (θ,φ) are related.

[0040] When the unit basis vectors are natural basis vectors, this equivalent expression is compact, but it should be remembered that the following approach (without this convenient notation) can be applied to any choice of model using orthogonal basis vectors in any domain. Other applications of the same underlying modeling technique would be different choices of orthogonal basis vectors in the time domain (Hermite polynomials, sine curves, etc.) or in domains other than the time domain (e.g., the frequency domain (e.g., via Fourier transform) or any other domain that naturally represents an HR filter).

[0041] Notice, (i.e., the estimated HR filter) is the result of the model evaluation specified in Equation (5) and should be similar to the measurement at the same location as (i.e., the HR filter). For the actual measurement, the test point (θ) is known. test ,φ test ), can compare h(θ) test ,φ test )and This is used to evaluate the quality of the model. If the model is considered accurate, it can also be used to generate estimates for some general points that are not among the points that have already been measured (i.e., ).

[0042] Also note that the equivalent matrix formula for equation (5) is:

[0043]

[0044] in

[0045] f(θ,φ) = the weighted row vector, length of which is...

[0046] f(θ,φ)=[F1(θ,φ),F2(θ,φ),…,F N [(θ,φ)],

[0047] = A set of N basis vectors, organized into rows of a matrix multiplied by columns, i.e.

[0048]

[0049] A complete binaural representation includes a left HR filter and a right HR filter for the left and right ears, respectively. Using two separate models of the form of equation (1), two distinct model parameters α representing the left and right HR filters are obtained. n,k The set of functions F. n (θ,φ) and basis functions e k The left HR filter model and the right HR filter model may often be different, but the two models are usually the same. Summary of the Invention

[0050] The three methods described above for inferring HR filters over a continuous domain of angles exhibit varying levels of computational complexity and perceived localization accuracy. In terms of computational complexity, directly using nearest neighbor points is the simplest, but requires densely sampled measurements of the HR filters, which are not readily available and result in large amounts of data. HR filter models offer the advantage of generating HR filters with point-like localization properties that change smoothly with DOA compared to methods that directly use measurements. These methods also represent HR filter sets in a more compact form (requiring less resources for transmission or storage, as well as program memory, when using these methods). However, these advantages come at the cost of computational complexity: the model must be evaluated to generate the HR filters before they can be used. Computational complexity is problematic for rendering systems with limited computational power, as it limits the number of audio objects that can be rendered, for example, in real-time audio scenes.

[0051] In spatial audio renderers, it is desirable to obtain, in real time, an estimated HR filter for any elevation-azimuth angle from a model evaluation equation such as Equation (5). To make this possible, the HR filter evaluation specified in Equation (5) needs to be performed very efficiently. The problem addressed in this disclosure is a good approximation of Equation (5) that can be evaluated efficiently. That is, this disclosure describes an optimization of the evaluation of an HR filter model of this type shown in Equation (5).

[0052] More specifically, this disclosure provides a mechanism for dynamically adjusting the precision of the estimated HR filters generated by the HRTF model: quantizing the model at a lower level of precision results in a corresponding reduction in computational complexity. Furthermore, the allocation of detail / complexity can be dynamically distributed among different parts of the generated HR filter pairs; the part of the filter pair that contributes most to convincing spatial awareness can be allocated a larger share of the computational load to allow for higher levels of detail, balanced with a smaller share of the computational load allocated to parts where the level of detail is less important.

[0053] In one aspect, a method for generating S head correlation (HR) filter components is provided. The estimated HR filter composed of the set of elements The method, where s = 1 to S, includes obtaining an alpha matrix (e.g., an N×K matrix), wherein the alpha matrix consists of S parts (e.g., S N×J sub-matrices, where J = K / S), each part of the alpha matrix corresponding to a different HR filter part among the S HR filter parts, each part of the alpha matrix consisting of N sub-vectors (e.g., N 1×J matrices), and each sub-vector including multiple scalar values ​​(e.g., J scalar values). The method also includes calculating each of the S HR filter parts separately, wherein for at least one of the S HR filter parts... calculate The steps include using the alpha matrix and... The corresponding portion shall not exceed the predetermined quantity q s To calculate the sub-vectors Where, q s Less than N (i.e., when calculating) Nq in this section is not used s (subvectors).

[0054] In another aspect, a method for filtering audio signals is provided. The method includes: obtaining an audio signal; generating an estimated head-related head-correlation (HR) filter according to any embodiment disclosed herein; and filtering the audio signal using the estimated HR filter.

[0055] In another aspect, a computer program including instructions is provided that, when executed by the processing circuitry of an audio rendering apparatus, causes the audio rendering apparatus to perform the methods described in any of the embodiments disclosed herein. In yet another aspect, a carrier containing a computer program is provided, wherein the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium.

[0056] On the other hand, a method for generating S head correlation (HR) filter components is provided. The estimated HR filter composed of the set of elements An audio rendering apparatus, wherein s = 1 to S. In one embodiment, the audio rendering apparatus is adapted to obtain an alpha matrix (e.g., an N×K matrix), wherein the alpha matrix consists of S parts (e.g., S N×J sub-matrices, where J = K / S), each part of the alpha matrix corresponding to a different HR filter part among S HR filter parts, each part of the alpha matrix consisting of N sub-vectors (e.g., N 1×J matrices), and each sub-vector including multiple scalar values ​​(e.g., J scalar values). The apparatus is also adapted to compute each of the S HR filter parts separately, wherein for at least one of the S HR filter parts... calculate The steps include using the alpha matrix and... The corresponding portion shall not exceed the predetermined quantity q s To calculate the sub-vectors Where, q s Less than N (i.e., when calculating) Nq in this section is not used s (subvectors).

[0057] In another aspect, an audio rendering apparatus for filtering audio signals is provided. In one embodiment, the audio rendering apparatus is adapted to acquire an audio signal, generate an estimated HR filter according to any of the methods disclosed herein, and filter the audio signal using the estimated HR filter.

[0058] In some embodiments, the audio rendering apparatus includes processing circuitry and a memory containing instructions executable by the processing circuitry.

[0059] The proposed embodiments described herein provide a flexible mechanism for reducing computational complexity when evaluating HR filters, for example, based on an HR filter model, by omitting components that do not contribute significant detail to the filter. Within this reduced complexity budget, computational complexity (and therefore filter accuracy) can be dynamically allocated to the parts of the HR filter pair that are most important for convincing spatial awareness. This enables efficient resource utilization in situations where the computational power of the rendering device is constrained (e.g., on mobile devices or in complex audio scenes with many rendered audio sources). Attached Figure Description

[0060] The accompanying drawings, which are included in and form part of this specification, illustrate various embodiments.

[0061] Figure 1 The image shows sound waves propagating toward the listener.

[0062] Figure 2 An example of the ITD and spectral cues of sound waves propagating toward the listener is shown.

[0063] Figure 3 The HR filter obtained by summing all basis vectors from the precise quantization of the model is compared with the HR filter calculated by optimization.

[0064] Figure 4 Three available basis vectors are shown.

[0065] Figure 5 This illustrates splitting the basis vectors into parts.

[0066] Figure 6The diagram shows the ordering of basis vectors based on energy.

[0067] Figure 7 The basis vectors that will be used in the estimation of the example filter are shown.

[0068] Figure 8 An example is shown of dividing K = 40 elements into S = 4 non-overlapping parts using discrete weighting values.

[0069] Figure 9 An example is shown where K = 40 elements are divided into S = 4 overlapping regions using consecutive weighted values.

[0070] Figure 10A This is a flowchart illustrating a process according to some embodiments.

[0071] Figure 10B This is a flowchart illustrating a process according to some embodiments.

[0072] Figure 11 An audio rendering unit according to an embodiment is shown.

[0073] Figure 12 A filter device according to some embodiments is shown. Detailed Implementation

[0074] As mentioned above, equation (6) is the equivalent matrix formula of equation (5). (The rest of the text is omitted.) Due to the angular dependence of and , equation (6) can be written as:

[0075]

[0076] in

[0077] = A synthesized HR function for one ear, containing filter taps,

[0078] =The weighted row vector of one ear, length

[0079] = The set of basis vectors for one ear, organized into rows in a matrix of rows multiplied by columns (also referred to as "alpha" in this paper).

[0080] The following algorithm is applicable to specific (θ,φ) pairs.

[0081] 1. Algorithm

[0082] This section describes an algorithm for calculating an optimized estimate of an HR filter (i.e., the position of a specific (θ,φ) object and one of its two ears).

[0083] Step 1: Divide matrix α into S parts (e.g., S = 8) (also called submatrices), where each part has N rows and multiple columns (e.g., S parts of equal length J, where J = K / S, but the S parts do not necessarily have the same length), assigning columns of consecutive ranges to different parts. For example, divide each basis vector αn into S subvectors, thus forming N×S subvectors (αn, ... n,s And each subvector consists of J columns (or "entries") (i.e., a set of J scalar values). In other words, each of the S parts of matrix α comprises N subvectors (e.g., part s of matrix α comprises subvectors: α 1,s ,α 2,s ,…,α N,s Matrices and vectors are represented in bold.

[0084]

[0085] Subvector α ns Represents the basis vector α n The sub-parts (or sub-vectors) of α (a summary of the different indexing schemes for α used in this document is given in the abbreviation section). In this embodiment, each sub-vector α ns It contains J columns and a single row (i.e., each subvector α) ns Containing J scalar values). Therefore, matrix α can be represented as an NxS matrix of subvectors (i.e., submatrices), as shown in equation (9), instead of being represented as an Nx1 matrix of basis vectors, as shown in equation (6a). Equivalently, matrix α can be represented as a 1xS matrix (vector) of submatrices, where each submatrix consists of N subvectors (i.e., submatrix s consists of the following subvectors: α 1s ,α 2s ,...,α Ns Each scalar value in α can be indexed in either of the following ways: using the row n and column k of the entire matrix as α. n,k Or, a portion of s, including rows n, a portion (submatrix) s, and columns j, are used as α. n,s,j That is, the relationship between k and the tuples s and j is: k = ((s-1) × J) + j.

[0086] α n,s The last entry in (i.e., α) n,s,(last) (Assuming part s is not the last part) and α n,(s+1) The first entry in (i.e., α) n,(s+1),1 ) are adjacent entries in the same row of the undivided matrix α. That is, if α n,s,J =α n,k Then α n,s+1,1 =α n,k+1 Similarly, α n,s,jand α (n+1),s,j These are adjacent entries in the same column of the undivided matrix α.

[0087] This partition of α is an alternative indexing scheme used in the following steps.

[0088] Step 2: Calculate the sub-vector α n,s Related importance metrics. For example, using energy E. ns As a measure:

[0089]

[0090] Where f n It is related to the basis vector α n The associated weight values. Other possible metrics are described below.

[0091] The sum of the n squared terms of each s of α is independent of (θ,φ) and can therefore be pre-computed (step 2).

[0092] Step 3: Determine f for the required (θ,φ) n .

[0093] Step 4: Based on the above equation (10), use the pre-calculated values ​​from step (2) and the stored results from step (3) to calculate the signal energy E of each sub-vector. ns .

[0094] Step 5: Calculate the sum of the signal energy (of all basis functions) for each part:

[0095]

[0096] Step 6: Determine the index sort M s It will part of E ns Sort the list of n values ​​of s from largest to smallest.

[0097] Step 7: For each section, determine the parts to be used for identification. The number of subvectors of part s is predetermined, where q s This indicates the pre-ordered quantity. This quantity (q) s The value in q is less than or equal to the total number of subvectors available within that section (i.e., N). s When =0, for Use zero values. This can be useful, for example, for efficiently truncating filters without changing the filter length. In one embodiment, where S = 8, q = [8,8,7,5,4,3,2,2] (i.e., q1 = 8, meaning that all 8 subvectors within part 1 are used to determine...), Furthermore, q8 = 2, which means that only 2 of the subvectors within part 8 are used to determine the position. ).

[0098] Step 8: By summing the signal energy values ​​calculated and stored in step (5) above, the calculation will be completed. Those used in the parts s The sum of the signal energies of each basis vector.

[0099] E′ s =∑ n E ns Where n = M1, M2, ..., M qs (12)

[0100] Step 9: Calculation The scaling value of each part s to illustrate the energy loss.

[0101]

[0102] Step 10: Use q, which represents the maximum energy contribution to part s within part s. s The subvectors (i.e., the sorted index M stored in step (7)) s The first q s (a few values), calculate an approximate value. Signal, represented as

[0103]

[0104] in

[0105] f′ is the selected subvector corresponding to the portion s (i.e., corresponding to the sorted index M). s The first q s The weight values ​​of the subvectors of each entry.

[0106] α′, the selected q within part s s The subvector (i.e., the one with the highest energy value E) ns q s (subvectors).

[0107] The number of floating-point multiplications in this step is reduced to q / N of the full model quantization (Equation (6)).

[0108] As a specific example of step 10, suppose:

[0109] q7 = 2,

[0110] For a subset s = 7, M1 = 3, and M2 = 5.

[0111] alpha n=3,s=7=[1.5,2.1,0.3,-0.8,-1.1]

[0112] alpha n=5,s=7 =[2.6,3.2,1.4,-1.9,-2.2]

[0113] f7 = [0.2, 0.7, 0.3, 0.1, 0.5, 0.8] (that is, f 3,7 =0.3 and f 5,7 =0.5),

[0114] p7 = 1.34, then

[0115]

[0116]

[0117]

[0118]

[0119]

[0120] Step 11: Cascade from adjacent parts To assemble HR filters The final estimate.

[0121]

[0122] Algorithm diagram

[0123] Figure 3 The HR filter obtained from the exact quantization of the model by summing all basis vectors is compared with the HR filter calculated using optimization. The steps taken for this specific calculation are described by [details omitted]. Figures 4 to 7 The series is shown. For this example calculation, S = 3 and there are N = 3 basis vectors and = 36 filter values. For step 7, the number of subvectors to be used for each part s is s =[2,2,1] (i.e., 1=2, etc.).

[0124] In the algorithm steps outlined above, alpha and f remain separate until step 10, where partial estimation... It is based on ′ s and' s The product of these matrices is calculated (Equation 14). This product is the sum of the products of the weighted values ​​and the basis vectors. Figures 4 to 7The corresponding data has been plotted, but with appropriate weighting factors applied. Although this doesn't match the algorithm steps, it's useful for illustration because it reveals which basis components have maximum values, thus aiding in the final calculation of the HR filter.

[0125] Figure 4 The following shows N=3 available basis vectors (i.e., basis vectors 401, 402, and 403) that can be used to construct HR filter estimates.

[0126] Figure 5 This illustrates splitting each of the three basis vectors into S = 3 parts (i.e., S sub-vectors of each of the N basis vectors), corresponding to step 1 of the algorithm. For example... Figure 5 As shown, the sub-vectors associated with the first part of the filter are sub-vectors 501, 502, and 502; the sub-vectors associated with the second part of the filter are sub-vectors 504, 505, and 506; and the sub-vectors associated with the third part of the filter are sub-vectors 507, 508, and 509.

[0127] Figure 6 The algorithm step 6 is shown, which sorts the subvectors within each corresponding part according to a metric (e.g., energy metric).

[0128] Figure 7 This shows the subvectors that will be used in the estimation of that part of the filter for each part. These unused subvectors (i.e., those located after ranking) are also shown. s Subsequent subvectors are discarded. Note that the discarded subvectors in this graph correspond to information that does not need to be computed, resulting in a reduction in the complexity of this method. Figure 7 As shown, for the first part of the filter, two sub-vectors associated with that part (i.e., sub-vectors 502 and 503) are used; for the second part of the filter, two sub-vectors associated with that part (i.e., sub-vectors 504 and 505) are used; and for the third part of the filter, only a single sub-vector associated with that part (i.e., sub-vector 508) is used.

[0129] Preprocessing

[0130] Step (2) in the above process can be performed independently of the object's location, and therefore can be performed as a preprocessing step instead of during rendering. This improves system performance by reusing computation results instead of unnecessarily repeating them. For real-time rendering applications, this preprocessing reduces the complexity of the time-critical parts of filter computation.

[0131] Partial overlap / windowing

[0132] The above division of the computation into non-overlapping parts is just one possible approach. These parts are non-overlapping and have a hard transition from one part to the next. In one embodiment, windowing is used to allow for a smooth transition / overlap between parts. This prevents artifacts from appearing in the resulting filter when there is a large change in the number of components used in two adjacent parts. Figure 8 (No overlap, some parts are different) and Figure 9 A simple example is shown in (overlap with simple window opening function).

[0133] Importance measurement

[0134] In steps 2, 4, 5, 8, and 9 above, signal energy is used as a measure of contribution to different subvectors. Other embodiments use different measures to determine the importance of filter components, including but not limited to: maximum value; and sum of absolute values. These measures have the property that they are simpler to compute than energy and increase with the signal energy, thus they can be used as proxies for energy calculation. Using such measures yields ranking behavior similar to that using energy calculation, but with reduced computational complexity.

[0135] For comparison, the energy metric and its scaling factor p are repeated below, as are the metric and scaling factor for the sum of the maximum and absolute values.

[0136] energy:

[0137]

[0138]

[0139]

[0140]

[0141] Maximum value:

[0142] MN ns =max(abs(f n α nsj ))=abs(f n )max(abs(α nsj ))

[0143]

[0144]

[0145]

[0146] Sum of absolute values:

[0147]

[0148]

[0149]

[0150]

[0151] Figure 10A This illustrates a method for generating S HR filter sections according to some embodiments. Head-related (HR) filters composed of a set of elements The flowchart of process 1000 is provided, where s = 1 to S. Process 1000 may begin at step s1002. Step s1002 includes: obtaining an alpha matrix (e.g., an N×K matrix), wherein the alpha matrix consists of S parts, and each part of the alpha matrix corresponds to a different HR filter part among the S HR filter parts (e.g., the first part of the alpha matrix corresponds to...). The second part of the alpha matrix corresponds to (etc.), each part of the alpha matrix consists of N sub-vectors (N>1, for example, N=8), and each sub-vector includes multiple scalar values. Step s1004 includes: calculating each of the S HR filter parts, wherein for at least one of the S HR filter parts... calculate The steps include using the alpha matrix and... The corresponding portion shall not exceed the predetermined quantity q s To calculate the sub-vectors Where, q s Less than N.

[0152] Figure 10B This is a flowchart illustrating a process 1050 for audio signal filtering according to some embodiments. Process 1050 includes step s1052, which includes obtaining an audio signal. Step s1054 includes generating an estimated head-related head-shift (HR) filter according to any embodiment disclosed herein. Step s1056 includes filtering the audio signal using the estimated HR filter.

[0153] Figure 11An audio rendering unit 1100 according to some embodiments is shown. The audio rendering unit 1100 includes a rendering unit 1102, an HR filter generator 1104, and an ITD generator 1106. In other embodiments, the ITD generator 1106 is optional (e.g., in some use cases, the ITD value is always zero). As described herein, the HR filter generator 1104 is used to generate estimated HR filters at any elevation and azimuth angles requested in real time by the rendering unit 1102. Similarly, the ITD generator generates ITDs at any elevation and azimuth angles requested in real time by the rendering unit 1102, which may require efficient evaluation of the ITDs from an ITD model already loaded into the unit, as described in reference [1].

[0154] Figure 12 This is a block diagram of an audio rendering apparatus 1200 for implementing an audio rendering unit 1100, according to some embodiments. That is, the audio rendering apparatus 1200 is operable to perform the processes disclosed herein. Figure 12As shown, the audio rendering device 1200 may include: a processing circuit (PC) 1202, which may include one or more processors (P) 1255 (e.g., a general-purpose microprocessor and / or one or more other processors, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.), which may coexist in a single enclosure or a single data center, or may be geographically distributed (i.e., the audio rendering device 1200 may be a distributed computing device); at least one network interface 1248, which includes a transmitter (Tx) 1 245 and receiver (Rx) 1247, enabling audio rendering apparatus 1200 to send data to and receive data from other nodes connected to network 110 (e.g., an Internet Protocol (IP) network), network interface 1248 (directly or indirectly) connected to network 110 (e.g., network interface 1248 may be wirelessly connected to network 110, in which case network interface 1248 is connected to an antenna arrangement); and storage unit (also referred to as a “data storage system”) 1208, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 1202 includes a programmable processor, computer program product (CPP) 1241 may be provided. CPP 1241 includes computer-readable medium (CRM) 1242 storing computer program (CP) 1243 including computer-readable instructions (CRI) 1244. CRM 1242 may be a non-transitory computer-readable medium, such as magnetic media (e.g., hard disk), optical media, memory devices (e.g., random access memory, flash memory), etc. In some embodiments, CRI 1244 of computer program 1243 is configured such that when executed by PC 1202, CRI causes audio rendering apparatus 1200 to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, audio rendering apparatus 1200 may be configured to perform the steps described herein without requiring code. That is, for example, PC 1202 may consist of only one or more ASICs. Therefore, the features of the embodiments described herein can be implemented in hardware and / or software.

[0155] The following is a summary of the various embodiments described herein:

[0156] A1. A method for generating a component consisting of S HR filter parts Head-related (HR) filters composed of a set of elements Method 1000 (see Figure 10AThe method includes obtaining an alpha matrix s1002 (e.g., an N×K matrix), wherein the alpha matrix consists of S parts, and each part of the alpha matrix corresponds to a different HR filter part among the S HR filter parts (e.g., the first part of the alpha matrix corresponds to...). The second part of the alpha matrix corresponds to (etc.), each part of the alpha matrix consists of N sub-vectors (N>1, for example N=8), and each sub-vector includes multiple scalar values; and each of the S HR filter parts in s1004 is calculated, wherein for at least one of the S HR filter parts, calculate The steps include using the alpha matrix and... The corresponding portion shall not exceed the predetermined quantity q s To calculate the sub-vectors Where, q s Less than N.

[0157] A2. The method according to embodiment A1, wherein the AND of the alpha matrix is ​​used. q in the corresponding part s Calculate using subvectors The method also includes: AND operations on the alpha matrix. For each subvector within the corresponding section, a metric is determined; and based on the determined metric, a metric is selected for computation. q s Subvectors.

[0158] A2b. The method according to embodiment A2, wherein determining the metric of the subvector includes: calculating the scalar value of the subvector using at least the subvector as input to the calculation.

[0159] A3. The method according to embodiment A2 or A2b, wherein determining the metric of the subvector includes determining an energy value based on the subvector and one or more weight values ​​associated with the subvector.

[0160] A4. The method according to embodiment A2 or A2b, wherein determining the metric of the subvector includes setting the metric of the subvector according to equation (17).

[0161] A5. The method according to embodiment A2 or A2b, wherein determining the metric of the subvector includes: setting the metric of the subvector according to equation (18).

[0162] A6. The method according to any one of embodiments A2 to A5, wherein q s=2, and use 2 subvectors to calculate The steps include: calculating w1×v11+w2×v12; and calculating 1×v21+w2×v22, where w1 is the AND of the alpha matrix. The weight value associated with the first subvector of the two subvectors in the corresponding part, w2 is the weight value associated with the second subvector of the two subvectors, v11 is the first scalar value in the first subvector, v21 is the second scalar value in the first subvector, v12 is the first scalar value in the second subvector, and v22 is the second scalar value in the second subvector.

[0163] A7. The method according to embodiment A6, wherein w1 = f1 × p s w2 = f2 × p s f1 is a predetermined weight value associated with the first sub-vector, f2 is a predetermined weight value associated with the second sub-vector, and p s It is the scaling factor associated with this section.

[0164] B1. A method for filtering audio signals 1050 (see...) Figure 10B The method includes: obtaining an audio signal s1052; generating a head-related head-related (HR) filter s1054 according to any one of embodiments A1 to A7; and filtering the audio signal s1056 using the estimated HR filter.

[0165] C1. A computer program 1243 including instructions 1244, which, when executed by the processing circuit 1202 of the audio rendering apparatus 1200, cause the audio rendering apparatus 1200 to perform the method according to any one of embodiments 1 to 9.

[0166] C2. A carrier comprising a computer program according to embodiment C1, wherein the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium 1242.

[0167] D1. An audio rendering apparatus 1200, the audio rendering apparatus 1200 being adapted to perform the method according to any of the embodiments described above.

[0168] E1. An audio rendering apparatus 1200, the apparatus 1200 comprising: a processing circuit 1202; and a memory 1242 containing instructions 1244 executable by the processing circuit, thereby enabling the apparatus to perform the method according to any of the embodiments described above.

[0169] Although various embodiments have been described herein, it should be understood that they are presented by way of example only and not by way of limitation. Therefore, the breadth and scope of this disclosure should not be limited to any of the exemplary embodiments described above. Furthermore, any combination of the foregoing elements with all possible variations is covered in this disclosure unless otherwise indicated or otherwise explicitly conflicted by the context.

[0170] Additionally, although the process described above and shown in the accompanying figures is presented as a series of steps, it is for illustrative purposes only. Therefore, it is conceivable that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.

[0171] abbreviation

[0172] A matrix of scalar weights used in the evaluation of the αHR filter model. N rows multiplied by K columns.

[0173] α (n,k) A single scalar entry in matrix α, indexed by row n and column k.

[0174] α n A row in matrix α. A vector of size 1 multiplied by K.

[0175] α n,s basis vector α n A subvector. The number of elements in the subvector corresponds to the number of columns in the part s.

[0176] α n,s,j A single scalar entry in matrix α, indexed by row n, part s, and column j within part s. In other words, a subvector α. n,s The j-th scalar entry in.

[0177] θ Angle of elevation

[0178] φ Azimuth

[0179] AR (Augmented Reality)

[0180] D / R ratio direct reverberation ratio

[0181] DOA arrival direction

[0182] FD frequency domain

[0183] FIR Finite Impulse Response

[0184] HR filter head correlation filter

[0185] HRIR head-related impulse response

[0186] HRTF Header-Related Transfer Functions

[0187] ILD (Inter-aural sound intensity difference)

[0188] IR impulse response

[0189] ITD Interaural Time Difference

[0190] MAA Minimum Hearing Angle

[0191] Mixed Reality (MR)

[0192] SAOC Spatial Audio Object Coding

[0193] TD time domain

[0194] VR Virtual Reality

[0195] XR Extended Reality

[0196] References:

[0197] [1]Mengqui Zhang and Erlendur Karlsson, USpatent application no.62 / 915,992, Oct.16, 2019.

Claims

1. A method (1050) for filtering audio signals, the method comprising: Obtain the (s1052) audio signal; Generate a head-related HR filter with (s1054) estimation; as well as The audio signal is filtered using the estimated HR filter (s1056). The method is characterized in that the estimated HR filter Composed of S HR filter sections The set consists of s = 1 to S, and the estimated HR filter (s1054) is generated by: The length of (s1002) is obtained as N basis vectors (α1, α2, ..., α N The alpha matrix of N basis vectors is organized into Travel The rows in the alpha matrix of the column; The alpha matrix is ​​divided into S parts, with columns of a continuous range assigned to different parts. Each part of the alpha matrix corresponds to a different HR filter part among the S HR filter parts. Each part consists of N sub-vectors, where the N sub-vectors are the rows of the alpha matrix corresponding to the part, and each sub-vector includes multiple scalar values, where the number of scalar values ​​in each sub-vector is equal to the number of columns in the corresponding part. Calculate (s1004) each of the S HR filter sections, wherein for at least one of the S HR filter sections... ,calculate The steps include using the alpha matrix with... q in the corresponding part s Calculate using subvectors , where q s Less than N, The method further includes: performing an AND operation on the alpha matrix. For each subvector within the corresponding part, determine the metric, and Select the metric based on the determined value for calculation. The q mentioned s Subvectors.

2. The method according to claim 1, wherein, Determining the metric for a subvector includes determining an energy value based on the subvector and one or more weight values ​​associated with the subvector.

3. The method according to claim 1, wherein, Determining the metric for a subvector includes determining the maximum value and one or more weight values ​​associated with the subvector.

4. The method according to claim 1, wherein, Determining the metric for a subvector includes determining the sum of the absolute values ​​associated with the subvector and one or more weight values.

5. The method according to any one of claims 2 to 4, wherein, q s =2, and Calculate using two sub-vectors The steps include: Calculate w1 × v11 + w2 × v12; and Calculate w1 × v21 + w2 × v22, where, w1 is the AND of the alpha matrix. The weight value associated with the first sub-vector of the two sub-vectors within the corresponding part. w2 is the weight value associated with the second sub-vector of the two sub-vectors. v11 is the first scalar value within the first sub-vector. v21 is the second scalar value within the first sub-vector. v12 is the first scalar value within the second sub-vector, and v22 is the second scalar value within the second subvector.

6. The method according to claim 5, wherein, w1 = f1 × p s , w2 = f2 × p s , f1 is a predetermined weight value associated with the first sub-vector. f2 is a predetermined weight value associated with the second sub-vector, and p s It is the scaling factor associated with the aforementioned portion.

7. The method according to claim 6, wherein, The scaling factor p s It is the sum of the signal energies of all basis functions in the aforementioned part and those used in the aforementioned part. The calculation is based on the relationship between the sum of the signal energies of the basis vectors.

8. A computer program product comprising instructions (1244) that, when executed by a processing circuit (1202) of an audio rendering apparatus (1200), cause the audio rendering apparatus (1200) to perform the method according to any one of claims 1 to 7.

9. An audio rendering apparatus (1200) comprising a processing circuit (1202) and a memory (1242), the memory containing instructions executable by the processing circuit, thereby enabling the audio rendering apparatus (1200) to: Obtain the (s1052) audio signal; Generate a head-related HR filter with (s1054) estimation; as well as The audio signal is filtered using the estimated HR filter (s1056). The device is characterized in that the estimated HR filter Composed of S HR filter sections The set consists of s = 1 to S, and the estimated HR filter (s1054) is generated by: The length of (s1002) is obtained as N basis vectors (α1, α2, ..., α N The alpha matrix of N basis vectors is organized into Travel The rows in the alpha matrix of the column; The alpha matrix is ​​divided into S parts, with columns of a continuous range assigned to different parts. Each part of the alpha matrix corresponds to a different HR filter part among the S HR filter parts. Each part of the alpha matrix consists of N sub-vectors, where each N sub-vector is a row of the alpha matrix corresponding to that part, and each sub-vector includes multiple scalar values, where the number of scalar values ​​in each sub-vector is equal to the number of columns in the corresponding part. Calculate (s1004) each of the S HR filter sections, wherein for at least one of the S HR filter sections... ,calculate The steps include using the alpha matrix with... q in the corresponding part s Calculate using subvectors , where q s Less than N, The audio rendering device (1200) can also be operated to: perform AND operations on the alpha matrix. For each subvector within the corresponding part, determine the metric, and Select the metric based on the determined value for calculation. The q mentioned s Subvectors.

10. The audio rendering apparatus (1200) according to claim 9, wherein, The memory also contains instructions that can be executed by the processing circuitry, thereby enabling the audio rendering apparatus (1200) to perform the method according to any one of claims 2 to 7.