Detecting anomalous values in head-related filter bank
By detecting outliers in the HR filter dataset, using SIOD and TIOD features, the spatial positioning inaccurate and rendering degradation problems caused by outliers in the HR filter group are solved, and higher quality spatial audio rendering is achieved.
Patent Information
- Application Number
- CN202280101650.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2025-06-13
AI Technical Summary
Outliers exist in the HR filter bank, resulting in inaccurate spatial positioning of the audio source and degradation in spatial audio rendering.
An outlier value detection framework is provided for detecting outliers in the HR filter data set. By extracting SIOD and TIOD characteristics, the standard for outliers is determined, and then the outlier values are detected and marked.
Effectively detect and eliminate outliers in the HR filter group, improve the spatial positioning accuracy of the audio source, and improve the quality of spatial audio rendering.
Smart Images

Figure CN120153670A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to detecting outliers in a head-related (HR) filter bank. Background Art
[0002] The human auditory system is equipped with two ears that capture sound waves propagating towards a listener. Figure 4 Shown is a sound wave propagating towards a listener from a direction of arrival (DoA) specified by a pair of elevation and azimuth angles in a spherical coordinate system. Along the propagation path towards us, each sound wave interacts with our upper torso, head, outer ear, and surrounding materials before reaching our left and right eardrums. This interaction causes temporal and spectral variations in the waveforms reaching the left and right eardrums, some of which are DoA-related. Our auditory system has learned to interpret these variations to infer various spatial characteristics of the sound wave itself and the acoustic environment in which the listener finds himself / herself. This ability is called spatial hearing, which involves how to evaluate the spatial cues embedded in the binaural signals (i.e., the sound signals in the left and right ear canals) to infer the location of the auditory event triggered by a sound event (physical sound source) and the acoustic characteristics caused by the physical environment (e.g., a small room, a tiled bathroom, an auditorium, a cave, etc.). By reintroducing the spatial cues that will result in the spatial perception of sound in the binaural signals, this human ability (spatial hearing) can then be used to create a spatial audio scene.
[0003] The main spatial cues include 1) angle-related cues: binaural cues (i.e., interaural level difference (ILD) and interaural time difference (ITD)) and monaural (or spectral) cues; 2) distance-related cues: intensity and direct-to-reverberant (D / R) energy ratio. The mathematical representation of the short-time (1 - 5 milliseconds) DoA-related temporal and spectral variations of the waveform is the so-called head-related (HR) filter. The frequency-domain (FD) representation of those filters is the so-called head-related transfer function (HRTF), and the time-domain (TD) representation is the head-related impulse response (HRIR). Figures 17A - 17E Shown is an example of the ITD and spectral cues of a sound wave propagating towards a listener. These four plots show the time-domain and frequency-domain responses of a pair of HR filters obtained at an elevation angle of 0 degrees and an azimuth angle of 40 degrees (data from the CIPIC database: subject-ID 28. This database is publicly available and can be accessed from the link https: / / www.ece.ucdavis.edu / cipic / spatial-sound / hrtf-data / ).
[0004] HR filters are typically estimated as the impulse response of a linear dynamic system from acoustic measurements that transform an original sound signal (input signal) into left and right ear signals (output signals), which can be measured at a set of predefined elevation and azimuth angles on a spherical surface at a constant radius from the listening subject (e.g., an artificial head, mannequin, or human subject) within the ear canal of the listening subject. Figure 18 Figure 1 depicts a simplified setup for binaural recording of HR filters. Such acoustic measurements require an acoustically well-treated environment (e.g., an anechoic chamber), a dedicated set of audio playback and acquisition equipment (e.g., speakers, in-ear microphones, amplifiers, sound cards, etc.), a source / listener positioning system, and a dedicated software package for playing, acquiring, and processing audio data.
[0005] The estimated (either through measurement or through numerical simulation) HR filters are typically provided as finite impulse response (FIR) filters. These FIR filters are typically several milliseconds long, where the time span of each filter can be divided into three consecutive time regions, namely the pre-active region, the active region, and the post-active region. In the pre-active region and the post-active region, the filter taps are zero or very close to zero due to estimation noise and contribute very little to binauralization. The active region contains the main part of the impulse response of the filter that represents the actual binauralization. At the start of the active region, there are typically strong oscillations near zero, and at the end of the active region, the oscillations gradually fade and drop to near zero values. In the Figures 17A - 17B plot of the time-domain response in Figure 2, the active region is roughly indicated by the dashed circle.
[0006] HR filter banks can be used:
[0007] directly in their original form by a binaural audio renderer, where spatial audio is rendered by filtering an audio source signal with a pair of HR filters located at the desired positions, or
[0008] for predictive modeling, where the resulting model is used to provide HR filters for a spatial audio renderer.
[0009] The spatial quality of the rendered spatial audio is largely determined by the HR filter bank used in the binaural renderer. Considerable effort has been invested in improving acoustic measurements to obtain high-quality HR filter banks. However, variability is an inherent part of HR filter measurements, where noise or measurement errors are inevitable.
[0010] In acoustic HR filter measurements, especially for human subjects, a special chair is usually designed, which has a structure with a headrest and a backrest that provides a reference position of the subject's head relative to the speaker(s), with the aim of minimizing unwanted head movement during the measurement. However, slight tilting of the subject's head or slight tilting of the vertical axis of rotation of the chair often occurs and results in misalignment between the speaker(s) and the head. This misalignment causes a time-of-arrival (TOA) shift of the signal at the eardrum, or in other words, an onset delay of the HR filter. The shift in the onset delay implies a discontinuity in the ITD between adjacent measurement points. For example, the ITD of the HR filter at an azimuth of 0 degrees is assumed to be 0. When misalignment occurs, the ITD deviates from 0. For an audio scenario where the source moves along a vertical line in front of the listener, even for a deviation as small as ±1 sample (e.g., 0.02 milliseconds at a sampling rate of 48 kHz), a renderer using an HR filter with such misalignment errors can perceive instability (left / right sway). An effective model-based method for correcting such misalignment errors in HR filters is discussed in WO 2022 / 223132.
[0011] Non-HR reflections from the mechanical setup are another source of error in HR filter measurements. The mechanical setup is always carefully designed to have a minimal impact on the incident acoustic waves. For example, the sides of the speaker are wrapped with acoustic absorption material, the support structure of the chair is covered with acoustic absorption material, and so on. However, non-HR reflections can still occur and can be captured in the recording. Such non-HR impulse responses appear in the resulting HR filter, which may disrupt the ILD cues at certain frequency bands, leading to a perceivable "auxiliary" source located somewhere other than the desired position. As discussed in WO 2021 / 074294, this type of error can be largely mitigated by linear regression methods. Summary of the Invention
[0012] In addition to the above two types of errors, the HR filter bank may also contain outlier HR filters (also known as "outliers") that deviate significantly from the overall pattern of the HR filters, and some outlier HR filters may not contain spatial cues for localization at all.
[0013] Outliers in the HR filter bank are caused by different reasons, such as mechanical problems in the measurement setup and malfunctioning activities during acoustic measurements. When the HR filter bank is obtained using low-cost devices (e.g., in a home environment), outliers may appear even more prominently.
[0014] Using an HR filter bank that includes such outliers can result in extremely poor spatial localization of the audio source, or cause other audible degradations in spatial audio rendering. If such outliers exist in the HR filter bank and this HR filter bank is used by a spatial audio renderer, the resulting binaural output will not be as expected. Further, if such an HR filter bank is used to train a predictive model, the outliers will result in an incorrect predictive model that generates defective filters. Accordingly, it is necessary to detect such outliers and remove them before the HR filter bank is used by a spatial audio renderer or provided to a modeling algorithm for learning.
[0015] Accordingly, embodiments of the present disclosure provide an outlier detection framework for detecting outliers in an HR filter dataset.
[0016] In one aspect of embodiments of the present disclosure, a method is provided. The method includes obtaining head-related HR filter data that indicates a set of HR filters for generating audio corresponding to temporal and / or spectral variations of sound waves. The method further includes determining a time interval for evaluating the HR filters included in the set of HR filters; and using the determined time interval to evaluate at least one HR filter included in the set of HR filters. Audio is generated based on the evaluation.
[0017] In another aspect, a computer program including instructions is provided, which when executed by a processing circuit, causes the processing circuit to perform the method of any one of the above embodiments.
[0018] In another aspect, a carrier containing the computer program of any one of the above embodiments is provided, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0019] In another aspect, a device is provided. The device is configured to obtain head-related HR filter data that indicates a set of HR filters for generating audio corresponding to temporal and / or spectral variations of sound waves. The device is further configured to determine a time interval for evaluating the HR filters included in the set of HR filters; and use the determined time interval to evaluate at least one HR filter included in the set of HR filters. Audio is generated based on the evaluation.
[0020] In another aspect, a device is provided. The device includes a memory and a processing circuit coupled to the memory, wherein the device is configured to perform the method of any one of the above embodiments.
[0021] Some embodiments of the present disclosure allow for the detection of outliers in an HR filter dataset without the need for a clean dataset for training. Additionally, the outlier detection method according to some embodiments is applicable to datasets that contain only normal data or both normal data and outliers. Finally, the outlier detection method according to some embodiments is insensitive to data changes (e.g., spatio-temporal changes, individual changes between different objects, and measurement-related changes between different measurement systems). BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings incorporated herein and forming a part of the specification illustrate various embodiments.
[0023] Figure 1 A system according to some embodiments is shown.
[0024] Figure 2A 、 Figure 2B 、 Figure 3A and Figure 3B The concept of an HR filter is shown.
[0025] Figure 4 A propagation vector indicating the propagation direction of a sound wave within a three-dimensional (3D) space is shown.
[0026] Figure 5 A set of head-related filters located on a 3D sphere is shown.
[0027] Figure 6 A device according to some embodiments is shown.
[0028] Figure 7A 、 Figure 7B and Figure 7C An exemplary response of an HR filter is shown.
[0029] Figure 8 A spatial deviation index (SIOD) curve is shown.
[0030] Figure 9A The SIOD curve is shown.
[0031] Figure 9B The SIOD curve is shown.
[0032] Figure 10A A temporal deviation index (TIOD) curve is shown.
[0033] Figure 10B The TIOD curve is shown.
[0034] Figure 11 A process according to some embodiments is shown.
[0035] Figure 12 A method for determining an initial starting point is shown.
[0036] Figure 13 An example of the frequency density at the peak position is shown.
[0037] Figure 14 Examples of the final start and end points are shown.
[0038] Figure 15 A process according to some embodiments is shown.
[0039] Figure 16 An apparatus according to some embodiments is shown.
[0040] Figures 17A - 17E Sound waves propagating towards the listener (which interact with the head and ears) and the resulting ITD are shown.
[0041] Figure 18 A simplified setup for HR binaural recording is shown. Detailed Description
[0042] Figure 1 An exemplary system 100 according to some embodiments is shown. System 100 includes headphones 106, an audio rendering unit 112, and a server 114. The server 114 is configured to transmit audio data 116 to the audio rendering unit 112 via a network 110. The network 110 can be a wired network or a wireless network. Alternatively or additionally, the network 110 can be a cloud through which the audio data 116 is transmitted from the server 114 to the audio rendering unit 112. In the present disclosure, audio data is defined as data for providing an audio experience to a listener after rendering (e.g., processing with one or more HR filters) as if the listener were in the three-dimensional (3D) space where one or more audio sources are located. The audio data includes audio samples of the source signal corresponding to one or more audio sources. In some embodiments, the audio data may additionally include HR filter information indicating the HR filter.
[0043] After receiving the audio data 116, the audio rendering unit 112 can generate a binaural audio signal and transmit the generated audio signal to the headphones 106. The headphones 106 are configured to generate audio based on the audio signal, thereby providing an audio (also referred to as spatial audio) experience to the user 102. In some embodiments, instead of the headphones 106, other audio generating devices such as a speaker array can be used. The number of speakers in the array can be any number greater than two.
[0044] In some embodiments, system 100 may optionally include an extended reality (XR) (e.g., virtual reality, mixed reality, or augmented reality) display head-mounted device 104. The XR display head-mounted device 104 may be configured to generate different views of a virtual reality (VR) environment based on the head orientation of user 102.
[0045] The XR display head-mounted device 104 may be communicatively coupled to the earphone 106. For example, the XR display head-mounted device 104 may detect the head orientation of user 102, and based on the detected head orientation of user 102, the XR display head-mounted device 104 may display different views of the VR environment and may trigger the audio rendering unit 112 to generate different audio signals such that user 102 may hear different audio based on the head orientation of user 102.
[0046] Figure 2A , Figure 2B , Figure 3A and Figure 3B illustrate the basic concept of HR filtering.
[0047] Figure 2A illustrates a sound wave 202 propagating in a first direction and reaching the right ear of user 102, while Figure 2B illustrates a sound wave 212 propagating in a second direction (different from the first direction) and reaching the right ear of user 102. As shown in Figure 2A and Figure 2B , depending on the direction of arrival (DoA) of the sound wave (relative to the center of user 102's head), the sound wave is diffracted and / or reflected in different ways (see the paths formed by the dashed arrows in Figure 2A and Figure 2B ). For simplicity of explanation, only reflections are shown in Figure 2A and Figure 2B .
[0048] The HR filter is used to generate an audio effect in which these different diffractions and reflections caused by different DoAs are factored. In other words, depending on the DoA of the sound wave, the sound wave undergoes different temporal and spectral changes before being perceived by user 102, and the mathematical representation of this temporal and spectral change is called the HR filter. Note that the reflection paths shown in Figure 2A and Figure 2B are provided for illustrative purposes only and may be different from the actual reflection paths in the real-world environment.
[0049] Figure 3A illustrates an exemplary time-domain response of the HR filter for the sound wave 202, and Figure 3BShows an exemplary time-domain response of the HR filter for the sound wave 212. As shown in the figure, due to the different temporal and spectral variations of the sound waves, the waveforms (including amplitude and time of arrival (TOA) (or start delay)) are different for sound waves 202 and 212. Note that the responses provided in Figure 3A and Figure 3B are only for showing several aspects of the influence of the HR filter, and thus may be different from the actual response.
[0050] As mentioned above, the temporal and spectral variations of the sound waves (also called "acoustic waves") caused by HR filtering change depending on the propagation direction of the sound waves.
[0051] In Figure 4 , the propagation vector 402 indicates the propagation direction of the sound wave within the 3D space defined by the three axes 412, 414, and 416. The propagation vector 402 can be defined using two angles - the azimuth and the elevation angle (θ). The azimuth is the angle between the axis 412 (e.g., the x-axis) and the projection vector 404, corresponding to the projection of the propagation vector 402 onto the plane formed by the axis 412 and the axis 414. The elevation angle (θ) is the angle between the propagation vector 402 and the projection vector 404.
[0052] Since the temporal and spectral variations of the sound wave can change depending on the azimuth and the elevation angle (θ), in some embodiments, various HR filters (which represent such temporal and spectral variations) are provided for different combinations of the azimuth and the elevation angle (θ).
[0053] Figure 5 Shows positions on a three-dimensional (3D) sphere around the user 102, which correspond to a set of HR filters (also called an HR filter bank) representing various temporal and spectral variations of the sound wave. The origin of the 3D sphere can correspond to the position of the user 102's head, and each point 502 can respectively correspond to the positions of the HR filters for the left ear and the right ear. As mentioned above, each HR filter can be associated with a specific combination of the azimuth and the elevation angle (θ).
[0054] The HR filter bank can be used to generate audio depending on the orientation of the user 102's head. For example, the HR filter 512 included in the HR filter bank can be used to generate audio corresponding to a first combination of the azimuth and the elevation angle (φ 1 , θ 1 ), while the HR filter 514 included in the HR filter bank can be used to generate sound corresponding to a second combination of the azimuth and the elevation angle (φ 2 , θ 2 ).
[0055] Typically, the HR filters included in an HR filter bank generate a common pattern of temporal and / or spectral variations of an acoustic wave. However, in some cases, some of the HR filters included in the HR filter bank generate temporal and / or spectral variations that deviate considerably from the common pattern. Such a set of HR filters showing a deviating pattern is referred to as an outlier or an outlier HR filter.
[0056] Using an HR filter bank that includes such outliers to generate an audio signal corresponding to an audio source may result in a very poor spatial localization of the audio source or cause other audible degradations in spatial audio rendering. Therefore, when generating an audio signal, it may be desirable to exclude such outliers or perform one or more additional correction processes to correct the incorrect performance of the outliers.
[0057] Thus, in some embodiments of the present disclosure, Figure 6 the outlier detector 600 shown in is used to detect one or more outliers within a set of HR filters of an HR filter data set.
[0058] Detecting outliers in a set of HR filters of an HR filter data set can be performed based on the characteristics of the set of HR filters and the identification of a typical active region, i.e., a specific time interval where the main part of the impulse response of the HR filter typically occurs. In the present disclosure, the identified typical active region of an HR filter is also referred to as the "valid region". More specifically, first, a set of features representing where the main part of the impulse response of the HR filter is located is calculated. Second, using the set of features, a set of criteria defining the start and end of a time frame that defines the boundary corresponding to the valid region is determined. Third, using the criteria, outliers are detected.
[0059] An HR filter that is an outlier and is correctly identified as an outlier is referred to as a true positive. An HR filter that is not an outlier but is incorrectly identified as an outlier is referred to as a false positive. An HR filter that is an outlier but is incorrectly identified as normal is referred to as a false negative. An HR filter that is not an outlier and is correctly identified as normal is referred to as a true negative. The goal of the embodiments of the present disclosure is to maximize true positives and true negatives while minimizing false positives and false negatives.
[0060] According to some embodiments, the information required to detect outliers is the HR filter data set and a set of specifications for outlier detection is the HR filter data set being evaluated, which is typically obtained by loading the original HR filter data set from an existing file into to obtain. is a set of specifications used by three modules included in the outlier detector 600.
[0061] Specify the methods and / or parameters required by the feature extraction module.
[0062] Specify the methods and / or parameters required by the criterion determination module.
[0063] Specify the methods and / or parameters required by the outlier identification module.
[0064] The outlier detection process performed by the outlier detector 600 involves various data variables and data representations. The symbolic notation for such data variables and data representations is explained as follows.
[0065] The general data structure is represented as a list of data sequences and / or a list of other data structures. The basic HR filter dataset is provided in the form of a data list which contains HR filters sampled at M elevation and azimuth angles {(θ[m], φ[m]): m = 1, …, M}, where θ and φ are the elevation and azimuth angles respectively, and m represents the index value.
[0066] θ = {θ[m]: m = 1, …, M} represents a series of elevation angles.
[0067] φ = {φ[m]: m = 1, …, M} represents a series of azimuth angles.
[0068] H l = {h l [m]: m = 1, ..., M} represents a set of left HR filters where h l [m] = [h l [1; m],... h l [n; m],..., h l [N l ; m] is a FIR filter of length N1, and n is the index of the filter tap at a certain moment.
[0069] H r = {h r [m]: m = 1, ..., M} represents a set of right HR filters, where h r [m] = [h r [1; m]..., h r [n; m],..., h r [N r ; m] is a FIR filter of length Nr, and n is the index of the filter tap at a certain moment. Usually, the lengths of the left and right filters are the same (meaning N l = N r )
[0070] In the present disclosure, for simplicity of representation, when subscripts and / or superscripts are not particularly required, they are omitted.
[0071] Referring back Figure 6 , the outlier detector 600 may include a feature extraction module 602, a criterion determination module 604, and an outlier identification module 606.
[0072] The outlier detector 600 is configured to detect outliers within an HR filter of an HR filter dataset based on identifying characteristics of the HR filter and identifying a valid region (i.e., a specific time interval in which the main part of the impulse response typically appears). The feature extraction module 602 is configured to extract a set of features corresponding to such characteristics of the HR filter (i.e., a set of features representing a specific time interval in which the main part of the impulse response typically appears).
[0073] Typically, a positive or negative peak of an HR filter indicates the location of the strongest impulse response of the HR filter. However, the location of such a peak is sensitive to noise. For example, in Figure 7A and Figure 7B , the left filter at elevation and azimuth angles (-100, 15) and (-80, 57) is contaminated with high DC offset type noise, and thus the peak does not appear where it normally should. Note that typically, the impulse response of an HR filter oscillates around zero (e.g., like the impulse response of the right HR filter shown in Figure 7C ). However, in Figure 7A and Figure 7B , the impulse response of the left HR filter increases from zero to a basic level and remains at that level. Thus, Figure 7A and Figure 7B almost the entire part of the left HR filter shown in
[0074] is affected by such high DC offset type noise. Thus, simply using the peak location as a feature for identifying the valid region (where the main part of the impulse response typically appears) of the HR filter may result in misidentification of the valid region and false detection of outliers.
[0075] The IOD is defined as the ratio of variance to the mean, where the mean is non-zero and is typically only used for positive statistics. However, the mean of the HR filter can be negative. Therefore, it is of great interest to determine whether there is an impulse response at a given moment, regardless of whether the impulse response is positive or negative. Accordingly, the IOD is modified to be the ratio of variance to the normalized L1 norm rather than the mean.
[0076] The SIOD can be defined at a given moment n for an HR filter bank, which can be either the left filter bank Hl or the right filter bank H r ). For example, the SIOD measurement can be calculated as follows:
[0077]
[0078] In this case, M′ = M. To calculate the SIOD curve for the entire filter bank including both the left and right filter banks, H = [H l , H r and M′ = 2M.
[0079] The SIOD measurement (also referred to simply as "SIOD") indicates the overall active region of all HR filters in the dataset. It is well known that for ipsilateral filters, the impulse response within the active region can be much stronger than that of the contralateral filters, and for ipsilateral filters, the starting point of the active region can be much earlier than that of the contralateral filters. The ideal SIOD curve is a smooth, right-skewed curve. It has a single unimodal maximum, and its value decreases sharply asymptotically when n becomes less than the n value of the maximum, and its value decreases slowly asymptotically when n becomes greater than the n value of the maximum.
[0080] For the HR filter, the "cup-shaped" region where the SIOD has a large value corresponds to the n region where the HR filter is most active. Figure 8 An example of an SIOD curve calculated for a diffuse-field (DF) equalized left HR filter according to the FABIAN database (https: / / depositonce.tu-berlin.de / handle / 11303 / 6153.4; note that the original filters in the FABIAN database are not diffuse-field DF equalized) is shown.
[0081] In Figure 8 's upper plot, the responses of the HR filters at five azimuth angles of 0 degrees (center), -30 degrees (right), -80 degrees (right), 30 degrees (left), and 80 degrees (left) on the horizontal plane are plotted. It can be seen that the "cup-shaped" region where the SIOD has a large value ( Figure 8 the circled part in the upper plot of Figure 8 ) corresponds to the n region where the active region of the HR filters in the dataset appears. The magnified part of the circled region is shown in
[0082] Figure 9A An example of the SIOD curve of the left HR filter in the HR filter dataset of Subject-19 in the Princeton database is shown, which contains outliers. Similarly, Figure 9B An example of the SIOD curve of the right HR filter in the HR filter dataset of Subject-19 in the Princeton database is shown, which contains outliers.
[0083] As Figure 8 、 Figure 9A and Figure 9B shown in, the maximum value of the SIOD curve is data-dependent. However, this correlation can be removed by min-max normalization such that the maximum value of SIOD is always 1 and the minimum value of SIOD is always 0.
[0084] TIOD can be defined for the HR filter h[m] at a certain moment n. For example, the TIOD measurement can be calculated as follows:
[0085]
[0086] TIOD is a robust measurement for detecting the main part of the impulse response of the HR filter. Figure 10A The TIOD curve of the left HR filter (from the FABIAN database) at an elevation angle of 0 degrees and an azimuth angle of 90 degrees is shown, where min-max normalization is applied to keep the TIOD values within the range of 0 and 1. Similarly, Figure 10B The TIOD curve of the left HR filter (from the FABIAN database) at an elevation angle of 0 degrees and an azimuth angle of -90 degrees is shown, where min-max normalization is applied to keep the TIOD values within the range of 0 and 1.
[0087] As Figure 8 shown in, the SIOD curve can have a steep taper on the left side, which indicates the start of the ipsilateral HR filter with the shortest start delay and the strongest impulse response in the filter bank. Therefore, the SIOD curve can be used to identify the starting point of the main part of the impulse response. Also as Figure 8 shown in, the SIOD curve can also have a long tail, and thus it is not easy to identify the point at which the main part of the impulse response of the contralateral HR filter with the longest start delay has started, because it is delayed compared to the ipsilateral filter and has lower energy / peak than the ipsilateral filter.
[0088] In contrast, since the TIOD curve indicates where the HR filter has low variance, the point identifying the end of the effective region is a relatively robust feature. Therefore, according to some embodiments, the feature extraction module 602 can be configured to extract from the dataset Extract SIOD and TIOD. In some embodiments, SIOD and / or TIOD can be specified in as features to be extracted from
[0089] After the feature extraction module 602 extracts the SIOD and TIOD values (corresponding to the SIOD curve and the TIOD curve), the criterion determination module 604 can determine a set of criteria for detecting outliers based on the extracted SIOD and TIOD values. This set of outlier detection criteria can define the effective region of the HR filter. The effective region is the region (time interval) within which the main part of the impulse response of any HR filter in the dataset should lie, i.e., covering the typical active region of the HR filter bank. When determining this set of outlier detection criteria, two steps can be performed.
[0090] The first step in determining this set of outlier detection criteria is to obtain an initial starting point and an initial ending point, which together define the initial effective region. Here, the initial starting point should be early enough to ensure that the ipsilateral filter is included in the initial effective region, and the initial ending point should be early enough to ensure that filters with an extremely long starting delay are excluded from the initial effective region.
[0091] The second step in determining this set of outlier detection criteria is to obtain a final starting point and a final ending point, which together define the final effective region. Here, the final starting point is determined based on the initial starting point, aiming to exclude filters with an extremely short starting delay. Similarly, the final ending point is determined based on the initial ending point, aiming to ensure that normal (typical) contralateral filters are included.
[0092] Thresholding is one of the efficient methods for determining the boundaries of the effective region. Some thresholds can be fixed constants specified in
[0093]
[0094]
[0095] specified in
[0095] The output of the outlier detector 600 is marked as It can be a set of labels that indicate for each HR filter in the dataset whether it is detected as an outlier, as shown in the following table.
[0096] HR Filter List Outlier (True / False) HR Filter #1 True HR Filter #2 False HR Filter #3 True
[0097] Alternatively, it can be a list of indices of the HR filters that have been detected as outliers. Thus, in the example provided in the following table, it can include a list of index values #1 and #3.
[0098] If no outliers have been detected, then it can be just an indicator indicating that no outliers have been found in the HR filter dataset.
[0099] Figure 11 FIG. 1100 shows a process 1100 for detecting outliers from a set of HR filters of an HR filter dataset, which will be explained with respect to Figure 1 FIG. 1100. The process 1100 can be executed by a server 114 configured to transmit audio data 116 to an audio renderer 112. Alternatively, the process 1100 can be executed by the audio renderer 112.
[0100] The process 1100 can start from step S1102. Step S1102 includes obtaining the input information required for outlier detection. For example, in the case where the process 1100 is executed by the audio renderer 112 or the server 114, the audio renderer 112 or the server 114 can obtain the input information by receiving the input information from another entity and / or by retrieving the input information from its memory. More specifically, in the case where the process 1100 is executed by the audio renderer 112, in one example, the input information can be included in the audio data 116.
[0101] The input information can include the HR filter dataset and a set of specifications for outlier detection is the HR filter dataset under evaluation, which is typically obtained by loading the original HR filter dataset from an existing file into to obtain. is a set of specifications used by three modules 602-606.
[0102] After obtaining the input information, the process 1100 can proceed to step S1104. Step S1104 includes extracting a set of eigenvalue Z, which indicates where the active regions of the set of HR filters of the HR filter dataset are located. Z contains according to the feature extraction settings The feature extraction method specified in extracts the feature values from
[0103] In some embodiments, the feature extraction setting indicates that SIOD and TIOD are the features to be extracted for outlier detection. In such embodiments, Z = {z S , z T}, where z S represents a set of SIOD values, and z T represents a set of TIOD values.
[0104] The SIOD values can be calculated separately for the left filter bank and the right filter bank . Therefore, z S can be expressed as:
[0105] z S = {SIOD l SIOD r} = {{SIOD l [n l : n l = 1,..., N l}, {SIOD r [n r : n r = 1,..., N r}.
[0106] Since the left and right sets of HR filters and basically follow the same pattern and usually have the same data structure (e.g., N l = N r = N), it is reasonable to calculate one SIOD curve for the combined set . z S can be expressed as:
[0107] z S = {SIOD} = {SIOD[n]: n = 1,…, N}.
[0108] As described above, in some embodiments, the set z S of SIOD values is calculated using equation (1). On the other hand, the set z T of TIOD values is calculated using equation (2).
[0109] z T can be expressed as:
[0110]
[0111] After a set of eigenvalues is extracted in step S1104, process 1100 may proceed to step S1106. Step S1106 includes determining a set of standard values for detecting outliers from a set of HR filters in the HR filter dataset . More specifically, in step S1106, using a set of eigenvalues Z (e.g., a set of SIOD values and a set of TIOD values), a set of standard values can be determined. The set of standard values indicates the start and end points of a valid region (i.e., a valid time interval) within which the main part of the impulse response of any HR filter in the dataset should lie (or typically lies).
[0112] The start point should be early enough so that the valid region includes / covers the ipsilateral filters with short start delays, and at the same time, the start point should be late enough so that the filters with ultra-short start delays are excluded from the valid region (not covered by the valid region). The end point should also be late enough so that the valid region includes / covers the contralateral filters with long start delays, and at the same time, the end point should be early enough so that the filters with ultra-long start delays are excluded from the valid region (not covered by the valid region).
[0113] As Figure 11 shown, the step of determining the outlier detection criteria includes two sub-steps: (1) obtaining an initial set of standard values and (2) obtaining a final set of standard values
[0114] Sub-step (1) - obtaining an initial set of standard values
[0115] The initial set of standard values may include two elements - where n s0 represents the initial start point, and n e0 represents the initial end point. The initial start point and the initial end point define the initial valid region.
[0116] The SIOD eigenvalue extracted in step S1104 can be used to estimate the initial start point n s0 . For example, using a thresholding method, n s0 can be obtained by the following formula:
[0117]
[0118] where represents the normalized cumulative sum of the SIOD values,
[0119]
[0120] the threshold η s0The smaller the value, the smaller the initial starting point n s0 will be. In some embodiments, a small η s0 value is recommended to minimize the false negative rate so that the ipsilateral HR filter cannot be detected as an outlier. The value of η can be within [0, 1]. s0 An example of η is 0.01.
[0121] Figure 12 An example of using the described thresholding method to determine the initial starting point is shown, where η s0 = 0.01. In Figure 12 , the FABIAN database was used. Here, given the threshold η s0 = 0.01, the initial starting point was found to be n s0 = 20.
[0122] After determining the initial starting point n s0 , the initial starting point n e0 can be determined as follows.
[0123] For the ipsilateral HR filter, the active region of the impulse response of the HR filter typically appears early, and for the contralateral HR filter, the active region of the impulse response of the HR filter typically appears late. The ipsilateral filter usually has a high signal-to-noise ratio (SNR). The position of the active region of the ipsilateral filter can be accurately represented by the peak position of its TIOD curve. However, the contralateral filter has a weak impulse response and can therefore be easily contaminated by noise. There may be peaks at a similar level caused by noise.
[0124] To capture such anomalies, for the contralateral filter, instead of using the position of a single TIOD maximum, averaging over the region where TIOD is greater than the threshold can be effective.
[0125] For example, let represent the list of peak positions of the TIOD curve. For the ipsilateral filter, For the contralateral filter, η n A typical value can be 0.5, given that the TIOD feature is min-max normalized and has values in [0, 1]. Let be the vector containing the unique elements in
[0126] The frequency density of the TIOD peak position can be estimated by the following formula
[0127]
[0128] is the count that appears in and Pr[i] represents the peak position of the frequency density. Figure 13 The curve in [[ ]] plots an example of such frequency density calculated from the FABIAN database. The starting and ending points of the curve correspond to the positions of the active regions of the ipsilateral and contralateral filters, respectively. To ensure exclusion of filters with extremely long start delays, an earlier initial end point is preferred.
[0129] Using a thresholding method, n can be obtained by the following formula e0 :
[0130]
[0131] η e0 is determined by calculating the q 0 th percentile of the frequency density. The larger the value of q 0 , the larger the threshold η e0 and the smaller the value of n e0 .
[0132] Figure 13 shows an example of determining the initial end point n e0 where the threshold η e0 is obtained from the 50th percentile of the frequency density of the active region positions (the FABIAN database is used in this example).
[0133] Sub-step (2) - Obtaining the final set of standard values
[0134] The principle for adjusting the initial start and end points is based on the fact that the position of the active region is likely to be a continuous function of direction and, therefore, is likely to be a contiguous set. In other words, the probability of a large difference between two adjacent elements in [[ ]] is low. Let where it should be less than the threshold η Δ . If then it is likely caused by an outlier. A larger η Δ value may result in a higher number of false positives, and a smaller η Δ value may result in a higher number of false negatives. η Δ = 2 or 3 seems to give a good balance.
[0135] However, this is not valid for spatially sparse sampled (especially azimuthally sparse sampled) HR filter banks, where The reason for being greater than the threshold is the large difference between the sampling azimuth angles. In this case, even if is greater than η Δ , for the typical peak positions, the frequency density is usually relatively high. This differentiates the HR filter with spatial sparse sampling from the outliers. Therefore, another condition regarding the frequency density is introduced, that is, Pr[i]<η p , where η p is determined by calculating the q 1 -th percentile of the frequency density. The larger the value of q 1 , the larger the threshold η p . A large threshold can lead to a high number of false negatives. The smaller the value of q 1 , the smaller the threshold η p . A small threshold can lead to a high number of false positives. The value of q 1 = 25 seems to give a good balance.
[0136] The final set of standard values includes the final starting point and the final ending point, which together define the valid region for identifying outliers. The final starting point n s can be obtained as follows.
[0137] Given obtain the list of peak positions represented by , and each peak position satisfies the following three conditions:
[0138] 1) The peak position is before the initial starting and ending points,
[0139] 2) The distance between the peak position and the next peak position is greater than the threshold η Δ ,
[0140] 3) The frequency density of the peak position is lower than the threshold η p , Pr[i]<η p .
[0141] The final starting point can be obtained through .
[0142] On the other hand, the final ending point n e can be obtained as follows.
[0143] Given the list of peak positions represented by can be obtained, and each peak position satisfies all three conditions:
[0144] 1) The peak position is after the initial starting and ending points,
[0145] 2) The distance between the peak position and the next peak position is greater than the threshold η Δ ,
[0146] 3) The frequency density at the peak position is lower than the threshold η p , Pr[i] < η p .
[0147] The final end point can be obtained by . The peak position is a point estimate corresponding to the point position where the strongest impulse response may occur. However, the main part of the impulse response appears within a certain time frame. Therefore, using the peak position without considering the impulse response distribution can lead to high false positives, where some normal (typical) contralateral filters may be identified as outliers.
[0148] Taking into account the impulse response distribution of the filter, the final end point can then be further updated by . is the average of the filter distribution, where the filter distribution n spread is calculated according to TIOD by the following formula
[0149]
[0150] Given that the TIOD feature is min-max normalized, the typical value of η spread is 0.1. Figure 14 Shows an example of the final start / end point, where typical values of the parameters are used, i.e., η Δ = 2, η spread = 0.1, and q 1 = 25. The FABIAN database is used in this example.
[0151] The parameters that can be used to determine the outlier detection criteria are summarized here. η s0 and q 0 are the parameters used to obtain the initial start / end point. η Δ , η spread and q 1 are the parameters used to obtain the final start / end point. In some embodiments, the parameters are specified (e.g., by the user) in .
[0152] As described above, the two sets of HR filters and follow substantially the same pattern. Therefore, in some embodiments, a common valid region is determined for both the left filter bank and the right filter bank. Then the set of standard values contains only two elements where n s is the start point of the valid region, and ne is the end of the valid region. Conversely, in other embodiments, the valid regions can be determined separately for the left and right filter banks. In such embodiments, the set of criteria can include four elements
[0153] After performing step s1106, process 1100 can proceed to step s1108. Step s1108 includes using a set of outlier criteria (i.e., ) determined in step s1106 to detect outliers from a set of HR filters in the HR filter dataset . Outlier detection methods specified in the specification can be used to detect outliers.
[0154] An exemplary simple method for detecting outliers is to locate the peak of the TIOD curve of each HR filter in the dataset. If the peak is outside the valid region defined by , the HR filter is marked as an outlier.
[0155] However, this method cannot detect noise, e.g., Figure 7B the left HR filter at (-80, -57) in
[0156]
[0157] . According to some embodiments, to detect such noise, the ratio of the energy within the valid region to the total energy of the HR filter can be used. For example, let γ[m] represent the energy ratio of the HR filter h[m], Figure 7C the left filter at (-70, 30) in
[0158]
[0159] In the above formula, the mean is removed, and the energy ratio is insensitive to DC offset. However, it may not be able to detect HR filters with an undesired DC offset, e.g., and for the right HR filter as such that the threshold for the ipsilateral filter is higher than that for the contralateral filter. In some embodiments, the threshold is (e.g., by the user) at Specified in the above embodiments, a single threshold is used for comparison of different ratios. However, in other embodiments, different thresholds may be provided for comparison with different ratios.
[0160] After detecting outliers in step S1108, process 1100 may proceed to step S1110. Step S1110 includes outputting a set of labels As described above, in some embodiments, the set of labels may indicate whether each HR filter included in the HR filter dataset is an outlier, while in other embodiments, the set of labels only includes a list of index values identifying outliers or only includes a list of index values identifying non-outliers.
[0161] Referring back Figure 1 , after obtaining a set of outlier labels, the set of outlier labels may be provided to audio rendering unit 112. For example, the set of outlier labels may be provided from server 114 to audio rendering unit 112, or may be directly detected by audio rendering unit 112 using process 1100 described above.
[0162] After obtaining a set of outlier labels, audio rendering unit 112 may use the set of outlier labels to generate an audio signal (to be provided to speakers 106 and 108). In one example, in the case where a set of HR filters used by audio rendering unit 112 to generate an audio signal includes outliers, audio rendering unit 112 may exclude the outliers from the set of HR filters and use the set of HR filters from which outliers have been excluded when generating the audio signal.
[0163] In another example, in the case where a set of HR filters used by audio rendering unit 112 to generate an audio signal includes outliers, audio rendering unit 112 may not use the set of HR filters to generate the audio signal and may use another set of HR filters that does not include any outliers. In some embodiments, it may be determined whether to use the set of HR filters including outliers to generate the audio signal based on the number of outliers. For example, in the case where the number of outliers is less than a threshold, audio rendering unit 112 may use the set of HR filters (with or without outliers), while in the case where the number of outliers is greater than or equal to the threshold, audio rendering unit 112 may use a different set of HR filters that does not include outliers or includes a number of outliers less than the threshold.
[0164] Figure 15Process 1500 according to some embodiments is shown. Process 1500 may start at step S1502. Step S1502 includes obtaining head-related HR filter data that indicates a set of HR filters for generating audio corresponding to temporal and / or spectral variations of sound waves. Step S1504 includes determining a time interval for evaluating the HR filters included in the set of HR filters. Step S1506 includes using the determined time interval to evaluate at least one HR filter included in the set of HR filters. Audio is generated based on the evaluation.
[0165] In some embodiments, process 1500 includes, based on the evaluation, performing one or more of the following: (i) marking the at least one HR filter for audio generation; (ii) selecting a different set of HR filters for audio generation; or (ii) not using the at least one HR filter for audio generation.
[0166] In some embodiments, process 1500 includes generating updated HR filter data indicating an updated set of HR filters, where the updated set of HR filters (i) includes the at least one marked HR filter, (ii) does not include the at least one marked HR filter, (iii) is a different set of HR filters, or (iv) does not include the at least one HR filter; and (i) storing and / or transmitting the updated HR filter data and / or (ii) using the updated HR filter data to generate audio.
[0167] In some embodiments, process 1500 includes using the obtained HR filter data to calculate a first set of eigenvalues and / or a second set of eigenvalues, where the eigenvalues included in the first set and / or the eigenvalues included in the second set indicate where a portion of the impulse response of the at least one HR filter is located in the time domain.
[0168] In some embodiments, the eigenvalues included in the first set and / or the eigenvalues included in the second set are a normalized measure of the deviation of the impulse response of the HR filters included in the set of HR filters and / or a normalized measure of the deviation of the impulse response of the HR filters.
[0169] In some embodiments, each eigenvalue included in the first set is a spatial deviation index SIOD, and each eigenvalue included in the second set is a temporal deviation index TIOD.
[0170] In some embodiments, each eigenvalue included in the first set is SIOD[n], where n is the index of the HR filter tap at a certain moment, and
[0171] where
[0172] M′ is the number of HR filters included in the set of HR filters or 1 / 2 of the number of HR filters included in the set of HR filters, m is the filter index of the HR filters included in the set of HR filters, and h[n; m] is the value of the HR filter having the index value m at the HR filter tap n.
[0173] In some embodiments, each eigenvalue included in the second set is TIOD[n], where n is the index of the HR filter tap at a certain moment, and
[0174] where
[0175] N is the length of the HR filters included in the set of HR filters, m is the filter index of the HR filters included in the set of HR filters, and h[n; m] is the value of the HR filter having the index value m at the HR filter tap n.
[0176] In some embodiments, determining the time interval includes: using the eigenvalues included in the first set to calculate the initial start time of the time interval, and / or using the eigenvalues included in the second set to calculate the initial end time of the time interval.
[0177] In some embodiments, the initial start time of the time interval is calculated based on the normalized cumulative sum of the eigenvalues included in the first set and a threshold for determining the initial start time.
[0178] In some embodiments, where
[0179] n s0 is the initial start time of the time interval, n is a positive integer, is the normalized cumulative sum of the eigenvalues included in the first set, η s0 is the threshold for determining the initial start time.
[0180] In some embodiments, the second set of eigenvalues (e.g., a set of eigenvalues including TIOD curve #1, TIOD curve #2,... etc.) includes multiple subgroups of eigenvalues (e.g., TIOD curve #1 corresponding to HR filter #1, TIOD curve #2 corresponding to HR filter #2,... etc.), each eigenvalue included in each subgroup of eigenvalues (e.g., TIOD value #1 of TIOD curve #1, TIOD value #2 of TIOD curve #1) is identified by an index value, in each subgroup of eigenvalues, the highest eigenvalue among the eigenvalues is identified by a peak index value, and the initial end time of the time interval is calculated based on the frequency density of the peak index values of the multiple subgroups.
[0181] In some embodiments, the initial end point n e0 is evaluated as
[0182] where
[0183] Pr[i] is the frequency density of the peak index value i, and the frequency density of the peak index value i indicates the number of times the peak index value i identifies the highest eigenvalue among the eigenvalues included in each subgroup of the eigenvalues of the plurality of subgroups. i is a positive integer, 1 ≤ i ≤ I, and I is the number of peak index values included in is a set of peak index values, η e0 is a threshold for determining the initial end time, and is the peak index value included in corresponding to the index of
[0184] In some embodiments, determining the time interval includes: obtaining a group of peak index values and calculating an adjusted start time of the time interval based on the initial start time of the time interval and the group of peak index values.
[0185] In some embodiments, each peak index value included in the group of peak index values satisfies the following conditions: the peak index value corresponds to a position before the initial end time of the time interval, the difference between the peak index value arranged in ascending order and the next peak index value within a set of peak index values is greater than the threshold, and the frequency density of the peak index value is lower than the threshold density value.
[0186] In some embodiments, the adjusted start time n s is calculated as where n s0 is the initial start time,[[]] is a set of peak index values, and is the group of peak index values.
[0187] In some embodiments, determining the time interval includes: obtaining a group of peak index values and calculating an adjusted end time of the time interval based on the initial end time of the time interval and the group of peak index values.
[0188] In some embodiments, each peak index value included in the group satisfies one or more of the following conditions: the peak index value indicates a position after the initial end time of the time interval, the difference between the peak index value arranged in ascending order and the next peak index value within a set of peak index values is greater than the threshold, and the frequency density of the peak index value is lower than the threshold density value.
[0189] In some embodiments, the adjusted end time ne is calculated as where n e0 is the initial end time, is a set of peak index values, and is the peak index value group.
[0190] In some embodiments, evaluating at least one HR filter included in the set of HR filters using the determined time interval includes: calculating a first energy amount of the at least one HR filter within the time interval; and calculating a total energy amount of the at least one HR filter.
[0191] In some embodiments, evaluating at least one HR filter included in the set of HR filters using the determined time interval includes: determining one or more ratios, each ratio being a ratio of the first energy amount to the total energy amount; and marking the at least one HR filter based on the determined one or more ratios.
[0192] In some embodiments, evaluating at least one HR filter included in the set of HR filters using the determined time interval includes: comparing each of the one or more ratios with at least one threshold; and marking the at least one HR filter if at least one of the one or more ratios is less than the at least one threshold.
[0193] In some embodiments, the at least one HR filter is associated with a specific angle, and the at least one threshold is determined based on the specific angle.
[0194] In some embodiments, evaluating at least one HR filter included in the set of HR filters using the determined time interval includes evaluating the m-th HR filter h[m] associated with the index value m,
[0195]
[0196] where γ is one of the one or more ratios, h[n,m] is the impulse response of the m-th HR filter at HR filter tap n, n s is the start time of the time interval, and n e is the end time of the time interval, and each of n and m is a positive integer.
[0197] In some embodiments, evaluating at least one HR filter included in the set of HR filters using the determined time interval includes evaluating the m-th HR filter h[m] associated with the index value m,
[0198]
[0199] where γ is one of the one or more ratios, h[n,m] is the impulse response of the m-th HR filter at tap n of the HR filter, n s is the start time of the time interval, and n e is the end time of the time interval, and each of n and m is a positive integer.
[0200] Figure 16 is a block diagram of a device 1600 for implementing the audio rendering unit 112 and / or the server 114 according to some embodiments. As Figure 16 shown, the device 1600 may include: a processing circuit (PC) 1602, which may include one or more processors (P) 1655 (e.g., general microprocessors and / or one or more other processors such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.), the processors may be co-located in a single enclosure or a single data center, or may be geographically distributed (i.e., the device 1600 may be a distributed computing device); at least one network interface 1648, each network interface 1648 including a transmitter (Tx) 1645 and a receiver (Rx) 1647 for enabling the device 1600 to transmit data to and receive data from other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network), the network interface 1648 being (directly or indirectly) connected to the network 110 (e.g., the network interface 1648 may be wirelessly connected to the network 110, in which case the network interface 1648 is connected to an antenna arrangement); and one or more storage units (also referred to as "data storage systems") 1608, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where the PC 1602 includes a programmable processor, a computer program product (CPP) 1641 may be provided. The CPP 1641 includes a computer readable medium (CRM) 1642 storing a computer program (CP) 1643 including computer readable instructions (CRI) 1644. The CRM 1642 may be a non-transitory computer readable medium such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., random access memory, flash memory), etc. In some embodiments, the CRI 1644 of the computer program 1643 is configured such that when executed by the PC 1602, the CRI causes the device 1600 to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, the device 1600 may be configured to perform the steps described herein without code. That is, for example, the PC 1602 may consist only of one or more ASICs. Thus, the features of the embodiments described herein may be implemented in hardware and / or software.
[0201] Although various embodiments are described herein, it should be understood that they are presented by way of example only and not by way of limitation. Accordingly, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments. Additionally, unless otherwise indicated herein or clearly contradicted by context, the present disclosure encompasses any combination of the above elements in all possible variations thereof.
[0202] Moreover, although the processes shown above and in the figures are shown as a series of steps, this is for illustrative purposes only. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be rearranged, and some steps may be performed in parallel.
Claims
1. A method (1500), comprising: obtaining (s1502) HR filter data indicative of a set of head-related (HR) filters for generating audio corresponding to temporal and / or spectral variations of sound waves; determining (s1504) a time interval for evaluating the HR filters included in the set of HR filters; and using the determined time interval to evaluate (s1506) at least one HR filter included in the set of HR filters, wherein audio is generated based on the evaluation.
2. The method according to claim 1, comprising: based on the evaluation, performing one or more of the following: (i) marking the at least one HR filter for audio generation; (ii) selecting a different set of HR filters for audio generation; or (ii) not using the at least one HR filter for audio generation.
3. The method according to claim 2, comprising: generating updated HR filter data indicative of an updated set of HR filters, wherein the updated set of HR filters (i) includes the at least one marked HR filter, (ii) does not include the at least one marked HR filter, (iii) is the different set of HR filters, or (iv) does not include the at least one HR filter; and (i) storing and / or transmitting the updated HR filter data, and / or (ii) using the updated HR filter data to generate audio.
4. The method according to at least one of claims 1 - 3, comprising: using the obtained HR filter data, calculating a first set of eigenvalues and / or a second set of eigenvalues, wherein the eigenvalues included in the first set and / or the eigenvalues included in the second set indicate where a part of the impulse response of the at least one HR filter is located in the time domain.
5. The method according to claim 4, wherein the eigenvalues included in the first set and / or the eigenvalues included in the second set are a normalized measurement of the deviation of the impulse response of the HR filters included in the set of HR filters and / or a normalized measurement of the deviation of the impulse response of the HR filter.
6. The method according to claim 4 or 5, wherein each eigenvalue included in the first set is a spatial deviation index (SIOD), and each eigenvalue included in the second set is a temporal deviation index (TIOD).
7. The method according to claim 6, wherein each eigenvalue included in the first set is SIOD[n], where n is the index of the HR filter tap at a certain moment, and wherein M′ is the number of HR filters included in the set of HR filters or 1 / 2 of the number of HR filters included in the set of HR filters, m is the filter index of the HR filters included in the set of HR filters, and h[n;m] is the value of the HR filter with index value m at the HR filter tap n.
8. The method according to claim 6 or 7, wherein Each eigenvalue included in the second group is TIOD[n], where n is the index of the HR filter tap at a certain moment, and wherein N is the length of the HR filters included in the group of HR filters, m is the filter index of the HR filters included in the group of HR filters, and h[n; m] is the value of the HR filter with index value m at the HR filter tap n.
9. The method according to at least one of claims 4-8, wherein, determining the time interval includes: using the eigenvalue included in the first group to calculate the initial start time of the time interval, and / or using the eigenvalue included in the second group to calculate the initial end time of the time interval.
10. The method according to claim 9, wherein, calculating the initial start time of the time interval based on the normalized cumulative sum of the eigenvalue included in the first group and a threshold for determining the initial start time.
11. The method according to claim 10, wherein wherein n s0 is the initial start time of the time interval n is a positive integer, is the normalized cumulative sum of the eigenvalue included in the first group, η s0 is the threshold for determining the initial start time.
12. The method according to at least one of claims 9-11, wherein the second group of eigenvalues (e.g., a group of eigenvalues including TIOD curve #1, TIOD curve #2,... etc.) includes multiple subgroups of eigenvalues (e.g., TIOD curve #1 corresponding to HR filter #1, TIOD curve #2 corresponding to HR filter #2,... etc.), each eigenvalue included in each subgroup of eigenvalues (e.g., TIOD value #1 of TIOD curve #1, TIOD value #2 of TIOD curve #1) is identified by an index value, in each subgroup of eigenvalues, the highest eigenvalue among the eigenvalues is identified by a peak index value, and calculating the initial end time of the time interval based on the frequency density of the peak index values of the multiple subgroups.
13. The method according to claim 12, wherein, Initial and terminal point n e0 Evaluated as wherein Pr[i] is the frequency density of the peak index value i, the frequency density of the peak index value i indicates the number of times the peak index value i identifies the highest eigenvalue among the eigenvalues included in each subgroup of the multiple subgroups of eigenvalues, i is a positive integer, 1 ≤ i ≤ I, I is the number of the peak index values included in is a set of the peak index values, η e0 is the threshold for determining the initial end time, and is the peak index value included in, which corresponds to the index of 14. The method according to at least one of claims 12-13, wherein, determining the time interval includes: Obtain a group of peak index values and calculating an adjusted start time of the time interval based on the initial start time of the time interval and the group of peak index values.
15. The method according to claim 14, wherein, each peak index value included in the group of peak index values satisfies the following conditions: the peak index value corresponds to a position before the initial end time of the time interval, the difference between the peak index value arranged in ascending order and the next peak index value within a group of peak index values is greater than a threshold, and the frequency density of the peak index value is lower than a threshold density value.
16. The method according to claim 14 or 15, wherein, The adjusted start time n s is calculated as wherein n s0 is the initial start time, is the set of peak index values, and is the peak index value group.
17. The method according to at least one of claims 13-16, wherein, Determining the time interval includes: Obtain a group of peak index values and Calculating an adjusted end time of the time interval based on the initial end time of the time interval and the group of peak index values.
18. The method according to claim 17, wherein, Each peak index value included in the group satisfies one or more of the following conditions: The peak index value indicates a position after the initial end time of the time interval, The difference between the peak index value and the next peak index value arranged in ascending order within a group of peak index values is greater than a threshold, and The frequency density of the peak index value is lower than a threshold density value.
19. The method according to claim 17 or 18, wherein, The adjusted end time n e is calculated as wherein n e0 is the initial end time, is the set of peak index values, and is the peak index value group.
20. The method according to any one of claims 1-19, wherein, Using the determined time interval to evaluate at least one HR filter included in the group of HR filters includes: Calculating a first amount of energy of the at least one HR filter within the time interval; and Calculating a total amount of energy of the at least one HR filter.
21. The method according to claim 20, wherein, Using the determined time interval to evaluate at least one HR filter included in the group of HR filters includes: Determining one or more ratios, each ratio being a ratio of the first amount of energy and the total amount of energy; and Marking the at least one HR filter based on the determined one or more ratios.
22. The method according to claim 21, wherein, Using the determined time interval to evaluate at least one HR filter included in the group of HR filters includes: Comparing each of the one or more ratios with at least one threshold; and Marking the at least one HR filter in the case where at least one of the one or more ratios is less than the at least one threshold.
23. The method according to claim 22, wherein The at least one HR filter is associated with a specific angle, and The at least one threshold is determined based on the specific angle.
24. The method according to claim 22, wherein Using the determined time interval to evaluate at least one HR filter included in the group of HR filters includes evaluating the m-th HR filter h[m] associated with the index value m, wherein, γ is one of the one or more ratios, h[n,m] is the impulse response of the m-th HR filter at tap n of the HR filter, n s is the start time of the time interval, and n e is the end time of the time interval, and Each of n and m is a positive integer.
25. The method according to any one of claims 22-24, wherein, Using the determined time interval to evaluate at least one HR filter included in the group of HR filters includes evaluating the m-th HR filter h[m] associated with the index value m, where γ is one of the one or more ratios, h[n,m] is the impulse response of the m-th HR filter at tap n of the HR filter, n s is the start time of the time interval, and n e is the end time of the time interval, and Each of n and m is a positive integer.
26. A computer program (1643) comprising instructions (1644) which, when executed by a processing circuit (1602), cause the processing circuit to perform the method according to any one of claims 1-25.
27. A carrier containing the computer program according to claim 26, wherein, The carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
28. An apparatus (1600) configured to: obtain (s1502) HR filter data indicative of a set of head-related (HR) filters for generating audio corresponding to temporal and / or spectral variations of sound waves; determine (s1504) a time interval for evaluating HR filters included in the set of HR filters; and evaluate (s1506) at least one HR filter included in the set of HR filters using the determined time interval, wherein audio is generated based on the evaluation.
29. The apparatus according to claim 28, wherein the apparatus is further configured to perform the method according to any one of claims 2-25.
30. An apparatus (1600), comprising: a memory (1641); and processing circuitry (1602) coupled to the memory, wherein the apparatus is configured to perform the method according to any one of claims 1-25.
Citation Information
Patent Citations
Modeling of the head-related impulse responses
WO2021074294A1
Error correction of head-related filters
WO2022223132A1