Binaural determination of direction to an audio object

By decomposing binaural signals into frequency bands and combining ILD and ITD cues, the method addresses the challenges of high complexity and accuracy in sound-source localization, enabling efficient and accurate sound direction tracking in audio systems.

WO2025193580A1PCT designated stage Publication Date: 2025-09-18DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/019121
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-14
Filing Date
2025-03-10
Publication Date
2025-09-18

Smart Images

  • Figure US2025019121_18092025_PF_FP_ABST
    Figure US2025019121_18092025_PF_FP_ABST
Patent Text Reader

Abstract

An audio-signal processing method directed at determining the direction to a sound source based on certain statistical characteristics of the corresponding binaural sound. In at least some examples, the method is configured to obtain at least two different azimuth-angle estimates using two different signal models, one representing an interaural level difference cue and the other representing an interaural time difference cue. The models are parameterized such that the computational complexity associated with computing the two azimuth-angle estimates is relatively low. The method is further configured to combine the two different azimuth-angle estimates to obtain a more-accurate azimuth-angle estimate. Beneficially, in at least some cases, directional position of a moving sound source can be accurately tracked in real time due to the relatively low computational complexity of the disclosed method.
Need to check novelty before this filing date? Find Prior Art

Description

BINAURAL DETERMINATION OF DIRECTION TO AN AUDIO OBJECT 1. Cross-Reference to Related Applications

[0001] This application claims the benefit of priority from International Patent Application No. PCT / CN2024 / 081445, filed 13 March 2024, U.S. Provisional Application No. 63 / 645,066, filed 9 May 2024, and European Patent Application No. 24175590.9, filed 14 May 2024, each of which is incorporated by reference herein in its entirety. 2. Field of the Disclosure

[0002] Various example embodiments relate to audio equipment and, more specifically but not exclusively, to binaural determination of the direction to an audio object. 3. Background

[0003] Binaural segregation and localization are important problems within computational auditory scene analysis and audio signal processing. Various technical solutions related to these problems find applications in simulating auditory perception, hearing prostheses, robust speech recognition, spatial sound reproduction, mobile robotics, etc. Binaural sound-source localization typically uses two microphones in a setup that emulates human-like hearing. Monaural and binaural cues can be extracted from the microphone signals and used to estimate the angular position (azimuth) of the sound source. However, in at least some cases, azimuth estimation in various acoustic environments is either not accurate enough or too computationally complex for at least some of the aforementioned applications. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS

[0004] Disclosed herein are various embodiments of an audio-signal processing method directed at determining the direction to a sound source based on certain statistical characteristics of the corresponding binaural sound. In at least some examples, the method is configured to obtain at least two different azimuth-angle estimates using two different signal models, one representing an interaural level difference (ILD) cue and the other representing an interaural time difference (ITD) cue. The models are parameterized such that the computational complexity associated with computing the two azimuth-angle estimates is relatively low. The method is further configured tocombine the two different azimuth-angle estimates to obtain a more-accurate azimuth-angle estimate. Beneficially, in at least some cases, directional position of a moving sound source can be accurately tracked in real time due to the relatively low computational complexity of the disclosed method.

[0005] According to an example embodiment, provided is an audio system for object-based audio, the audio system comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: decompose a left binaural signal and a right binaural signal corresponding to an audio object into band signals corresponding to a plurality of audio frequency bands; and, for a selected band of the plurality of audio frequency bands, obtain a set of statistical characteristics of a corresponding pair of the band signals; compute a first azimuth- angle estimate based on a first subset of the set of statistical characteristics and further based on a model representing an ILD cue; compute a second azimuth-angle estimate based on a different second subset of the set of statistical characteristics and further based on a model representing an ITD cue; and compute a third azimuth-angle estimate using a weighted sum of the first and second azimuth-angle estimates.

[0006] According to another example embodiment, provided is a method of determining a direction to an audio object, the method comprising: decomposing a left binaural signal and a right binaural signal corresponding to the audio object into band signals corresponding to a plurality of audio frequency bands; and, for a selected band of the plurality of audio frequency bands, obtaining a set of statistical characteristics of a corresponding pair of the band signals; computing a first azimuth-angle estimate based on a first subset of the set of statistical characteristics and further based on a model representing an ILD cue; computing a second azimuth-angle estimate based on a different second subset of the set of statistical characteristics and further based on a model representing an ITD cue; and computing a third azimuth-angle estimate using a weighted sum of the first and second azimuth-angle estimates.

[0007] According to yet another example embodiment, provided is a non-transitory computer- readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method of determining the direction to an audio object.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:

[0009] FIG. 1 is a schematic diagram illustrating interaural differences of time and intensity for an ideal spherical head according to one example.

[0010] FIG. 2 is a block diagram illustrating an audio system in which various embodiments can be practiced according to some examples.

[0011] FIG. 3 is a block diagram illustrating a source direction detection block used in the audio system of FIG. 2 according to various examples.

[0012] FIGS. 4A-4B graphically illustrate an approximation that can be used in the ILD estimation component used in the source direction detection block of FIG. 3 according to some examples.

[0013] FIG. 5 is a flowchart illustrating a method of determining the azimuth angle to a sound source implemented in the audio system of FIG. 2 according to various examples.

[0014] FIGS. 6A-6E graphically illustrates azimuth-angle determination results obtained with the method of FIG. 5 according to one example.

[0015] FIG. 7 is a block diagram illustrating a computing device, one or more instances of which can be used in the audio system of FIG. 2 according to various examples. DETAILED DESCRIPTION

[0016] FIG. 1 is a schematic diagram illustrating interaural differences of time and intensity for an ideal spherical head 102 subjected to sound waves 101 received from a distant sound source according to one example. In the example shown, the sound waves 101 arrive earlier in time at the right ear R than at the left ear L. This time-of-arrival difference is referred to as the interaural time difference (ITD). The sound intensity difference between the ipsilateral right ear R and the contralateral left ear L is referred to as the interaural level difference (ILD). The ILD is caused bythe “shadowing” effect of the head 102, which prevents some of the incoming sound energy from reaching the contralateral left ear L.

[0017] In various examples, the ITD and ILD cues are effective in complementary ranges of frequencies of acoustic environments. More specifically, the ILDs are more pronounced at frequencies greater than approximately 1.5 kHz because the head 102 is large compared to the wavelengths of the incoming sound waves 101 at those frequencies, thereby producing substantial sound reflections. The ITDs, on the other hand, are relatively similar for all frequencies. However, for periodic sounds, the ITDs can be decoded unambiguously only at frequencies for which the maximum physically possible ITD is smaller than one half of the period of the waveform at that frequency. Since for a typical human head 102 the maximum possible ITD is about 660 μs, the ITDs are generally more useful at frequencies lower than about 1.5 kHz. For pure tones, the interaural phase delay (IPD) can also be used due to the straightforward one-to-one mapping between the ITDs and the IPDs of pure tones (also see Eq. (33)).

[0018] In some cases, for a spherical head 102 having a uniform surface, the ITD, ^, corresponding to the sound waves 101 can be approximated using Eq. (1) as follows: ^= ^2^ / ^^ sin ^ (1)where r is the radius of the head 102; c is the speed of sound; and ^ is the azimuth angle. The approximation corresponding to Eq. (1) can be relatively accurate for frequencies lower than approximately 500 Hz but breaks down for higher frequencies. For frequencies higher than about 2 kHz, the ITD corresponding to the sound waves 101 can be approximated using Eq. (2) as follows: ^= ^^ / ^^^^ + sin ^^ (2)In some examples, the radius r is about 9 cm.

[0019] The ILD values are not as well predicted by analytical models as the ITD values. In some examples, the ILDs can be obtained by applying suitable signal processing to microphone signals captured by two microphones corresponding to the left ear L and the right ear R, respectively. It has been found empirically that the ILDs tend to depend on the azimuth angle of the sound source, frequency, and distance from the sound sources to the head 102. For example, the ILDs produced by distant sound sources can become as large as 25 dB in magnitude at relatively high frequencies, and the ILDs can be even greater when the sound source is relatively close to one of the ears L and R.

[0020] FIG. 2 is a block diagram illustrating an audio system 200 in which various embodiments can be practiced according to some examples. The system 200 includes a microphone array 210 having microphones 2021 and 2022 that are separated by a nonzero distance 2r. For illustration purposes and without any implied limitations, the first microphone 2021 is also referred as the left (L) microphone, and the second microphone 2022is also referred as the right (R) microphone. In the example shown, the microphone array 210 receives a sound wave 201 from a sound source S. The azimuth angle ^ to the sound source S is defined in accordance with the definition illustrated in FIG. 1. In some examples, the sound source S is movable with respect to the microphone array 210. In such examples, the azimuth angle ^ may change over time. In various additional examples, the audio environment corresponding to the audio system 200 may include two or more separate and distinct sound sources variously positioned with respect to the microphone array 210. The term ‘audio object’ may be used herein to refer to one or more sound sources that may be perceived to emanate from a particular physical location or locations in the listening environment during playback.

[0021] In some examples, the microphone array 210 may be absent, and audio signals 212 may be provided to the audio system 200 by a suitable external source. In some configurations of the audio system 200, such an external source can be a memory in which the previously recorded audio signals 212 are stored. In some other configurations of the audio system 200, such an external source can be a binaural renderer or virtualizer configured to transform other audio into binaural audio. For illustration purposes and without any implied limitations, example embodiments are described below in reference to the configuration illustrated in FIG. 2, wherein the audio signals 212 are generated with the microphone array 210. Based on the provided description, a person of ordinary skill in the pertinent art will be able to make and use other embodiments of the audio system 200 without any undue experimentation.

[0022] The left and right audio signals 212 captured by the microphones 2021 and 2022, respectively, are subjected to spectral decomposition in a QMF transform block 220. Herein, the acronym QMF stands for Quadrature Mirror Filter. In one example, the QMF transform block 220 includes a QMF filter bank configured to decompose the audio signals 212 into a plurality of QMF band signals 222, each representing a specific frequency range of the original signals 212. The term “quadrature” refers to the phase shift relationship between the bands, and the term “mirror” denotes the symmetrical nature of the filter responses about the Nyquist frequency. An important propertyof the frequency band signals 222 obtained with the QMF filter bank of the QMF transform block 220 is that these signals enable an (almost) perfect signal reconstruction as well as a good separation between the band signals (no spectral leakage) without any audible aliasing distortions. In various examples, the QMF filter bank of the QMF transform block 220 can be implemented using a plurality of finite impulse response (FIR) filters and / or infinite impulse response (IIR) filters. In some examples, the number (N+1) of QMF band signals 222 is in the range between 8 and 64.

[0023] The band signals 222 are applied to a source direction detection block 230 configured to perform signal processing directed at determining the azimuth angle ^ corresponding to the sound source S. In some examples, the source direction detection block 230 determines the value of the azimuth angle ^ on a per-band basis. In some cases, the per-band azimuth angle values can be combined into a single azimuth angle value, e.g., via weighted sum averaging. Various examples of signal processing performed in the block 230 are described in more detail below, e.g., in reference to FIGS. 3-7.

[0024] An output signal 232 generated by the source direction detection block 230 provides the determined azimuth-angle value(s) to a downstream processing block 240. The downstream processing block 240 may also receive one or more additional inputs 238, e.g., including the original signals 212 and / or the band signals 222. The processing operations of the downstream processing block 240 depend on the specific embodiment and / or application. For example, one application can be directed at separating the audio object corresponding to the sound source S from binaural audio or at guiding postprocessing. Another application can be directed at up-mixing binaural content to object-based content or multi-channel content for re-rendering of the resulting combined audio scene. It should be noted here that the number of different applications of the binaural audio technology is growing fast, with the corresponding audio systems being configured to handle a larger variety of virtualized signals on wearable or mobile audio devices with the goal of providing better immersive and / or personalized experience to the users. As such, the target platforms for being connected with or incorporating the source direction detection block 230 may include a wide range of devices, from portable devices to large speaker system installations. Beneficially, the computational complexity of the signal-processing algorithm implemented in the source direction detection block 230 is relatively low, which makes it especially attractive for wearable devices.

[0025] Output audio signals 242 generated by the downstream processing block 240 based on the signals 232 and 238 are directed to an audio rendering component 250. In various examples, the audio rendering component 250 operates to render and play back the audio content represented by the audio signals 242. In various examples, the audio rendering component 250 may include any professional or consumer-grade audio system, such as a home theater (e.g., including an A / V receiver, a soundbar, a Blu-ray player, etc.), one or more E-media devices (e.g., a computer, a tablet, a mobile phone equipped with headphones and / or speakers, etc.), a TV set, and a sound reproduction system. In some examples, the audio rendering component 250 provides an audio environment for playback of audio or audio / visual content using a plurality of speakers and suitable playback devices. In some examples, the audio rendering component 250 may represent any environment in which a listener is experiencing playback of the audio content, such as a cinema, a concert hall, an outdoor theater, a home theater or room, a listening booth, a car, a game console, a headset device, a public address (PA) system, or other audio playback environment.

[0026] FIG. 3 is a block diagram illustrating the source direction detection block 230 according to various examples. The input to the source direction detection block 230 includes the band signals 222 having a first set of band signals 222L corresponding to the microphone 2021 and a second set of band signals 222R corresponding to the microphone 2022. Each of the sets 222L and 222R includes (N+1) individual band signals {222Ln} and {222Rn}, respectively, where n is an index in the range [0, N]. For each matching pair of the band signals 222Lnand 222Rn, the source direction detection block 230 has a corresponding band direction detection block 300n. For brevity, only one of the blocks 300nis explicitly shown in FIG. 3, i.e., the band direction detection block 3000. In various examples, other blocks 300n(not explicitly shown) may have a structure that is similar to the shown structure of the block 3000.

[0027] The block 3000 includes a smoothing component 310 configured to receive the band signals 222L0and 222R0. In one example, the smoothing component 310 is configured to apply the following smoothing operation to each of the band signals 222L0 and 222R0: ^^^^^^^^^^ = ^^^^^^^^^^^^^^^^ − 1^ + ^1 − ^^^^^^^^^^^^ (3)where ^^^^^^^and i is an index. Each of the band signals 222L0 and 222R0 provides a respective stream of the input values ^^^^. The streams of corresponding output values ^^^^^^^^^^generated by the smoothing component 310 form smoothed band signals 312L and 312R.

[0028] The block 3000also includes a statistical estimation component 320 configured to receive the smoothed band signals 312L and 312R. In one example, using these signals, the statistical estimation component 320 operates to compute a binaural cross-channel product M and a combined binaural channel energy T as follows: ^= ^^∗ (4)^ = ^^∗ + ^^∗ (5)where L and R denote the smoothed band signals 312L and 312R, respectively, and the “*” symbol denotes conjugation. The computed streams of M and T values are then provided to an ILD estimation component 330 and an IPD estimation component 340 as indicated in FIG. 3.

[0029] In one example, the ILD estimation component 330 performs the following operations: (1) Getting a head-related transfer function (HRTF) coefficient ^^for this band from a lookup table (LUT). In some examples, such a LUT can be based on Table 1. Table 1: HRTF Coefficient ^ Band # 10 11 12 13 14 15 16 17 18 19 8 38(2) Computing the parameter ^ #^ = $ .(3) Computing the azimuth-angleas ^%,'() = arcsin ^−^^^^^ − .1 + ^^ / ^^ if |^^^^^ −.1 + ^^ / ^| < 1.(4)the azimuth angle estimate as ^%,'() = arcsin ^−^^^^^ + .1 + ^^ / ^^.

[0030] In one example, the IPD estimation component 340 performs the following operations: (1) Getting a band-based constant value 2%from a lookup table. (2) Calculating the azimuth-angle estimate as ^%,'3) = 2%atan ^^^

[0031] The block 3000also includes a mixer 350 configured to combine the azimuth-angle values ^%,'()and ^%,'3)into a combined azimuth-angle value ^%for this frequency band. In some examples, the mixer 350 operates to compute ^%as follows: ^% = ^%^%,'$) + ^1 − ^%^^%,'() (6)where ^%is a band-specific weighting coefficient.

[0032] Example implementations of various components of the band direction detection block 300n(where n=0, 1, …, N) are described in more detail below. ILD Cue

[0033] The ILD cues are processed in the block 300n using the respective ILD estimation component 330. Some embodiments of the ILD estimation component 330 may benefit from at least some features described in International Patent Application Publication No. WO2022182943A1, which is incorporated herein by reference in its entirety.

[0034] A z-domain representation of an ILD filter is given by Eq. (7) as follows: 56 = % 8% ;9^ ^ 7 9:<78<9:;9 (7)where 5^6^ is the filter’sare filter coefficients. Eq. (7) can also be rewritten by factoring out the coefficient a0and redefining the coefficients a1, b0, and b1. The resulting expression for the filter’s transfer function is as follows: %;95^6^ = 78%9:(8)A frequency-domain the substitution 6 = >?@, whichproduces the following expressions: 5A>?@B = %78%9C;DE= %78%9"^^@8<7%7"^^@8<7%9 + G <7%7^HI@J%9^HI@(9)From Eq.domain is provided by: =K / + <7%7^HI@J%9^HI@ / where L =

[0035] Using a different parametrization of Eq. (10), one arrives at the following expression for|5|:5 = √OF8 / OP"^^ F F F F| | @8P "^^ @8Q ^HI @$ (11)whereV= S= + UTST (13)W = UTST − S= (14)^ = 1 + 2U / T^XYL + UT (15)

[0036] Using still another parametrization of Eq. (11), one arrives at the following expression for |5|: .Z 8Z ^HI[8Z ^ F|5| = 7 9 F HI [$ (16)where\T = (U + S^XYL) (17)\= = 42 / (1 − ^XYL)(U + S^XYL) (18)The parameters a and b transfer functioncoefficient UT of Eq. (8) as follows:U = 1 + U / T (20)S = 2UT (21)The parameter β used in Eq. (19) is tied to the radius r of the head 102 (FIG. 1). The parameters k and β used in Eqs. (18)-(19) are interrelated as follows: 2= / a8 / (22)The parameters k and β are also related to the z-domain transfer function coefficients a0, b0, and b1of Eq. (8) as follows: UT = aJ / a8 / (23)ST = 1 + 2Y^`^ (24)S= = UT − 2Y^`^ (25)

[0037] As can be seen from Eqs. (15) and (17)-(19), the coefficients \T, \=, \ / and ^ used in the expression for the amplitude frequency response |5| given by Eq. (16) are frequency dependent. For a selected fixed frequency, the coefficients \T, \=, \ / and ^ are constant. As such, the above parameterization enables a representation of the frequency response |5| as a square rootof a quadratic equation, in which the variable is the sine of the azimuth angle ^ (see Eq. (16)). Note that Eq. (16) beneficially lends itself to programming that causes the corresponding computations to have a relatively low computational complexity.

[0038] FIGS. 4A-4B graphically illustrate an approximation that can be implemented in the ILD estimation component 330 according to some examples. More specifically, the illustrated approximation reduces the expression for the coefficient \ / given by Eq. (19) to the following expression: \= 42^^1 / / − ^XYL^ (26)In other words, the second term, 2 / _ / Y^` / L, inthe righthand side of Eq. (19).

[0039] FIG. 4A graphically illustrates an effect of the above approximation on the modeled amplitude response for the ipsilateral ear. Therein, a curve 402 graphically shows the amplitude response modeled with Eq. (16) in which the coefficient \ / is given by Eq. (19). A curve 412 graphically shows the amplitude response modeled with Eq. (16) in which the coefficient \ / is given by Eq. (26). The difference between the curves 402 and 412 is relatively small, which indicates that the approximation represented by Eq. (26) is relatively accurate in at least some examples.

[0040] FIG. 4B graphically illustrates an effect of the above approximation on the modeled amplitude response for the contralateral ear. Therein, a curve 404 graphically shows the amplitude response modeled with Eq. (16) in which the coefficient \ / is given by Eq. (19). A curve 414 graphically shows the amplitude response modeled with Eq. (16) in which the coefficient \ / is given by Eq. (26). While being larger than the difference between the curves 402 and 412 of FIG. 4A, the difference between the curves 404 and 414 is still sufficiently small for the approximation represented by Eq. (26) to be sufficiently accurate for the contralateral ear as well, in at least some examples.

[0041] With the coefficient \ / given by Eq. (26), the amplitude frequency response|5|can be approximated as follows: .(<8 F( ) F= %"^^@8 / b =J"^^@ ^HI[)By parametrizing Eq.expression for |5|:|5| = I8^^HI[<I (28)wherec = 22 / (1 − ^XYL) (29)` = U + S^XYL (30)Based on the above, the contralateral ear amplitude frequency response|5|"^I^d<can be approximated as follows: |5| ^^< = 1 + HI["^I^d I (31)The corresponding frequency response |5|He^Hcan beobtained from Eq. (31) by changing the sign of Y^`^, which leads to the following expression: |5| ^^HI[He^H = 1 − I (32)For the approximation m, n are fixed for a fixedfrequency and fixed size of the head 102. ITD Cue

[0042] The ITD cues are processed in the block 300n using the respective IPD estimation component 340. As already indicated above, the ITD is caused by a relative difference in the effective travel distance from the sound source S to the two ears L, R or to the two microphones 2021 and 2022. The ITD can also be represented using the group delay or IPD. The ITD and IPD on a frequency band are related as: f^g(!") = '3)(hi)hi (33)where !"is the center frequencythe ITD can be represented by a single value for all frequency bands rather than by different respective values for different frequency bands.

[0043] In some examples, the IPD can be extracted from the audio data as follows:fjg = atan (k^k^kkk∗) (34)where the top bar denotes time averaging. The azimuth angle estimate ^%,'3)is computed from theIPD as follows:^%,'3) = 2%fjg = 2%atan (k^k^kkk∗) (35)where 2% is band-of the scaling coefficient2% used in the IPD estimation component 340 according to some examples.Table 2: Band-Dependent Scaling Coefficient lmBand # 0 2 4 5 6 7 8 9 10 !"(Hz) 47 141 234 281 328 422 469 516 656 9

[0044] The ITD effect can be observed more readily by a listener on lower frequencies and is more likely to be applied on lower frequency bands than on higher frequency bands in a virtualizer adapted for generating stereo audio. There are at least two reasons for these aspects of the model representing the ITD cue, one reason being that the head-shadow effect will be weakened for higher frequencies, and the other reason being that the phase will be wrapped around 2M when the frequency is higher, thereby making it harder to simulate or disambiguate the time difference based solely on the phase difference. The phase wrapping does not occur when the following condition is met: 2f^g^<n! < 1 (36)For example, for f^g^<n= 80 ms, the maximum nonwrapped frequency is about 625 Hz. For an example QMF band structure shown in Table 2, the highest nonwrapped band is the 10thband. Signal Model

[0045] The source direction detection block 230 operates to process the plurality of frequency band signals 222 generated by the QMF transform block 220 (see FIG. 2). Each of the band signals 222 is a two-channel signal, and these two channels are denoted as Lb and Rb, respectively, in the signal model presented below. For higher frequency bands, the ILD is the dominant feature for rendering because the IPD on those bands wraps one or more times and the amplitude response of the HRTF is a major contributor to binaural listening. For those bands, the presented signal model does not rely on the ITD. The expressions for the band’s channels Lb and Rb are as follows: ^% = ℎpY + qp (37)^% = ℎdY + qd (38)where Y is the source object; ℎpand ℎdare amplitude responses of the HRTF for the left and right channels; and qpand qdare diffuse signals for the left and right channels.

[0046] Eqs. (37) and (38) are approximations that are valid under the following set of assumptions: rsYYt = u / (39)= = = 0 (40)where E denotes the matrix of the band’schannels. Considering response and assuming that the left ear is the contralateral ear, the covariance matrix of Eq. (42) can be expressed as follows: (1 + ^^HI[ / / / ^^HI[ / / =I ) u + g [1 − { I |]u =~^== ^= / ^method: min rs(~ℎpY ℎpY + qℎ ^ − ^ ~ p+ ^) / t (44)By solving thethe following expression for the steer matrix W: ℎ / / / é p u^ / / ℎpℎdu^ / / Thisthe statistical estimation component 320, e.g., as follows: √#F8^F8^ #where M and T are given byenergy difference of the leftand right smoothed signals expressed as:^ = ^^∗ − ^^∗ (47)

[0048] From Eq. (43) for the covariance matrix, the ratio M / T can be expressed as:# $= ^9F8^F9^99J^FF = IFJ^F^HIF[ / ^I^HI[ ≝ ^^ (48)I Let^ = ^^ be the HRTF above equation can berewritten as the following quadratic equation with respect to Y^`^: Y^` / ^ + 2^ / ^^^Y^`^ − ^^ = 0 (49)By solving this quadratic equation, one finds the following expression for Y^`^: Y^`^ = −^^^^^ ∓ .1 + ^^ / ^ (50)The valid value range for in Eq. (50) is resolvedas: arcsin ^−^^^^ − .1 + ^ / ^^, ^! |^ ^^ − .1 + ^ / ^| < 1^ ^ ^ ^ ^ ^'() = ^ (51)arcsin ^−^^^^^ + .1 + ^^ / ^^, ^! |^^^^^ + .1 + ^^ / ^| < 1where ^^is the HRTF related coefficient which is a respective constant for each band; and ^^=#$ (also see Eq. (48)). In at least some examples, Eq. (51) can be used to program the ILD estimation components 330 of the different band direction detection blocks 300nof the source direction detection block 230.

[0049] Finally, ^%,'3)can be obtained from Eqs. (4) and (35): ^%,'3) = 2%atan ^^^ (52)In at least some examples, Eq. (52) can be used to program the IPD estimation component 340 in the different band direction detection blocks 300nof the source direction detection block 230. Mixer Operations

[0050] As already indicated above, the ITD and ILD have different respective levels of reliability for different frequency ranges. The respective mixers 350 of the different band direction detection blocks 300nare configured to use different respective pairs of weighting coefficients that track the respective levels of reliability of the estimated ITD and ILD values. In one example, the mixer 350 is configured to compute the value of ^%for the corresponding band b as follows: ^% = ^%^%,'$) + ^1 − ^%^^%,'() (53)where ^%is the confidence value for the corresponding band b. In some examples, the confidence values ^%for different bands can be assigned to different bands based on the following expression: ^^>J% , S < S^^where S^^is a threshold value. Example Workflow

[0051] FIG. 5 is a flowchart illustrating a method 500 of determining the azimuth angle to an audio object implemented in the audio system 200 according to some examples. In some examples, the method 500 can be executed by one or more computing devices programmed to perform various operations of the QMF transform block 220 and the source direction detection block 230. An example computing device 700 suitable for these purposes is described below in reference to FIG. 7.

[0052] The method 500 includes the computing device receiving the L and R channel signals generated with an array of microphones (in a block 502). In some examples, the signals received in the block 502 are the L and R channels of the audio signals 212 captured by the microphones 2021and 2022. In some other examples, the signals received in the block 502 are retrieved from a memory in which those signals (e.g., analogous to the audio signals 212) were previously saved. In some cases, the signal retrieval is performed via a communication link, e.g., via a network.

[0053] The method 500 also includes the computing device decomposing each of the received L and R channel signals into a respective plurality of band signals (in a block 504). Operations of the block 504 include applying a QMF transform to the received L and R channel signals to generate a first set of QMF band signals (e.g., 222L, FIG. 3) corresponding to the L channel and a second set of band signals (e.g., 222R, FIG. 3) corresponding to the R channel. In some examples, operations of the block 504 include an optional operation of rebounding the QMF band signals to reduce the total number of bands. Such rebounding is performed such that the first and second sets of rebounded- band signals represent the same band configuration so that matching L-R pairs of band signals can be provided for subsequent processing.

[0054] The method 500 also includes the computing device performing operations of blocks 5060-506N. In some examples, the blocks 5060-506N are executed in parallel, as indicated in FIG. 5. In some other examples, the blocks 5060-506N are executed in a sequence (not explicitly shown). Each of the blocks 5060-506Nis configured to process a respective one matching L-R pair of band signals. An example of such matching pair corresponding to the block 5060is the pair of band signals 222L0 and 222R0 shown in FIG. 3. For illustration purposes and without any implied limitations, only operations of the block 506nare explicitly shown in FIG. 5 and described below.Based on the provided description, a person of ordinary skill in the pertinent art will be able to implement the corresponding operations of all of the blocks 5060-506Nwithout any undue experimentation.

[0055] The block 506n optionally includes the computing device applying one or more smoothing operations to each of the band signals 222Lnand 222Rn(in a block 508). A nonlimiting example of such smoothing operations of the block 508 is represented by Eq. (3). The smoothed band signals 312L and 312R shown in FIG. 3 provide an example of the smoothed signals generated in the block 508.

[0056] The block 506nalso includes the computing device obtaining a set of statistical characteristics of the smoothed band signals (in a block 510). In some examples, the set of statistical characteristics includes the binaural cross-channel product M, combined binaural channel energy T, and binaural channel energy difference Y computed in accordance with Eqs. (4), (5), and (47), respectively.

[0057] The block 506n also includes the computing device obtaining first and second azimuth- angle estimates (in a block 512). Operations of the block 512 include computing the first azimuth- angle estimate based on a first subset of the set of statistical characteristics obtained in the block510. In some examples, these computations include computing the parameter ^ #^ = $ and finding asolution to the quadratic equation given in Eq. (49). Operations of the block 512include computing the second azimuth-angle estimate based on a different second subset of the set of statistical characteristics obtained in the block 510. In some examples, these computations include computing the estimate based on the binaural cross-channel product M in accordance with Eq. (52).

[0058] The block 506n also includes computing a third azimuth-angle estimate (in a block 514). In some examples, the third azimuth-angle estimate is obtained by computing a weighted sum of the first and second azimuth-angle estimates computed in the block 512. These computations can be performed, e.g., in accordance with Eq. (53), with the weights being determined based on Eq. (54).

[0059] The method 500 also includes the computing device computing a fourth azimuth-angle estimate (in a block 516). In some examples, the fourth azimuth-angle estimate is obtained by computing a weighted sum of some or all of the third azimuth-angle estimates computed in the blocks 5060-506N. In other words, the fourth azimuth-angle estimate is obtained by combining theazimuth-angle estimates corresponding to different bands. In some examples, the block 516 is optional and may be omitted.

[0060] The method 500 also includes the computing device outputting the computed azimuth- angle estimates (in a block 518). Such outputting may include transmitting the computed azimuth- angle estimates to the device or system configured to implement the downstream processing block 240. In various examples, the transmitted estimates may include the third azimuth-angle estimates computed in the blocks 5060-506N and / or the fourth azimuth-angle estimate computed in the block 516. The method 500 is terminated after the output operations of the block 518 are completed.

[0061] As indicated above, in some examples of the method 500, some or all of the QMF bands 222 are rebounded to reduce the total number of bands processed in the source direction detection block 230. Table 3 illustrates an example configuration in which such QMF-bands rebounding reduces the total number of bands down to 20 bands. In this example, the sampling rate is 48 kHz. Table 3: Example Rebounded Band Configuration Band # 0 1 2 3 4 5 6 7 8 9 8 8

[0062] FIGS. 6A-6E graphically illustrates azimuth-angle determination results obtained with the method 500 according to one example. In the example shown, the sound source S emits a bongos sound while rotating around the microphone array 210 at a constant angular speed. The angular range of rotation is from 0.5M to −0.5M. The QMF bands 222 are rebounded into the twenty bands described in Table 3. The threshold value used in the computations corresponding to Eq. (54) is S^^=6.

[0063] Each of the individual band graphs presented in FIGS. 6A-6E shows a respective reference azimuth angle of the sound source S as a function of time and a corresponding azimuth- angle estimate obtained with the method 500. The respective line types used for plotting thereference and the estimate in the graphs are indicated in the legends. Each reference is a straight line with the fixed slope representing the constant angular speed of the sound source S. For most of the individual bands, the azimuth-angle estimates obtained with the method 500 follow the reference line fairly accurately, thereby demonstrating excellent performance of the method 500. For the first and last bands, deviations of the azimuth-angle estimates from the reference line are larger than for the other bands due to the lower energy density of the bongos sound in those bands, which causes the signal-to-noise ratio to be relatively low and affects the accuracy of the azimuth-angle estimates accordingly. Example Hardware

[0064] FIG. 7 is a block diagram illustrating a computing device 700 according to various examples. In some examples, the computing device 700 is configured to perform at least some operations of the method 500. In some examples, two or more instances of the computing device 700 are used in the audio system 200.

[0065] The computing device 700 of FIG. 7 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 700 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 702 and one or more storage devices 704). Additionally, in various embodiments, the computing device 700 may not include one or more of the components illustrated in FIG. 7, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 700 may not include a display device 710, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 710 may be coupled.

[0066] The computing device 700 includes a processing device 702 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic datafrom registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. In various embodiments, the processing device 702 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.

[0067] The computing device 700 also includes a storage device 704 (e.g., one or more storage devices). In various embodiments, the storage device 704 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memory devices, solid- state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 704 may include memory that shares a die with the processing device 702. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 704 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device 702), cause the computing device 700 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.

[0068] The computing device 700 further includes an interface device 706 (e.g., one or more interface devices 706). In various embodiments, the interface device 706 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 700 and other computing devices. For example, the interface device 706 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 700. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 706 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi(IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 706 for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 706 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, circuitry included in the interface device 706 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 706 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.

[0069] In some embodiments, the interface device 706 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 706 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 706 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 706 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 706 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 706 may be dedicated to wireless communications, and a second set of circuitry of the interface device 706 may be dedicated to wired communications.

[0070] The computing device 700 also includes battery / power circuitry 708. In various embodiments, the battery / power circuitry 708 may include one or more energy storage devices (e.g.,batteries or capacitors) and / or circuitry for coupling components of the computing device 700 to an energy source separate from the computing device 700 (e.g., to AC line power).

[0071] The computing device 700 also includes a display device 710 (e.g., one or multiple individual display devices). In various embodiments, the display device 710 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.

[0072] The computing device 700 also includes additional input / output (I / O) devices 712. In various embodiments, the I / O devices 712 may include one or more data / signal transfer interfaces, audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc.

[0073] Depending on the specific embodiment, various components of the interface devices 706 and / or I / O devices 712 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 706 and / or I / O devices 712 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 702 and / or the storage device 704. In some additional examples, the interface devices 706 and / or I / O devices 712 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 702 and / or the storage device 704 into an analog form suitable for being transmitted through a communication channel.

[0074] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGs. 1-7, provided is an audio system for object-based audio, the audio system comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: decompose a left binaural signal and a right binaural signal corresponding to an audio object into band signals corresponding to a plurality of audio frequency bands; and, for a selected band of the plurality of audio frequency bands, obtain a set of statistical characteristics of a corresponding pair of the bandsignals; compute a first azimuth-angle estimate based on a first subset of the set of statistical characteristics and further based on a model representing an ILD cue; compute a second azimuth- angle estimate based on a different second subset of the set of statistical characteristics and further based on a model representing an ITD cue; and compute a third azimuth-angle estimate using a weighted sum of the first and second azimuth-angle estimates.

[0075] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGs. 1-7, provided is a method of determining a direction to an audio object, the method comprising: decomposing a left binaural signal and a right binaural signal corresponding to the audio object into band signals corresponding to a plurality of audio frequency bands; and, for a selected band of the plurality of audio frequency bands, obtaining a set of statistical characteristics of a corresponding pair of the band signals; computing a first azimuth-angle estimate based on a first subset of the set of statistical characteristics and further based on a model representing an ILD cue; computing a second azimuth- angle estimate based on a different second subset of the set of statistical characteristics and further based on a model representing an ITD cue; and computing a third azimuth-angle estimate using a weighted sum of the first and second azimuth-angle estimates.

[0076] In some embodiments of the above method, the decomposing comprises applying a QMF transform to the left and right binaural signal.

[0077] In some embodiments of any of the above methods, the decomposing further comprises rebounding at least a subset of QMF frequency bands obtained via the QMF transform to reduce a total number of frequency bands.

[0078] In some embodiments of any of the above methods, the obtaining comprises: smoothing a left channel of the selected frequency band to obtain a first smoothed signal; and smoothing a right channel of the selected frequency band to obtain a second smoothed signal.

[0079] In some embodiments of any of the above methods, the set of statistical characteristics is selected from the group consisting of: a cross product of the first and second smoothed signals; a combined energy of the first and second smoothed signals; and an energy difference of the first and second smoothed signals.

[0080] In some embodiments of any of the above methods, the first subset of the set of statistical characteristics includes a cross product of the first and second smoothed signals and a combined energy of the first and second smoothed signals; and wherein the different second subset of the set of statistical characteristics includes the combined energy of the first and second smoothed signals.

[0081] In some embodiments of any of the above methods, weights used in the weighted sum are band dependent.

[0082] In some embodiments of any of the above methods, the weight assigned to the second azimuth-angle estimate exponentially decreases with an increase of the center frequency when the center frequency of the selected band is smaller than a threshold value; and wherein the weight assigned to the second azimuth-angle estimate is zero when a center frequency of the selected band is greater than the threshold value.

[0083] In some embodiments of any of the above methods, said computing the first azimuth- angle estimate includes finding a solution to a quadratic equation with respect to a sine of the azimuth angle.

[0084] In some embodiments of any of the above methods, coefficients of the quadratic equation are determined based on the first subset of the set of statistical characteristics and further based on a head-related transfer function (HRTF).

[0085] In some embodiments of any of the above methods, said computing the second azimuth- angle estimate includes: computing an arctangent value corresponding to a cross product of the first and second smoothed signals; and applying a band-dependent scaling coefficient to the computed arctangent value.

[0086] In some embodiments of any of the above methods, the method further comprises computing a fourth azimuth-angle estimate using a weighted sum of the third azimuth-angle estimates corresponding to a plurality of different selected bands of the plurality of audio frequency bands.

[0087] In some embodiments of any of the above methods, the method further comprises outputting the third azimuth-angle estimate or the fourth azimuth-angle estimate to a computing device configured to perform downstream processing of the audio object.

[0088] In some embodiments of any of the above methods, the computing device implements a virtualizer configured to cause a wearable headset to generate stereo sound corresponding to the audio object.

[0089] In some embodiments of any of the above methods, the method further comprises outputting the third azimuth-angle estimate to a computing device configured to perform downstream processing of the audio object.

[0090] In some embodiments of any of the above methods, the downstream processing includes modifying an audio scene representing object-based content by incorporating the audio object into the audio scene.

[0091] In some embodiments of any of the above methods, the method further comprises generating sound corresponding to the modified audio scene with an audio-rendering device.

[0092] In some embodiments of any of the above methods, the method further comprises: generating the left binaural signal using a first microphone; and generating the right binaural signal using a second microphone, wherein the first and second microphones are separated by a nonzero distance.

[0093] In some embodiments of any of the above methods, the azimuth angle is an angle between the direction to a sound source corresponding to the audio object and a line connecting the first and second microphones.

[0094] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.

[0095] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processesherein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0096] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0097] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0098] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

[0099] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.

[0100] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.

[0101] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non- transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.

[0102] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.

[0103] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.

[0104] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.

[0105] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor areseparate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”

[0106] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.

[0107] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”

[0108] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.

[0109] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.

[0110] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not beconstrued to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.

[0111] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0112] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.

[0113] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

CLAIMS What is claimed is:

1. A method of determining a direction to an audio object, the method comprising: decomposing a left binaural signal and a right binaural signal corresponding to the audio object into band signals corresponding to a plurality of audio frequency bands; and for a selected band of the plurality of audio frequency bands, obtaining a set of statistical characteristics of a corresponding pair of the band signals; computing a first azimuth-angle estimate based on a first subset of the set of statistical characteristics and further based on a model representing an interaural level difference (ILD) cue; computing a second azimuth-angle estimate based on a different second subset of the set of statistical characteristics and further based on a model representing an interaural time difference (ITD) cue; and computing a third azimuth-angle estimate using a weighted sum of the first and second azimuth-angle estimates.

2. The method of claim 1, wherein the decomposing comprises applying a quadrature mirror filter (QMF) transform to the left and right binaural signal.

3. The method of claim 1 or 2, wherein the decomposing further comprises rebounding at least a subset of QMF frequency bands obtained via the QMF transform to reduce a total number of frequency bands.

4. The method of any preceding claim, wherein the obtaining comprises: smoothing a left channel of the selected frequency band to obtain a first smoothed signal; and smoothing a right channel of the selected frequency band to obtain a second smoothed signal.

5. The method of claim 4, wherein the set of statistical characteristics is selected from the group consisting of: a cross product of the first and second smoothed signals; a combined energy of the first and second smoothed signals; and an energy difference of the first and second smoothed signals.

6. The method of claim 4, wherein the first subset of the set of statistical characteristics includes a cross product of the first and second smoothed signals and a combined energy of the first and second smoothed signals; and wherein the different second subset of the set of statistical characteristics includes the combined energy of the first and second smoothed signals.

7. The method of any preceding claim, wherein weights used in the weighted sum are band dependent; wherein the weight assigned to the second azimuth-angle estimate exponentially decreases with an increase of the center frequency when the center frequency of the selected band is smaller than a threshold value; and wherein the weight assigned to the second azimuth-angle estimate is zero when a center frequency of the selected band is greater than the threshold value.

8. The method of any preceding claim, wherein said computing the first azimuth-angle estimate includes finding a solution to a quadratic equation with respect to a sine of the azimuth angle.

9. The method of claim 8, wherein coefficients of the quadratic equation are determined based on the first subset of the set of statistical characteristics and further based on a head-related transfer function (HRTF).

10. The method of claim 4 or any claim dependent thereon, wherein said computing the second azimuth-angle estimate includes: computing an arctangent value corresponding to a cross product of the first and second smoothed signals; andapplying a band-dependent scaling coefficient to the computed arctangent value.

11. The method of any preceding claim, further comprising computing a fourth azimuth-angle estimate using a weighted sum of the third azimuth-angle estimates corresponding to a plurality of different selected bands of the plurality of audio frequency bands.

12. The method of any preceding claim, further comprising outputting the third azimuth-angle estimate or the fourth azimuth-angle estimate to a computing device configured to perform downstream processing of the audio object.

13. The method of claim 12, wherein the computing device implements a virtualizer configured to cause a wearable headset to generate stereo sound corresponding to the audio object.

14. The method of any preceding claim, further comprising outputting the third azimuth angle estimate to a computing device configured to perform downstream processing of the audio object.

15. The method of claim 14, wherein the downstream processing includes modifying an audio scene representing object-based content by incorporating the audio object into the audio scene.

16. The method of claim 15, further comprising generating sound corresponding to the modified audio scene with an audio rendering device.

17. The method of any preceding claim, further comprising: generating the left binaural signal using a first microphone; and generating the right binaural signal using a second microphone, wherein the first and second microphones are separated by a nonzero distance.

18. The method of any preceding claim, wherein the azimuth angle is an angle between the direction to a sound source corresponding to the audio object and a line connecting the first and second microphones.

19. An audio system for object-based audio, the audio system comprising:at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: decompose a left binaural signal and a right binaural signal corresponding to an audio object into band signals corresponding to a plurality of audio frequency bands; and for a selected band of the plurality of bands: obtain a set of statistical characteristics of a corresponding pair of the band signals; compute a first azimuth-angle estimate based on a first subset of the set of statistical characteristics and further based on a model representing an ILD cue; compute a second azimuth-angle estimate based on a different second subset of the set of statistical characteristics and further based on a model representing an ITD cue; and compute a third azimuth-angle estimate using a weighted sum of the first and second azimuth-angle estimates.

20. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1 to 18.

Citation Information

Patent Citations

  • Virtualizer for binaural audio

    WO2022182943A1

  • Sound outputting apparatus and method of controlling the same

    US20120051553A1

  • Binaural signal post-processing

    WO2022133128A1