Spatially preserved binaural headtracking for a wearable device

WO2026177931A1PCT designated stage Publication Date: 2026-08-27DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/014919
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-07-10
Filing Date
2026-02-11
Publication Date
2026-08-27

Smart Images

  • Figure US2026014919_27082026_PF_FP_ABST
    Figure US2026014919_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A method for controlling a wearable device to output audio includes receiving input binaural audio configured to be output by the wearable device. The method includes receiving data indicative of an orientation and / or a movement of a head of a user of the wearable device from a sensor of the wearable device. The method includes determining a virtual listening space based on the data. The method includes determining a head-related transfer function (HRTF) based on the virtual listening space. The method includes rendering, based on the orientation and / or the movement of the head of the user, output binaural audio to be output by the wearable device. The output binaural audio may be rendered based on the input binaural audio and based on the HRTF.
Need to check novelty before this filing date? Find Prior Art

Description

SPATIALLY PRESERVED BINAURAL HEADTRACKING FOR A WEARABLE DEVICE CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 841,894, filed July 10, 2025 and International Patent Application No.PCT / CN2025 / 078092, filed February 192025, each incorporated by reference in their entireties.TECHNICAL FIELD

[0002] The present disclosure relates generally to outputting binaural audio with a wearable device that is worn on a head of a user. More specifically, the present disclosure relates to outputting the binaural audio based on input binaural audio in a manner that spatially preserves the intention of the content creator of the input binaural audio.SUMMARY

[0003] Wearable devices (e.g., earbuds, headphones, smart glasses, other head-mounted displays, etc.) are often equipped with one or more sensors and one or more electronic processors. For example, a wearable device may include an inertial measurement unit (IMU) configured to provide real-time data from a gyroscope(s) and / or an accelerometer(s) to indicate a status / orientation of the wearable device. This real-time data may be processed / analyzed to determine the orientation and / or physical movement of a head of a user wearing the wearable device. Using such orientation and movement information of the user, spatial computational audio can be output by the wearable device, for example, by executing one or more algorithms in the electronic processor(s) of the wearable device.

[0004] Executing such an algorithm(s) improves the user’s experience by providing a virtual listening space that is relatively static to the real-world environment surrounding the user rather than being static to the wearable device itself. In other words, when the user uses a wearable device equipped with such an algorithm(s), the sound will be perceived by the user as if the sound is generated from the real-world environment, which enhances the listening experience provided to the user by making the listening experience more immersive.

[0005] Binaural audio is audio that aims to replicate a way that a human hears sound by simulating spatial cues of human ears. In some instances, binaural audio may provide a moreimmersive listening experience compared to, for example, stereo audio that may merely provide a sense of directionality by separating sound into left and right channels. Binaural audio may be produced through binaural recording or may be rendered from object / channel -based audio with a head-related transfer function (HRTF). A spatial audio experience combines binaural audio and sensor data (e.g., IMU sensor data of a head-mounted device) to make the audio from the headmounted device sound like it is fixed to the real-world rather than fixed to head as mentioned previously herein. Such audio output may create a sound scape for the listener and may use a headtracking feature of the head-mounted / wearable device.

[0006] In many instances, spatial audio is rendered from multi-channel audio or object-based audio, so that during processing, each channel or object may be rendered considering head measured information like angles and position. Combined binaural audio is then output for consumption by the user. However, for wearable devices that receive input binaural audio, the input binaural audio may be professional generated content (i.e., professional binaural audio as explained herein) such that the audio channels are well-located to a specific position. In such instances, because wearable devices are usually located at the end of an audio processing chain, a wearable device applying an additional HRTF process on the input binaural audio may cause a technological problem of a double-processing issue. This double-processing issue distorts perceptual object localization in the input binaural audio and damages the intention of a creator of the input binaural audio (e.g., professional binaural audio content creator object-based intention). This distortion and damaged intention of the creator results in a lower quality listening experience for the user.

[0007] To address the above-noted technological problem of the double-processing issue, the systems, methods, and devices described herein perform headtracking of a user wearing a wearable device as well as processing of input binaural audio that preserves a spatial effect of the input binaural audio to maintain the intention of the creator of the input binaural audio. The disclosed systems, methods, and devices allow for binaural audio to be output by a wearable device in accordance with an orientation and / or movement of a user’s head as measured by a sensor(s) of the wearable device while reducing or preventing the distortion and damaged intention of output audio that may be caused by the double-processing issue described above.

[0008] In one aspect of the present disclosure, there is provided a method for controlling a wearable device to output audio. The method may include receiving, with an electronic processor, input binaural audio configured to be output by the wearable device. The method may include receiving, with the electronic processor and from a sensor of the wearable device,data indicative of an orientation and / or a movement of a head of a user of the wearable device. The method may include determining, with the electronic processor, a virtual listening space based on the data. The method may include determining, with the electronic processor, a head-related transfer function (HRTF) based on the virtual listening space. The method may include rendering, with the electronic processor and based on the orientation and / or the movement of the head of the user, output binaural audio to be output by the wearable device. The output binaural audio may be rendered based on the input binaural audio and based on the HRTF.

[0009] In addition to any combination of features described above, the method may include extracting, with the electronic processor, a center channel from the input binaural audio. The virtual listening space may include the center channel. The output binaural audio may be rendered based on the center channel from the input binaural audio.

[0010] In addition to any combination of features described above, the method may include performing, with the electronic processor, pre-cancelation of residual power in each of a virtual right speaker and a virtual left speaker included in the virtual listening space. Pre-cancelation may be performed after extracting the center channel and before rendering the output binaural audio.

[0011] In addition to any combination of features described above, the HRTF may include a combined HRTF. In addition to any combination of features described above, the method may include determining, with the electronic processor, a first HRTF based on the virtual listening space. The method may include determining, with the electronic processor, a normalized HRTF based on the first HRTF. The normalized HRTF may correspond to a zero-movement position of the head of the user. The method may include determining, with the electronic processor, an all-pass response HRTF based on the first HRTF. The all-pass response HRTF may maintain a phase of the first HRTF and unify a magnitude of the first HRTF to a residual power for the virtual left speaker and the virtual right speaker. The method may include combining, with the electronic processor, the normalized HRTF and the all-pass response HRTF to generate the combined HRTF.

[0012] In addition to any combination of features described above, the method may include determining, with the electronic processor, an inverse angular attenuation for each ear of the user based on the data. Determining the inverse angular attenuation for a left ear of the user may include determining a first coefficient for the virtual right speaker to the left ear. Determining the inverse angular attenuation for a right ear of the user may include determining a secondcoefficient for the virtual left speaker to the right ear. Determining the HRTF may include determining the HRTF based on the inverse angular attenuation for each ear of the user.

[0013] In addition to any combination of features described above, the first coefficient may be based on a first angle between the virtual right speaker and the left ear, and the second coefficient may be based on a second angle between the virtual left speaker and the right ear.

[0014] In addition to any combination of features described above, the output binaural audio may be unchanged with respect to the input binaural audio in response to the data indicating that the orientation of the head of the user corresponds to a zero-movement position. The output binaural audio may be different than the input binaural audio in response to the data indicating that the orientation of the head of the user corresponds to a non-zero-movement position.

[0015] In addition to any combination of features described above, a virtual sound scape of the input binaural audio may be preserved based on the orientation and / or the movement of the head of the user.

[0016] In addition to any combination of features described above, determining the HRTF may include selecting, with the electronic processor, the HRTF from among a plurality of HRTFs stored in a memory.

[0017] In addition to any combination of features described above, the method may include determining, with the electronic processor and based on the data from the sensor of the wearable device, that the orientation of the head of the user has changed from a first orientation to a second orientation. In response to determining that the orientation of the head of the user has changed from a first orientation to a second orientation, the method may include determining, with the electronic processor, a second virtual listening space based on the data. The second virtual listening space may be different than the virtual listening space. The method may include determining, with the electronic processor, a second HRTF based on the second virtual listening space. The second HRTF may be different than the HRTF. The method may include rendering, with the electronic processor and based on the second orientation of the head of the user, second output binaural audio to be output by the wearable device. The second output binaural audio may be rendered based on the input binaural audio and based on the second HRTF.

[0018] In addition to any combination of features described above, the wearable device may include one of a set of earbuds, headphones, smart glasses, and another head-mounted display.

[0019] In another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method.

[0020] In another aspect of the present disclosure, there is provided a wearable device, wherein the wearable device includes a computing apparatus, including: an electronic processor; and a memory storing instructions, which when executed by the electronic processor, cause the computing apparatus to perform the method.

[0021] Other aspects of the embodiments will become apparent by consideration of the detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] FIG. 1 illustrates a wearable device being worn on a head of a user, according to some example instances.

[0023] FIG. 2 is a hardware block diagram of the wearable device of FIG. 1, according to some example instances.

[0024] FIG. 3 illustrates a block diagram of a spatially preserved binaural headtracking (SPBH) device implemented by the electronic processor of the wearable device of FIGS. 1 and 2, according to some example instances.

[0025] FIG. 4 illustrates a more detailed block diagram of the SPBH device of FIG. 3, according to some example instances.

[0026] FIG. 5 A illustrates a virtual speaker configuration with a virtual left speaker, a virtual right speaker, and a center channel, according to some example instances.

[0027] FIG. 5B illustrates head movement of a user’s head and an angle between a left ear of the user and an inverse right speaker vector, according to some example instances.

[0028] FIG. 6 illustrates a flowchart of a method for controlling the wearable device to output audio, according to some example instances.DETAILED DESCRIPTION

[0029] The following description sets forth exemplary methods, parameters, and the like. It should be recognized, however, that such description is not intended as a limitation on the scope of the present disclosure but is instead provided as a description of example embodiments.

[0030] Although the following description uses terms “first,” “second,” etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first touch could be termed a second touch, and, similarly, a second touch could be termed a first touch, without departing from the scope of the various described embodiments. The first touch and the second touch are both touches, but they are not the same touch.

[0031] The terminology used in the description of the various described embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various described embodiments and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “includes,” “including,” “comprises,” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0032] The term “if’ is, optionally, construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” is, optionally, construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event],” depending on the context. The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific termsused herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0033] FIG. 1 illustrates a wearable device 105 being worn on a head 110 of a user 115 according to some example instances. The wearable device 105 is illustrated as a set of headphones. In other instances, the wearable device 105 may alternatively be embodied as a set of earbuds, smart glasses, other head-mounted displays such as virtual reality or augmented reality headsets, etc. As shown in FIG. 1, the wearable device 105 may be communicatively coupled to an external device 120 such as a smart phone, a tablet, a laptop computer, a desktop computer, or the like. As shown in FIG. 1, the wearable device 105 may be communicatively coupled to the external device 120 via a communicative connection 125 such as a wired communicative connection 125. In other instances, the wearable device 105 may be communicatively coupled to the external device 120 via a wireless communicative connection (e.g., via a short-range wireless connection such as Bluetooth™ and / or a longer-range wireless connection). In some instances, the wearable device 105 is communicatively coupled to additional or alternative devices such as other external devices located nearby the wearable device 105 and / or other devices located remotely from the wearable device 105 (e.g., a server, a cloud-based device, etc.).

[0034] FIG. 2 is a hardware block diagram of the wearable device 105 according to some example instances. In the example illustrated in FIG. 2, the wearable device 105 includes an electronic processor 205 (for example, a microprocessor or other electronic device). The electronic processor 205 includes input and output interfaces (not shown) and is electrically coupled to a memory 210, a network interface 215, speakers 220 (e.g., one or more audio output devices), and one or more orientation and / or movement sensors 225. In some instances, the wearable device 105 includes fewer or additional components in configurations different from that illustrated in FIG. 2 and / or optionally combines two or more components shown in FIG. 2. For example, the wearable device 105 may additionally include one or a combination of a microphone(s), a camera(s), and a display screen. As another example, the electronic processor 205 may include multiple electronic processors 205 within the wearable device 105 (e.g., a main electronic processor 205 of the wearable device 105 and an electronic processor included in an IMU) that together function to control various aspects of the wearable device 105. In some instances, the wearable device 105 is implemented within a distributed system including one or more components located in different devices. For example, in some embodiments, the wearable device 105 (e.g., the electronic processor 205) includes local hardware components andone or more external hardware components (e.g., one or more electronic processors of the external device 120). In other words, the electronic processor 205 may include any one or a combination of electronic processors located within a single device (e.g., the wearable device 105) or distributed among various devices and / or systems. For example, the electronic processor 205 may include a first electronic processor 205 of the wearable device 105, a second electronic processor of the external device 120 that is configured to communicate with the wearable device 105, or both the first electronic processor 205 and the second electronic processor. Thus, in the claims, if an apparatus or system is claimed, for example, as including an electronic processor or other element configured in a certain manner, for example, to make multiple determinations, the claim or claim element should be interpreted as meaning one or more electronic processors (or other element) where any one of the one or more electronic processors (or other element) is configured as claimed, for example, to make some or all of the multiple determinations. To reiterate, those electronic processors and processing may be distributed within a single device or across multiple devices.

[0035] In some instances, the wearable device 105 performs functionality other than the functionality described below. The various components of the wearable device 105 described herein are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0036] The memory 210 may include read only memory (ROM), random access memory (RAM), other non-transitory computer-readable media, or a combination thereof. The electronic processor 205 is configured to receive instructions and data from the memory 210 and execute, among other things, the instructions. In particular, the electronic processor 205 executes instructions stored in the memory 210 to perform the methods described herein.

[0037] The network interface 215 sends data to and receives data from other devices (e.g., the external device 120, cloud-based devices, server devices, etc.). In some instances, the network interface 215 includes one or more transceivers for wirelessly communicating with the other devices. Alternatively or in addition, the network interface 215 may include a connector or port for receiving a wired connection to one or more of the other devices, such as connector to receive a coaxial cable. The electronic processor 205 may receive data (for example, audio data and / or the like) from one or more other devices (e.g., the external device 120 and / or another device) via the network interface 215. The electronic processor 205 may output the data (e.g., audio / audio data) via the speakers 220 (and / or via a display screen in instances where the wearable device 105 includes the display screen). In some instances, the speakers 220 include aleft speaker and right speaker. In some instances, the speakers 220 may include additional and / or alternative speakers 220. In some instances, the wearable device 105 may also receive power from another device (e.g., the external device 120) via the network interface 215 to power components of the wearable device 105 (e.g., the electronic processor 205). In some instances, the wearable device 105 may include a separate power connection to receive power from another device and / or may include its own power supply such as a rechargeable or replaceable battery.

[0038] The orientation and / or movement sensor(s) 225 may include one or more sensors that are configured to collect and / or determine data indicative of an orientation (e.g., three-dimensional orientation including pitch, roll, yaw, and / or the like) and / or a movement of the head 110 of the user 115 of the wearable device 105 (i.e., head orientation / movement data). For example, the orientation and / or movement sensor(s) 225 may include an inertial measurement unit (IMU) that may include one or more accelerometers and / or gyroscopes. The head orientation / movement data may be an output of any one or a combination of sensors of the IMU (e.g., raw data from one or more sensors). Additionally or alternatively, the head orientation / movement data may be an output of the IMU itself. For example, the IMU may include an electronic processor configured to analyze raw data from one or more sensors and make one or more determinations to generate head orientation / movement data that is then provided to the electronic processor 205. In some instances, the head orientation / movement data includes absolute orientation data, velocity data, acceleration data, magnetic field data, and / or the like. In some instances, the head orientation / movement data includes linear and angular data, acceleration data from one or more accelerometers, angular velocity data from one or more gyroscopes, magnetometer data, and / or the like. In some instances, the electronic processor 205 is configured to receive the data indicative of the orientation and / or the movement of the head 110 of the user 115 of the wearable device 105 (i.e., the head orientation / movement data) from the orientation and / or movement sensor(s) 225.

[0039] In some instances, the external device 120 may include the same or similar components shown in FIG. 2 with respect to the wearable device 105. Such components of the external device 120 may be similar to the components described above with respect to the wearable device 105 and may perform similar general functions. In some instances, the external device 120 includes fewer or additional components in configurations different from that illustrated in FIG. 2 and / or optionally combines two or more components shown in FIG. 2. For example, the external device 120 may include a display screen such as a touchscreen configured to display a user interface and / or receive user inputs from a user. As another example, theexternal device 120 may include one or more microphones. In some instances, one or more components of the external device 120 (e.g., the speaker(s), the network interface, and / or the like) may be different than that of the wearable device 105.

[0040] FIG. 3 illustrates a block diagram 300 of a spatially preserved binaural headtracking (SPBH) device 305 implemented by the electronic processor 205 of the wearable device 105 according to some example instances. As shown in FIG. 3, the SPBH device 305 receives input binaural audio 310, for example, from the external device 120 or from another device. The SPBH device 305 also receives orientation and / or movement data 315 (i.e., head orientation / movement data 315) from the orientation and / or movement sensor(s) 225 (e.g., IMU data from an IMU) of the wearable device 105. As explained in greater detail herein, the SPBH device 305 is configured to generate output binaural audio 320 (e.g., head tracked output binaural audio 320) based on the input binaural audio 310 and the head orientation / movement data 315. In some instances, the SPBH device 305 is implemented by the electronic processor(s) 205 of the wearable device 105 such that any type of input binaural audio that is received by the wearable device 105 is rendered to head tracked output binaural audio that is output based on an orientation and / or movement of the user’s head 110. In some instances, the SPBH device 305 is additionally or alternatively implemented by the electronic processor(s) of the external device 120 in conjunction with the head orientation / movement data 315 from the wearable device 105 that is communicatively coupled to the external device 120. When the SPBH device 305 is described herein as performing a certain task or action, it should be understood that the electronic processor(s) that implement the SPBH device 305 is (are) the hardware that performs the certain task or action. While the tasks or actions performed by the SPBH device 305 may be described herein as being performed by the electronic processor 205 of the wearable device 105, other electronic processors (e.g., an electronic processor(s) of the external device 120) may additionally or alternatively perform the tasks or actions in some instances as explained previously herein.

[0041] FIG. 4 illustrates a more detailed block diagram of the SPBH device 305 according to some example instances. In some instances, the SPBH device 305 receives the input audio 310, x (e.g., input binaural audio 310, x) that is configured to be output by the wearable device 105. As indicated previously herein, the input audio includes input binaural audio. In some instances, binaural audio is audio produced through binaural recording (e.g., using two microphones positioned to approximate the position of ears on a human head, each microphone associated with generating respective ear-specific signals (i.e., a left ear signal and a right ear signal). Insome instances, binaural audio is audio produced by rendering non-binaural audio (e.g., object / channel -based audio) with filters that mimic how sound interacts with human anatomy (e.g., a head-related transfer function). In either case, binaural audio includes a pair of earspecific signals (i.e., a left ear signal and a right ear signal) that include the appropriate spatial cues for reproduction by transducers at or near a listener’s ears (e.g., a near-field speaker proximate to the left ear and a near-field speaker proximate to the right ear). The process of generating binaural audio from non-binaural audio (e.g., object / channel -based audio) is often referred to as binaural rendering, as discussed above, or binauralization. When binaural audio is reproduced by near-field speakers such as headphones, it generally provides a more detailed and realistic spatial representation of sound compared to corresponding stereo audio reproduced on the same near-field speakers. In some instances, binaural audio may provide a more immersive listening experience compared to, for example, stereo audio. In some instances, stereo audio is audio that may merely provide a sense of directionality by separating sound into left and right channels. In some instances, stereo audio is audio designed for far-field playback (e.g., loudspeaker playback) as opposed near-field playback (e.g., headphone playback), and does not provide as precise spatial perception as binaural audio. Overall, binaural audio may be considered a specific form of two-channel audio that is distinct from stereo audio both in terms of how the audio is produced and / or its content (e.g., presence or absence of spatial cues introduced via the process of binauralization or binaural recording).

[0042] In some instances, the input binaural audio 310 includes professional binaural audio that was rendered with an object-based file. In some instances, there may be different types of binaural audio such as professional binaural audio and stereo binaural audio. For example, professional binaural audio includes binaural audio that was rendered with an object-based file (e.g., Dolby Atmos™ audio). On the other hand, stereo binaural audio may be rendered by up-mixing stereo audio and is not professional binaural audio because it was not rendered with an object-based file.

[0043] As indicated in FIGS. 3 and 4, the SPBH device 305 may also receive data indicative of an orientation and / or a movement of a head 110 of a user 115 of the wearable device 105 (i.e., head orientation and / or movement data) from one or more sensors 225 of the wearable device 105.

[0044] In some instances, the electronic processor 205 implementing the SPBH device 305 extracts a center channel (at block 405) from the input binaural audio 310, x. The center extraction 405 may improve localization during headtracking. In some instances, the extractedcenter channel is a balanced part of a left channel and a right channel with respect to both power and phase of the input binaural audio 310. Equations 1-3 below may represent different channels of the input binaural audio 310, where xcrepresents the input binaural signal 310, c is the channel number where c = {0,1,2} that represents channel ID = {left, right, center}, b indicates the bth band of a signal, and k represents the extraction coefficient. In some instances, the SPBH device 305 operates in a complex quadrature mirror filter (CQMF) domain. For example, a CQMF analysis may be executed by the electronic processor 205 to convert the input audio 310 into the CQMF domain such that other processing actions shown in FIG. 4 may be completed in CQMF domain. For example, the input audio 310 is CQMF analyzed into CQMF bands b and blocks.Equation 1: x̃0(b) = k0(b)x0(b)Equation 2: %{(b) = k1(b')x1(b)Equation 3: %J(b) = (1 — fc0(b))x0(b) + (1 — fc1(b))x1(b)

[0045] The center part / channel may be linearly extracted from the left channel and the right channel during center extraction 405 as part of the spatially preserved rendering process performed by the SPBH device 305. In some instances, the signal x after center extraction 405 is represented by Equation 4 below where the signals are transposed.Equation 4: x̃ = [x̃0, x̃1, x̃2]T

[0046] In some instances, the electronic processor 205 implementing the SPBH device 305 performs pre-cancelation (at block 410) of residual power in each of a virtual right speaker and a virtual left speaker included in a virtual listening space. In some instances and as shown in FIG.4, the pre-cancelation 410 is performed after extracting the center channel (at block 410) and before rendering the output binaural audio 320 by a Tenderer 415. In other instances, the pre-cancelation 410 is performed at other times during audio processing. In some instances, the pre-cancelation 410 may not be performed (i.e., the pre-cancelation 410 is optional and may not be performed by the SPBH device 305 in some instances).

[0047] In some instances, the electronic processor 205 implements a HRTF processing device 420 as part of the SPBH device 305. The HRTF processing device 420 may performsome or all of its actions (e.g., any combination of any one or more of its actions) before, during, or after center extraction 405 and pre-cancelation 410 are performed. The HRTF processing device 420 may determine a virtual listening space based on the data indicative of the orientation and / or the movement of the head 110 of a user 115 of the wearable device 105 as explained in greater detail herein. Additionally, the HRTF processing device 420 may determine a HRTF (e.g., a combined HRTF) based on the virtual listening space as explained in greater detail herein.

[0048] The HRTF determined by the HRTF processing device 420 may be provided to the renderer 415. The electronic processor 205 implementing the renderer 415 may be configured to render, based on the orientation and / or the movement of the head 110 of the user 115, output binaural audio to be output by the wearable device 105. Accordingly, the output binaural audio 320 is rendered by the renderer 415 based on the input binaural audio 310 (e.g., x) and based on the HRTF (e.g., h) determined by the HRTF processing device 420. For example, the virtual listening space may include the center channel (e.g., x̃') that was extracted from the input binaural audio 310 (e.g., x). Continuing this example, the output binaural audio 320 may be rendered based on the center channel (e.g., x̃') from the input binaural audio 310.

[0049] Referring back to the HRTF processing device 420, the HRTF processing device 420 may perform an inverse angular attenuation 425 with respect to an amount of sound from one channel (e.g., virtual speaker) that is heard by an ear of the user 115 that is opposite / contralateral to the channel. For example, the inverse angular attenuation 425 for a virtual right speaker R (see FIG. 5A) indicates how much sound from the right virtual speaker R is heard by the left ear of the user 110 (i.e., the opposite / contralateral ear with respect to the virtual right speaker R). Similarly, the inverse angular attenuation 425 for a virtual left speaker L (see FIG. 5 A) indicates how much sound from the virtual left speaker L is heard by the right ear of the user 110 (i.e., the opposite / contralateral ear with respect to the virtual left speaker L). The electronic processor 205 performing the inverse angular attenuation 425 may determine an inverse angular attenuation for each ear of the user based on the head orientation and / or movement data 315. In some instances, determining the inverse angular attenuation for a left ear of the user 115 includes determining a first coefficient for a virtual right speaker to the left ear. In some instances, determining the inverse angular attenuation for a right ear of the user 115 includes determining a second coefficient for a virtual left speaker to the right ear. In some instances, the HRTF processing device 420 is configured to determine the HRTF based on the inverse angular attenuation 425 for each ear of the user 115. In some instances, the first coefficient is based on afirst angle between the virtual right speaker and the left ear. In some instances, the second coefficient is based on a second angle between the virtual left speaker and the right ear.

[0050] In some instances, to maintain the intention of creator of the input binaural audio 310, the output binaural audio 320 should not be changed unless the user’s head 110 changes positions. In other words, in some instances, the output binaural audio 320 is unchanged with respect to the input binaural audio 310 in response to the head orientation and / or movement data indicating that the orientation of the head 110 of the user 115 corresponds to a zero-movement position. Conversely, in some instances, the output binaural audio 320 is different than the input binaural audio 310 in response to the head orientation and / or movement data 315 indicating that the orientation of the head 110 of the user 115 corresponds to a non-zero-movement position. In some instances, the zero-movement position corresponds to rotation / movement of the user’s head 110 below a rotation / movement threshold value or position within a range associated with a stationary user head position (e.g., ±1 degree of rotation / movement, ±5 degrees of rotation / movement, ±10 degrees of rotation / movement, or the like). In some instances, the nonzero-movement position corresponds to rotation / movement of the user’s head 110 greater than or equal to a rotation / movement threshold value or position within a range associated with a stationary user head position (e.g., ±1 degree of rotation / movement, ±5 degrees of rotation / movement, ±10 degrees of rotation / movement, or the like). In some instances, a zeromovement position may correspond to an original orientation / position of the user’s head 110 (e.g., the user 115 looking straightforward with their sight line parallel to the ground). In some instances, any orientation / position besides the original orientation / position (e.g., outside of a rotation / movement threshold from the original orientation / position) may be considered to be a non-zero-movement position (e.g., the user 115 rotating their head and / or moving their body to look in another direction).

[0051] In some instances, a head rotation matrix Rh is used during the inverse angular attenuation calculation 425. The electronic processor 205 may calculate a coefficient for attenuation on the contralateral ear of each channel. Specifically, during the inverse angular attenuation calculation, the electronic processor 205 calculates a coefficient for a side virtual speaker (non-center) to a contralateral ear (e.g., left speaker to right ear and right speaker to left ear). In some instances, the calculations for the two pairs are symmetrical.

[0052] FIG. 5A illustrates a virtual speaker configuration 505 with a virtual left speaker L, a virtual right speaker R, and a center channel C. FIG. 5B illustrates a use case 510 includinghead movement of the user’s head 110 and an angle between a left ear of the user 115 and an inverse right speaker vector.

[0053] In some instances, an original orientation uo of the head 110 of the user 115 may be represented by Equation 5 below. Equation 5 below represents a user’s head 110 facing the x-axis (see FIG. 5A).Equation 5: u0= [1,0, 0]T

[0054] In some instances, only angles are considered so the ears and the virtual speakers may be placed on an X-Y plane, and distance may be considered a unit distance. In some instances, the right speaker vector sris represented by Equation 6 below, and the inverse right speaker vectoris represented by Equation 7 below, where 9s represents a speaker angle that may change based on a vertical head tilt angle of the user 115. These vectors are shown in FIG.5 A.Equation 6: sr= [cos0s,sin0s, 0]TEquation 7:= [— cos6s, —sin0s, 0]T

[0055] In some instances, the original left ear vector eio is represented by Equation 8 below, where 6eis the ear angle of the left ear 515.Equation 8: el0= [cos0e, — sin0e, 0]T

[0056] In some instances, the left ear vector ei after head rotation is represented by Equation 9 below, where Rh is the head rotation matrix mentioned previously herein. The left ear vector ei is shown in FIG. 5B with respect to the left ear 515.Equation 9: e;=

[0057] In some instances, the angle 9abetween the inverse right speaker and the left ear 515 is represented by Equation 10 below. The angle 9a is also shown in FIG. 5B.Equation 10: 0a= acos (ez■

[0058] In some instances, the electronic processor 205 is configured to calculate a non-linear coefficient such that when the inverse speaker vector and the contralateral ear angle are closeenough to each other (e.g., within a threshold value), there is full attenuation (e.g., to prevent sound being output from a sound source in the virtual listening space that is opposite to the contralateral ear from being heard by the contralateral ear). In some instances, attenuation is applied to a HRTF cross-talk path that is from a virtual object to a contralateral ear to reduce influence of an opposite direction sound source during head-tracking (e.g., for the zero state when there is no head rotation). On the other hand, when the inverse speaker vector and the contralateral ear angle are greater than a second threshold value apart from each other, the attenuation goes to zero (e.g., there is no attenuation such that sound being output from the sound source in the virtual listening space that is facing the contralateral ear, for example due to head rotation, is fully provided to the contralateral ear). In some instances, when the head rotation is far from the original / zero state, a cross-talk path may be applied to construct a consistent listening space. In some instances, the right-speaker-to-left-ear coefficient go is represented by Equation 11 below, where 9T is an included angle threshold and A is a constant attenuation factor. In some instances, 0T is the angular threshold for an included angle when the cross-talk path is leveraged to build the sound scene.Equation 11: g0=0, else

[0059] Similarly, in some instances, the left-speaker-to-right-ear coefficient gi is represented by Equation 12 below, where 6b= acos (er■ sj).Equation 12: g1=I 0, elser00r01r02

[0060] In one example where A = 1, 0e= 0T= TT / 2, and Rh= r10r11r12, then e, =r20r21r22— [r01, r11,r21]T, the coefficients go and gi may be respectively represented by simplified equations 13 and 14 below. In some instances, A is an attenuation factor to control the overall attenuation gain, 0eis the ear angle, 0Tis the angular threshold for an included angle, Rh is the rotation matrix of the user’s head 110, and e / is the rotated left ear vector.Equation 13: ga=+'<l®”1<l®’1( 0, else|rnsin0s- roicos0s|, if |0a| < |0T|Equation 14:0, else

[0061] In some instances, the HRTF processing device 420 determines a first HRTF based on the virtual listening space. In some instances, a scene constructor of the HRTF processing device 420 generates the virtual listening space based on the head orientation / movement data from the sensor(s) 225 of the wearable device 105. For example, the virtual listening space is controllable through head rotation, object distance, and other parameters. The scene constructor may output parameters for a HRTF generator. The HRTF generator may include and / or utilize a HRTF library 430 that includes sets of HRTFs for different angles and distances. Based on the head orientation / movement data, the electronic processor 205 implementing the HRTF processing device 420 may allocate a virtual space to each channel. Additionally, the electronic processor 205 is configured to read a HRTF from the HRTF library as h. An initial state for HRTF when no head movement may be referred to as h0.

[0062] As explained above, in some instances, the electronic processor 205 implementing the HRTF processing device 420 determines a virtual listening space based on the head orientation / movement data (e.g., using a scene constructor). The electronic processor 205 may also determine the first HRTF based on the virtual listening space. In some instances, the HRTF may be determined (e.g., selected and / or calculated) in any one or a combination of different manners. For example, the electronic processor 205 may select a HRTF (e.g., from a library of stored HRTFs that may be included in the memory 210 or accessible via communication with another device) based on data corresponding to the user 115 (e.g., demographic data including but not limited to age, gender assigned at birth, geographic data, etc.; physiological data including but not limited to weight, height, head size, ear shape / pitch / di stance, etc.; and / or the like). In some instances, the electronic processor 205 selects the first HRTF and modifies the first HRTF based on the data corresponding to the user 115. In some instances, the electronic processor 205 generates the HRTF using a generative model and the data corresponding to the user 115. In some instances, at least some portions of the HRTF selection and / or generation process are performed by the external device 120 and / or other devices (e.g., a server, a cloud-based computing device, etc.). In such instances, the first HRTF and / or the determinations related to the first HRTF that were made by the other devices may be transmitted to the electronic processor 205 of the wearable device 105.

[0063] As explained above, in some instances, the electronic processor 205 implementing the HRTF processing device 420 determines the first HRTF based on the headorientation / movement data of the user 115. In some instances, the orientation and / or the movement of the head 110 of the user 115 includes a predicted orientation and / or movement that represents an estimate of an orientation and / or movement of the user’s head at a time in the future. For example, by analyzing audio that will be output by the wearable device 105, the electronic processor 205 may determine that it is likely that the user’s head 110 will turn in a certain direction at a certain time (e.g., in response to a loud sound coming from a certain direction). The electronic processor 205 may estimate the orientation and / or movement of the user’s head at a time in the future (e.g., when the loud sound audio is output) based on the analysis of the audio indicating when the loud sound will be output and where in the virtual listening space the loud sound will be output. This example is merely one example of a predicted orientation and / or movement. The electronic processor 205 may generate other predicted orientations and / or movements in other instances based on audio analysis and / or head orientation / movement data of the user 115.

[0064] In some instances, the electronic processor 205 implementing the HRTF processing device 420 determines a normalized HRTF h based on the first HRTF h, where the normalized HRTF corresponds to the zero-movement position of the head 110 of the user 115. In other words, a normalized HRTF h is the relative response to zero movement. For band b, the normalized HRTF h may be represented by Equation 15 below.Equation 15: h(b) =

[0065] In some instances, normalizing the first HRTF may prevent / avoid the doubleprocessing issue described previously herein. Specifically, because the input binaural audio 310 is binaural audio to which a HRTF has been applied before receipt by the wearable device 105, a difference between the original HRTF (i.e., the first HRTF) and a rotated HRTF is applied during the normalization procedure to result in proper headtracking and sound output of the input binaural audio 310 on the wearable device 105.

[0066] In addition to determining a normalized HRTF h(b), in some instances, the HRTF processing device 420 also determines an all-pass response HRTF h (at block 435) based on the first HRTF h. In some instances, the all-pass response HRTF h maintains a phase of the first HRTF h and unifies a magnitude of the first HRTF h to a residual power for the virtual left speaker and the virtual right speaker. In some instances, a speaker-to-ear response may be defined as a matrix H. In some instances, the rendered output binaural audio after speaker-to-earrendering may be represented by Equation 16 below, where x is the separated input signal and x is the assumed rendered singal according to the HRTF matrix H.Equation 16: x = Hx

[0067] In some instances, an ideal matrix for zero movement is represented by Equation 17 below, where x = x.Equation 17: Ho= [J *]

[0068] However, the zero response causes significant fluctuation during head rotation around zero degrees and may be incompatible with a HRTF system for the whole listening space. For example, the zero response may create a valley in both amplitude response and phase response of the HRTFs during head rotation. Accordingly, it is a goal of the present disclosure to create an all-pass response where power is low such that it is less perceivable to the user 115. Additionally, residual power Prfrom a contralateral ear could be eliminated during precancelation 410 as explained in greater detail herein. In some instances, Equation 17 below represents an all-pass response HRTF h for all bands.Equation 18: / i(b) =

[0069] According to Equation 18, the phase of the first HRTF h is maintained, and the magnitude of the first HRTF h is unified to a residual power Prfor each band b.

[0070] The influence of residual power Pris quite small. Nevertheless, in some instances, pre-cancelation 410 can be used to further reduce or eliminate the residual power Pras indicated by the following equations in which Ho is a matrix of speaker-to-ear response at a zero movement position of the user’s head 110, the all-pass response from the left speaker to the right ear is denoted as h.Lr(b) and from the right speaker to the left ear is denoted as / iRZ(b). In Equations 20 and 21, x0(b) is the rendered left channel signal, x1(b) is the rendered right channel signal,is the separated left signal, %i(6) is the separated right signal, andis the separated center signal.^RZ(^)Equation 19: Ho= 1, W) 1 1Equation 20: x0(b) = x0(b) + hRiWx^b) + x2(b)Equation 21: x1(b) = hLr(b)x0(b) + %1(b) + x2(b)

[0071] In some instances, to make x0(b) = x0(b) and xx(b) = x^b) (which makes the output audio signal 320, x0(b), xx(b) when there is no head rotation the same as the input audio signal 310, x0(b), x1(b)), let the canceled signal be x̃0'(b) and x̃1'(b), and it is desirable to make Equations 22 and 23 true.Equation 22: x^' (b) + hRZ(b)xf'(b) = x^(b)Equation 23:hLr(b)x^' (b) + %7'(b) = %7(b)

[0072] In some instances, after deduction to calculate two pre-cancellated signals that satisfy Equation 22 and 23, a pre-cancelation signal may be determined using Equations 24 and 25 below.Equation 24: xHW =Equation 25: xl'(b) =1 - hLr(b)hRl(b)

[0073] In some instances, the HRTF processing device 420 includes a combiner 440 implemented by the electronic processor 205. The electronic processor 205 implementing the combiner 440 is configured to combine the normalized HRTF h and the all-pass response HRTF h to generate a combined HRTF h. As indicated in FIG. 4, in some instances, the combiner 440 takes the inverse angular attenuation g into account when determining the combined HRTF h. In other words, determining the combined HRTF h includes determining the combined HRTF h. based on the inverse angular attenuation g for each ear of the user. In some instances, the combiner 440 generates the combined HRTF h. according to the following equations, where ĥLL(b) is a final combined HRTF for a transmission path from left speaker to left ear, h̃LL(b) is a normalized HRTF for a transmission path from left speaker to left ear, ĥLr(b) is an all-pass response HRTF for a transmission path from left speaker to right ear, ĥRL(b) is an all-pass response HRTF for a transmission path from right speaker to left ear, h̃Rr(b) is a final combinedHRTF for a transmission path from left speaker to right ear, hLr(b is a normalized HRTF for a transmission path from left speaker to right ear, hRr(b) is a final combined HRTF for a transmission path from right speaker to right ear, hRr(b) is a normalized HRTF for a transmission path from right speaker to right ear, hRZ(b) is a final combined HRTF for a transmission path from right speaker to left ear, hRZ(b) is a normalized HRTF for a transmission path from right speaker to left ear, hcz(b) is a final combined HRTF for a transmission path from center speaker to left ear, hcz(b) is a normalized HRTF for a transmission path from center speaker to left ear, iCr(b) is a final combined HRTF for a transmission path from center speaker to right ear, and hCr(b>) is a normalized HRTF for a transmission path from center speaker to right ear.Equation 26: hLZ(b) = hLZ(b)Equation 27: iRr(b) = (1 - ^i) iLr(b) + gQh^r(b~)Equation 28: hRr(b) = hRr(b')Equation 29: / iRZ(b) = (1 - g0)hRl(b) +Equation 30: / icz(b) = / icz(b)Equation 31: hCr(b) = / iCr(b)

[0074] In some instances, Equation 26 means a final left-speaker-to-left-ear HRTF is a normalized HRTF. In some instances, Equation 27 means a final left-speaker-to-right-ear HRTF is a combination of the normalized HRTF and the all-pass response HRTF according to attenuation gains. In some instances, Equation 30 means a center-speaker-to-left-ear HRTF is also a normalized HRTF.

[0075] In some instances, the renderer 415 is configured to use the combined HRTF h. to render the output binaural audio 320. For example, the renderer 415 may be configured to render the output binaural audio 320 based on the input binaural audio 310, %' (e.g., after center extraction 405 and pre-cancelation 410) and based on the combined HRTF h. By rendering the output binaural audio 320 based on the combined HRTF h, the output binaural audio 320 isrendered based on the orientation and / or the movement of the head 110 of the user 115.Equation 32 below may represent the output binaural audio 320, y in some instances, where y(b) is the output binaural audio 320, y for each band b. The combined HRTF h. for each band of each channel may be multiplied by a respective input binaural audio signal for each band to result in the output binaural audio 320, y.hLi(b) hRi(b)E y(b) = hCi(b)quation 32:hLrW hRr(b) hc'rW.*2(6).

[0076] In some instances, rendering the output binaural audio 320 based on the input binaural audio 310, x' (e.g., after center extraction 405 and pre-cancelation 410) and based on the combined HRTF h. results in a virtual sound scape of the input binaural audio 310 being preserved based on the orientation and / or the movement of the head 110 of the user 115. For example, rendering of the output binaural audio 320 may be adjusted in response to the orientation of the user’s head 110 changing via movement of the user’s head 110 and / or body.

[0077] In some instances, the electronic processor 205 implementing the SPBH device 305 may be configured to determine that the orientation of the head 110 of the user 115 has changed from a first orientation (for which a first virtual listening space and first HRTF are determined as explained herein) to a second orientation based on the data from the sensor(s) 225 of the wearable device 105. In response to determining that the orientation of the head 110 of the user 115 has changed from a first orientation to a second orientation, the electronic processor 205 may be configured to determine a second virtual listening space based on the head orientation and / or movement data 315 data. The second virtual listening space may be different than the first virtual listening space due to the movement / change in orientation of the user’s head 110. Also in response to determining that the orientation of the head 110 of the user 115 has changed from a first orientation to a second orientation, the electronic processor 205 may be configured to determine a second HRTF based on the second virtual listening space. The second HRTF may be different than the first HRTF based on the difference(s) between the listening spaces that is caused by the movement / change in orientation of the user’s head 110. The electronic processor 205 may also be configured to render, based on the second orientation of the head of the user, second output binaural audio 320 to be output by the wearable device 105. The second output binaural audio 320 may be rendered based on the input binaural audio 310 and based on the second HRTF.

[0078] FIG. 6 illustrates a flowchart of a method 600 for controlling the wearable device 105 to output audio according to some example instances. The method 600 is described as being performed by the electronic processor 205 implementing the SPBH device 305, but other electronic processors may be involved in at least portions of the method 600 as described previously herein. The method 600 includes actions performed by the SPBH device 305 as previously explained herein. While a particular order of processing steps, message receptions, and / or message transmissions is indicated in FIG. 6 as an example, timing and ordering of such steps, receptions, and transmissions may vary where appropriate without negating the purpose and advantages of the examples set forth in detail throughout the remainder of this disclosure.

[0079] At block 605, the electronic processor 205 receives the input audio 310 that is configured to be output by the wearable device 105. In some instances, the input audio 310 includes input binaural audio. In some instances, the input audio 310 may be received from the external device 120 or from another device such as a server, a cloud-based device, etc.

[0080] At block 610, the electronic processor 205 receives, from a sensor(s) 225 of the wearable device 105, data indicative of an orientation and / or a movement of a head of a user of the wearable device 105 (i.e., head orientation / movement data). In some instances, the electronic processor 205 implements the HRTF processing device 420 explained previously herein to receive the head orientation / movement data.

[0081] At block 615, the electronic processor 205 determines a virtual listening space based on the head orientation / movement data as explained previously herein.

[0082] At block 620, the electronic processor 205 determines a head-related transfer function (HRTF) (i.e., a first HRTF) based on the virtual listening space. In some instances, the electronic processor 205 loads / selects and / or generates the HRTF according to head movement of the user 115 (as indicated by the head orientation / movement data) and a virtual speaker location (e.g., of a virtual right speaker and a virtual left speaker) within / defined by the virtual listening space.

[0083] In some instances, the electronic processor 205 generates an all-pass HRTF for side-to-side transmission paths of output audio 320 (e.g., a left-speaker-to-right-ear path and a right-speaker-to-left-ear path). In some instances, the electronic processor 205 generates a normalized HRTF. In some instances, the generation of the all-pass HRTF and / or the normalized HRTF is based on the first HRTF. In some instances, the electronic processor 205 determines attenuation factors for the side-to-side transmission paths of the output audio 320 (e.g., to attenuate audiowhen the user’s head 110 is turned away from a sound source such as one of the virtual left speaker and the virtual right speaker). In some instances, the electronic processor 205 combines the normalized HRTF and the all-pass HRTF to generate a combined HRTF to be used to render the output audio 320.

[0084] At block 625, the electronic processor 205 renders, based on the orientation and / or the movement of the head 110 of the user 115, the output binaural audio 320 to be output by the wearable device 105. The output binaural audio 320 may be rendered based on the input binaural audio 310 and based on the HRTF (e.g., the combined HRTF).

[0085] As indicated in FIG. 6, in some instances, the method 600 may repeat as additional input audio 310 is received to be output by the wearable device 105.

[0086] The foregoing description, for purpose of explanation, has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The embodiments were chosen and described in order to best explain the principles of the techniques and their practical applications. Others skilled in the art are thereby enabled to best utilize the techniques and various embodiments with various modifications as are suited to the particular use contemplated.

[0087] Although the disclosure and examples have been fully described with reference to the accompanying drawings, it is to be noted that various changes and modifications will become apparent to those skilled in the art. Such changes and modifications are to be understood as being included within the scope of the disclosure and examples as defined by the claims.

[0088] Various features and advantages are set forth in the following claims.

Claims

CLAIMSWhat is claimed is:

1. A method for controlling a wearable device to output audio, the method comprising: receiving, with an electronic processor, input binaural audio configured to be output by the wearable device;receiving, with the electronic processor and from a sensor of the wearable device, data indicative of an orientation and / or a movement of a head of a user of the wearable device;determining, with the electronic processor, a virtual listening space based on the data; determining, with the electronic processor, a head-related transfer function (HRTF) based on the virtual listening space; andrendering, with the electronic processor and based on the orientation and / or the movement of the head of the user, output binaural audio to be output by the wearable device, wherein the output binaural audio is rendered based on the input binaural audio and based on the HRTF.

2. The method of claim 1, further comprising extracting, with the electronic processor, a center channel from the input binaural audio, wherein the virtual listening space includes the center channel;wherein the output binaural audio is rendered based on the center channel from the input binaural audio.

3. The method of claim 2, further comprising performing, with the electronic processor, pre-cancelation of residual power in each of a virtual right speaker and a virtual left speaker included in the virtual listening space, wherein pre-cancelation is performed after extracting the center channel and before rendering the output binaural audio.

4. The method of claim 3, wherein the HRTF includes a combined HRTF, and further comprising:determining, with the electronic processor, a first HRTF based on the virtual listening space;determining, with the electronic processor, a normalized HRTF based on the first HRTF, wherein the normalized HRTF corresponds to a zero-movement position of the head of the user;determining, with the electronic processor, an all-pass response HRTF based on the first HRTF, wherein the all-pass response HRTF maintains a phase of the first HRTF and unifies amagnitude of the first HRTF to a residual power for the virtual left speaker and the virtual right speaker; andcombining, with the electronic processor, the normalized HRTF and the all-pass response HRTF to generate the combined HRTF.

5. The method of claim 4, further comprising:determining, with the electronic processor, an inverse angular attenuation for each ear of the user based on the data, wherein determining the inverse angular attenuation for a left ear of the user includes determining a first coefficient for the virtual right speaker to the left ear, and wherein determining the inverse angular attenuation for a right ear of the user includes determining a second coefficient for the virtual left speaker to the right ear;wherein determining the HRTF includes determining the HRTF based on the inverse angular attenuation for each ear of the user.

6. The method of claim 5, wherein the first coefficient is based on a first angle between the virtual right speaker and the left ear, and wherein the second coefficient is based on a second angle between the virtual left speaker and the right ear.

7. The method of any one of the preceding claims, wherein the output binaural audio is unchanged with respect to the input binaural audio in response to the data indicating that the orientation of the head of the user corresponds to a zero-movement position; andwherein the output binaural audio is different than the input binaural audio in response to the data indicating that the orientation of the head of the user corresponds to a non-zeromovement position.

8. The method of any one of the preceding claims, wherein a virtual sound scape of the input binaural audio is preserved based on the orientation and / or the movement of the head of the user.

9. The method of any one of the preceding claims, wherein determining the HRTF includes selecting, with the electronic processor, the HRTF from among a plurality of HRTFs stored in a memory.

10. The method of any one of the preceding claims, further comprising:determining, with the electronic processor and based on the data from the sensor of the wearable device, that the orientation of the head of the user has changed from a first orientation to a second orientation; andin response to determining that the orientation of the head of the user has changed from a first orientation to a second orientation:determining, with the electronic processor, a second virtual listening space based on the data, wherein the second virtual listening space is different than the virtual listening space,determining, with the electronic processor, a second HRTF based on the second virtual listening space, wherein the second HRTF is different than the HRTF, and rendering, with the electronic processor and based on the second orientation of the head of the user, second output binaural audio to be output by the wearable device, wherein the second output binaural audio is rendered based on the input binaural audio and based on the second HRTF.

11. The method of any one of the preceding claims, wherein the wearable device includes one of a set of earbuds, headphones, smart glasses, and another head-mounted display.

12. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any one of claims 1-11.

13. The wearable device, wherein the wearable device includes a computing apparatus, comprising:an electronic processor; anda memory storing instructions, which when executed by the electronic processor, cause the computing apparatus to perform the method of any one of claims 1-11.