Outputting binaural audio with a wearable device

WO2026177929A1PCT designated stage Publication Date: 2026-08-27DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/014915
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-07-10
Filing Date
2026-02-11
Publication Date
2026-08-27

Smart Images

  • Figure US2026014915_27082026_PF_FP_ABST
    Figure US2026014915_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A method for controlling a wearable device to output audio includes analyzing two-channel input audio to determine a first probability that indicates whether the input audio includes binaural audio. The method includes performing binaural separation of the input audio to determine an audio scene center of the binaural audio, and performing stereo separation of the input audio to extract three-channel audio from the input audio. The method includes combining the audio scene center and the three-channel audio to generate combined audio such that (i) the audio scene center is weighted based on the binaural steering signal corresponding to the first probability and (ii) the three-channel audio is weighted based on a stereo steering signal corresponding to an opposite probability of the first probability. The method includes rendering, based on an orientation and / or a movement of a head of a user, output binaural audio to be output by the wearable device.
Need to check novelty before this filing date? Find Prior Art

Description

D24130W001OUTPUTTING BINAURAL AUDIO WITH A WEARABLE DEVICE CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims the benefit of priority from U.S. Provisional Patent Application No. 63,841,888, filed on July 10, 2025 and International Patent Application No. PCT / CN2025 / 078134, filed on February 19, 2025, each of which is hereby incorporated by reference in their entireties.TECHNICAL FIELD

[0002] The present disclosure relates generally to outputting binaural audio with a wearable device that is worn on a head of a user.SUMMARY

[0003] Wearable devices (e.g., earbuds, headphones, smart glasses, other head-mounted displays, etc.) are often equipped with one or more sensors and one or more electronic processors. For example, a wearable device may include an inertial measurement unit (IMU) configured to provide real-time data from a gyroscope(s) and / or accelerometer(s) to indicate a status / orientation of the wearable device. This real-time data may be processed / analyzed to determine the orientation and / or physical movement of a head of a user wearing the wearable device. Using such orientation and movement information of the user, spatial computational audio can be output by the wearable device, for example, by executing one or more algorithms in the electronic processor(s) of the wearable device.

[0004] Executing such an algorithm(s) improves the user’s experience by providing a virtual listening space that is relatively static to the real-world environment surrounding the user rather than being static to the wearable device itself. In other words, when the user uses a wearable device equipped with such an algorithm(s), the sound will be perceived by the user as if the sound is generated from the real-world environment, which enhances the listening experience provided to the user by making the listening experience more immersive.

[0005] The algorithm(s) that may be used to provide the immersive listening experience described above may be referred to as a virtualizer (e.g., a lightweight intelligent virtualizer, also known as a LIV) that receives two-channel audio as an input and generates binaural audio to beD24130W001 output by the wearable device. In some instances, the computational complexity of the virtualizer is low enough to be executed by an electronic processor(s) of a wearable device.

[0006] Because wearable devices are the devices that actually output audio to be consumed / heard by a user (e.g., with speakers), wearable devices are usually located at the end of an audio processing chain. Accordingly, there is a possibility that input audio received by the wearable device may already be binaural audio. If an additional binaural rendering process is applied by a wearable device to input binaural audio received by the wearable device, a technological problem of a double-processing issue arises. This double-processing issue distorts perceptual object localization in the input binaural audio and damages the intention of a creator of the input binaural audio (e.g., professional binaural audio content creator object-based intention). This distortion and damaged intention of the creator results in a lower quality listening experience for the user.

[0007] To address the above-noted technological problem of the double-processing issue, the systems, methods, and devices described herein (e.g., the virtualizer implemented by the electronic processor of the wearable device) perform intelligent processing that handle stereo input audio, binaural input audio, and combinations thereof. The disclosed systems, methods, and devices allow for binaural audio to be output by a wearable device in accordance with an orientation and / or movement of a user’s head as measured by a sensor(s) of the wearable device regardless of whether the input audio received by the wearable device is stereo audio, binaural audio, or a combination of stereo audio and binaural audio. Thus, the disclosed systems, methods and devices allow the wearable device to provide an enhanced / immersive listening experience for any two-channel input audio that is received by the wearable device (e.g., an immersive listening experience without distortion and without damaging the intention of the creator of binaural audio content that is included in the input audio).

[0008] In one aspect of the present disclosure, there is provided a method for controlling a wearable device to output audio. The method may include receiving, with an electronic processor, input audio configured to be output by the wearable device. The input audio may include two-channel input audio. The method may include analyzing, with the electronic processor, the input audio to determine a first probability that indicates a likelihood whether the input audio includes binaural audio. The method may include generating, with the electronic processor, a binaural steering signal based at least in part on the first probability. The method may include performing, with the electronic processor, binaural separation of the input audio to determine an audio scene center of the binaural audio. The method may include performing,D24130W001 with the electronic processor, stereo separation of the input audio to extract three-channel audio from the input audio. The method may include combining, with the electronic processor, the audio scene center and the three-channel audio to generate combined audio. The audio scene center and the three-channel audio may be combined such that (i) the audio scene center is weighted based on the binaural steering signal corresponding to the first probability and (ii) the three-channel audio is weighted based on a stereo steering signal corresponding to an opposite probability of the first probability. The method may include receiving, with the electronic processor and from a sensor of the wearable device, data indicative of an orientation and / or a movement of a head of a user of the wearable device. The method may include rendering, with the electronic processor and based on the orientation and / or the movement of the head of the user, output binaural audio to be output by the wearable device. The rendering may be based on the combined audio and the data indicative of the orientation and / or the movement of the head of the user of the wearable device.

[0009] In addition to any combination of features described above, the method may include determining, with the electronic processor, a virtual listening space based on the data. The method may include determining, with the electronic processor, a first head-related transfer function (HRTF) based on the virtual listening space. The method may include determining, with the electronic processor, a normalized HRTF based on a zero-rotation state of the virtual listening space. The method may include mixing, with the electronic processor, the first HRTF and the normalized HRTF to generate a mixed HRTF. The first HRTF and the normalized HRTF may be mixed such that (i) the normalized HRTF is weighted based on the binaural steering signal corresponding to the first probability and (ii) the first HRTF is weighted based on the stereo steering signal corresponding to the opposite probability of the first probability.Rendering the output binaural audio may include rendering the output binaural audio based on the mixed HRTF and a reverb generated with a feedback delay network.

[0010] In addition to any combination of features described above, the method may include analyzing, with the electronic processor, the input audio to determine a second probability that indicates a likelihood whether the input audio includes professional binaural audio. The professional binaural audio may have been rendered with an object-based file. Rendering the output binaural audio may include rendering the output binaural audio based on the second probability and a direct-to-late rate that is configured to control a perceived distance of objects in the virtual listening space.D24130W001

[0011] In addition to any combination of features described above, analyzing the input audio to determine the second probability may include receiving, with the electronic processor, metadata associated with the input audio. The metadata may indicate that the input audio includes professional binaural audio. Analyzing the input audio to determine the second probability may also include determining, with the electronic processor, that the second probability is 100% based on the metadata.

[0012] In addition to any combination of features described above, one or both of combining the audio scene center and the three-channel audio to generate combined audio and rendering the output binaural audio may include utilizing the metadata to at least partially maintain content creator object-based intention indicated by the metadata.

[0013] In addition to any combination of features described above, the method may include converting, with the electronic processor, the input audio into a complex quadrature mirror filter (CQMF) domain. Performing the binaural separation, performing the stereo separation, combining the audio scene center and the three-channel audio to generate the combined audio, and rendering the output binaural audio may be completed in the CQMF domain. The method may also include synthesizing, with the electronic processor, the output binaural audio from the CQMF domain to a time domain for output by the wearable device.

[0014] In addition to any combination of features described above, analyzing the input audio to determine the first probability may include analyzing the input audio using a machine learning model to classify audio content into two or more types corresponding to respective classifiers. The machine learning model may be trained from audio features including inter-channel phase difference (ICPD) and inter-channel level difference (ICLD).

[0015] In addition to any combination of features described above, the respective classifiers for the input audio may be determined using a parallel structure or a cascade structure.

[0016] In addition to any combination of features described above, analyzing the input audio to determine the first probability may include receiving, with the electronic processor, metadata associated with the input audio. The metadata may indicate that the input audio includes binaural audio. Analyzing the input audio to determine the first probability may also include determining, with the electronic processor, that the first probability is 100% based on the metadata.D24130W001

[0017] In addition to any combination of features described above, the wearable device may include one of a set of earbuds, headphones, smart glasses, and another head-mounted display.

[0018] In another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method.

[0019] In another aspect of the present disclosure, there is provided a wearable device, wherein the wearable device includes a computing apparatus, including: an electronic processor; and a memory storing instructions, which when executed by the electronic processor, cause the computing apparatus to perform the method.

[0020] Other aspects of the embodiments will become apparent by consideration of the detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0021] FIG. 1 illustrates a wearable device being worn on a head of a user, according to some example instances.

[0022] FIG. 2 is a hardware block diagram of the wearable device of FIG. 1, according to some example instances.

[0023] FIG. 3 illustrates a block diagram of a virtualizer implemented by an electronic processor of the wearable device of FIGS. 1 and 2, according to some example instances.

[0024] FIG. 4 illustrates a more detailed block diagram of the virtualizer of FIG. 3, according to some example instances.

[0025] FIGS. 5 A and 5B respectively illustrate a parallel structure and a cascade structure for a set of classifiers of a binaural detector of the virtualizer of FIG. 3, according to some example instances.

[0026] FIG. 6 illustrates a high-level schematic diagram of a steerer of the virtualizer of FIG.3, according to some example instances.

[0027] FIG. 7 illustrates a block diagram of a separator of the virtualizer of FIG. 3, according to some example instances.D24130W001

[0028] FIG. 8 illustrates an assumption of encoding a channel-based signal that may be used by a stereo separator of the separator of the virtualizer of FIG. 3, according to some example instances.

[0029] FIG. 9 illustrates a block diagram of a rendering device of the virtualizer of FIG. 3, according to some example instances.

[0030] FIG. 10 illustrates an example sound scene of a three-object system, according to some example instances.

[0031] FIG. 11 illustrates a flowchart of a method for controlling the wearable device to output audio, according to some example instances.DETAILED DESCRIPTION

[0032] The following description sets forth exemplary methods, parameters, and the like. It should be recognized, however, that such description is not intended as a limitation on the scope of the present disclosure but is instead provided as a description of example embodiments.

[0033] Although the following description uses terms “first,” “second,” etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first touch could be termed a second touch, and, similarly, a second touch could be termed a first touch, without departing from the scope of the various described embodiments. The first touch and the second touch are both touches, but they are not the same touch.

[0034] The terminology used in the description of the various described embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various described embodiments and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “includes,” “including,” “comprises,” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.D24130W001

[0035] The term “if’ is, optionally, construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” is, optionally, construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event],” depending on the context. The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0036] FIG. 1 illustrates a wearable device 105 being worn on a head 110 of a user 115 according to some example instances. The wearable device 105 is illustrated as a set of headphones. In other instances, the wearable device 105 may alternatively be embodied as a set of earbuds, smart glasses, other head-mounted displays such as virtual reality or augmented reality headsets, etc. As shown in FIG. 1, the wearable device 105 may be communicatively coupled to an external device 120 such as a smart phone, a tablet, a laptop computer, a desktop computer, or the like. As shown in FIG. 1, the wearable device 105 may be communicatively coupled to the external device 120 via a communicative connection 125 such as a wired communicative connection 125. In other instances, the wearable device 105 may be communicatively coupled to the external device 120 via a wireless communicative connection (e.g., via a short-range wireless connection such as Bluetooth™ and / or a longer-range wireless connection). In some instances, the wearable device 105 is communicatively coupled to additional or alternative devices such as other external devices located nearby the wearable device 105 and / or other devices located remotely from the wearable device 105 (e.g., a server, a cloud-based device, etc.).

[0037] FIG. 2 is a hardware block diagram of the wearable device 105 according to some example instances. In the example illustrated in FIG. 2, the wearable device 105 includes an electronic processor 205 (for example, a microprocessor or other electronic device). The electronic processor 205 includes input and output interfaces (not shown) and is electrically coupled to a memory 210, a network interface 215, speakers 220 (e.g., one or more audio outputD24130W001 devices), and one or more orientation and / or movement sensors 225. In some instances, the wearable device 105 includes fewer or additional components in configurations different from that illustrated in FIG. 2 and / or optionally combines two or more components shown in FIG. 2. For example, the wearable device 105 may additionally include one or a combination of a microphone(s), a camera(s), and a display screen. As another example, the electronic processor 205 may include multiple electronic processors 205 within the wearable device 105 (e.g., a main electronic processor 205 of the wearable device 105 and an electronic processor included in an IMU) that together function to control various aspects of the wearable device 105. In some instances, the wearable device 105 is implemented within a distributed system including one or more components located in different devices. For example, in some embodiments, the wearable device 105 (e.g., the electronic processor 205) includes local hardware components and one or more external hardware components (e.g., one or more electronic processors of the external device 120). In other words, the electronic processor 205 may include any one or a combination of electronic processors located within a single device (e.g., the wearable device 105) or distributed among various devices and / or systems. For example, the electronic processor 205 may include a first electronic processor 205 of the wearable device 105, a second electronic processor of the external device 120 that is configured to communicate with the wearable device 105, or both the first electronic processor 205 and the second electronic processor. Thus, in the claims, if an apparatus or system is claimed, for example, as including an electronic processor or other element configured in a certain manner, for example, to make multiple determinations, the claim or claim element should be interpreted as meaning one or more electronic processors (or other element) where any one of the one or more electronic processors (or other element) is configured as claimed, for example, to make some or all of the multiple determinations. To reiterate, those electronic processors and processing may be distributed within a single device or across multiple devices.

[0038] In some instances, the wearable device 105 performs functionality other than the functionality described below. The various components of the wearable device 105 described herein are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0039] The memory 210 may include read only memory (ROM), random access memory (RAM), other non-transitory computer-readable media, or a combination thereof. The electronic processor 205 is configured to receive instructions and data from the memory 210 and execute,D24130W001 among other things, the instructions. In particular, the electronic processor 205 executes instructions stored in the memory 210 to perform the methods described herein.

[0040] The network interface 215 sends data to and receives data from other devices (e.g., the external device 120, cloud-based devices, server devices, etc.). In some instances, the network interface 215 includes one or more transceivers for wirelessly communicating with the other devices. Alternatively or in addition, the network interface 215 may include a connector or port for receiving a wired connection to one or more of the other devices, such as connector to receive a coaxial cable. The electronic processor 205 may receive data (for example, audio data and / or the like) from one or more other devices (e.g., the external device 120 and / or another device) via the network interface 215. The electronic processor 205 may output the data (e.g., audio / audio data) via the speakers 220 (and / or via a display screen in instances where the wearable device 105 includes the display screen). In some instances, the speakers 220 include a left speaker and right speaker. In some instances, the speakers 220 may include additional and / or alternative speakers 220. In some instances, the wearable device 105 may also receive power from another device (e.g., the external device 120) via the network interface 215 to power components of the wearable device 105 (e.g., the electronic processor 205). In some instances, the wearable device 105 may include a separate power connection to receive power from another device and / or may include its own power supply such as a rechargeable or replaceable battery.

[0041] The orientation and / or movement sensor(s) 225 may include one or more sensors that are configured to collect and / or determine data indicative of an orientation (e.g., three-dimensional orientation including pitch, roll, yaw, and / or the like) and / or a movement of the head 110 of the user 115 of the wearable device 105 (i.e., head orientation / movement data). For example, the orientation and / or movement sensor(s) 225 may include an inertial measurement unit (IMU) that may include one or more accelerometers and / or gyroscopes. The head orientation / movement data may be an output of any one or a combination of sensors of the IMU (e.g., raw data from one or more sensors). Additionally or alternatively, the head orientation / movement data may be an output of the IMU itself. For example, the IMU may include an electronic processor configured to analyze raw data from one or more sensors and make one or more determinations to generate head orientation / movement data that is then provided to the electronic processor 205. In some instances, the head orientation / movement data includes absolute orientation data, velocity data, acceleration data, magnetic field data, and / or the like. In some instances, the head orientation / movement data includes linear and angular data, acceleration data from one or more accelerometers, angular velocity data from one or moreD24130W001 gyroscopes, magnetometer data, and / or the like. In some instances, the electronic processor 205 is configured to receive the data indicative of the orientation and / or the movement of the head 110 of the user 115 of the wearable device 105 (i.e., the head orientation / movement data) from the orientation and / or movement sensor(s) 225.

[0042] In some instances, the external device 120 may include the same or similar components shown in FIG. 2 with respect to the wearable device 105. Such components of the external device 120 may be similar to the components described above with respect to the wearable device 105 and may perform similar general functions. In some instances, the external device 120 includes fewer or additional components in configurations different from that illustrated in FIG. 2 and / or optionally combines two or more components shown in FIG. 2. For example, the external device 120 may include a display screen such as a touchscreen configured to display a user interface and / or receive user inputs from a user. As another example, the external device 120 may include one or more microphones. In some instances, one or more components of the external device 120 (e.g., the speaker(s), the network interface, and / or the like) may be different than that of the wearable device 105.

[0043] FIG. 3 illustrates a block diagram 300 of a virtualizer 305 implemented by the electronic processor 205 of the wearable device 105 according to some example instances. As shown in FIG. 3, the virtualizer 305 receives two-channel input audio 310, for example, from the external device 120 or from another device. The virtualizer 305 also receives orientation and / or movement data 315 (i.e., head orientation / movement data 315) from the orientation and / or movement sensor(s) 225 of the wearable device 105. As explained in greater detail herein, the virtualizer 305 is configured to generate output binaural audio 320 based on the two-channel input audio 310 and the head orientation / movement data 315. In some instances, the virtualizer 305 is implemented by the electronic processor(s) 205 of the wearable device 105 such that any type of two-channel audio that is received by the wearable device 105 is rendered to binaural audio that is output based on an orientation and / or movement of the user’s head 110. In some instances, the virtualizer 305 is additionally or alternatively implemented by the electronic processor(s) of the external device 120 in conjunction with the head orientation / movement data 315 from the wearable device 105 that is communicatively coupled to the external device 120. When the virtualizer 305 is described herein as performing a certain task or action, it should be understood that the electronic processor(s) that implement the virtualizer 305 is (are) the hardware that performs the certain task or action. While the tasks or actions performed by the virtualizer 305 may be described herein as being performed by the electronic processor 205 ofD24130W001 the wearable device 105, other electronic processors (e.g., an electronic processor(s) of the external device 120) may additionally or alternatively perform the tasks or actions in some instances as explained previously herein.

[0044] FIG. 4 illustrates a more detailed block diagram of the virtualizer 305 according to some example instances. In some instances, the virtualizer 305 receives the input audio 310 that is configured to be output by the wearable device 105. As indicated previously herein, the input audio includes two-channel input audio.

[0045] In some instances, the virtualizer 305 operates in a complex quadrature mirror filter (CQMF) domain. For example, as shown at block 405 of FIG. 4, a CQMF analysis may be executed by the electronic processor 205 to convert the input audio 310 into the CQMF domain such that other processing actions shown in FIG. 4 may be completed in CQMF domain. For example, the input audio 310 is CQMF analyzed into CQMF bands and blocks.

[0046] In some instances, the virtualizer 305 includes a binaural detector 410 that is configured to determine a likelihood whether the input audio 310 includes a first type of audio (e.g., binaural audio). For example, the binaural detector 410 may determine whether the input audio 310 includes the first type of audio versus a second type of audio (e.g., stereo audio). In some instances, binaural audio is audio produced through binaural recording (e.g., using two microphones positioned to approximate the position of ears on a human head, each microphone associated with generating respective ear-specific signals (i.e., a left ear signal and a right ear signal). In some instances, binaural audio is audio produced by rendering non-binaural audio (e.g., object / channel -based audio) with filters that mimic how sound interacts with human anatomy (e.g., a head-related transfer function). In either case, binaural audio includes a pair of ear-specific signals (i.e., a left ear signal and a right ear signal) that include the appropriate spatial cues for reproduction by transducers at or near a listener’s ears (e.g., a near-field speaker proximate to the left ear and a near-field speaker proximate to the right ear). The process of generating binaural audio from non-binaural audio (e.g., object / channel -based audio) is often referred to as binaural rendering, as discussed above, or binauralization. When binaural audio is reproduced by near-field speakers such as headphones, it generally provides a more detailed and realistic spatial representation of sound compared to corresponding stereo audio reproduced on the same near-field speakers. In some instances, binaural audio may provide a more immersive listening experience compared to, for example, stereo audio. In some instances, stereo audio is audio that may merely provide a sense of directionality by separating sound into left and right channels. In some instances, stereo audio is audio designed for far-field playback (e.g.,D24130W001 loudspeaker playback) as opposed near-field playback (e.g., headphone playback), and does not provide as precise spatial perception as binaural audio. Overall, binaural audio may be considered a specific form of two-channel audio that is distinct from stereo audio both in terms of how the audio is produced and / or its content (e.g., presence or absence of spatial cues introduced via the process of binauralization or binaural recording).

[0047] In some instances, the electronic processor 205 implementing the binaural detector 410 analyzes the input audio 310 to determine a first probability that indicates a likelihood whether the input audio 310 includes binaural audio. The electronic processor 205 implementing the binaural detector 410 may also analyze the input audio 310 to determine an opposite probability of the first probability, where the opposite probability indicates a likelihood whether the input audio 310 includes stereo audio. For example, the first probability may represent the probability that the input audio 310 includes binaural audio while the opposite probability (i.e., one minus the first probability) may represent the probability that the input audio 310 includes stereo audio. As a specific example, when the first probability is 70%, the opposite probability may be 30%.

[0048] In some instances, the binaural detector 410 classifies the input audio 310 into only two types: binaural audio and stereo audio as explained previously herein. In other instances, the binaural detector 410 classifies the input audio into additional types / categories and / or sub-types / sub-categories. For example, the electronic processor 205 implementing the binaural detector 410 analyzes the input audio 310 to determine a second probability that indicates a likelihood whether the input audio 310 includes professional binaural audio that was rendered with an object-based file. In some instances, there may be different types of binaural audio such as professional binaural audio and stereo binaural audio. For example, professional binaural audio includes binaural audio that was rendered with an object-based file (e.g., Dolby Atmos™ audio). On the other hand, stereo binaural audio may be rendered by up-mixing stereo audio and is not professional binaural audio because it was not rendered with an object-based file.

[0049] In some instances, performing one or more of the classifications of types of input audio 310 allows the virtualizer 305 to maintain the intention of the creator of binaural audio (e.g., content creator object-based intention of professional binaural audio) when such intention exists while allowing stereo audio to be converted to binaural audio to be provide a more immersive listening experience for the user 115.D24130W001

[0050] In some instances, analyzing the input audio 310 to determine the first probability includes receiving, with the electronic processor 205, metadata associated with the input audio 310. The metadata may indicate that the input audio 310 includes binaural audio and / or professional binaural audio. The electronic processor 205 may determine (i) that the first probability that indicates the likelihood of whether the input audio 310 includes binaural audio is 100% based on the metadata and / or (ii) that the second probability that indicates the likelihood of whether the input audio 310 includes professional binaural audio is 100% based on the metadata. In other words, metadata received with the input audio 310 may indicate whether the input audio is binaural audio and / or professional binaural audio.

[0051] In some instances, analyzing the input audio 310 to determine the first probability includes analyzing the input audio 310 using a machine learning model to classify audio content into two or more types corresponding to respective classifiers. In some instances, the machine learning model is trained from audio features including, but not limited to, inter-channel phase difference (ICPD) and inter-channel level difference (ICLD). In some instances, the binaural detector 410 includes several classifiers, which may complete binary classification or multi class classification. In some instances, the binaural detector 410 implements classifiers with pretrained adaptive boosting (AdaBoost) models. The raw features on which the classification may be based come from the ICPD and the ICLD. In some instances, for a CQMF domain signal, for a bth band signal xb ch0and xb chl, the original ICPD is represented by Equations 1 and 2 below, where blk is the block number representing sequence in the time domain, b represents the band, chO represents the left channel, and chi represents the right channel. In some instances, x^ is the cross-channel conjugate multiplication of the bth band. In some instances, in Equation 2, 3 is the imaginary part of the cross-channel conjugate multiplication x^ and 31 is the real part of the cross-channel conjugate multiplication x^. In some instances, APis a calculated ICPD.Equation 1.Equation 2: AP= atan

[0052] In some instances, the original ICLD is represented by Equations 3 and 4 below, where b again represents the bth band, ch represents channel number, and blk again is the block number representing sequence in the time domain. In some instances, Ech is the energy of the channel ch. In some instances, ALis a calculated ICLD.D24130W001Equation 3: Ech= (xbiChi btk)2Equation 4: &

[0053] With further statistical estimation on the above-noted raw features, features are extracted for training and inference by the machine learning model. In some instances, the binaural detector 410 may provide a classification between stereo audio and binaural audio. In some instances, the binaural detector 410 may provide additional classifications such as further classifying binaural audio into professional binaural audio and stereo binaural audio as described previously herein. In some instances, there are two structures for a classifier set: a parallel structure and a cascade structure. In other words, the respective classifiers for the input audio 310 may be determined using a parallel structure or a cascade structure. These structures may provide confidence scores co, ci, ... cnfor each classification as indicated in FIGS. 5A and 5B that illustrate a parallel structure 505 and a cascade structure 510, respectively.

[0054] In some instances, the binaural detector 410 does not use classifiers to generate the confidence scores. Rather, the confidence scores may be generated from parameter settings of the wearable device 105. For example, the confidence scores may indicate in a binary manner whether input audio 310 is determined to be binaural or stereo (i.e., not binaural). In some instances, a flag may be set or cleared to indicate whether the input audio 310 is determined to be binaural or stereo. For example, an is binaural flag may be set to “1” in response to the input audio 310 being determined to be binaural and may be set to “0” in response to the input audio 310 being determined to be stereo (i.e., not binaural). In some instances, the is binaural flag may be set by the electronic processor 205 based on metadata received with the input audio 310 that indicates a characterization / type of the input audio 310. In some instances, the confidence scores are used by a steerer 415 to generate one or more steering signals for processing of the input audio 310. For example, the steerer 415 (e.g., a steering model) implemented by the electronic processor 205 may be configured to synthesize confidence scores from the binaural detector 410 and provide coefficients that make up a steering signal for other components of the virtualizer 305 such as a separator 420 and / or rendering device 435. For example, the confidence scores may be synthesized by generating a smoothed confidence score using leaky integration. In some instances, the electronic processor 205 is configured to generate a binaural steering signal Sb based at least in part on the first probability that indicates the likelihood whether the input audio 310 includes binaural audio.D24130W001

[0055] FIG. 6 illustrates a high-level schematic diagram of the steerer 415 and its inputs 605 and outputs 610 according to some example instances. As shown in FIG. 6, the steerer 415 receives the confidence scores from the binaural detector 410 as inputs 605 and generates steering signals Sb and spas outputs 610, where sbis a binaural steering signal and spis a professional binaural steering signal. In some instances, the steerer 415 may not generate the professional binaural steering signal sp. In some instances, the steerer 415 may generate additional and / or alternative steering signals not shown in FIG. 6. For example, the steerer 415 may generate a stereo steering signal corresponding to an opposite probability of the first probability that indicates the likelihood whether the input audio 310 includes binaural audio. In some instances, the stereo steering signal may be determined by subtracting the binaural steering signal Sb from one (i.e., one minus Sb). In some instances, the steerer 415 may receive additional, fewer, and / or alternative inputs than the confidence scores shown in FIG. 6 (e.g., depending on how many classifications are made by the binaural detector 410).

[0056] As shown in FIG. 4, the virtualizer 305 also includes a separator 420 configured to separate two-channel input audio 310 that has been converted into the CQMF domain (at block 405) to, for example, more than two channels (e.g., multi-channel audio that includes multiple channels and / or objects). In some instances, the separator 420 includes two distinct separators: a binaural separator 425 and a stereo separator 430. In some instances, the electronic processor 205 that implements the binaural separator 425 performs binaural separation of the input audio 310 to determine an audio scene center of binaural audio included in the input audio 310 as explained in greater detail below. In some instances, the electronic processor 205 that implements the stereo separator 430 performs stereo separation of the input audio 310 to extract three-channel audio from the input audio 310 as explained in greater detail below. In some instances, both of the binaural separator 425 and the stereo separator 430 perform distinct separation steps on the input audio 310 regardless of a type of the input audio. In other words, different separation methods are used on the input audio 310 for binaural audio included in the input audio 310 and for stereo audio included in the input audio 310. In other instances, only one of the two above-noted distinct separation steps is performed on the input audio 310 (e.g., when the type of input audio 310 is 100% known, for example, based on metadata received with the input audio 310).

[0057] FIG. 7 illustrates a block diagram of the separator 420 according to some example instances. In some instances, the separations of the input audio 310 result in separate multichannel audio signals 705 and 710 as shown in FIG. 7. In some instances, the separate multi-D24130W001 channel audio signals 705 and 710 are combined according to steering information (e.g., one or more steering signals) from the steerer 415 as indicated in FIG. 7. For example, during combination of the two signals 705 and 710 by the electronic processor 205, the multi-channel audio signal 705 from the binaural separator 425 may be weighted based on the binaural steering signal Sb from the steerer 415 and corresponding to the first probability that indicates the likelihood whether the input audio 310 includes binaural audio. Continuing this example, during combination of the two signals 705 and 710 by the electronic processor 205, the multi-channel audio signal 710 from the stereo separator 430 may be weighted based on a stereo steering signal corresponding to an opposite probability of the first probability (i.e., one minus Sb) as indicated in FIG. 7. In other words, in some instances, the electronic processor 205 combines the audio scene center (of the binaural audio of the input audio 310 determined during binaural separation) and the three-channel audio (extracted from the stereo audio of the input audio 310 during stereo separation) to generate combined audio 715. During combination of the two signals 705 and 710 by the electronic processor 205, the audio scene center and the three-channel audio are combined such that (i) the audio scene center is weighted based on the binaural steering signal Sb corresponding to the first probability and (ii) the three-channel audio is weighted based on a stereo steering signal corresponding to an opposite probability of the first probability (i.e., one minus Sb). In some instances, combining the audio scene center and the three-channel audio to generate the combined audio 715 in the weight manner explained above includes utilizing metadata (e.g., object-based metadata from professional binaural audio) received with the input audio 310 to at least partially maintain content creator object-based intention indicated by the metadata. As indicated in FIG. 4, in some instances, performing the binaural separation 425, performing the stereo separation 430, and combining the audio scene center and the three-channel audio to generate the combined audio 715 may be completed in CQMF domain. In some instances, the multi-channel audio signals 705 and 710 are dynamically combined based on the steering signal Sb.

[0058] With respect to more specific functionality of the binaural separator 425, in some instances, the binaural separator 425 is configured to extract the audio scene center from binaural audio included in the input audio 310 (i.e., the binaural separator 425 is based on a method of center extraction). The audio scene center may improve localization after a rendering process that is described in greater detail below. For example, in an audio scene, there are natural center-panned objects including, but not limited to, speech in movie scenes, vocal and some instruments in a song, etc. If these natural center-panned objects are split to symmetric virtual speakers rather than to a center speaker, the localization of these natural center-pannedD24130W001 objects will be inaccurate and make the objects sound like a plane rather than a shaped object. Thus, extracting the audio scene center may improve localization by preventing these natural center-panned objects from be split into symmetric virtual speakers. In some instances, the binaural separator 425 may work directly in CQMF domain and may save complexity with banding according to a critical band. For example, the banding technique may calculate and apply coefficients according to several adjacent CQMF bands rather than on each CQMF band. In some instances, the banding of CQMF bands follows equivalent rectangular bandwidth (ERB) bands. For example, the CQMF bands number for 48kHz input may be 77 while the ERB band number may be 20. In some instances, the binaural separator 425 calculates one or more statistical values from a banded signal as indicated in Equations 5-9 below, where xch brepresents the channel ch and band b of the input audio 310. In some instances, m is the start band of the kth critical band and n is the end band. In some instances, £krepresents left energy,krepresents right energy, Ckrepresents center energy, Skrepresents side energy, and £7?* represents left-right conjugate multiplicationEquation 5:&Equation 6:k= t>=m \xi,b IEquation 7:&Equation 8:Equation 9: £££ =1b=mxo,bxi,b*

[0059] In some instances, the statistical values may be smoothed. In some instances, several rates may be calculated with the factors determined in Equations 5-9 as indicated in Equations 10-13 below, where a is a constant value (e.g., a typical value for a is 0.489 although other values may alternatively be used). In some instances, the rates are calculated from a differential signal to a sum signal to represent the energy distribution for each band. In some instances, Rikrepresents a rate for a left side for the kth ERB band, Rr krepresents a rate for the right side for the kth ERB band, Rfkrepresents a rate for a front side for the kth ERB band, and 6krepresents a phase difference between the left side and the right side of the kth band. In some instances, in Equation 13 the angle operation is equivalent to arctan(I / R), similar to Equation 2.D24130W001Equation 10:Equation 11:Equation 12:Equation 13: 0k= angle (£7?*)

[0060] Non-linearizing a ratio (e.g., the energy ratio) and introducing phase between the left and right ERB bands by the electronic processor 205 provides coefficients for an extraction matrix as indicated in Equations 14 and 15 below, where a, b and c are linear factors of the coefficients and c is a non-linear factor for phase. For example, typical values for a, b, and c are a=0.262, b=0.27, c= 0.113, and a typical value for e is 2. These values are examples and other values may be used in other instances. In some instances, mo kis the coefficient for separation for a left channel of the kth ERB bands, and ml kis the coefficient for separation for a right channel of the kth ERB bands.Equation 14:Equation 15: ml k= [a — b(Rrk3+ ) + c(Rfk3+ E)](1 — 0fc)e

[0061] In some instances, the binaural separator 425 determines an extraction matrix t represented by Equation 16 below.Equation 16:

[0062] In some instances, each element in the critical band may be up-mixed with a matrix as indicated in Equation 17 below (e.g., through a multiplication between the extraction matrix Mkand original input signals xo band xl b).>Equation 17:D24130W001

[0063] With respect to more specific functionality of the stereo separator 430, in some instances, the stereo separator 430 extracts more channels (i.e., multi-channel audio) from the two-channel input audio 310 (i.e., the stereo separator 430 is based on a method of channel decoding). In some instances, the stereo separator 430 assumes the input audio 310 to be encoded from a three-channel system. For example, FIG. 8 illustrates an assumption of encoding a channel-based signal. FIG. 8 includes a three-channel audio scene that includes three virtual objects: a left object L, a right object R, and a center object C. FIG. 8 also illustrates a speaker angle 9 from a line of symmetry about which the left object and right object are symmetrical with each other. In some instances, the stereo signal is assumed to be matrix encoded from the three-channel signal. In some instances, the encoding gain for the left object L is based on the speaker angle 9. In some instances, there is a corresponding speaker angle 9 between the right object R and the center object C on which an encoding gain for the right channel R is based. Equation 18 represents the encoding gain gi(0 for the left object L, and Equation 19 represents the encoding gain gr(fT) for the right object R.Equation 18:Equation 19:

[0064] From the encoding gain, the stereo separator 430 may use a matrix decoding method. In some instances, the decoding follows a procedure explained below on each band. First, the stereo separator 430 may calculate a statistical estimation on the input signal x using Equation 20 below. In some instances, C is the cross-channel conjugated multiplication for the bth band, D is the difference between self-conjugated multiplication of channels for the bth band, and A is the sum between self-conjugated multiplication of channels for the bth band. As described previously herein, b may represent the band, xo may represent the input signal corresponding to the left channel, and xi may represent the input signal corresponding to the right channel.Equation 29."

[0065] In some instances, the stereo separator 430 smooths the statistical estimations to get Cb, T)b, <Ab, which represent the smoothed signals of Cb, T>b, and <Jlb. Second, the stereo separator 430 may estimate an encoding angle for this band using Equation 21 below.D24130W001 Equation 21: 0b= angle(T>b,Cb)

[0066] Third, the stereo separator 430 may construct a decoding matrix M as indicated by Equation 22 below.>

[0067] Fourth, the stereo separator 430 may decode the input Sb as indicated in Equation 23 below and may determine a residual signal Rb as indicated in Equation 24 below. For each band, it may be assumed that the band includes two signals: a steered signal Sb and a diffuse signal (i.e., a residual signal Rb. The steered signal is a signal having direction and is encoded with the encoding matrix. The diffuse signal is a signal that spreads across virtual speakers. Extracting the steered signal may be referred to as matrix decoding (which is opposite to encoding). In some instances, the decoding process decodes two-channel audio to three-channel audio.&

[0068] Fifth, the stereo separator 430 may pan the signal with a function represented by Equation 25 below, where 0 = 0S, 02= 0 or 0 = 0, 02= 0Sdepending on where the estimated 0bis located, and where 0Sis a parameter speaker angle which is adjustable. In some instances, 0 and 02are the angles for the decoding matrix, and 0bis the estimated steer signal angle for the bth band (e.g., 0bmay be -n to JC).

[0069] In some instances, a panning matrix Pb generated by the stereo separator 430 is represented by Equation 26 below.D24130W001

[0070] Sixth, the stereo separator 430 may decorrelate the residual signal Rb with a 3x2 decorrelator matrix Db and spread the result into channels. In some instances, the separated signal from the stereo separator 430 is represented by Equation 27 below.<Equation 27:>

[0071] In some instances, the output of binaural separator 425 and the stereo separator 430 is dynamically combined based on the binaural steer signal Sb as explained previously herein. In some instances, the combined output 715 of the separator 420 is represented by Equation 28 below, where the separated binaural channels / data is weighted based on the binaural steering signal Sb corresponding to the first probability and where the separated stereo channels / data is weighted based on the stereo steering signal (i.e., one minus Sb) corresponding to an opposite probability of the first probability.Equation 28:

[0072] In alternate instances, the separator 420 may be omitted from the virtualizer 305 so that the two-channel input audio 310 may be processed directly by a rendering device 435 that is described in greater detail below. In other alternate instances, the stereo separator 430 may be based on a method of center extraction that is similar to the center extraction method performed by the binaural separator 425. In such instances, the separator 420 may include a single separator in the separator 420 that performs center extraction on both binaural audio included in the input audio 310 and stereo audio included in the input audio 310.

[0073] With reference to FIG. 4, the electronic processor 205 implements the rendering device 435 to render the output binaural audio 320 to be output by the wearable device 105D24130W001 based on the orientation and / or the movement of the head 110 of the user 115. In some instances, the rendering of the output binaural audio 320 is based on the combined audio 715 and the data indicative of the orientation and / or the movement of the head 110 of the user 115 of the wearable device 105. The rendering of the output binaural audio 320 may additionally or alternatively be based on one or more steering signals from the steerer 415, for example, as explained in greater detail below.

[0074] FIG. 9 illustrates a block diagram of the rendering device 435 that may be implemented by the electronic processor 205 according to some example instances. As shown in both of FIGS. 4 and 9, the rendering device 435 may include a head-related transfer function (HRTF) engine 440 and a Tenderer 445. Also as indicated in both of FIGS. 4 and 9, the rendering device 435 may receive a plurality of inputs including, but not limited to, the combined audio 715, steering signals from the steerer 415 (e.g., Sb, sp, and / or the like), head orientation / movement data from the sensor(s) 225 of the wearable device 105, and a reverb r, 905 generated with a feedback delay network (FDN) 450. Based on these inputs, the rendering device 435 may output / render the output binaural audio 320. In some instances, the rendering device 435 may receive more or fewer inputs and / or alternate inputs.

[0075] As shown in FIG. 9, the HRTF engine 440 may include a universal HRTF generator 910 that is configured to generate complex valued HRTF coefficients for each CQMF band. For a binaural audio input, the HRTF coefficients are normalized to avoid the double processing problem explained previously herein. The reverb r, 905 generated from the FDN may be taken into account during rendering by the Tenderer 445 to account for room simulation (e.g., sound vibrations within a simulated room). The two steering signals Sb and spof FIG. 9 respectively represent binaural steering and professional content steering and are used to control HRTF mixing and rendering. As shown in FIG. 9, the rendering device 435 receives the combined signal x , 715 (e.g., from the separator 420) as input, and the rendering device 435 outputs binaural audio y, 320.

[0076] In some instances, a scene constructor 915 of the HRTF engine generates a virtual listening space based on the head orientation / movement data from the sensor(s) 225 of the wearable device 105. For example, the virtual listening space is controllable through head rotation, object distance, and other parameters. The scene constructor 915 may output parameters for the HRTF generator 910.D24130W001

[0077] FIG. 10 illustrates an example sound scene of a three-object system according to some example instances. In FIG. 10, a head yaw rotation is y, a speaker angle is 0S, and a distance for left (L), right (R) and center (C) channels is dt, drand dc, respectively. In some instances, a rotation matrix of head rotation of the head 110 of the user 115 may be represented by Rh, and an original location for object coordinates may be L = [Lo, LltL2, ... ] . For example,-djsin (0S)if Lorepresents the left channel,dzcos (0S) if the rotated angles for the three channels 0are L' = RhL. These calculations describe how the original virtual objects are represented in a coordinate system and how a head rotation (e.g., headtracking) affects how the virtual objects are represented.

[0078] In some instances, the HRTF generator 910 may provide a pair of complex valued coefficients h(b, L') = { / q(b, L'), hr(b, L')} for the bth band, where h corresponds to the original object, hi is the object-to-left-ear transfer function, and h, is the object-to-right-ear transfer function. In some instances, the HRTF generator 910 provides transfer functions from one object to two ears. For CQMF representation, there may be complex coefficients for each band. The location of the objects is provided to the HRTF generator 910 to generate the desired coefficients. For separated stereo audio, the coefficients may be directly applied, but for separated binaural audio, the coefficients are normalized (by a normalizer 920 of FIG. 9) to prevent / avoid the double-processing issue described previously herein. For example, for stereo audio, the coefficients may be directly applied because a HRTF was not previously applied to the stereo audio. On the other hand, for binaural audio, the binaural audio has already had an original HRTF applied. Thus, for headtracking, a difference between the original HRTF and a rotated HRTF is applied using the normalization procedure performed by the normalizer 920 of FIG. 9. In some instances, the normalization is based on a zero-rotation state of the virtual scene. For a given virtual scene parameter (including speaker angle, distance, etc.), reference virtual scene locations are recorded, in which the head rotation is zero. The normalization may be represented by Equation 29 below, where h is the normalized HRTF.Equation 29:

[0079] In some instances, the values are complex values, and the magnitude response and phase response are normalized together. Then, as shown in FIG. 9, a HRTF mixer 925 may mixD24130W001 a normalized HRTF and an original HRTF together based on the steering signal Sb. The mixed coefficient may be represented by Equation 30 below.Equation 30:

[0080] The rendering device 435 applies the HRTF coefficient on corresponding signal bands that are added together. In some instances, a direct-to-late rate rai is introduced by the rendering device 435 to control a perceived distance of objects with respect to the virtual room. The reverb amount r may combine the direct-to-late rate and a steering signal from the steerer 415 according to Equation 31 below, where aris the reverb amount and spis the professional binaural steering signal.Equation 31: ar= sp(1 — rd;)

[0081] The Tenderer 445 may then apply the reverb amount to the combined audio 715 to render the output binaural audio 320 for each of two speakers 220 of the wearable device 105 (e.g., a left speaker 220 and a right speaker 220) in accordance with Equations 32 and 33 below, where yi is the final left channel output and j', is the final right channel output.Equation 32:Equation 33:

[0082] As explained above with respect to FIG. 9, in some instances, the electronic processor 205 implementing the rendering device 435 determines a virtual listening space based on the head orientation / movement data (e.g., using a scene constructor 915). The electronic processor 205 may also determine a first HRTF based on the virtual listening space. In some instances, the HRTF may be determined (e.g., selected and / or calculated) in any one or a combination of different manners. For example, the electronic processor 205 may select a HRTF (e.g., from a library of stored HRTFs that may be included in the memory 210 or accessible via communication with another device) based on data corresponding to the user 115 (e.g., demographic data including but not limited to age, gender assigned at birth, geographic data, etc.; physiological data including but not limited to weight, height, head size, ear shape / pitch / distance, etc.; and / or the like). In some instances, the electronic processor 205 selects the first HRTF and modifies the first HRTF based on the data corresponding to the user 115. In some instances, the electronic processor 205 generates the HRTF using a generativeD24130W001 model and the data corresponding to the user 115. In some instances, at least some portions of the HRTF selection and / or generation process is performed by the external device 120 and / or other devices (e.g., a server, a could-based computing device, etc.). In such instances, the first HRTF and / or the determinations related to the first HRTF that were made by the other devices may be transmitted to the electronic processor 205 of the wearable device 105.

[0083] In some instances and as explained previously herein, the electronic processor 205 determines a normalized HRTF based on a zero-rotation state of the virtual listening space. In some instances, the zero-rotation state corresponds to rotation / movement of the user’s head 110 below a rotation / movement threshold value or position within a range associated with a stationary user head position (e.g., ±1 degree of rotation / movement, ±5 degrees of rotation / movement, ±10 degrees of rotation / movement, or the like).

[0084] In some instances, the electronic processor 205 mixes the first HRTF and the normalized HRTF to generate a mixed HRTF such that (i) the normalized HRTF is weighted based on the binaural steering signal Sb corresponding to the first probability that indicates the likelihood of whether the input audio 310 includes binaural audio and (ii) the first HRTF is weighted based on the stereo steering signal (i.e., one minus Sb) corresponding to the opposite probability of the first probability. In some instances and as explained previously herein, rendering the output binaural audio 320 includes rendering the output binaural audio 320 based on the mixed HRTF and a reverb r, 905 generated with the FDN 450.

[0085] In instances in which a professional binaural steering signal spis generated by the steerer 415 to indicate the second probability that indicates the likelihood of whether the input audio 310 includes professional binaural audio, the electronic processor 205 may be configured to render the output binaural audio 320 based on the second probability and a direct-to-late rate rai that is configured to control a perceived distance of objects in the virtual listening space.

[0086] The rendering device 435 is configured to to render the output binaural audio 320 in a manner that at least partially maintains content creator object-based intention of professional binaural audio. Accordingly, when metadata is provided with the input audio 310 to indicate that the input audio 310 includes professional binaural audio, the electronic processor 205 may render the output binaural audio 320 by utilizing the metadata to at least partially maintain content creator object-based intention indicated by the metadata.

[0087] In some instances, the electronic processor 205 is configured to render the output binaural audio 320 in the CQMF domain. In some instances, at block 455, the electronicD24130W001 processor 205 is configured to synthesize the output binaural audio 320 from the CQMF domain to a time domain for output by the wearable device 105 (e.g., for output by the speakers 220 of the wearable device 105).

[0088] FIG. 11 illustrates a flowchart of a method 1100 for controlling the wearable device 105 to output audio according to some example instances. The method 1100 is described as being performed by the electronic processor 205 implementing the virtualizer 305, but other electronic processors may be involved in at least portions of the method 1100 as described previously herein. The method 1100 includes actions performed by the virtualizer 305 as previously explained herein. While a particular order of processing steps, message receptions, and / or message transmissions is indicated in FIG. 11 as an example, timing and ordering of such steps, receptions, and transmissions may vary where appropriate without negating the purpose and advantages of the examples set forth in detail throughout the remainder of this disclosure.

[0089] At block 1105, the electronic processor 205 receives the input audio 310 that is configured to be output by the wearable device 105. In some instances, the input audio 310 includes two-channel input audio. In some instances, the input audio 310 may be received from the external device 120 or from another device such as a server, a cloud-based device, etc.

[0090] At block 1110, the electronic processor 205 analyzes the input audio 310 to determine a first probability that indicates a likelihood whether the input audio 310 includes binaural audio. In some instances, the electronic processor 205 implements the binaural detector 410 explained previously herein to analyze the input audio 310 and to determine the first probability.

[0091] At block 1115, the electronic processor 205 generates a binaural steering signal Sb based at least in part on the first probability. In some instances, the electronic processor 205 implements the steerer 415 explained previously herein to generate the binaural steering signal Sb.

[0092] At block 1120, the electronic processor 205 performs binaural separation of the input audio 310 to determine an audio scene center of the binaural audio included in the input audio 310. In some instances, the electronic processor 205 implements the binaural separator 425 explained previously herein to perform the binaural separation.

[0093] At block 1125, the electronic processor 205 performs stereo separation of the input audio 310 to extract three-channel audio from the input audio (e.g., extract three-channel audioD24130W001 from the stereo audio included in the input audio 310). In some instances, the electronic processor 205 implements the stereo separator 430 explained previously herein to perform the stereo separation.

[0094] At block 1130, the electronic processor 205 combines the audio scene center and the three-channel audio to generate the combined audio 715. In some instances, the audio scene center and the three-channel audio are combined such that (i) the audio scene center is weighted based on the binaural steering signal Sb corresponding to the first probability and (ii) the three-channel audio is weighted based on a stereo steering signal (one minus Sb) corresponding to an opposite probability of the first probability.

[0095] At block 1135, the electronic processor 205 receives, from a sensor(s) 225 of the wearable device 105, data indicative of an orientation and / or a movement of a head of a user of the wearable device 105 (i.e., head orientation / movement data). In some instances, the electronic processor 205 implements the rendering device 435 that includes the HRTF engine 440 explained previously herein to receive the head orientation / movement data.

[0096] At block 1140, the electronic processor 205 renders, based on the orientation and / or the movement of the head of the user, output binaural audio 320 to be output by the wearable device 105. In some instances, the rendering is based on the combined audio 715 and the data indicative of the orientation and / or the movement of the head 110 of the user 115 of the wearable device 105. In some instances, the electronic processor 205 implements the Tenderer 445 within the rendering device 435 explained previously herein to render the output binaural audio 320.

[0097] As indicated in FIG. 11, in some instances, the method 1100 may repeat as additional input audio 310 is received to be output by the wearable device 105.

[0098] The foregoing description, for purpose of explanation, has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The embodiments were chosen and described in order to best explain the principles of the techniques and their practical applications. Others skilled in the art are thereby enabled to best utilize the techniques and various embodiments with various modifications as are suited to the particular use contemplated.D24130W001

[0099] Although the disclosure and examples have been fully described with reference to the accompanying drawings, it is to be noted that various changes and modifications will become apparent to those skilled in the art. Such changes and modifications are to be understood as being included within the scope of the disclosure and examples as defined by the claims.

[0100] Various features and advantages are set forth in the following claims.

Claims

D24130W001CLAIMSWhat is claimed is:

1. A method for controlling a wearable device to output audio, the method comprising: receiving, with an electronic processor, input audio configured to be output by the wearable device, wherein the input audio includes two-channel input audio;analyzing, with the electronic processor, the input audio to determine a first probability that indicates a likelihood whether the input audio includes binaural audio;generating, with the electronic processor, a binaural steering signal based at least in part on the first probability;performing, with the electronic processor, binaural separation of the input audio to determine an audio scene center of the binaural audio;performing, with the electronic processor, stereo separation of the input audio to extract three-channel audio from the input audio;combining, with the electronic processor, the audio scene center and the three-channel audio to generate combined audio, wherein the audio scene center and the three-channel audio are combined such that (i) the audio scene center is weighted based on the binaural steering signal corresponding to the first probability and (ii) the three-channel audio is weighted based on a stereo steering signal corresponding to an opposite probability of the first probability;receiving, with the electronic processor and from a sensor of the wearable device, data indicative of an orientation and / or a movement of a head of a user of the wearable device; and rendering, with the electronic processor and based on the orientation and / or the movement of the head of the user, output binaural audio to be output by the wearable device, wherein the rendering is based on the combined audio and the data indicative of the orientation and / or the movement of the head of the user of the wearable device.

2. The method of claim 1, further comprising:determining, with the electronic processor, a virtual listening space based on the data; determining, with the electronic processor, a first head-related transfer function (HRTF) based on the virtual listening space;determining, with the electronic processor, a normalized HRTF based on a zero-rotation state of the virtual listening space; andmixing, with the electronic processor, the first HRTF and the normalized HRTF to generate a mixed HRTF, wherein the first HRTF and the normalized HRTF are mixed such that (i) the normalized HRTF is weighted based on the binaural steering signal corresponding to theD24130W001 first probability and (ii) the first HRTF is weighted based on the stereo steering signal corresponding to the opposite probability of the first probability;wherein rendering the output binaural audio includes rendering the output binaural audio based on the mixed HRTF and a reverb generated with a feedback delay network.

3. The method of claim 2, further comprising analyzing, with the electronic processor, the input audio to determine a second probability that indicates a likelihood whether the input audio includes professional binaural audio, wherein the professional binaural audio was rendered with an object-based file;wherein rendering the output binaural audio includes rendering the output binaural audio based on the second probability and a direct-to-late rate that is configured to control a perceived distance of objects in the virtual listening space.

4. The method of claim 3, wherein analyzing the input audio to determine the second probability includes:receiving, with the electronic processor, metadata associated with the input audio, wherein the metadata indicates that the input audio includes professional binaural audio; and determining, with the electronic processor, that the second probability is 100% based on the metadata.

5. The method of claim 4, wherein one or both of combining the audio scene center and the three-channel audio to generate combined audio and rendering the output binaural audio includes utilizing the metadata to at least partially maintain content creator object-based intention indicated by the metadata.

6. The method of any one of the preceding claims, further comprising:converting, with the electronic processor, the input audio into a complex quadrature mirror filter (CQMF) domain, wherein performing the binaural separation, performing the stereo separation, combining the audio scene center and the three-channel audio to generate the combined audio, and rendering the output binaural audio are completed in the CQMF domain; andsynthesizing, with the electronic processor, the output binaural audio from the CQMF domain to a time domain for output by the wearable device.D24130W001 7. The method of any one of the preceding claims, wherein analyzing the input audio to determine the first probability includes analyzing the input audio using a machine learning model to classify audio content into two or more types corresponding to respective classifiers, wherein the machine learning model is trained from audio features including inter-channel phase difference (ICPD) and inter-channel level difference (ICLD).

8. The method of claim 7, wherein the respective classifiers for the input audio are determined using a parallel structure or a cascade structure.

9. The method of any one of the preceding claims, wherein analyzing the input audio to determine the first probability includes:receiving, with the electronic processor, metadata associated with the input audio, wherein the metadata indicates that the input audio includes binaural audio; and determining, with the electronic processor, that the first probability is 100% based on the metadata.

10. The method of any one of the preceding claims, wherein the wearable device includes one of a set of earbuds, headphones, smart glasses, and another head-mounted display.

11. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any one of claims 1-10.

12. The wearable device, wherein the wearable device includes a computing apparatus, comprising:an electronic processor; anda memory storing instructions, which when executed by the electronic processor, cause the computing apparatus to perform the method of any one of claims 1-10.