Methods, devices, and systems for rendering spatialized binaural audio
The binaural audio device enhances spatialization by processing binaural downmixed signals with head motion, ensuring accurate spatial cues and immersive audio experience despite listener movement.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BOSE CORP
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-23
AI Technical Summary
Binaural audio recordings break when the listener moves, failing to maintain the spatialization and spatial cues intended by the original recording, especially with binaural downmixes on wearable devices.
A binaural audio device processes binaural downmixed signals to extract a center channel and residual signals, applies spatialization based on head motion, and adjusts frequency-dependent time delays to ensure accurate spatialization and recombination of audio signals for left and right ears.
Maintains spatialization and spatial cues as the listener moves, providing an authentic immersive audio experience by dynamically adjusting audio signals based on head motion.
Smart Images

Figure US2026011236_23072026_PF_FP_ABST
Abstract
Description
METHODS, DEVICES, AND SYSTEMS FORRENDERING SPATIALIZED BINAURAL AUDIOCross-Reference to Related Applications
[0001] This application claims priority to U.S. Patent Application Serial No. 63 / 745,746, filed on January 15, 2025, and titled “METHODS, DEVICES, AND SYSTEMS FOR RENDERING SPATIALIZED BINAURAL AUDIO,” which application is herein incorporated in its entirety.Field of the Disclosure
[0002] This disclosure relates to methods, devices, and systems for rendering spatialized binaural audio, and more particularly, to modifying binaural audio according to motion of a binaural audio device.Background
[0003] Audio content may be provided in immersive, spatialized, three-dimensional, and multi-channel audio formats. A binaural version of immersive audio content may also be provided for wearable audio devices. The binaural version of the immersive audio content attempts to incorporate spatialization information corresponding to audio as it would have been heard by ears of a static head in a three-dimensional environment.Summary
[0004] Some immersive or spatialized audio systems attempt to create a listening experience that is indistinguishable from sitting in the “sweet-spot” between a pair of out-loud stereo speakers. For example, these systems may attempt to replicate the listening experience of an audio engineer listening to monitor speakers in a professional recording studio. Using a model of sound propagation, room acoustics, and real-time knowledge of head motion of the listener, these systems synthesize sound waves that would have occurred at the ears of the listener had they been listening to an out-loud system. This is desirable because sitting between a pair of stereo speakers is precisely how stereo content was intended to be experienced.
[0005] Recently, there has been rapid growth in binaural audio formats. Originally, binaural recordings were created by recording signals from microphones placed in ears of a dummy head (or sometimes a real person). Because the binaural recording captures sound at the ears of the listener, when the binaural recording is played back by a wearable audio device, theDocket No. BOSE-143WO / RS-25-443-WOsound waves at the ear are the same as what was originally heard, therefore the binaural recording should sound the same as the original live experience. This result remains true assuming that the listener remains in the same spot where the recording was made and keeps their head perfectly still. Once the listener changes their head position or head orientation, the audio present at their ears is very different, and the illusion of a binaural recording breaks. In other words, the binaural recording can easily “collapse” to sound like it is inside the head of the listener and sound nothing like the original live experience in terms of sound location and spaciousness. Accordingly, the methods, devices, and systems described below attempt to combine the binaural recording with spatialization technology for an improved listening experience.
[0006] Usage of binaural audio has been growing rapidly, in part due to multi-channel audio formats such as Dolby Atmos® offering a “binaural downmix” for wearable audio devices. This binaural downmix converts stereo audio content or multi-channel audio content with three or more audio channels (left channel, right channel, center channel, etc.) into two audio channels, a left binaural audio channel and a right binaural audio channel. This binaural downmix is available on most audio content platforms, such as Apple Music®, Tidal®, and Amazon Music®. Further, the binaural downmix attempts to add all of the spatial cues that would have occurred if the recording was captured using an unmoving, static dummy head. As the binaural downmix has already been processed with some spatialization, rendering them by processing with further spatialization to account for head movement will fail to provide accurate spatialization.
[0007] The present disclosure is generally directed to methods, devices, and systems for providing spatialized, binaural audio. The audio is provided by a binaural audio device capable of presenting separate audio signals to left and right ears independently. The binaural audio device may be a pair of earbuds, audio headphones, audio sunglasses, a pair of open earbuds, an automotive nearfield headrest, a pair of out-loud speakers with static or dynamic crosstalk cancellation, etc. In some particular examples, the binaural audio device may be a wearable audio device (such as a pair of earbuds, audio headphones, etc.) with left and right acoustic transducers. The binaural audio device receives a binaural downmixed signal from an external device, such as a smartphone. The binaural audio device processes the binaural downmixed signal to generate left and right binaural signals. The binaural audio device then extracts a center channel signal from the left and right binaural signals and also generates left and right residual signals. A spatialization transformer is applied to the center channel signal to spatialize the center channel signal according to a motion signal provided by a motion sensor. Frequency-Docket No. BOSE-143WO / RS-25-443-WOdependent time delays and / or equalization may be applied to the left and right residual signals to ensure proper recombination with the center channel signal. The left acoustic transducer then renders audio based on the time delayed left residual signal and the spatialized center channel signal, while the right acoustic transducer renders audio based on the time delayed right residual signal and the spatialized center channel signal.
[0008] Generally, in one aspect, a method for rendering output audio is provided. The method includes (1) generating a center channel signal, a left residual channel, and a right residual channel based on a left binaural signal and a right binaural signal; (2) generating, via a spatialization transformer, a spatialized center signal based on the center channel signal; and (3) rendering the output audio based on the spatialized center signal, the left residual signal, and the right residual signal.
[0009] According to an example, the output audio comprises left output audio and right output audio. The left output audio may be based on the spatialized center signal and the left residual signal. The right output audio may be based on the spatialized center signal and the right residual signal. The left output audio may be rendered by a left acoustic transducer. The right output audio may be rendered by a right acoustic transducer.
[0010] According to an example, the method further includes (1) generating a delayed left signal based on the left residual signal and a first time delay and (2) generating a delayed right signal based on the right residual signal and a second time delay. The output audio is rendered further based on the delayed left signal and the delayed right signal.
[0011] According to an example, the method further includes adjusting the spatialization transformer based on a motion signal. The motion signal may be provided by a motion sensor.
[0012] According to an example, the left residual signal is generated based on the center channel signal and the left binaural signal. The left residual signal may be generated by subtracting the center channel signal from the left binaural signal.
[0013] According to an example, the right residual signal is generated based on the center channel signal and the right binaural signal. The right residual signal may be generated by subtracting the center channel signal from the right binaural signal.
[0014] According to an example, the center channel signal may be generated by multiplying a sum of the left binaural signal and the right binaural signal by a predetermined center gain constant. The center channel signal, the left residual signal, and the right residual signal may be generated via linear or non-linear source and / or channel separation.
[0015] Generally, in another aspect, a binaural audio device is provided. The binaural audio device includes a controller. The controller is configured to: (1) generate a center channelDocket No. BOSE-143WO / RS-25-443-WOsignal, a left residual signal, and a right residual signal based on a left binaural signal and a right binaural signal; (2) generate, via a spatialization transformer, a spatialized center signal based on the center channel signal; and (3) render output audio based on the spatialized center signal, the left residual signal, and the right residual signal.
[0016] According to an example, the output audio comprises left output audio and right output audio. The left output audio may be based on the spatialized center signal and the left residual signal. The right output audio may be based on the spatialized center signal and the right residual signal. The binaural audio device may further include a left acoustic transducer configured to render the left output audio and a right acoustic transducer configured to render the right output audio.
[0017] According to an example, the controller is further configured to: (1) generate a delayed left signal based on the left residual signal and a first time delay; and (2) generate a delayed right signal based on the right residual signal and a second time delay. The output audio may be rendered further based on the delayed left signal and the delayed right signal.
[0018] According to an example, the controller is further configured to adjust the spatialization transformer based on a motion signal.
[0019] According to an example, the binaural audio device further includes a motion sensor configured to provide the motion signal.
[0020] According to an example, the left residual signal is generated based on the center channel signal and the left binaural signal. The left residual signal may be generated by subtracting the center channel signal from the left binaural signal.
[0021] According to an example, the right residual signal is generated based on the center channel signal and the right binaural signal. The right residual signal may be generated by subtracting the center channel signal from the right binaural signal.
[0022] According to an example, the center channel signal may be generated by multiplying a sum of the left binaural signal and the right binaural signal by a predetermined center gain constant. The center channel signal, the left residual signal, and the right residual signal may be generated via linear or non-linear source and / or channel separation.
[0023] In various implementations, a processor or controller can be associated with one or more storage media (generically referred to herein as “memory,” e.g., volatile and non-volatile computer memory such as ROM, RAM, PROM, EPROM, and EEPROM, floppy disks, compact disks, optical disks, magnetic tape, Flash, OTP-ROM, SSD, HDD, etc.). In some implementations, the storage media can be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform at least some of the functionsDocket No. BOSE-143WO / RS-25-443-WOdiscussed herein. Various storage media can be fixed within a processor or controller or can be transportable, such that the one or more programs stored thereon can be loaded into a processor or controller so as to implement various aspects as discussed herein. The terms “program” or “computer program” are used herein in a generic sense to refer to any type of computer code (e.g., software or microcode) that can be employed to program one or more processors or controllers.
[0024] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.
[0025] Other features and advantages will be apparent from the description and the claims.Brief Description of the Drawings
[0026] In the drawings, like reference characters generally refer to the same parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the various embodiments.
[0027] FIG. 1 illustrates a system for rendering output audio including a binaural audio device and an external device, in accordance with the present disclosure.
[0028] FIG. 2 is a schematic view of a binaural audio device, in accordance with the present disclosure.
[0029] FIG. 3 is a functional block diagram of aspects of the binaural audio device, in accordance with the present disclosure.
[0030] FIG. 4 is a flowchart illustrating a method for rendering output audio according to the present disclosure.Detailed Description
[0031] The present disclosure is generally directed to methods, devices, and systems for providing spatialized, binaural audio. The audio is provided by a binaural audio device capable of presenting separate audio signals to left and right ears independently. The binaural audio device may be a pair of earbuds, audio headphones, audio sunglasses, a pair of open earbuds, an automotive nearfield headrest, a pair of out-loud speakers with static or dynamic crosstalk cancellation, etc. In some particular examples, the binaural audio device may be a wearableDocket No. BOSE-143WO / RS-25-443-WOaudio device (such as a pair of earbuds, audio headphones, etc.) with left and right acoustic transducers. The binaural audio device receives a binaural downmixed signal from an external device, such as a smartphone. The binaural audio device processes the binaural downmixed signal to generate left and right binaural signals. The binaural audio device then extracts a center channel signal from the left and right binaural signals and also generates left and right residual signals. A spatialization transformer is applied to the center channel signal to spatialize the center channel signal according to a motion signal provided by a motion sensor. Frequencydependent time delays and / or equalization may be applied to the left and right residual signals to ensure proper recombination with the center channel signal. The left acoustic transducer then renders audio based on the time delayed left residual signal and the spatialized center channel signal, while the right acoustic transducer renders audio based on the time delayed right residual signal and the spatialized center channel signal.
[0032] The following description should be read in view of FIGS. 1-4.
[0033] The term “wearable audio device,” as used in this application, in addition to including its ordinary meaning or its meaning known to those skilled in the art, is intended to mean a device that fits around, on, in, or near an ear (including open-ear audio devices worn on the head or shoulders of a user) and that radiates acoustic energy into or towards the ear. Wearable audio devices are sometimes referred to as headphones, earphones, earpieces, headsets, earbuds, or sport headphones, and can be wired or wireless. A wearable audio device includes an acoustic driver to transduce audio signals to acoustic energy. The acoustic driver can be housed in an earcup. While some of the figures and descriptions following can show a single wearable audio device, having a pair of earcups (each including an acoustic driver) it should be appreciated that a wearable audio device can be a single stand-alone unit having only one earcup. Each earcup of the wearable audio device can be connected mechanically to another earcup or headphone, for example by a headband and / or by leads that conduct audio signals to an acoustic driver in the earcup or headphone. A wearable audio device can include components for wirelessly receiving audio signals. A wearable audio device can include components of an active noise reduction (ANR) system. Wearable audio devices can also include other components such as a microphone so that they can function as a headset. While the non-limiting example of FIG. 1 depicts the binaural audio device 100 as audio headphones, the binaural audio device 100 described below may be any of the aforementioned types of devices.
[0034] The term “head related transfer function” or acronym “HRTF” is intended to be used broadly herein to reflect any manner of calculating, determining, or approximating head relatedDocketNo. BOSE-143WO / RS-25-443-WOtransfer functions. For example, a head related transfer function as referred to herein may be generated or selected specific to each user, e.g., taking into account that user’s unique physiology (e.g., size and shape of the head, ears, nasal cavity, oral cavity, etc.). Alternatively, a generalized head related transfer function may be generated or selected that is applied to all users, or a plurality of generalized head related transfer functions may be generated that are applied to subsets of users (e.g., based on certain physiological characteristics that are at least loosely indicative of that user’s unique head related transfer function, such as age, gender, head size, ear size, or other parameters). In one embodiment, certain aspects of the head related transfer function may be accurately determined, while other aspects are roughly approximated (e.g., accurately determines the inter-aural delays, but coarsely determines the magnitude response). In various examples, a number of HRTF’s may be stored, e.g., in a memory, and selected for use relative to a determined angle of arrival of a virtual acoustic signal.
[0035] FIG. 1 illustrates an example system 10 for rendering output audio. The system 10 of FIG. 1 includes a binaural audio device 100 and an external device 200. In the non-limiting example of FIG. 1, the binaural audio device 100 is shown as a wearable audio device embodied as audio headphones, while the external device 200 is shown as a smartphone. However, the binaural audio device 100 may be any audio device capable of presenting separate audio signals to left and right ears independently, such as a pair of earbuds, audio headphones, audio sunglasses, a pair of open earbuds, an automotive nearfield headrest, a pair of out-loud speakers with static or dynamic crosstalk cancellation, etc. Similarly, the external device 200 may be any device capable of providing audio data to the binaural audio device 100, such as a personal computer, tablet computer, portable music player, audio speaker, etc. In the example of FIG.1, the external device 200 transmits a binaural downmix signal 202 to the binaural audio device 100. The binaural downmix signal 202 may be transmitted via any suitable wired connection or wireless protocol (such as Bluetooth, Wi-Fi, etc.). The binaural downmix signal 202 may be generated by an audio content platform (such as Spotify®). The binaural downmix signal 202 is generated by converting a spatialized, multi-channel audio signal (corresponding to an immersive audio format such as Dolby Atmos®) into a binaural format for a binaural audio device 100. The binaural downmix signal 202 includes left channel audio to be rendered to a left ear of a listener and right channel audio to be rendered to a right ear of the listener. Accordingly, the binaural downmix signal 202 attempts to translate spatialized audio with more than two channels (such as left, right, front, rear, front left, front right, etc.) into two binaural channels. The spatialization of the binaural downmix signal 202 corresponds to a static head in a virtual environment. As will be described below, the binaural downmix signal 202 must beDocket No. BOSE-143WO / RS-25-443-WOfurther processed to use head motion data of a user to provide a more authentic immersive experience.
[0036] FIG. 2 is a schematic representation of the binaural audio device 100 of FIG. 1. As shown in the non-limiting example of FIG. 2, the binaural audio device 100 includes a controller 101. The controller 101 includes a processor 155 for processing data, a memory 175 for storing data, and a transceiver 185 for transmitting and / or receiving data via an antenna 187. For example, the transceiver 185 may be used to receive the binaural downmix signal 202 transmitted by the external device 200 as shown in FIG. 1. The binaural audio device 100 further includes a left acoustic transducer 103, a right acoustic transducer 105, and a motion sensor 131. The left and right acoustic transducers 103, 105 are used to provide output audio to the left and right ears of the listener, respectively. For example, the left acoustic transducer 103 may be arranged in a left earcup to be worn over the left ear, while the right acoustic transducer 105 may be arranged in a right earcup to be worn over the right ear. The motion sensor 131 is used to track the movement of the head of the listener, and may be arranged at any practical location on the binaural audio device 100.
[0037] FIG. 3 is a functional block diagram illustrating various aspects of the controller 101 of the binaural audio device 100. In particular, these aspects may be executed by the processor 155 of the controller 101 and / or stored in the memory 175 of the controller 101. As shown in FIG. 3, the transceiver 185 of the controller 101 receives the binaural downmix signal 202. The transceiver 185 passes the binaural downmix signal 202 to a binaural channel splitter 111. The binaural channel splitter Ill is configured to convert the single binaural downmix signal 202 received by the transceiver 185 into a left binaural signal 102 and a right binaural signal 104. In previous systems, the left binaural signal 102 would be rendered to the left ear of the listener and the right binaural signal 104 would be rendered to the right ear of the listener.
[0038] The left and right binaural signals 102, 104 are then provided to the center extractor 113 to generate a center channel signal 106. The center channel signal 106 represents the “front audio” content of the left and right binaural signals 102, 104. Accordingly, the center channel signal 106 may represent a virtual audio source arranged in front of the listener in a virtual space. The center channel signal 106 may be determined based on the left binaural signal 102, the right binaural signal 104, a predetermined center gain constant 122, and / or a predetermined side gain constant 124. In some examples, the center channel signal 106 may be calculated by dividing the sum of the left and right channel signals 102, 104 by 2. In other examples, the center channel signal 106 may be calculated by dividing the sum of the left and right binauralDocket No. BOSE-143WO / RS-25-443-WOsignals 102, 104 by the difference between the predetermined center gain constant 122 and the predetermined side gain constant 124.
[0039] A left residual calculator 115 determines a left residual signal 108 based on the left binaural signal 102 and the center channel signal 106. In some examples, the left residual signal 108 is calculated by subtracting the center channel signal 106 from the left binaural signal 102. Similarly, a right residual calculator 117 determines a right residual signal 110 based on the right binaural signal 104 and the center channel signal 106. In some examples, the right residual signal 110 is calculated by subtracting the center channel signal 106 from the right binaural signal 104.
[0040] As can be seen in FIG. 3, the center extractor 113, the left residual calculator 115, and the right residual calculator 117 may be considered to be aspects of a channel extractor 161. Broadly, the channel extractor 161 is configured to generate signals corresponding to three channels (center channel signal 106, left residual signal 108, and right residual signal 110) based on two input signals (left and right binaural signals 102, 104). However, the channel extractor 161 could instead generate the center channel signal 106, the left residual signal 108, and the right residual signal 110 through other means, such as via a linear or non-linear method for source and / or channel separation. For example, the center channel signal 106, the left residual signal 108, and the right residual signal 110 could be generated by an upmixer implemented by the channel extractor 161 such that the upmixer receives the left and right binaural signals 102, 104.
[0041] The center channel signal 106 is then spatialized via a spatialization transformer 119. The spatialization transformer 119 may implement one or more head related transfer functions (HRTFs), to generate a spatialized center signal 112. These HRTFs may be adjusted according to a motion signal 126 provided by a motion sensor 131. In some examples, the motion sensor 131 may be embodied as one or more accelerometers, gyroscopes, or inertial measurement units (IMUs). The motion signal 126 may be used to determine the location and / or orientation of the binaural audio device 100.
[0042] Processing the center channel signal 106 with the spatialization transformer 119 introduces a time delay, which may cause the spatialized center signal 112 to be out-of-sync with the left and right residual signals 108, 110. Accordingly, to compensate for the delay introduced by the spatialization transformer 119, the left residual signal 108 may be processed by a first delay 121 to generate a delayed left signal 114, and the right residual signal 110 may be processed by a second delay 123 to generate a delayed right signal 116. The timing of the first and second delay 121, 123 may be programmed to correspond to the processing delayDocket No. BOSE-143WO / RS-25-443-WOintroduced by the spatialization transformer 119. Accordingly, the delay time of the first and second delay 121, 123 may be approximately equal. Thus, the delays 121, 123 ensure proper recombination of the delayed left and right signals 114, 116 with the spatialized center signal 112. In some examples, the first and second delay 121, 123 may be frequency dependent. In further examples, the first and second delay 121, 123 may implement equalization.
[0043] The delayed left signal 114 and the spatialized center signal 112 are then provided to a left driver 125. The left driver 125 combines the delayed left signal 114 and the spatialized center signal 112 to generate a left driver signal 118. The left driver signal 118 is then provided to the left acoustic transducer 103. As previously described, the left acoustic transducer 103 may be positioned in or near the left ear of the listener, such as within a left earcup or a left earbud. The left acoustic transducer 103 renders output audio (which may be referred to as “left” output audio) based on the left driver signal 118. Accordingly, the left output audio includes aspects of both spatialized front audio as well as downmixed left binaural audio. Similarly, the delayed left signal 114 and the spatialized center signal 112 are then provided to a left driver 125. The right driver 127 combines the delayed right signal 116 and the spatialized center signal 112 to generate a right driver signal 120. The right driver signal 120 is then provided to the right acoustic transducer 105. As previously described, the right acoustic transducer 105 may be positioned in or near the right ear of the listener, such as within a right earcup or a right earbud. The right acoustic transducer 105 renders output audio (which may be referred to as “right” output audio) based on the right driver signal 120. Accordingly, the right output audio includes aspects of both spatialized front audio as well as downmixed right binaural audio. Notably, the spatialized front audio rendered by both acoustic transducers 103, 105 will change according to the motion of the head of the listener as represented by the motion signal 126. By contrast, the rendered audio corresponding to the residual signals 108, 110 will remain static despite any head movements by the listener.
[0044] As described above, the center channel signal 106 (M for “middle” or “mid” channel signal), may be determined based on the left and right binaural signals 102, 104 (L and / ?). Further, the left and right residual signals 108, 110 (Lrand Rr) may be generated as shown below.
[0046] Lr= L - M, (2)
[0047] Rr= R - M. (3)Docket No. BOSE-143WO / RS-25-443-WO
[0048] Accordingly, if the residual signals 108, 110 are expressed in terms of the left and right binaural signals 102, 104, then:>
[0051] A similar approach is to pass the left and right binaural signals 102, 104 through a mid-side matrix, extract the center channel signal 106, and then express new left and right signals in terms of L and R. The “side” component (which may refer to either the left or right side) may be annotated as S.
[0053] Although the mid-side matrix is not orthonormal, L and R can be recovered later. First, the side component will be extracted, and then moved back to a left-right space. This gives the following:
[0055] Note that Lnewand Rneware identical to Lrand Rr. They also have the unfavorable perspective of being exactly out of phase with one another. To work with mid-side processing in a more generic way, the predetermined center gain constant 122 (gm) and the predetermined side gain constant 124 (gs) applied to M and S may be assumed to be generic. Variables gmand gswill represent gains applied to the mid and side components, respectively. This takes the form of:
[0056] M = [l |] H'" “] ["] (8)- newl LI U t) gs] L j J
[0057] which, when expressed in terms of L and R takes the form of:
[0060] If gm= 0 and gs= 1, the original Lnewand Rnewmay be recovered. With this new form, however, the issue of the residual signals 108, 110 being exactly out-of-phase with one another may be addressed. For example, gm= 2 / 3 and gs= 1 / 3 results in:J R
[0061] Lnew= - + ~, (11)D J
[0062] Rnew= - + ~, (12)Docket No. BOSE-143WO / RS-25-443-WO
[0063] which starts to move in the direction of Lnewbeing mostly the left binaural signal 102 and Rnewbeing mostly the right binaural signal 104. Because the “residual” approach is the same, a more generic approach can be applied as well. This is as follows:
[0064] M = gm(L + R , (13)
[0065] Lr= L - M, (14)
[0066] Rr= R — M. (15)
[0067] If dm = 1 / 2, equations 1-3 may be used to derive Lrand Rr. If gm< 1 / 2, the residual signals 108, 110 predominantly correspond to their appropriate binaural signals 102, 104, with a little bit of bleed. If gm> 1 / 2, the opposite binaural signal 102, 104 is the predominant component in the listener’s ear. For example, for gm= 1, left residual signal 108 is — R and right residual is — L. Last, if gm= 0, the original binaural signals 102, 104 are recovered with no extra mid (M) component.
[0068] Upon real-world testing, it was determined that setting gmto about -8 dB (« 0.4) resulted in a number of benefits. First, the left side content is not exactly out of phase with the right side content. Second, the mid content is not overly predominant (relative to the side content). Third, this setting aids in counter rotation cases since the mid content is carried with the side content that does not rotate. Fourth, this example setting preserves a sense that the center channel signal 106 is externalized — that is, when using the appropriate spatialization.
[0069] FIG. 4 is an example flow chart of a method 900 for rendering output. Referring to FIGS. 1-4, the method 900 includes, in step 902, generating a center channel signal 106, a left residual signal 108, and a right residual signal 110 based on a left binaural signal 102 and a right binaural signal 104.
[0070] The method 900 further includes, in step 904, generating, via a spatialization transformer 119, a spatialized center signal 112 based on the center channel signal 106.
[0071] The method 900 further includes, in step 906, rendering the output audio based on the spatialized center signal 112, the left residual signal 108, and the right residual signal 110.
[0072] According to an example, the method 900 further includes, in optional step 908, generating a delayed left signal 114 based on the left residual signal 108 and a first delay 121. The method 900 further includes, in optional step 910, generating a delayed right signal 116 based on the right residual signal 110 and a second delay 123. The output audio is rendered further based on the delayed left signal 114 and the delayed right signal 116.
[0073] According to an example, the method further includes, in optional step 912, adjusting the spatialization transformer 119 based on a motion signal 126.Docket No. BOSE-143WO / RS-25-443-WO
[0074] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0075] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0076] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified.
[0077] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.”
[0078] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified.
[0079] It should also be understood that, unless clearly indicated to the contrary, in any methods claimed herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.Docket No. BOSE-143WO / RS-25-443-WO
[0080] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively.
[0081] The above-described examples of the described subject matter can be implemented in any of numerous ways. For example, some aspects may be implemented using hardware, software, or a combination thereof. When any aspect is implemented at least in part in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single device or computer or distributed among multiple device s / computers .
[0082] The present disclosure may be implemented as a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0083] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0084] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to anDocket No. BOSE-143WO / RS-25-443-WOexternal computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0085] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user's computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some examples, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0086] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to examples of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0087] The computer readable program instructions may be provided to a processor of a, special purpose computer, or other programmable data processing apparatus to produce aDocket No. BOSE-143WO / RS-25-443-WOmachine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram or blocks.
[0088] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0089] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various examples of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0090] Other implementations are within the scope of the following claims and other claims to which the applicant may be entitled.
[0091] While various examples have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the examples described herein. More generally, those skilled in the art will readily appreciate thatDocket No. BOSE-143WO / RS-25-443-WOall parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific examples described herein. It is, therefore, to be understood that the foregoing examples are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, examples may be practiced otherwise than as specifically described and claimed. Examples of the present disclosure are directed to each individual feature, system, article, material, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, and / or methods, if such features, systems, articles, materials, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
Claims
Docket No. BOSE-143WO / RS-25-443-WOClaimsWhat we claim is:
1. A method for rendering output audio, comprising:generating a center channel signal, a left residual signal, and a right residual signal based on a left binaural signal and a right binaural signal;generating, via a spatialization transformer, a spatialized center signal based on the center channel signal; andrendering the output audio based on the spatialized center signal, the left residual signal, and the right residual signal.
2. The method for rendering output audio of claim 1, wherein the output audio comprises left output audio and right output audio.
3. The method for rendering output audio of claim 2, wherein the left output audio is based on the spatialized center signal and the left residual signal, and wherein the right output audio is based on the spatialized center signal and the right residual signal.
4. The method for rendering output audio of claim 2, wherein the left output audio is rendered by a left acoustic transducer, and wherein the right output audio is rendered by a right acoustic transducer.
5. The method for rendering output audio of claim 1, further comprising:generating a delayed left signal based on the left residual signal and a first time delay; andgenerating a delayed right signal based on the right residual signal and a second time delay, wherein the output audio is rendered further based on the delayed left signal and the delayed right signal.
6. The method for rendering output audio of claim 1, further comprising adjusting the spatialization transformer based on a motion signal.Docket No. BOSE-143WO / RS-25-443-WO7. The method for rendering output audio of claim 6, wherein the motion signal is provided by a motion sensor.
8. The method for rendering output audio of claim 1, wherein the left residual signal is generated based on the center channel signal and the left binaural signal.
9. The method for rendering output audio of claim 8, wherein the left residual signal is generated by subtracting the center channel signal from the left binaural signal.
10. The method for rendering output audio of claim 1, wherein the right residual signal is generated based on the center channel signal and the right binaural signal.
11. The method for rendering output audio of claim 10, wherein the right residual signal is generated by subtracting the center channel signal from the right binaural signal.
12. The method for rendering output audio of claim 1, wherein the center channel signal is generated by multiplying a sum of the left binaural signal and the right binaural signal by a predetermined center gain constant.
13. The method for rendering output audio of claim 1, wherein the center channel signal, the left residual signal, and the right residual signal are generated via linear source and / or channel separation.
14. The method for rendering output audio of claim 1, wherein the center channel signal, the left residual signal, and the right residual signal are generated via non-linear source and / or channel separation.
15. A binaural audio device, comprising a controller configured to:generate a center channel signal, a left residual signal, and a right residual signal based on a left binaural signal and a right binaural signal;generate, via a spatialization transformer, a spatialized center signal based on the center channel signal; andrender output audio based on the spatialized center signal, the left residual signal, and the right residual signal.Docket No. BOSE-143WO / RS-25-443-WO16. The binaural audio device of claim 15, wherein the output audio comprises left output audio and right output audio.
17. The binaural audio device of claim 16, wherein the left output audio is based on the spatialized center signal and the left residual signal, and wherein the right output audio is based on the spatialized center signal and the right residual signal.
18. The binaural audio device of claim 16, further comprising a left acoustic transducer configured to render the left output audio and a right acoustic transducer configured to render the right output audio.
19. The binaural audio device of claim 15, wherein the controller is further configured to:generate a delayed left signal based on the left residual signal and a first time delay; andgenerate a delayed right signal based on the right residual signal and a second time delay, wherein the output audio is rendered further based on the delayed left signal and the delayed right signal.
20. The binaural audio device of claim 15, wherein the controller is further configured to adjust the spatialization transformer based on a motion signal.
21. The binaural audio device of claim 20, further comprising a motion sensor configured to provide the motion signal.
22. The binaural audio device of claim 15, wherein the left residual signal is generated based on the center channel signal and the left binaural signal.
23. The binaural audio device of claim 22, wherein the left residual signal is generated by subtracting the center channel signal from the left binaural signal.
24. The binaural audio device of claim 15, wherein the right residual signal is generated based on the center channel signal and the right binaural signal.Docket No. BOSE-143WO / RS-25-443-WO25. The binaural audio device of claim 24, wherein the right residual signal is generated by subtracting the center channel signal from the right binaural signal.
26. The binaural audio device of claim 15, wherein the center channel signal is generated by multiplying a sum of the left binaural signal and the right binaural signal by a predetermined center gain constant.
27. The binaural audio device of claim 15, wherein the center channel signal, the left residual signal, and the right residual signal are generated via linear source and / or channel separation.
28. The binaural audio device of claim 15, wherein the center channel signal, the left residual signal, and the right residual signal are generated via non-linear source and / or channel separation.