Generation of interactive audio content

By performing beamforming and spatial matching on audio channel recordings from multiple devices and combining them with timbre matching, the method generates a perceptually-matched audio stream that provides an immersive, interactive binaural audio experience.

WO2025111240A1PCT designated stage expired Publication Date: 2025-05-30DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/056469
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2024-11-19
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Generating and rendering immersive, user-generated audio content that is spatially aware and dynamically adjusts based on the listener's head movement is challenging due to differences in device characteristics between multiple audio capture devices.

Method used

The method involves obtaining audio channel recordings from multiple devices, performing beamforming and spatial matching to generate a spatially-processed set of recordings, and then combining these with timbre matching to create a perceptually-matched audio stream that can be rendered as interactive binaural audio content.

Benefits of technology

This approach allows for seamless blending of audio content from multiple devices, providing an immersive experience by accurately simulating spatial audio and adapting to the listener's head orientation in real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024056469_30052025_PF_FP_ABST
    Figure US2024056469_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for generating audio streams for immersive audio content are provided. In some embodiments, the techniques involve obtaining a first set of audio channel recordings from a first audio capture device and a second set of audio channel recordings from a second audio capture device. The techniques may involve performing beamforming using the first set of audio channel recordings to generate a spatially-processed set of audio channel recordings. The techniques may involve performing spatial matching and timbre matching on the spatially-processed set of audio channel recordings using information associated with the second set of audio channel recordings to generate a matched set of audio channel recordings. The techniques may involve combining the matched set of audio channel recordings with the second set of audio channel recordings to generate a perceptually-matched audio stream.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATION OF INTERACTIVE AUDIO CONTENTCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to US provisional applications Nos. 63 / 612,281, filed December 19, 2023, and 63 / 684,338, filed on August 16, 2024, and PCT Application No. PCT / CN2023 / 132873, filed on November 21, 2023, all of which are incorporated herein by reference in their entirety.TECHNICAL FIELD

[0002] This disclosure pertains to systems, methods, and media for generation of interactive audio content.BACKGROUND

[0003] Content creators are increasingly creating sophisticated user-generated content. Such user-generated content may be immersive, e.g., such that audio content is to be rendered based on an orientation of a listener’ s head and which dynamically changes responsive to movement of the listener’s head. However, generating and rendering such immersive content is difficult.NOTATION AND NOMENCLATURE

[0004] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.

[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).

[0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.

[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.SUMMARY

[0008] Techniques for generating audio streams for immersive audio content are provided. Various embodiments described herein provide methods, systems and devices that apply the described techniques for generating audio streams.

[0009] In some other embodiments, a method may involve obtaining a first set of audio channel recordings from a first audio capture device and a second set of audio channel recordings from a second audio capture device. The method may further involve performing beamforming using the first set of audio channel recordings to generate a spatially-processed set of audio channel recordings associated with the first audio capture device. The method may further involve performing spatial matching and timbre matching on the spatially -processed set of audio channel recordings associated with the first audio capture device using information associated with the second set of audio channel recordings from the second audio capture device to generate a matched set of audio channel recordings associated with the first audio capture device. The method may further involve combining the matched set of audio channel recordings associated with the first audio capture device with the second set of audio channel recordings from the second audio capture device to generate a perceptually-matched audio stream.

[0010] In some examples, the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

[0011] In some examples, the information associated with the second set of audio channel recordings comprises a modified binaural signal generated based on a direction of arrival of the second set of audio channel recordings. In some examples, the modified binaural signal is generated based on the direction of arrival and a head-related transfer function.

[0012] In some examples, performing the spatial matching comprises adjusting levels of the spatially-processed set of audio channel recordings. In some examples, adjusting the levels of the spatially-processed set of audio channel recordings comprises setting the levels based on levels associated with the second set of audio channel recordings. In some examples, the levels associated with the second set of audio channel recordings comprise level differences between channels.

[0013] In some examples, performing timbre matching comprises performing dynamic equalization on the spatially-processed set of audio channel recordings based on the second set of audio channel recordings.

[0014] In some examples, the method may further involve providing the perceptually-matched audio stream to an audio playback device.

[0015] According to some embodiments, a method of rendering and presenting binaural audio content may involve receiving, at an audio playback device, a perceptually-matched audio stream, wherein the perceptually-matched audio stream comprises a matched set of audio channel recordings corresponding to audio content recorded by a first audio capture device, and a second set of audio channel recordings corresponding to audio content recorded by a second audio capture device, and wherein the matched set of audio channel recordings have been spatially matched and timbre matched based on information associated with the second set of audio channel recordings. The method may further involve obtaining headtracking data from one or more sensors disposed in or on a set of earbuds or headphones paired with the audio playback device. The method may further involve mixing the perceptually-matched audio stream based on the headtracking data to generate rendered binaural audio content. The method may further involve presenting the rendered binaural audio content via the set of earbuds or headphones.

[0016] In some examples, mixing the perceptually-matched audio stream based on the headtracking data comprises generating mixing weights based on head orientation of a user in a horizontal plane and applying the mixing weights to perceptually-matched audio stream.

[0017] In some examples, the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

[0018] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.

[0019] According to some embodiments, an audio capture device may comprise one or more processors and one or more processor-readable media. The one or more processor-readable media may store instructions which, when executed by the one or more processors, cause the device to obtain a first set of audio channel recordings from the audio capture device and a second set of audio channel recordings from a second audio capture device. The instructions may cause the device to perform beamforming using the first set of audio channel recordings to generate a spatially-processed set of audio channel recordings associated with the audio capture device. The instructions may cause the device to perform spatial matching and timbre matching on the spatially-processed set of audio channel recordings associated with the audio capture device using information associated with the second set of audio channel recordings from the second audio capture device to generate a matched set of audio channel recordings associated with the audio capture device. The instructions may cause the device to combine the matched set of audio channel recordings associated with the audio capture device with the second set of audio channel recordings from the second audio capture device to generate a perceptually-matched audio stream.

[0020] According to some other embodiments, an audio playback device may comprise one or more processors and one or more processor-readable media. The one or more processor-readable media may store instructions which, when executed by the one or more processors, cause the device to receive a perceptually-matched audio stream, wherein the perceptually-matched audio stream comprises a matched set of audio channel recordings corresponding to audio contentrecorded by a first audio capture device, and a second set of audio channel recordings corresponding to audio content recorded by a second audio capture device, and wherein the matched set of audio channel recordings have been spatially matched and timbre matched based on information associated with the second set of audio channel recordings. The instructions may cause the device to obtain headtracking data from one or more sensors disposed in or on a set of earbuds or headphones paired with the audio playback device. The instructions may cause the device to mix the perceptually-matched audio stream based on the headtracking data to generate rendered binaural audio content. The instructions may cause the device to present the rendered binaural audio content via the set of earbuds or headphones.

[0021] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.

[0022] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a block diagram depicting example components of an example audio capture device and an example audio playback device in accordance with some embodiments.

[0024] Figures 2A and 2B illustrate example implementations of spatial processing and perceptual matching as implemented on an audio capture device in accordance with some embodiments.

[0025] Figure 3 is a block diagram of an example system for rendering interactive binaural content in accordance with some embodiments.

[0026] Figure 4A is a flowchart of an example process for generating matched audio channel recordings across multiple audio capture devices in accordance with some embodiments.

[0027] Figure 4B is a flowchart of an example process for rendering binaural audio content as interactive audio content in accordance with some embodiments.

[0028] Figure 5A shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.

[0029] Figure 5B illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.

[0030] Figure 5C illustrates a schematic block diagram of an example CPU implemented in the device architecture of Figure 5B that may be used to implement various aspects of the present disclosure.

[0031] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION OF EMBODIMENTS

[0032] Content creators are increasingly creating sophisticated user-generated content. Such user-generated content may be immersive, e.g., such that audio content is to be with audio objects rendered and perceived to be at particular spatial locations within a three-dimensional frame surrounding the listener. However, generating and rendering such immersive content is difficult. For example, content creators may create user-generated content using multiple capture devices. In one example, a content creator may use a first audio capture device, such as a mobile phone, to capture audio content using one or more (e.g., two, three, etc.) microphones of the mobile phone. The content creator may simultaneously use a second audio capture device, such as a set of earbuds, headphones, smart glasses, AR / VR headset, etc. with microphones disposed in or on the second audio capture device to capture audio content. Use of these multiple audio capture devices may allow interactive audio content to be generated. As used herein, “interactive” audio content refers to audio content that is rendered such that audio objects may be perceived as being in particular spatial locations, and where the spatial perception is adjusted in real-time or near real-time based on the head orientation of a listener (e.g., as a listener moves their head). However, because the multiple audio capture devices have different device characteristics (e.g., due to frequency characteristics of the microphones, microphone locations, etc.), the audio content captured by each audio capture device may be substantially different in perceptual characteristics (e.g., spatialcharacteristics, timbre, etc.). Accordingly, listening to interactive audio content captured by audio capture devices having different characteristics may lead to a jarring experience for the listener.

[0033] Disclosed herein are techniques for performing perceptual matching on audio content recorded by multiple audio capture devices. Generally, the techniques disclosed herein are discussed with regard to a first audio capture device that is a mobile phone, and a second audio capture device that is a set of earbuds or headphones with microphones disposed in or on the earbuds or headphones, however, this is merely one example. Other examples of audio capture devices include a tablet computer, a laptop computer, an AR / VR headset, smart glasses, an audio conferencing or video conferencing system, etc.

[0034] Briefly stated, techniques for obtaining a first set of audio channel recordings from a first audio capture device and a second set of audio channel recordings from a second audio capture device are described. The techniques may involve performing beamforming using the first set of audio channel recordings to generate a spatially -processed set of audio channel recordings. The techniques may involve performing spatial matching and timbre matching on the spatially- processed set of audio channel recordings using information associated with the second set of audio channel recordings to generate a matched set of audio channel recordings. The techniques may involve combining the matched set of audio channel recordings with the second set of audio channel recordings to generate a perceptually-matched audio stream.

[0035] As described below in more detail below in connection with Figure 1, the audio content captured by a first audio capture device may be spatially matched and timbre matched to the audio content captured by a second audio capture device. For example, levels of the audio content may be adjusted based on inter-channel level differences in order to spatially match the audio content of the two devices. As another example, spectral shape may be matched for the audio content in order to match timbre characteristics of the two devices. Note that as used herein, perceptual matching may involve spatial matching and timbre matching. By performing such perceptual matching, the spatial and timbre characteristics of audio content captured by multiple devices may be aligned such that when rendered as interactive binaural content by an audio playback device, the content from the two capture devices is seamlessly blended to create an immersive experience.

[0036] Figure 1 is a block diagram depicting example components of an example audio capture device 102 and an example audio playback device 110 in accordance with some embodiments. Audio capture device 102 includes a spatial processing block 106 and a perceptual matching block 108. Audio playback device 110 includes a mixing block 112.

[0037] Spatial processing block 106 receives, as input, a set of audio channel recordings 104. Audio channel recordings 104 may comprise N audio channels (e.g., recorded from microphones of a mobile phone) and two binaural channels (e.g., recorded from microphones of a set of earbuds). Spatial processing block 106 generates at least a beamformed signal 107 which is provided as input to perceptual matching block 108. Perceptual matching block 108 generates, as output, a perceptually matched immersive audio stream 109. Audio playback device 110 is configured to receive the perceptually matched immersive audio stream 109, which may be received as input by mixing block 112. Mixing block 112 may optionally receive listener behavior information 114. Mixing block 112 may generate, as output, an interactive binaural audio output 115.

[0038] Referring to audio capture device 102, as illustrated, audio capture device 102 may obtain a set of audio channel recordings 104. The audio channel recordings may be obtained from multiple audio capture devices, such as a mobile phone, a set of earbuds with microphones disposed in or on the earbuds, a set of smart glasses with microphones disposed in or on the smart glasses, etc. In the example shown in Figure 1, audio channel recordings 104 includes two binaural channels captured from microphones of a set of earbuds, and N channels captured from microphones of a mobile phone, where the mobile phone corresponds to audio capture device 102. The audio channel recordings 104 are provided to spatial processing block 106. As described below in more detail in connection with Figures 2A and 2B, spatial processing block 106 is configured to generate at least a beamformed signal that simulates head rotation while capturing the audio channel recordings. In some embodiments, beamforming may be performed based on the N channels recorded using the mobile device. As described below in more detail in connection with Figures 2A, spatial processing block 106 may additionally generate a modified binaural signal based on a direction of arrival (DOA) estimation and a head-related transfer function (HRTF) that additionally simulates head rotation while capturing the audio channel recordings.

[0039] The output 107 of spatial processing block 106, which may include the beamformed signal and optionally a modified binaural signal (or an un-modified version of the binaural signal) may be provided as input to perceptual matching block 108. Perceptual matching block 108 is configured to match the binaural signal (or the modified binaural signal) and the beamformed signal. In particular, as described below in connection with Figures 2A and 2B, perceptual matching may involve spatial matching between the beamformed signal and the binaural signal (or the modified binaural signal). Perceptual matching may additionally or alternatively involve timbre matching. In general, timbre matching may involve dynamic equalization such that thespectrum of the beamformed signal after timbre matching is similar in shape to that of the binaural signal (or the modified binaural signal).

[0040] The perceptually matched audio stream 109 generated by audio capture device 102 may be received by an audio playback device 110. Note that audio playback device 110 may receive the perceptually matched audio stream directly from audio capture device 102 or via an intermediary device, such as a server or cloud device. Audio playback device 110 may use mixing block 112 to render the perceptually matched audio stream as interactive binaural content based on listener behavior information 114. For example, listener behavior information 114 may include listener head orientation information. More detailed techniques for rendering the audio content are shown in and described below in connection with Figures 3 and 4B.

[0041] As described above, an audio capture device may perform spatial processing and perceptual matching based on the spatial processing. In some embodiments, spatial processing may involve estimating a DOA associated with a subset of the audio channels, e.g., binaural audio content captured by a pair of earbuds. The spatial processing may involve generating a modified binaural signal, where the modified binaural signal is generated based on the estimated DOA. For example, the modified binaural signal may be generated by performing an HRTF adjustment using the estimated DOA. As a more particular example, the HRTF adjustment may simulate the binaural signal as captured at a set of predetermined head rotation angles to generate a spatial- processed binaural recording. Examples of predetermined head rotation angles include 90 degrees to the left, 90 degrees to the right, 80 degrees to the left, 80 degrees to the right, or the like. In some embodiments, a subset of audio channels (e.g., the N audio channels captured by a mobile phone) may be processed using beamforming. For example, fixed beamforming may be performed to generate a spatially-processed set of audio channels to simulate fixed beams pointing to predetermined angles. The predetermined angles may include 0 degrees to the front and 180 degrees to the back (e.g., in front of the listener’s head and behind the listener’s head), or the like.

[0042] Perceptual matching may be performed based on the spatially processed audio signals. For example, in some embodiments, perceptual analysis may be performed based on a modified binaural signal and the spatially-processed N audio channels (e.g., obtained by a mobile phone). The perceptual analysis may identify spatial and timbre features of the audio channels obtained by the two audio capture devices. Spatial matching may be performed to spatially match the N audio channels to the binaural audio channels, and timbre matching may be performed to match the spectral characteristics (e.g., shape) of the N audio channels to the binaural audio channels. Thespatially matched and timbre matched N audio channels and binaural audio channels may then be multiplexed to form the perceptually matched immersive audio stream.

[0043] Figure 2A illustrates an example implementation of spatial processing and perceptual matching as implemented on an audio capture device 102 in accordance with some embodiments. As illustrated, audio capture device 102 may include a spatial processing block 202 and a perceptual matching block 210. Spatial processing block 202 may include a beamforming block 204, a spatial analysis block 206, and a head-related transfer function (HRTF) adjustment block 208. Perceptual matching block 210 may include a perceptual analysis block 212, a spatial matching block 214, a timbre matching block 216, and a mixing block 218.

[0044] Beamforming block 204 of spatial processing block 202 receives, as input, N audio channels (e.g., from audio channel recordings 104) and generates a beamformed signal 205 as output. Spatial analysis block 206 of spatial processing block 202 receives two binaural channels and the N audio channels as input and generates an output 207 representing an estimated direction of arrival (DOA) associated with the audio signal. The output 207 of spatial analysis block 206 is provided to HRTF adjustment block 208. HRTF adjustment block 208 generates a modified binaural signal 209 based on the output 207 of spatial analysis block 206. The output 205 of beamforming block 204 is provided to perceptual analysis block 212 and spatial matching block 214 of perceptual matching block 210. The output 209 of HRTF adjustment block 208 is also provided to perceptual analysis block 212. Perceptual analysis block 212 is configured to generate spatial information and / or perceptual information 213 associated with the beamformed signal 205 generated by beamforming block 204 and / or the modified binaural signal 209 generated by HRTF adjustment block 208. The output 213 of perceptual analysis block 212 is optionally provided to spatial matching block 213 and timbre matching block 216. Spatial matching block 214 is configured to modify the beamformed signal based on spatial information received from perceptual analysis block 212 to generate a modified signal 215. Timbre matching block 216 receives modified signal 215 generated by spatial matching block 214 and is configured to perform timbre matching on the spatially matched beamformed signal to generate a timbre matched signal 217. The timbre matched signal 217 generated by timbre matching block 216 is provided to mixing block 218. Mixing block 218 also receives, as input, the two binaural audio channels (e.g., from audio channel recordings 104). Mixing block 218 generates, as output, a perceptually matched immersive audio stream 109.

[0045] Referring to spatial processing block 202, spatial processing block 202 may be configured to receive N audio channels from a first audio capture device (e.g., a mobile phone) and two binaural channels from a second audio capture device (e.g., from microphones disposed in or on a pair of earbuds or headphones). Spatial analysis block 206 may utilize the N audio channels and the two binaural channels to estimate a DOA associated with the audio signal which is represented in output 207. HRTF adjustment block 208 may be configured to generate a modified binaural signal 209 based on the DOA. In particular, the modified binaural signal 209 may be a simulated signal that simulates a set of predetermined head rotation angles. In other words, the modified binaural signal may simulate a listener associated with the audio capture device rotating their head as the binaural audio channels are captured. Beamforming block 204 may be configured to generate a spatially -processed signal 205 associated with the first audio capture device, for example, to simulate fixed beams pointing at predetermined angles, such as straight in front of and / or behind a listener associated with the first audio capture device, 45 degrees left and right, or the like. In the case of beams that are straight in front of and / or behind the listener, such beams may be used to simulate left and right head rotation.

[0046] The beamformed signal 205 and the modified binaural signal 209 may be provided to perceptual matching block 210. In particular, perceptual analysis block 212 may be configured to determine spatial characteristics and / or perceptual characteristics 213 associated with the beamformed signal and the modified binaural signal. Spatial matching block 214 may be configured to modify the beamformed signal based on the spatial characteristics to generate modified signal 215. In particular, in some embodiments, spatial matching block 214 may adjust the beamformed signal to be closer in direction to the direction associated with the modified binaural signal. To perform spatial matching, inter-aural level differences may be adjusted. For example, level differences may be adjusted to match those of the modified binaural signal. Because spatial cues are generally perceived based on inter-aural level differences at relatively high frequencies, level differences may be adjusted on a frequency-dependent basis for frequency channels above a predetermined frequency (e.g., above 1000 Hz, above 2000 Hz, etc.).

[0047] Timbre matching 216 may be configured to adjust timbre characteristics associated with the spatially-matched beamformed signal based on timbre characteristics of the modified binaural signal to generate timbre matched signal 217. Timbre matching may involve performing a dynamic equalization process such that a spectrum of the spatially-matched beamformed signal is similar in shape to the spectrum of the modified binaural signal.

[0048] The output (e.g., timbre matched signal 217) of timbre matching block 216 is considered the spatially and perceptually matched version of the set of audio channels from the first audio capture device (e.g., a mobile phone), where matching has been performed relative to the audio content captured by the second audio capture device (e.g., a set of earbuds or headphones). The spatially and perceptually matched audio channels associated with the first audio capture device are multiplexed with the modified binaural signal by mixing block 218 to form the perceptually matched immersive audio stream 109. The audio stream may be saved, transmitted to a playback device, etc. Note that, in some embodiments, the audio stream may be packaged with corresponding video content.

[0049] In the implementation shown in and described above in connection with Figure 2A, spatial matching is performed by generating a modified binaural signal that is in turn generated based on an estimated DOA. Estimating the DOA for the binaural signal can be computationally intensive. Accordingly, in some embodiments, spatial and timbre matching may be performed using the binaural signal without generating a modified binaural signal based on the DOA. For example, in some embodiments, spatial matching may be performed using tuning parameters specific to the first audio capture device such that the N audio channels captured using microphones of the first audio capture device (e.g., the mobile phone) are adjusted based on the tuning parameters. As a more particular example, known average level differences of a beamformed signal may be used to adjust a beamformed signal, rather than adjusting the beamformed signal based on the modified binaural signal, as shown in and described above in connection with Figure 2A. Perceptual matching may then be performed on the spatially-processed beamformed signal similar to what is described above in connection with Figure 2A.

[0050] Figure 2B illustrates another example implementation of spatial processing and perceptual matching as implemented on an audio capture device 102 in accordance with some embodiments. In the example shown in Figure 2B, spatial matching is performed without determining an estimated DOA of the binaural signal. As illustrated, audio capture device 102 may include a spatial processing block 252 and a perceptual matching block 260. Spatial processing block 262 may include a beamforming block 204. Perceptual matching block 260 may include a perceptual analysis block 262, a spatial matching block 214, a timbre matching block 216, and a mixing block 218.

[0051] Beamforming block 204 of spatial processing block 252 receives, as input, N audio channels (e.g., from audio channel recordings 104) and generates a beamformed signal 205 asoutput. The beamformed signal 205 of beamforming block 204 is provided as input to both perceptual analysis block 262 and spatial matching block 214 of perceptual matching block 260. Perceptual analysis block 262 is configured to generate spatial information and / or perceptual information 263 associated with the beamformed signal 205 generated by beamforming block 204 and / or the two binaural audio channels. The output 263 of perceptual analysis block 262 is optionally provided to timbre matching block 216. Spatial matching block 214 is configured to modify the beamformed signal 205 to generate a spatially matched signal 215. Timbre matching block 216 receives spatially matched signal 215 of spatial matching block 214 and is configured to perform timbre matching to generate a timbre matched signal 217. The timbre matched signal 217 is provided to mixing block 218. Mixing block 218 also receives, as input, the two binaural audio channels. Mixing block 218 generates, as output, a perceptually matched immersive audio stream 109.

[0052] Referring to spatial processing block 252, spatial processing block 252 is configured to receive N audio channels recorded using a first audio capture device (e.g., a mobile phone) and two binaural channels recorded using a second audio capture device (e.g., a set of earbuds or headphones). Unlike what is shown in and described above in connection with Figure 2A, a modified binaural signal is not generated based on an estimated DOA of the binaural signal. As illustrated in Figure 2B, beamforming block 204 generates a beamformed signal 205, which is provided as input to perceptual matching block 260.

[0053] Unlike what is shown in and described above in connection with Figure 2A, perceptual analysis block 262 does not receive a modified binaural signal, and instead receives the original binaural signal and the beamformed signal 205. Spatial matching block 214 performs spatial matching based on tuning parameters and / or device information associated with the first audio capture device to generate a spatially matched signal 215. The tuning parameters and / or the device information may include information such as relative locations of microphones of the first audio capture device. Rather than spatially matching the beamformed signal 205 to a representation or version of the binaural signal as shown in and described above in connection with Figure 2A, spatial matching block 214 may use the tuning parameters and / or device information to adjust level differences in the beamformed signal. Timbre matching and mixing may be performed by timbre matching block 216 and mixing block 218, similar to what is shown in and described above in connection with Figure 2A. In particular, timbre matching block 216 generates a timbre matched signal 217. Mixing block 218 receives the timbre matched signal 217 and generates perceptually matched immersive audio stream 109 as an output.

[0054] In some embodiments, an audio playback device may render the perceptually matched audio stream to generate an interactive binaural audio output. In particular, the interactive binaural audio input may dynamically adapt to the listener’s head orientation such that audio objects are rendered spatially in a manner that is responsive to the listener’s head orientation. The interactive binaural audio output may be rendered by mixing the N audio channels captured by the first audio capture device and the two binaural audio channels captured by the second audio capture device based on the listener’s head orientation. Note that the listener’s head orientation may be determined based on one or more sensors (e.g., one or more accelerometers, one or more gyroscopes, or any combination thereof) disposed in or on earbuds or headphones worn by the listener. Mixing may involve applying weights to different audio channels based on the angle of the listener’s head, e.g., within a horizontal plane.

[0055] By way of example, the interactive binaural audio output, represented herein as Y(n), where n represents the audio frame index, may be determined by:T(n) = X(n) (n)

[0056] Because Y(n) is a binaural output (i.e., with two audio channels), where Y is an L x 2 matrix, and L is the number of samples in an audio frame. In the equation given above, X(n) is an L x K matrix of audio samples of K perceptually matched channels. Note that K generally represents the total number of audio channels recorded in total by the first audio capture device and the second audio capture device. In the example in which the second audio capture device (e.g., a pair of earbuds) captures two audio channels, and the first audio capture device (e.g., a mobile phone) captures N audio channels (as shown in and described above in connection with Figures 1, 2A, and 2B), K is N + 2. In an example in which N = 2 and in which the second audio capture device captures two audio channels, K is 4. The weight matrix, represented in the equation above as W(n), is a Kx 2 matrix, where the entries are dependent on the listener’ s head orientation. In one example, the weights in weight matrix W(n) may depend on the angle 0n, which represents the angle of the listener’s head in the horizontal plane. In such an example, and in which K is 4, an example weight matrix is:W (n) = max (

[0057] Figure 3 is a block diagram of an example system 300 for rendering interactive binaural content in accordance with some embodiments. System 300 may be implemented as part of an audio playback device (e.g., audio playback device 110 of Figure 1). System 300 includes a mixing block 302. Mixing block 302 may receive, as input, a perceptually matched audio stream 109 (e.g., as generated by an audio capture device). Mixing block 302 may additionally receive, as input, head orientation data 304. Mixing block 302 may generate, as output, interactive binaural audio content 115 by applying head orientation data 304 to render the perceptually matched audio stream. Mixing block 302 may receive head orientation data 304, e.g., via one or more sensors disposed in or on earbuds or headphones operatively coupled to the audio playback device (e.g., via a wireless or wired communication channel). Mixing block 302 may then apply a weight matrix to audio channels included in the perceptually matched audio stream 301 to generate interactive binaural audio content 115. Because the weight matrix is dependent on head orientation data 304 (e.g., as described above), application of the weight matrix to the audio channels causes the resulting binaural audio content to be dynamically adjusted based on the listener’s head orientation.

[0058] Figure 4A is a flowchart of an example process 400 for generating a perceptually- matched audio stream in accordance with some embodiments. Blocks of process 400 may be executed by one or more processors of an audio capture device, such as a mobile phone, a tablet computer, etc. Note that the audio capture device may be paired with a second audio capture device such as a pair of earbuds, a set of smart glasses, etc. An example of a control system that may be used to execute blocks of process 400 is control system 510 of Figure 5A.

[0059] Process 400 includes various blocks or steps as illustrated by blocks 402, 404, 406 and 408. The blocks of process 400 may arranged differently from what is illustrated in Figure 4A. For example, some blocks may be executed in an order other than what is shown, while in other examples two or more blocks of process 400 may be omitted, replaced, and / or split into additional blocks without departing from the present disclosure. Additionally, various blocks may be executed substantially in parallel. An example process 400 may commence at 402.

[0060] At 402, process 400 can obtain a first set of audio channel recordings from a first audio capture device and a second set of audio channel recordings from a second audio capture device. Examples of the first audio capture device and the second audio capture device include a mobile phone, a tablet computer, a set of earbuds or headphones with microphones disposed in or on the earbuds or headphones, a pair of smart glasses with microphones disposed in or on the smart glasses, an AR / VR headset, etc. Note that the first audio capture device and the second audiocapture device may be paired with each other such that a communication channel operatively couples the two devices (e.g., via a wireless communication channel such as BLUETOOTH). In some embodiments, the first audio capture device may perform blocks of process 400. The total number of audio channel recordings is generally referred to herein as K, with N audio channel recordings recorded by the first audio capture device. In the examples shown in and described above in connection with Figures 1, 2A, and 2B, the second audio capture device records two binaural audio channels, such that K = N + 2. For example, the second audio capture device may be a set of earbuds with one microphone disposed in each earbud.

[0061] At 404, process 400 can perform beamforming using the first set of audio channel recordings to generate a spatially-processed set of audio channel recordings associated with the first audio capture device. For example, as described above in connection with Figures 2A and 2B, the beamformed signal may simulate beams extending to predetermined angles, such as 0 degrees and 180 degrees (e.g., straight in front of and / or behind a user associated with capture of the audio content), and / or any other predetermined angles.

[0062] At 406, process 400 can perform spatial matching and timbre matching on the spatially- processed set of audio channel recordings associated with the first audio capture device using information associated with the second set of audio channel recordings from the second audio capture device to generate a matched set of audio channel recordings associated with the first audio capture device. For example, spatial matching may be performed by matching spatial characteristics of the audio channel recordings associated with the first audio captured device to spatial characteristics of the audio channel recordings of the second audio capture device. As another example, timbre matching may be performed by matching timbre characteristics of the audio channel recordings of the first audio capture device to timbre characteristics of the audio channel recordings of the second audio capture device. As illustrated in and described above in connection with Figure 2A, in some embodiments, spatial matching may be performed on a modified binaural signal generated based on the audio channel recordings of the second audio capture device. For example, the modified binaural signal may simulate head rotation, e.g., to the left and / or to the right. As illustrated in and described above in connection with Figure 2B, in some embodiments, spatial matching may be performed based on device information associated with the first audio capture device, such as relative locations of the microphones of the first audio capture device. In some implementation, spatial matching may involve adjusting levels based on level differences, e.g., to account for spatial perception based on inter-aural level differences. In some embodiments, timbre matching may involve dynamic equalization of at least a portion of thespectrum of the audio channel recordings associated with the first audio capture device to match a shape of the spectrum of the audio channel recordings associated with the second audio capture device.

[0063] At 408, process 400 can combine the matched set of audio channel recordings associated with the first audio capture device with the second set of audio channel recordings from the second audio capture device to generate a perceptually matched audio stream. For example, the matched set of audio channel recordings may be multiplexed with the second set of audio channel recordings. Note that in some embodiments, the perceptually matched audio stream may be encoded in any suitable manner. The perceptually matched audio stream may be stored, transmitted to another device (e.g., a server device, a cloud device, an audio playback device, etc.). The perceptually matched audio stream may be stored in connection with video content.

[0064] Figure 4B is a flowchart of an example process 450 for rendering and playback of a perceptually matched audio stream in accordance with some embodiments. Blocks of process 450 may be executed by one or more processors of an audio playback device, such as a mobile phone, a tablet computer, etc. An example of a control system that may be used to execute blocks of process 450 is control system 510 of Figure 5A.

[0065] Process 450 includes various blocks or steps as illustrated by blocks 452, 454, 456, and 458. The blocks of process 450 may be arranged differently from what is illustrated in Figure 4B. For example, some blocks may be executed in an order other than what is shown, while in other examples two or more blocks of process 450 may be omitted, replaced, and / or split into additional blocks without departing from the present disclosure. Additionally, various blocks may be executed substantially in parallel. An example process 450 may commence at 452.

[0066] At 452, process 450 receives, at an audio playback device, a perceptually matched audio stream. The perceptually matched audio stream may include a matched set of audio channel recordings corresponding to audio content recorded by a first audio capture device, and a second set of audio channel recordings corresponding to audio content recorded by a second audio capture device. For example, the first audio capture device may be a mobile phone, and the second audio capture device may be a set of earbuds, headphones, smart glasses, AR / VR headset, etc. paired with the mobile phone, as shown in and described above in connection with Figures 1, 2A, and 2B. The matched set of audio channel recordings may have been spatially matched and timbre matched based on the information associated with the second set of audio channel recordings. Forexample, matching may have been performed by the first audio capture device, as shown in and described above in connection with Figures 1, 2A, 2B, and 4A.

[0067] At 454, process 450 can obtain headtracking data from one or more sensors disposed in or on a set of earbuds or headphones paired with the audio playback device. For example, the sensors may include one or more accelerometers, gyroscopes, or any combination thereof. The headtracking data may indicate an orientation of the listener’ s head. For example, orientation may be indicated by rotation around one or more axes (e.g., X, Y, and / or Z axes), yaw, pitch and roll angles, or the like. Note that the head orientation may be determined by the audio playback device paired with the earbuds or headphones. For example, the earbuds or headphones may transmit, via a wireless or wired communication channel, the sensor data to the audio playback device, and the audio playback device may determine the head orientation based on the received sensor data.

[0068] At 456, process 450 can mix the perceptually matched audio stream based on the headtracking data to generate rendered binaural audio content. For example, as shown in and described above in connection with Figure 3, process 450 can determine weights of a weight matrix based on an angle of the listener’s head with respect to one or more planes or axes, e.g., the horizontal plane. The weights may then be applied to the matched set of audio channel recordings and the second set of audio channel recordings (which may total K channels in total, as described above in connection with Figure 4A) to generate two binaural audio channels that are determined based on the headtracking data.

[0069] At 458, process 450 can present the rendered binaural audio content via the set of earbuds or headphones. For example, process 450 can transmit instructions to play the rendered binaural audio content.

[0070] Process 450 can loop back to 454 and can obtain updated headtracking data. In this way, process 450 can loop through blocks 454-458 such that the perceptually matched audio stream is rendered and presented by the audio playback device as interactive binaural content based on the headtracking data (e.g., the listener’s head orientation).

[0071] Figure 5A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 5A are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 500 may be configured for performing at least someof the methods disclosed herein. In some implementations, the apparatus 500 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0072] According to some alternative implementations the apparatus 500 may be, or may include, a server. In some such examples, the apparatus 500 may be, or may include, an encoder. Accordingly, in some instances the apparatus 500 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 500 may be a device that is configured for use in “the cloud,” e.g., a server.

[0073] In this example, the apparatus 500 includes an interface system 505 and a control system 510. The interface system 505 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 505 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may, in some examples, pertain to one or more software applications that the apparatus 500 is executing.

[0074] The interface system 505 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.

[0075] The interface system 505 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 505 may include one or more wireless interfaces. The interface system 505 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 505 may include one or more interfaces between the control system 510 and a memory system, such as the optional memory system 515 shown in Figure 5A. However, the control system 510 may include a memory systemin some instances. The interface system 505 may, in some implementations, be configured for receiving input from one or more microphones in an environment.

[0076] The control system 510 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0077] In some implementations, the control system 510 may reside in more than one device. For example, in some implementations a portion of the control system 510 may reside in a device within one of the environments depicted herein and another portion of the control system 510 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 510 may reside in a device within one environment and another portion of the control system 510 may reside in one or more other devices of the environment. For example, a portion of the control system 510 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 510 may reside in another device that is implementing the cloudbased service, such as another server, a memory device, etc. The interface system 505 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.

[0078] In some implementations, the control system 510 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 510 may be configured for implementing methods of generating matched sets of audio channel recordings across different audio capture devices, rendering interactive audio content, or the like.

[0079] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 515 shown in Figure 5A and / or in the control system 510. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, filter audio signals to reduce noise, determine filter coefficients to perform such filtering, etc. The software may, forexample, be executable by one or more components of a control system such as the control system 510 of Figure 5A.

[0080] In some examples, the apparatus 500 may include the optional microphone system 520 shown in Figure 5A. The optional microphone system 520 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 500 may not include a microphone system 520. However, in some such implementations the apparatus 500 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 510. In some such implementations, a cloud-based implementation of the apparatus 500 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 510.

[0081] According to some implementations, the apparatus 500 may include the optional loudspeaker system 525 shown in Figure 5A. The optional loudspeaker system 525 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 500 may not include a loudspeaker system 525. In some implementations, the apparatus 500 may include headphones. Headphones may be connected or coupled to the apparatus 500 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).

[0082] Figure 5B illustrates a schematic block diagram of an example device architecture 501 (in this example, an apparatus 501) that may be used to implement various aspects of the present disclosure. The apparatus 501 of Figure 5B is an instance of the apparatus 500 of Figure 5 A. Architecture 501 includes but is not limited to servers and client devices, systems, etc., which may be configured to perform the methods that are described with reference to any or all of Figures 4A and / or 4B. As shown, the architecture 501 includes central processing unit (CPU) 541, which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 542 or a program loaded from, for example, storage unit 548 to random access memory (RAM) 543. The CPU 541 may be, for example, an electronic processor 541. In these examples, the CPU 541 is an instance of the control system 510 of Figure 5A and the ROM 542 and RAM 543 are instances of the memory system 515. In RAM 543, the data required when CPU 541 performs the various processes is also stored, as required. CPU 541, ROM 542, and RAM 543 are connected to one another via bus 544. Input / output (I / O) interface 545 is alsoconnected to bus 544. The bus 544 and the I / O) interface 545 are instances of the interface system 505 of Figure 5A.

[0083] The following components are connected to VO interface 545: input unit 546, that may include a keyboard, a mouse, or the like; output unit 547 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 548 including a hard disk, or another suitable storage device; and communication unit 549 including a network interface card such as a network card (e.g., wired or wireless).

[0084] In some implementations, input unit 546 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0085] In some implementations, output unit 5947 include systems with various number of speakers. Output unit 547 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0086] In some embodiments, communication unit 549 is configured to communicate with other devices (e.g., via a network). Drive 550 is also connected to VO interface 545, as required. Removable medium 551, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 550, so that a computer program read therefrom is installed into storage unit 548, as required. A person skilled in the art would understand that although apparatus 501 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0087] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 549, and / or installed from the removable medium 551, as shown in Figure 5B.

[0088] Figure 5C illustrates a schematic block diagram of an example CPU 541 implemented in the device architecture 501 of Figure 5B that may be used to implement various aspects of the present disclosure. The CPU 541 includes an electronic processor 560 and a memory 561. Theelectronic processor 560 is electrically and / or communicatively connected to the memory 561 for bidirectional communication. The memory 561 may store capture processing software 562 and / or playback processing software 563. The memory 561 may be, for example, a ROM, a RAM, or another non-transitory computer readable medium. The electronic processor 560 may implement the capture processing software 562 stored in the memory 561 to perform, among other things, the method 400 of Figure 4A. Additionally, the electronic processor 560 may implement the playback processing software 563 stored in the memory 561 to perform, among other things, the method 450 of Figure 4B4B.

[0089] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 541 in combination with other components of Figure 5B), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0090] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0091] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a varietyof operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0092] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.

[0093] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):

[0094] EEE 1. A method of generating audio streams for immersive audio content, the method comprising: obtaining a first set of audio channel recordings from a first audio capture device and a second set of audio channel recordings from a second audio capture device (104, 402)', performing beamforming using the first set of audio channel recordings to generate a spatially- processed set of audio channel recordings associated with the first audio capture device (404)', performing spatial matching and timbre matching on the spatially -processed set of audio channel recordings associated with the first audio capture device using information associated with the second set of audio channel recordings from the second audio capture device to generate a matched set of audio channel recordings associated with the first audio capture device (406); and combining the matched set of audio channel recordings associated with the first audio capture device with thesecond set of audio channel recordings from the second audio capture device to generate a perceptually-matched audio stream (408).

[0095] EEE 2. The method of EEE 1, wherein the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

[0096] EEE 3. The method of any one of EEEs 1 or 2, wherein the information associated with the second set of audio channel recordings comprises a modified binaural signal generated based on a direction of arrival of the second set of audio channel recordings.

[0097] EEE 4. The method of EEE 3, wherein the modified binaural signal is generated based on the direction of arrival and a head-related transfer function.

[0098] EEE 5. The method of any one of EEEs 1-4, wherein performing the spatial matching comprises adjusting levels of the spatially-processed set of audio channel recordings.

[0099] EEE 6. The method of EEE 5, wherein adjusting the levels of the spatially-processed set of audio channel recordings comprises setting the levels based on levels associated with the second set of audio channel recordings.

[0100] EEE 7. The method of EEE 6, wherein the levels associated with the second set of audio channel recordings comprise level differences between channels.

[0101] EEE 8. The method of any one of EEEs 1-7, wherein performing timbre matching comprises performing dynamic equalization on the spatially -processed set of audio channel recordings based on the second set of audio channel recordings.

[0102] EEE 9. The method of any one of EEEs 1-8, further comprising providing the perceptually-matched audio stream to an audio playback device.

[0103] EEE 10. A method of rendering and presenting binaural audio content, the method comprising: receiving, at an audio playback device (110), a perceptually-matched audio stream, wherein the perceptually-matched audio stream comprises a matched set of audio channel recordings corresponding to audio content recorded by a first audio capture device, and a second set of audio channel recordings corresponding to audio content recorded by a second audio capture device, and wherein the matched set of audio channel recordings have been spatially matched and timbre matched based on information associated with the second set of audio channel recordings(452) obtaining headtracking data (304) from one or more sensors disposed in or on a set of earbuds or headphones paired with the audio playback device (454); mixing the perceptually- matched audio stream based on the headtracking data to generate rendered binaural audio content (456); and presenting the rendered binaural audio content via the set of earbuds or headphones (458).

[0104] EEE 11. The method of EEE 10, wherein mixing the perceptually-matched audio stream based on the headtracking data comprises generating mixing weights based on head orientation of a user in a horizontal plane and applying the mixing weights to perceptually-matched audio stream.

[0105] EEE 12. The method of any one of EEEs 10 or 11, wherein the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

[0106] EEE 13. An apparatus configured for implementing the method of any one of EEEs 1- 12.

[0107] EEE 14. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of EEEs 1-12.

[0108] EEE 15. An audio capture device (102) comprising: one or more processors (541); and one or more processor-readable media (561 ) storing instructions which, when executed by the one or more processors, cause the device to: obtain a first set of audio channel recordings from the audio capture device and a second set of audio channel recordings from a second audio capture device (104, 402); perform beamforming using the first set of audio channel recordings to generate a spatially-processed set of audio channel recordings associated with the audio capture device (404); perform spatial matching and timbre matching on the spatially -processed set of audio channel recordings associated with the audio capture device using information associated with the second set of audio channel recordings from the second audio capture device to generate a matched set of audio channel recordings associated with the audio capture device (406); and combine the matched set of audio channel recordings associated with the audio capture device with the second set of audio channel recordings from the second audio capture device to generate a perceptually- matched audio stream (408).

[0109] EEE 16. The audio capture device of EEE 15, wherein the audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

[0110] EEE 17. The audio capture device of any one of EEEs 15 or 16, wherein the information associated with the second set of audio channel recordings comprises a modified binaural signal generated based on a direction of arrival of the second set of audio channel recordings.

[0111] EEE 18. An audio playback device (110) comprising: one or more processors (541 ); and one or more processor-readable media (561 ) storing instructions which, when executed by the one or more processors, cause the device to : receive a perceptually-matched audio stream, wherein the perceptually-matched audio stream comprises a matched set of audio channel recordings corresponding to audio content recorded by a first audio capture device, and a second set of audio channel recordings corresponding to audio content recorded by a second audio capture device, and wherein the matched set of audio channel recordings have been spatially matched and timbre matched based on information associated with the second set of audio channel recordings (452); obtain headtracking data (304) from one or more sensors disposed in or on a set of earbuds or headphones paired with the audio playback device (454); mix the perceptually-matched audio stream based on the headtracking data to generate rendered binaural audio content (456); and present the rendered binaural audio content via the set of earbuds or headphones (458).

[0112] EEE 19. The audio playback device of EEE 18, wherein mixing the perceptually- matched audio stream based on the headtracking data comprises generating mixing weights based on head orientation of a user in a horizontal plane and applying the mixing weights to perceptually- matched audio stream.

[0113] EEE 20. The audio playback device of any one of EEEs 18 or 19, wherein the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

[0114] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.

Claims

CLAIMS1. A method of generating audio streams for immersive audio content, the method comprising: obtaining a first set of audio channel recordings from a first audio capture device and a second set of audio channel recordings from a second audio capture device (104, 402)', performing beamforming using the first set of audio channel recordings to generate a spatially-processed set of audio channel recordings associated with the first audio capture device (404); performing spatial matching and timbre matching on the spatially-processed set of audio channel recordings associated with the first audio capture device using information associated with the second set of audio channel recordings from the second audio capture device to generate a matched set of audio channel recordings associated with the first audio capture device (406); and combining the matched set of audio channel recordings associated with the first audio capture device with the second set of audio channel recordings from the second audio capture device to generate a perceptually-matched audio stream (408).

2. The method of claim 1, wherein the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

3. The method of any one of claims 1 or 2, wherein the information associated with the second set of audio channel recordings comprises a modified binaural signal generated based on a direction of arrival of the second set of audio channel recordings.

4. The method of claim 3, wherein the modified binaural signal is generated based on the direction of arrival and a head-related transfer function.

5. The method of any one of claims 1-4, wherein performing the spatial matching comprises adjusting levels of the spatially-processed set of audio channel recordings.

6. The method of claim 5, wherein adjusting the levels of the spatially-processed set of audio channel recordings comprises setting the levels based on levels associated with the second set of audio channel recordings.

7. The method of claim 6, wherein the levels associated with the second set of audio channel recordings comprise level differences between channels.

8. The method of any one of claims 1-7, wherein performing timbre matching comprises performing dynamic equalization on the spatially-processed set of audio channel recordings based on the second set of audio channel recordings.

9. The method of any one of claims 1-8, further comprising providing the perceptually-matched audio stream to an audio playback device.

10. A method of rendering and presenting binaural audio content, the method comprising: receiving, at an audio playback device (110), a perceptually-matched audio stream, wherein the perceptually-matched audio stream comprises a matched set of audio channel recordings corresponding to audio content recorded by a first audio capture device, and a second set of audio channel recordings corresponding to audio content recorded by a second audio capture device, and wherein the matched set of audio channel recordings have been spatially matched and timbre matched based on information associated with the second set of audio channel recordings (452 ; obtaining headtracking data (304) from one or more sensors disposed in or on a set of earbuds or headphones paired with the audio playback device (454)', mixing the perceptually-matched audio stream based on the headtracking data to generate rendered binaural audio content (456 ; and presenting the rendered binaural audio content via the set of earbuds or headphones (458).

11. The method of claim 10, wherein mixing the perceptually-matched audio stream based on the headtracking data comprises generating mixing weights based on head orientation of a user in a horizontal plane and applying the mixing weights to perceptually-matched audio stream.

12. The method of any one of claims 10 or 11, wherein the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

13. An apparatus configured for implementing the method of any one of claims 1-12.

14. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of claims 1-1215. An audio capture device (102) comprising: one or more processors (541 ); and one or more processor-readable media (561 ) storing instructions which, when executed by the one or more processors, cause the device to: obtain a first set of audio channel recordings from the audio capture device and a second set of audio channel recordings from a second audio capture device (104, 402)', perform beamforming using the first set of audio channel recordings to generate a spatially- processed set of audio channel recordings associated with the audio capture device (404)', perform spatial matching and timbre matching on the spatially-processed set of audio channel recordings associated with the audio capture device using information associated with the second set of audio channel recordings from the second audio capture device to generate a matched set of audio channel recordings associated with the audio capture device (406); and combine the matched set of audio channel recordings associated with the audio capture device with the second set of audio channel recordings from the second audio capture device to generate a perceptually-matched audio stream (408).

16. The audio capture device of claim 15, wherein the audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

17. The audio capture device of any one of claims 15 or 16, wherein the information associated with the second set of audio channel recordings comprises a modified binaural signal generated based on a direction of arrival of the second set of audio channel recordings.

18. An audio playback device (110) comprising: one or more processors (541 ); and one or more processor-readable media (561 ) storing instructions which, when executed by the one or more processors, cause the device to : receive a perceptually-matched audio stream, wherein the perceptually-matched audio stream comprises a matched set of audio channel recordings corresponding to audio content recorded by a first audio capture device, and a second set of audio channel recordings corresponding to audio content recorded by a second audio capture device, and wherein thematched set of audio channel recordings have been spatially matched and timbre matched based on information associated with the second set of audio channel recordings (452); obtain headtracking data (304) from one or more sensors disposed in or on a set of earbuds or headphones paired with the audio playback device (454)', mix the perceptually-matched audio stream based on the headtracking data to generate rendered binaural audio content (456); and present the rendered binaural audio content via the set of earbuds or headphones (458).

19. The audio playback device of claim 18, wherein mixing the perceptually-matched audio stream based on the headtracking data comprises generating mixing weights based on head orientation of a user in a horizontal plane and applying the mixing weights to perceptually- matched audio stream.

20. The audio playback device of any one of claims 18 or 19, wherein the first audio capture device is a mobile device and the second audio capture device is a set of earbuds or headphones with one or more microphones disposed in or on the set of earbuds or headphones.

Citation Information

Patent Citations

  • Distributed Audio Capture and Mixing

    US20180310114A1

  • Crosstalk cancellation and adaptive binaural filtering for listening system using remote signal sources and on-ear microphones

    US20230319488A1