Transmission of interactive audio content
By encoding audio content into distinct streams with metadata for device compatibility, the method addresses the challenge of rendering interactive audio across various devices, ensuring both compatible and non-compatible devices can render appropriate audio content.
Patent Information
- Application Number
- PCT/US2025/035025
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-16
- Filing Date
- 2025-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies face challenges in packaging interactive audio content that requires rendering based on a listener's head orientation, as it is difficult to ensure compatibility with both compatible and non-compatible audio playback devices.
A method involving encoding audio content from multiple microphones into distinct streams, with metadata containing sensor and device information, allowing both compatible and non-compatible devices to render appropriate audio content, where compatible devices can reconstruct and render interactive audio based on headtracking data.
Enables both compatible and non-compatible devices to effectively render audio content, with compatible devices providing interactive immersive experiences and non-compatible devices rendering binaural audio, using a single audio package.
Smart Images

Figure US2025035025_02012026_PF_FP_ABST
Abstract
Description
TRANSMISSION OF INTERACTIVE AUDIO CONTENTCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from International Patent Application No. PCT / CN2024 / 101030, filed on 24 June 2024, and U.S. Provisional Application Ser. No. 63 / 684,144, filed on 16 August 2024, each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] This disclosure pertains to systems, methods, and media for transmission of interactive audio content.BACKGROUND
[0003] Content creators are increasingly creating sophisticated user-generated content. Such user-generated content may be interactive, e.g., such that audio content is to be rendered based on an orientation of a listener’s head and which dynamically changes responsive to movement of the listener’ s head. However, packaging such interactive audio content can be difficult.NOTATION AND NOMENCLATURE
[0004] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.
[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
[0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implementsa decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.
[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.SUMMARY
[0008] Techniques are described for processing immersive audio content into an audio package to facilitate rendering on a variety of devices, where the device may either be capable or incapable of rendering the immersive audio content. Various embodiments described herein provide methods, systems, and devices that apply the described techniques for processing the immersive audio content into an audio package.
[0009] In some embodiments, a method for transmitting audio content may involve capturing audio content from a plurality of microphones using at least one audio capture device. The method may further involve encoding a subset of audio content channels associated with a subset of the plurality of microphones to form a first audio stream and encoding the remaining subset of audio content channels associated with the remaining microphones of the plurality of microphones to form a second audio stream. The method may further involve determining sensor data associated with the at least one audio capture device and device information associated with the at least one audio capture device. The method may further involve storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream. The method may further involve generating an audio package comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non- interactive audio content.
[0010] In some examples, the method may further involve transmitting the audio package to a user device configured to render and playback audio content based on the first audio stream and the second audio stream.
[0011] In some examples, the sensor data comprises data indicative of an orientation of the at least one audio capture device during capture of the audio content.
[0012] In some examples, the device information comprises information indicative of relative microphone placement of the plurality of microphones.
[0013] In some examples, the at least one audio capture device comprises a set of earbuds or headphones with microphones disposed in or on the set of earbuds or headphones, and a mobile device. In some examples, the first audio stream comprises audio content captured using microphones of the set of earbuds or headphones. In some examples, the sensor data and device information are associated with the mobile device.
[0014] In some examples, the metadata further comprises beamforming coefficients usable to render the audio content as spatially interactive audio content by an audio playback device.
[0015] According to some embodiments, a method for presenting audio content may involve receiving, at an audio playback device, an audio package comprising a first audio stream and a metadata stream, wherein metadata of the metadata stream is usable to decode and reconstruct a second audio stream and perform spatially interactive rendering and playback of the first audio stream and the second audio stream. The method may further involve decoding the first audio stream to generate a first set of audio channels and generating a second set of audio channels associated with the second audio stream. The method may further involve recovering a full set of audio channels based on the first set of audio channels and the second set of audio channels using the metadata. The method may further involve obtaining headtracking data based on sensor data associated with a set of earbuds or headphones paired with the audio playback device. The method may further involve rendering the full set of audio channels based on the headtracking data and using sensor data encoded in the metadata to generate rendered audio content. The method may further involve presenting the rendered audio content using the set of earbuds or headphones.
[0016] In some examples, the metadata comprises sensor data and device information associated with an audio capture device that recorded the first set of audio channels and / or the second set of audio channels, and wherein the sensor data and / or the device information are used to render thefull set of audio channels. In some examples, the device information is used to determine beamforming coefficients to render the full set of audio channels.
[0017] In some examples, recovering the full set of audio channels comprises performing an inverse channel transform on the second set of audio channels using mixing coefficients and combining the first set of audio channels and a result of the inverse channel transform to obtain the full set of audio channels.
[0018] In some examples, rendering the full set of audio channels comprises performing beamforming using the second set of audio channels and the headtracking data to generate a beamformed second set of audio channels. In some examples, beamforming coefficients used to perform the beamforming are included in the metadata. In some examples, the method may further involve performing mixing of the first set of audio channels and the beamformed second set of audio channels to generate the rendered audio content.
[0019] In some embodiments, an audio capture device is described. An audio capture device may comprise: one or more processors; and one or more processor-readable media storing instructions. The instructions, when executed by the one or more processors, may cause performance of capturing audio content from a plurality of microphones. The instructions may further cause performance of encoding a subset of audio content channels associated with a subset of the plurality of microphones to form a first audio stream and encoding the remaining subset of audio content channels associated with the remaining microphones of the plurality of microphones to form a second audio stream; The instructions may further cause performance of determining sensor data associated with the at least one audio capture device and device information associated with the at least one audio capture device. The instructions may further cause performance of storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream. The instructions may further cause performance of generating an audio package comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non-interactive audio content.
[0020] In some embodiments, an audio playback device is described. An audio playback device may comprise: one or more processors; and one or more processor-readable media storing instructions. The instructions, when executed by the one or more processors, may cause performance of receiving an audio package comprising a first audio stream and a metadata stream,wherein metadata of the metadata stream is usable to decode and reconstruct a second audio stream and perform spatially interactive rendering and playback of the first audio stream and the second audio stream. The instructions may cause performance of decoding the first audio stream to generate a first set of audio channels and generating a second set of audio channels associated with the second audio stream. The instructions may cause performance of recovering a full set of audio channels based on the first set of audio channels and the second set of audio channels using the metadata. The instructions may cause performance of obtaining headtracking data based on sensor data associated with a set of earbuds or headphones paired with the audio playback device. The instructions may cause performance of rendering the full set of audio channels based on the headtracking data and using sensor data encoded in the metadata to generate rendered audio content. The instructions may cause performance of presenting the rendered audio content using the set of earbuds or headphones.
[0021] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.
[0022] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0023] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a block diagram of an example system for generating a package usable to generate interactive audio content in accordance with some embodiments.
[0025] Figure 2A is a block diagram of an example system associated with a non-compatible device for rendering non-interactive audio content in accordance with some embodiments.
[0026] Figure 2B is a block diagram of an example system associated with a compatible device for rendering and playback of interactive audio content in accordance with some embodiments.
[0027] Figure 3 is a block diagram of an example implementation of a channel recovery block in accordance with some embodiments.
[0028] Figure 4 is a block diagram of an example implementation of a spatial processing block of a compatible audio playback device in accordance with some embodiments.
[0029] Figure 5A is a flowchart of an example process for generating an audio package usable to render and playback interactive audio content in accordance with some embodiments.
[0030] Figure 5B is a flowchart of an example process for rendering and playback of interactive audio content in accordance with some embodiments.
[0031] Figure 6A shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0032] Figure 6B illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.
[0033] Figure 6C illustrates a schematic block diagram of an example CPU implemented in the device architecture of Figure 6B that may be used to implement various aspects of the present disclosure.
[0034] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION OF EMBODIMENTS
[0035] Content creators are increasingly creating sophisticated user-generated content. Such user-generated content may be interactive, e.g., such that audio content is to be rendered based on an orientation of a listener’s head and which dynamically changes responsive to movement of thelistener’ s head. However, packaging such interactive audio content can be difficult. For example, the audio content may be captured by multiple audio capture devices. As a more particular example, audio content may be captured using microphones on a user device such as a mobile phone or tablet computer, as well as using microphones disposed on a wearable device, such as a set of earbuds or headphones or a pair of smart glasses. However, it may be difficult to package this audio content in a manner such that the audio content can be played back by compatible devices capable of rendering interactive immersive audio content, as well as by non-compatible devices that are not capable rendering interactive audio content. Note that, as used herein a “compatible device” or a “compatible audio playback device” is an audio playback device that is capable of rendering audio content based on listener head orientation to provide an interactive immersive audio listening experience. A “non-compatible device” is an audio playback device that is not capable rendering audio content based on listener head orientation to provide an interactive immersive audio listening experience. As used herein, “interactive immersive” (which may be used interchangeably herein with “interactive” content) audio content refers to audio content that is rendered to provide a perception of spatial location of audio objects, where the spatial location is rendered and / or adjusted as the listener moves their head.
[0036] Briefly stated, techniques for transmitting audio content are provided. In some embodiments, the techniques may involve capturing audio content from a plurality of microphones. The techniques may further involve encoding a subset of audio content channels to form a first audio stream and encoding the remaining subset of audio content channels to form a second audio stream. The techniques may further involve determining sensor data and device information. The techniques may further involve storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream. The techniques may further involve generating an audio package comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non-interactive audio content.
[0037] As described above, an audio capture device may generate an audio package or a transmission file. The audio package or transmission file may include a first audio stream associated with two audio channels which may be rendered and presented by a non-compatible device as non-interactive binaural audio content. The audio package or transmission file may additionally include a metadata stream. The metadata stream may be parsed (e.g., by an audioplayback device) to obtain a second audio stream, which may be decoded to generate the remaining number of audio channels captured by the audio capture device(s). For example, in an instance in which two audio channels are captured from two microphones disposed on a set of earbuds, and in which N-2 audio channels are captured form N-2 microphones of a mobile phone, the first audio stream may correspond to the two audio channels associated with the microphones of the earbuds, and the second audio stream may correspond to the N-2 microphones of the mobile phone. The metadata stream may include metadata usable to reconstruct the full N channels captured by the one or more audio capture devices, as well as sensor data usable to render the audio content as interactive immersive binaural audio content based on sensor data associated with an audio playback device. A non-compatible device may discard the metadata stream and play back the first audio stream, whereas a compatible device may unpack and decode the metadata stream in order to render and play back interactive immersive content. Accordingly, using the techniques disclosed herein, a single audio package or transmission file may be used by both compatible and non-compatible devices.
[0038] Figure 1 illustrates an example system 100 for generating a package usable to generate interactive audio content in accordance with some embodiments. As illustrated, system 100 includes capture processing block 104, audio encoding blocks 106a and 106b, a metadata compiler block 112, and a packaging block 114. Capture processing block 104 is configured to receive, as input, N microphone audio inputs 102. Capture processing block 104 generates, as an output, N-2 audio channels, and the remaining two audio channels. The two audio channels are provided as input to audio encoding block 106a, and the N-2 audio channels are provided as input to audio encoding block 106b. Metadata compiler 112 receives the output of audio encoding block 106b as input, as well as device information 108 and sensor data 110. Metadata compiler 112 generates, as an output, a metadata stream based on the output of audio encoding block 106b, device information 108, and sensor data 110. The output of metadata compiler 112 and audio encoding block 106a are provided as input to packaging block 114. Packaging block 114 generates, as an output, a transmission file 116.
[0039] As illustrated, microphone audio inputs 102 from one or more audio capture devices may be obtained. In some embodiments, microphone audio inputs 102 may be obtained from multiple microphones of a single audio capture device, such as a mobile phone, tablet computer, microphones disposed in a pair of smart glasses or AR / VR headset, etc. In some embodiments, microphone audio inputs 102 may be obtained from multiple microphones across multiple audio capture devices, such as a mobile phone and a set of earbuds or headphones with microphones,mobile phone and a pair of smart glasses, a tablet computer and a set of earbuds or headphones with microphones, etc. In general, the total number of microphone audio inputs is generally referred to herein as N.
[0040] Microphone audio inputs 102 may be processed by capture processing block 104. Capture processing block 104 may perform various processing techniques, such as synchronizing the microphone inputs in time, adjusting timbre and / or level such that the microphone audio inputs are similar to one another, etc. Capture processing block 104 may perform a channel transform to generate two channels that may be used to provide binaural, non-interactive playback for audio playback devices that are not capable of rendering interactive audio content, and may generate N- 2 channels corresponding to the remaining audio channels. The channel transform may be a frequency dependent mix.
[0041] The two channels generated by capture processing block 104 may be encoded by audio encoding block 106a as a first audio stream, and the N-2 audio channels generated by capture processing block 104 may be encoded by audio encoding block 106b as a second audio stream. Each of encoding block 106a and encoding block 106b may perform audio encoding using, e.g., a predetermined codec to generate the first audio stream and the second audio stream, respectively. Audio encoding may involve compressing the audio data corresponding to the audio channels in a manner specified by the codec to generate the corresponding audio stream. Metadata compiler 112 may store the second audio stream in conjunction with metadata encoding device information 108 and sensor data 110. For example, in some embodiments, metadata compiler 112 may receive device information 108 and sensor data 110 and aggregate device information 108 and / or sensor data 110 in conjunction with timestamp information (e.g., a time at which device information 108 and / or sensor data 110 was obtained), or the like, and may store the device information 108 and sensor data 110 as metadata along with metadata representing the second audio stream. Device information 108 may include information about the one or more audio capture device that recorded audio content associated with the second audio stream (e.g., the N-2 channels), such as the type of device, whether the device was operating in landscape mode or portrait mode, the relative placement and / or distances between microphones, etc. Sensor data 110 may include device orientation and / or movement information associated with the one or more audio capture device that recorded the audio content associated with the second audio stream. Sensor data 110 may include data from one or more accelerometers and / or gyroscopes. Note that metadata compiler 112 generates a metadata stream that includes the second audio stream, and metadata associated with device information 108 and sensor data 1 10.
[0042] The first audio stream and the metadata stream may be packaged as a single transmission file 116 by packaging block 114. Packaging block 114 may be configured to generate a single file 116 that includes both the first audio stream and the metadata stream as components of the single file 116. The transmission file may be stored (e.g., in memory of one of the one or more audio capture devices that obtained at least a subset of microphone audio inputs 102, such as in memory of a mobile phone), may be transmitted (e.g., to a server or cloud device, to another user device, etc.), or the like.
[0043] Because the first audio stream (comprising the audio content associated with two channels) and the metadata stream (comprising the audio content associated with N-2 channels) are packaged in one transmission file but are distinct within the transmission file, a compatible device capable of rendering interactive audio content is able to decode and reconstruct the full N channels and perform spatial processing and rendering using the metadata in the transmission file. A non-compatible device that is not capable of rendering interactive audio content is capable of utilizing the transmission file to render and playback binaural audio content using the two audio channels of the first audio stream. The non-compatible device may discard (e.g., not use) the information associated with the metadata stream. Accordingly, the transmission file is usable by both compatible and non-compatible devices.
[0044] Figure 2A illustrates an example system 200 associated with a non-compatible device for rendering non-interactive audio content in accordance with some embodiments. System 200 includes an unpacking block 202, an audio decoding block 204, and a playback processing block 206. Unpacking block 202 receives, as input, transmission file 116. Unpacking block 202 generates, as an output, a first audio stream associated with two binaural audio channels. The first audio stream is provided as input to audio decoding block 204. Audio decoding block 204 generates, as an output, two decoded audio channels. The two decoded audio channels are provided as input to playback processing block 206.
[0045] Transmission file 116 is received by an audio playback device (e.g., a mobile phone, a tablet computer, a desktop computer, smart glasses, an AR / VR headset, etc.), which may be a non- compatible device. As described above in connection with Figure 1, transmission file 116 includes a first audio stream (comprising audio content associated with two audio channels) and a metadata stream (comprising metadata usable to decode N-2 audio channels and device and sensor information). Unpacking block 202 is configured to unpack the transmission file 116, e.g., to obtain the first audio stream and the metadata stream. For example, unpacking block 202 mayidentify and / or retrieve the first audio stream and the metadata stream from transmission file 1 16. The first audio stream is decoded by audio decoding block 204 to obtain the two decoded audio channels. For example, in some embodiments, decoding block 204 may perform decoding techniques specified by a codec used to encode the first audio stream to perform a reversal of the operations used to encode the first audio stream in order to generate the two decoded audio channels. Playback processing block 206 is configured to render and play back the two decoded audio channels as non-interactive binaural audio output. For example, in some implementations, playback processing block 206 may be configured to mix the audio channels and cause the mixed audio to be played back by a pair of associated earbuds and / or headphones.
[0046] Figure 2B illustrates an example system 250 associated with a compatible device for rendering interactive audio content in accordance with some embodiments. System 250 includes an unpacking block 252, an audio decoding block 204, a metadata parser 254, an audio decoding block 256, an information processing block 258, a channel recovery block 260, a spatial processing block 264, and a playback processing block 266. Unpacking block 252 receives, as input, transmission file 116. Unpacking block 252 generates, as output, a first audio stream that is provided to audio decoding block 204 as input, and a metadata stream that is provided as input to metadata parser 254. Audio decoding block 204 generates two decoded audio channels which are provided as input to channel recovery block 260. Metadata parser 254 generates a second audio stream that is provided to audio decoding block 256 and device information and sensor data which are provided as input to information processing block 258. Audio decoding block 256 generates N-2 decoded audio channels. The two decoded audio channels and the N-2 decoded audio channels are provided to channel recovery block 260. Channel recovery block 260 may optionally receive, as input, information from information processing block 258. Channel recovery block 260 reconstructs the original N channels. The original N channels that are output from channel recovery block 260 are provided as input to spatial processing block 264. Spatial processing block 264 may optionally take as input information from information processing block 258 and / or sensor data 262. Spatial processing block 264 may generate rendered interactive audio content. The rendered interactive audio content may be provided to playback processing block 266.
[0047] As illustrated, transmission file 116 is received by a compatible audio playback device (e.g., a mobile phone, a tablet computer, a desktop computer, smart glasses, an AR / VR headset, etc.). As described above in connection with Figure 1, transmission file 116 includes a first audio stream (comprising audio content associated with two audio channels) and a metadata stream (comprising metadata usable to decode N-2 audio channels and device and sensor information).
[0048] Unpacking block 252 is configured to unpack transmission file 116, e.g., to obtain the first audio stream and the metadata stream. For example, unpacking block 252 may identify and / or retrieve the first audio stream and the metadata stream from transmission file 116. The first audio stream is decoded by audio decoding block 204 to obtain two decoded audio channels. Similar to what is described above, audio decoding block 204 may reverse an encoding process as specified by a codec to decode the first audio stream. Metadata parser 254 is configured to parse the metadata stream. In particular, metadata parser 254 may obtain the second audio stream from the metadata stream and provide the second audio stream to audio decoding block 256, which is configured to obtain the decoded N-2 audio channels. Audio decoding block 256 may be configured to reverse an encoding process used to encode the second audio stream, e.g., as specified by a codec used to encode the second audio stream.
[0049] Metadata parser 254 may be configured to provide the device information and sensor data metadata to information processing block 258. Information processing block 258 may provide the device information to channel recovery block 260. For example, in some embodiments, information processing block 258 may identify timestamp information associated with the device information and / or sensor data metadata, may transform the device information and / or sensor data metadata into a format usable by channel recover block 260 and / or spatial processing block 264, or the like. Channel recovery block 260 is configured to take, as input, the two audio channels associated with the first audio stream and the N-2 audio channels associated with the second audio stream and perform an inverse channel transform to reconstruct the original N audio channels. More detailed techniques associated with channel recovery are shown in and described below in connection with Figure 3.
[0050] Information processing block 258 is also configured to provide the device information and the sensor data captured by the audio capture device and encoded in metadata to spatial processing block 264. Spatial processing block 264 is configured to utilize the device information and sensor data encoded in the metadata along with sensor data 262 to perform spatial rendering of the N audio channels in accordance with sensor data 262. Note that sensor data 262 is sensor data associated with the audio playback device, for example, indicative of a head orientation of a listener of the compatible audio playback device. Sensor data 262 may be obtained from one or more accelerometers, gyroscopes, etc. disposed in or on earbuds or headphones associated with (e.g., paired with) the compatible audio playback device. Accordingly, spatial processing block 264 may render the N audio channels in an interactive manner based on the listener’s head orientation as binaural audio content. In particular, the N audio channels may be mixed based onthe listener’s head orientation to dynamically generate interactive immersive binaural audio content. The compatible audio playback device may perform audio playback of the rendered audio content using playback processing block 266.
[0051] As described above, to generate the transmission file, the audio capture device may perform a channel transform to transform the N-2 channels before storing the audio content associated with the N-2 channels as metadata. A compatible audio playback device (e.g., as described above in connection with Figure 2B) may perform an inverse channel transform (e.g., using channel recovery block 260 of Figure 2B) to obtain the original N-2 channels, and therefore, to obtain the original N channels when combined with the two binaural channels. In some embodiments, the inverse channel transform may be implemented using mixing coefficients. The mixing coefficients may be stored as part of the device information that is encoded as metadata within the transmission file. In some embodiments, the channel transform may be implemented as a full rank matrix operation, and the inverse channel transform may be the inverse of the full rank matrix operation. Note that, in some embodiments, the mixing coefficients may be determined such that the two binaural channels form a suitable stereo signal In such instances, the mixing coefficients may be determined based on relative microphone locations (e.g., distance of microphones from one another) and / or tuning parameters. The mixing coefficients may be determined such that the mixing coefficients satisfy the constraint of being represented by a full rank matrix operation, which enables the inverse channel transform to be performed.
[0052] Figure 3 is a block diagram of an example implementation of a channel recovery block 260 in accordance with some embodiments. Channel recovery block 260 includes an inverse channel transform block 304. Channel recovery block 260 may receive, as input, two binaural audio channels and N-2 audio channels. The N-2 audio channels may be provided as input to inverse channel transform block 304. Inverse channel transform block 304 may optionally take, as input, mixing coefficients 302. Inverse channel transform block 304 may perform an inverse channel transform to recover the original N-2 channels. Channel recovery block 260 may concatenate the output of inverse channel transform block 304 (e.g., the original N-2 channels) with the two binaural channels to generate, as output, the recovered original N channels.
[0053] As described above, a channel recovery block may be implemented on a compatible audio playback device. As illustrated, channel recovery block 260 may be configured to take, as input, two decoded binaural channels and N-2 decoded channels. The N-2 decoded channels may be provided to inverse channel transform block 304. Inverse channel transform block 304 may usemixing coefficients 302 to perform an inverse channel transform on the N-2 decoded channels to recover the original N-2 channels. Note that the mixing coefficients may be stored as part of the device information encoded as metadata in the transmission file generated by the audio capture device, as described above in connection with Figure 1.
[0054] As described above, after reconstructing the original N channels, a compatible audio playback device may render the audio content interactively as interactive binaural audio content. In particular, the audio content may be rendered based on a head orientation of the listener such that audio objects are rendered at perceived spatial locations based on the head orientation. Rendering may be performed based on headtracking data, or sensor data, which may be obtained from one or more sensors disposed in or on earbuds or headphones worn by the listener. Note that the earbuds or headphones may be paired with the audio playback device such that the audio playback device may utilize the sensor data or headtracking data to render the audio content, which may then be played back by the audio playback device.
[0055] Figure 4 is a block diagram of an example implementation of a spatial processing block 264 of a compatible audio playback device in accordance with some embodiments. Spatial processing block 264 may include a beamforming block 402 and a mixing block 404. As illustrated, spatial processing block 264 may take, as input, the reconstructed N audio channels (e.g., as reconstructed by channel recovery block 260 of Figure 2B and Figure 3). The N-2 audio channels may be provided to beamforming block 402. Beamforming block 402 may be configured to perform beamforming to simulate head rotation to generate a beamformed signal. The output of beamforming block 402 (e.g., the beamformed signal) may be provided as input to mixing block 404. Beamforming block 402 and / or mixing block 404 may optionally take, as input, headtracking data 406. Mixing block 404 may generate, as output, a head tracked binaural output.
[0056] With respect to beamforming block 402, the simulated head rotation may simulate the user capturing the audio content rotating their head while capturing the audio content. Such simulation is performed to allow, during audio playback, adjustment of the audio content responsive to movement of the listener’s head by utilizing the simulated beamformed output corresponding to the movement of the listener’s head. Beamforming block 402 may utilize beamforming coefficients, which may be based on the device orientation. Note that device orientation may be part of the device information which is stored as metadata within the transmission file by the audio capture device. In some embodiments, the beamforming coefficients may be determined by the compatible audio playback device based on the device orientation.Alternatively, in some embodiments, the beamforming coefficients may be determined by the audio capture device and stored as metadata within the transmission file.
[0057] With respect to mixing block 404, the beamformed signal may be mixed with the two binaural channels by mixing block 404. Mixing may be performed using headtracking data 406, which may be obtained using one or more sensors (e.g., accelerometers, gyroscopes, etc.) disposed in or on earbuds or headphones associated with the audio playback device. In some embodiments, mixing block 404 may perform cross-fading when mixing the audio channels. The output of mixing block 404 is a head-tracked binaural output that is interactive and dynamic based on headtracking data 406.
[0058] In some embodiments, mixing may be performed by applying weights to different audio channels to generate two binaural signals. By way of example, the mixing may be performed using:Y(ri) = X(ri)W(ri)
[0059] In the equation given above, Y(n ) corresponds to the output binaural audio content and is an L x 2 matrix, X( n ) corresponds to the full set of reconstructed audio channels and is an L x N matrix (where N represents the full number of audio channels), and W(n) is an Ax 2 weight matrix. Note that L represents the number of audio samples in an audio frame, and n represents the index of the audio frame. In one example, mixing may be performed based on headtracking data representative of the listener’s head orientation. For example, if 6nrepresents the head orientation of the listener’s head in the horizontal plane, the weight matrix may be represented as:
[0060] Using the weights determined based on listener head orientation, mixing may be performed to dynamically generate a mixed binaural audio signal based on listener head orientation.
[0061] Figure 5A is a flowchart of an example process 500 for generating a transmission file (sometimes referred to herein as “an audio package”) usable by both compatible audio playback devices (to render interactive audio content) and non-compatible audio playback devices (to render non-interactive audio content) in accordance with some embodiments. Blocks of process 500 maybe executed by one or more processors and / or control systems of an audio capture and playback device. An example of such a control system is control system 610 of Figure 6 A. In some embodiments, blocks of process 500 may be executed in an order other than what is shown in Figure 5A. In some embodiments, two or more blocks of process 500 may be executed substantially in parallel. In some embodiments, one or more blocks of process 500 may be omitted.
[0062] Process 500 can begin at block 502 by capturing audio content from a plurality of microphones using at least one audio capture device (e.g., by capture processing block 104 of Figure 1). For example, the audio capture device may be a mobile phone, a tablet computer, etc. In some embodiments, the microphones may be disposed on two audio capture devices, such as a pair of earbuds and a mobile phone, a pair of smart glasses and a mobile phone, etc. In general, the number of microphones is referred to herein as N, which generate a corresponding N audio channels. In one example, two audio channels may be captured using two microphones of a pair of earbuds (e.g., one microphone disposed in each earbud of the pair), and N-2 (e.g., two, three, four, etc.) microphones may be disposed in a device such as a mobile phone, tablet computer, etc. paired with the earbuds.
[0063] At 504, process 500 can encode a subset of audio content channels associated with a subset of the plurality of microphones to form a first audio stream and encode the remaining subset of audio content channels associated with the remaining microphones to form a second audio stream (e.g., via audio encoding block 106a and / or audio encoding block 106b of Figure 1). In the example shown in and described above in connection with Figure 1 , the first audio stream may include two encoded channels (e.g., two binaural channels) and the second audio stream may include the remaining N-2 channels. Note that, in some embodiments, the first audio stream may include audio content captured by a first device, and the second audio stream may include audio content captured by a second device. For example, in an instance in which the first audio stream comprises two audio channels, the two audio channels may be those associated with two microphones of a pair of earbuds, and the second audio stream may comprise the remaining audio channels associated with the microphones of another device, such as a paired mobile phone.
[0064] At 506, process 500 may determine sensor data associated with the at least one audio capture device and device information associated with the at least one audio capture device (e.g., device information 108 and sensor data 110 of Figure 1). In some embodiments, the sensor data may indicate an orientation of the at least one audio capture device, such as whether the audio capture device was in landscape mode or portrait mode when capturing the content. In someembodiments, the device information may include a device type associated with the at least one audio capture device, relative placement of microphones with respect to one another (e.g., distance between the microphones, etc.), or the like. In some embodiments, the device information may include beamforming coefficients which may be used by a compatible audio playback device to perform beamforming to simulate head rotation, as described above in connection with Figure 4.
[0065] At 508, process 500 can store data associated with the second audio stream, the sensor data, and the device information as metadata, where the metadata includes data usable to decode and reconstruct the second audio stream (e.g. , by metadata compiler 112 of Figure 1 ). For example, the metadata may include mixing coefficients which may be used to reconstruct the A- 2 audio channels associated with the second audio stream, as shown in and described above in connection with Figure 3.
[0066] At 510, process 500 can generate an audio package comprising the first audio stream and the metadata (e.g., by audio packaging block 116 of Figure 1). The audio package may be usable by both a compatible audio playback device capable of rendering interactive audio content and by a non-compatible audio playback device that is not capable of rendering interactive audio content. For example, as shown in and described above in connection with Figure 2A, a non-compatible audio playback device may discard the metadata and render binaural audio content associated with the first audio stream in a non-interactive manner. Conversely, as shown in and described above in connection with Figure 2B, a compatible audio playback device may parse the metadata to reconstruct the full N channels, and subsequently render the reconstructed audio content based on headtracking data associated with the audio playback device.
[0067] Figure 5B is a flowchart of an example process 550 for rendering interactive audio content in accordance with some embodiments. Blocks of process 550 may be executed by one or more processors and / or control systems of one or more audio playback devices, such as a mobile phone, a tablet computer, a desktop computer, a laptop computer, smart glasses, an AR / VR headset, etc. In some embodiments, earbuds or headphones may be coupled to and / or paired with (e.g., via a wireless connection such as BLUETOOTH) the audio playback device. An example of a control system is control system 610 shown in Figure 6A. In some embodiments, blocks of process 550 may be executed in an order other than what is shown in Figure 5B. In some embodiments, two or more blocks of process 550 may be executed substantially in parallel. In some embodiments, one or more blocks of process 550 may be omitted.
[0068] Process 550 can begin at block 552 by receiving, at an audio playback device (e.g., by unpacking block 252 of Figure 2B), an audio package (sometimes referred to herein as “a transmission file,” as shown in and described above in connection with Figure 1). The audio package may comprise a first audio stream and a metadata stream, where metadata of the metadata stream is usable to decode and reconstruct a second audio stream and to perform spatially interactive rendering and playback of the first audio stream and the second audio stream. An example of such an audio package is the transmission file shown in and described above in connection with Figure 1. The audio package may be received via a wired or wireless communication, may be downloaded from a website or application (e.g., a social media application), streamed from a website or application, or the like.
[0069] At 554, process 550 can decode the first audio stream and the metadata stream to generate a first set of audio channels and to generate a second set of audio channels associated with the second audio stream. For example, as shown in and described above in connection with Figure 2B, the first audio stream may correspond to two binaural channels (e.g., captured using microphones disposed in or on a pair of earbuds, headphones, smart glasses, etc. The second set of audio channels may comprise audio content recorded by one or more microphones of an audio capture device such as a mobile phone, a tablet computer, etc. The first audio stream may be decoded by unpacking the audio package to retrieve the first audio stream and decoding the first audio stream using an audio decoder to obtain the first set of audio channels (e.g., by audio decoding block 204 of Figure 2B). The metadata stream may be unpacked and / or parsed to obtain the second audio stream (e.g., by metadata parser 254 of Figure IB), which may then be decoded to generate the second set of audio channels (e.g., by audio decoding block 256 of Figure IB).
[0070] At 556, process 550 can recover a full set of audio channels based on the first set of audio channels and the second set of audio channels using the metadata (e.g., by channel recovery block 260). For example, process 550 can obtain metadata indicative of device information and / or sensor data associated with the audio capture device that recorded and / or captured the audio content associated with the first audio stream and / or the second audio stream by parsing the metadata of the metadata stream. The full set of audio channels may be recovered by, e.g., performing an inverse channel transform on the second set of audio channels, as shown in and described above in connection with Figure 3. The inverse channel transform may utilize the device information stored in the metadata. The full set of audio channels may include an inverse channel transform of the second set of audio channels concatenated with the first set of audio channels.
[0071] At 558, process 550 can obtain headtracking data based on sensor data associated with a set of earbuds or headphones paired with the audio playback device. The headtracking data may be obtained based on sensor data from one or more accelerometers and / or gyroscopes disposed in or on the set of earbuds or headphones. The headtracking data may indicate an orientation of the listener’s head. Note that the headtracking data may be dynamic and indicate movement of the listener’s head, e.g., as they rotate their head up, down, left, right, any combination of roll, pitch, and yaw, or any combination thereof.
[0072] At 560, process 550 can render the full set of audio channels based on the headtracking data and using sensor data encoded in the metadata to generate rendered audio content (e.g., by spatial processing block 264 of Figure IB). For example, the sensor data encoded in the metadata may be used to perform beamforming techniques. The beamforming techniques may be used to, e.g., simulate head rotation of a listener. The headtracking data may be used to mix the first set of audio channels and the second set of audio channels in accordance with a head orientation of the listener to generate interactive binaural audio content. An example of techniques for rendering the full set of audio channels in accordance with the headtracking data is described above in connection with Figure 4.
[0073] At 562, process 550 can present the rendered audio content using the set of earbuds or headphones (e.g., by playback processing block 266 of Figure IB). For example, process 550 can play back the rendered audio content via the set of earbuds or headphones. Note that in some embodiments, process 550 can loop back to 558 to obtain additional headtracking data. In this way, process 550 can dynamically present the interactive audio content in accordance with the listener’s head orientation and / or movements.
[0074] Figure 6A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 6A are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 600 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 600 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.
[0075] According to some alternative implementations the apparatus 600 may be, or may include, a server. In some such examples, the apparatus 600 may be, or may include, an encoder.Accordingly, in some instances the apparatus 600 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 600 may be a device that is configured for use in “the cloud,” e.g., a server.
[0076] In this example, the apparatus 600 includes an interface system 605 and a control system 610. The interface system 605 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 605 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may, in some examples, pertain to one or more software applications that the apparatus 600 is executing.
[0077] The interface system 605 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0078] The interface system 605 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 605 may include one or more wireless interfaces. The interface system 605 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 605 may include one or more interfaces between the control system 610 and a memory system, such as the optional memory system 615 shown in Figure 6A. However, the control system 610 may include a memory system in some instances. The interface system 605 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0079] The control system 610 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0080] In some implementations, the control system 610 may reside in more than one device. For example, in some implementations a portion of the control system 610 may reside in a device within one of the environments depicted herein and another portion of the control system 610 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 610 may reside in a device within one environment and another portion of the control system 610 may reside in one or more other devices of the environment. For example, a portion of the control system 610 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 610 may reside in another device that is implementing the cloud- based service, such as another server, a memory device, etc. The interface system 605 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.
[0081] In some implementations, the control system 610 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 610 may be configured for implementing methods of generating a package usable to generate and play interactive audio content, or the like.
[0082] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 615 shown in Figure 6A and / or in the control system 610. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, filter audio signals to reduce noise, determine filter coefficients to perform such filtering, etc. The software may, for example, be executable by one or more components of a control system such as the control system 610 of Figure 6 A.
[0083] In some examples, the apparatus 600 may include the optional microphone system 620 shown in Figure 6A. The optional microphone system 620 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 600 may not include a microphone system 620. However, in some suchimplementations the apparatus 600 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 610. In some such implementations, a cloud-based implementation of the apparatus 600 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 610.
[0084] According to some implementations, the apparatus 600 may include the optional loudspeaker system 625 shown in Figure 6A. The optional loudspeaker system 625 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 600 may not include a loudspeaker system 625. In some implementations, the apparatus 600 may include headphones. Headphones may be connected or coupled to the apparatus 600 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).
[0085] Figure 6B illustrates a schematic block diagram of an example device architecture 601 (in this example, an apparatus 601) that may be used to implement various aspects of the present disclosure. The apparatus 601 of Figure 6B is an instance of the apparatus 600 of Figure 6 A. Architecture 601 includes but is not limited to servers and client devices, systems, etc., which may be configured to perform the methods that are described with reference to any or all of Figures 5A and / or 5B. As shown, the architecture 601 includes central processing unit (CPU) 641, which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 642 or a program loaded from, for example, storage unit 648 to random access memory (RAM) 643. The CPU 641 may be, for example, an electronic processor 641. In these examples, the CPU 641 is an instance of the control system 610 of Figure 6A and the ROM 642 and RAM 643 are instances of the memory system 615. In RAM 643, the data required when CPU 641 performs the various processes is also stored, as required. CPU 641, ROM 642, and RAM 643 are connected to one another via bus 644. Input / output (I / O) interface 645 is also connected to bus 644. The bus 644 and the I / O interface 645 are instances of the interface system 605 of Figure 6 A.
[0086] The following components are connected to I / O interface 645: input unit 646, that may include a keyboard, a mouse, or the like; output unit 647 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 648 including a hard disk, or another suitable storage device; and communication unit 649 including a network interface card such as a network card (e.g., wired or wireless).
[0087] In some implementations, input unit 646 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0088] In some implementations, output unit 647 include systems with various number of speakers. Output unit 647 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0089] In some embodiments, communication unit 649 is configured to communicate with other devices (e.g., via a network). Drive 650 is also connected to I / O interface 645, as required. Removable medium 651, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 650, so that a computer program read therefrom is installed into storage unit 648, as required. A person skilled in the art would understand that although apparatus 601 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alterations all fall within the scope of the present disclosure.
[0090] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 649, and / or installed from the removable medium 651, as shown in Figure 6B.
[0091] Figure 6C illustrates a schematic block diagram of an example CPU 641 implemented in the device architecture 601 of Figure 6B that may be used to implement various aspects of the present disclosure. The CPU 641 includes an electronic processor 660 and a memory 661. The electronic processor 660 is electrically and / or communicatively connected to the memory 661 for bidirectional communication. The memory 661 stores encoding software 662 and decoding software 663. The memory 661 may be, for example, a ROM, a RAM, or another non-transitory computer readable medium. The electronic processor 660 may implement the encoding software 662 stored in the memory 661 to perform, among other things, the method 500 of Figure 5 A and / or the method 550 of Figure 5B. Additionally, the electronic processor 660 may implement the decoding software 663 stored in the memory 661 to perform, among other things, the methods that are described with reference to any or all of Figure 6.
[0092] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 641 in combination with other components of Figure 6B), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0093] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0094] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine- readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0095] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer programcodes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0096] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0097] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.
[0098] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.
[0099] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):EEE 1. A method for transmitting audio content, the method comprising: capturing audio content from a plurality of microphones using at least one audio capture device (502); encoding a subset of audio content channels associated with a subset of the plurality of microphones to form a first audio stream and encoding the remaining subset of audio content channels associated with the remaining microphones of the plurality of microphones to form a second audio stream (504)-, determining sensor data (110) associated with the at least one audio capture device and device information (108) associated with the at least one audio capture device (506)-. storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream (508)-, and generating an audio package (116) comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non-interactive audio content (510).EEE 2. The method of EEE 1 , further comprising transmitting the audio package to a user device configured to render and playback audio content based on the first audio stream and the second audio stream.EEE 3. The method of any one of EEEs 1 or 2, wherein the sensor data comprises data indicative of an orientation of the at least one audio capture device during capture of the audio content.EEE 4. The method of any one of EEEs 1-3, wherein the device information comprises information indicative of relative microphone placement of the plurality of microphones.EEE 5. The method of any one of EEEs 1-4, wherein the at least one audio capture device comprises a set of earbuds or headphones with microphones disposed in or on the set of earbuds or headphones, and a mobile device.EEE 6. The method of EEE 5, wherein the first audio stream comprises audio content captured using microphones of the set of earbuds or headphones.EEE 7. The method of any one of EEEs 5 or 6, wherein the sensor data and device information are associated with the mobile device.EEE 8. The method of any one of EEEs 1-7, wherein the metadata further comprises beamforming coefficients usable to render the audio content as spatially interactive audio content by an audio playback device.EEE 9. A method for presenting audio content, the method comprising: receiving, at an audio playback device, an audio package (116) comprising a first audio stream and a metadata stream, wherein metadata of the metadata stream is usable to decode and reconstruct a second audio stream and perform spatially interactive rendering and playback of the first audio stream and the second audio stream (552); decoding the first audio stream to generate a first set of audio channels and generating a second set of audio channels associated with the second audio stream (554)-, recovering a full set of audio channels based on the first set of audio channels and the second set of audio channels using the metadata (556); obtaining headtracking data (406) based on sensor data associated with a set of earbuds or headphones paired with the audio playback device (558); rendering the full set of audio channels based on the headtracking data and using sensor data encoded in the metadata to generate rendered audio content (560); and presenting the rendered audio content using the set of earbuds or headphones (562).EEE 10. The method of EEE 9, wherein the metadata comprises sensor data and device information associated with an audio capture device that recorded the first set of audio channels and / or the second set of audio channels, and wherein the sensor data and / or the device information are used to render the full set of audio channels.EEE 11. The method of EEE 10, wherein the device information is used to determine beamforming coefficients to render the full set of audio channels.EEE 12. The method of any one of EEEs 9-11 , wherein recovering the full set of audio channels comprises performing an inverse channel transform on the second set of audio channels using mixing coefficients and combining the first set of audio channels and a result of the inverse channel transform to obtain the full set of audio channels.EEE 13. The method of any one of EEEs 9-12, wherein rendering the full set of audio channels comprises performing beamforming using the second set of audio channels and the headtracking data to generate a beamformed second set of audio channels.EEE 14. The method of EEE 13, wherein beamforming coefficients used to perform the beamforming are included in the metadata.EEE 15. The method of any one of EEEs 13 or 14, further comprising performing mixing of the first set of audio channels and the beamformed second set of audio channels to generate the rendered audio content.EEE 16. An apparatus configured for implementing the method of any one of EEEs 1-15.EEE 17. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of EEEs 1-15.EEE 18. An audio capture device, comprising: one or more processors (641 ); and one or more processor-readable media ( 661 ) storing instructions which, when executed by the one or more processors, cause performance of: capturing audio content from a plurality of microphones (502); encoding a subset of audio content channels associated with a subset of the plurality of microphones to form a first audio stream and encoding the remaining subset of audio content channels associated with the remaining microphones of the plurality of microphones to form a second audio stream (504); determining sensor data (110) associated with the at least one audio capture device and device information (108) associated with the at least one audio capture device (506); storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream (508); andgenerating an audio package (116) comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non-interactive audio content (510).EEE 19. The audio capture device of EEE 18, wherein the instructions further cause performance of transmitting the audio package to a user device configured to render and playback audio content based on the first audio stream and the second audio stream.EEE 20. The audio capture device of any one of EEEs 18 or 19, wherein the sensor data comprises data indicative of an orientation of the audio capture device during capture of the audio content.EEE 21. An audio playback device comprising: one or more processors (641 ); and one or more processor-readable media ( 661 ) storing instructions which, when executed by the one or more processors, cause performance of: receiving an audio package (116) comprising a first audio stream and a metadata stream, wherein metadata of the metadata stream is usable to decode and reconstruct a second audio stream and perform spatially interactive rendering and playback of the first audio stream and the second audio stream (552); decoding the first audio stream to generate a first set of audio channels and generating a second set of audio channels associated with the second audio stream (554)-, recovering a full set of audio channels based on the first set of audio channels and the second set of audio channels using the metadata (556); obtaining headtracking data (406) based on sensor data associated with a set of earbuds or headphones paired with the audio playback device (55<S); rendering the full set of audio channels based on the headtracking data and using sensor data encoded in the metadata to generate rendered audio content (560); and presenting the rendered audio content using the set of earbuds or headphones (562).EEE 22. The audio playback device of EEE 21, wherein the metadata comprises sensor data and device information associated with an audio capture device that recorded the first set of audio channels and / or the second set of audio channels, and wherein the sensor data and / or the device information are used to render the full set of audio channels.EEE 23. The audio playback device of any one of EEEs 21 or 22, wherein rendering the full set of audio channels comprises performing beamforming using the second set of audio channels and the headtracking data to generate a beamformed second set of audio channels.
Claims
CLAIMS1. A method for transmitting audio content, the method comprising: capturing audio content from a plurality of microphones using at least one audio capture device (502); encoding a subset of audio content channels associated with a subset of the plurality of microphones to form a first audio stream and encoding the remaining subset of audio content channels associated with the remaining microphones of the plurality of microphones to form a second audio stream (504); determining sensor data (110) associated with the at least one audio capture device and device information (108) associated with the at least one audio capture device (506); storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream (508); and generating an audio package (116) comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non-interactive audio content (510).
2. The method of claim 1 , further comprising transmitting the audio package to a user device configured to render and playback audio content based on the first audio stream and the second audio stream.
3. The method of any one of claims 1 or 2, wherein the sensor data comprises data indicative of an orientation of the at least one audio capture device during capture of the audio content.
4. The method of any one of claims 1-3, wherein the device information comprises information indicative of relative microphone placement of the plurality of microphones.
5. The method of any one of claims 1-4, wherein the at least one audio capture device comprises a set of earbuds or headphones with microphones disposed in or on the set of earbuds or headphones, and a mobile device.
6. The method of claim 5, wherein the first audio stream comprises audio content captured using microphones of the set of earbuds or headphones.
7. The method of any one of claims 5 or 6, wherein the sensor data and device information are associated with the mobile device.
8. The method of any one of claims 1 -7, wherein the metadata further comprises beamforming coefficients usable to render the audio content as spatially interactive audio content by an audio playback device.
9. A method for presenting audio content, the method comprising: receiving, at an audio playback device, an audio package (116) comprising a first audio stream and a metadata stream, wherein metadata of the metadata stream is usable to decode and reconstruct a second audio stream and perform spatially interactive rendering and playback of the first audio stream and the second audio stream (552); decoding the first audio stream to generate a first set of audio channels and generating a second set of audio channels associated with the second audio stream (554)', recovering a full set of audio channels based on the first set of audio channels and the second set of audio channels using the metadata (556); obtaining headtracking data (406) based on sensor data associated with a set of earbuds or headphones paired with the audio playback device (558); rendering the full set of audio channels based on the headtracking data and using sensor data encoded in the metadata to generate rendered audio content (560); and presenting the rendered audio content using the set of earbuds or headphones (562).
10. The method of claim 9, wherein the metadata comprises sensor data and device information associated with an audio capture device that recorded the first set of audio channels and / or the second set of audio channels, and wherein the sensor data and / or the device information are used to render the full set of audio channels.
11. The method of claim 10, wherein the device information is used to determine beamforming coefficients to render the full set of audio channels.
12. The method of any one of claims 9-11, wherein recovering the full set of audio channels comprises performing an inverse channel transform on the second set of audio channels using mixing coefficients and combining the first set of audio channels and a result of the inverse channel transform to obtain the full set of audio channels.
13. The method of any one of claims 9-12, wherein rendering the full set of audio channels comprises performing beamforming using the second set of audio channels and the headtracking data to generate a beamformed second set of audio channels.
14. The method of claim 13, wherein beamforming coefficients used to perform the beamforming are included in the metadata.
15. The method of any one of claims 13 or 14, further comprising performing mixing of the first set of audio channels and the beamformed second set of audio channels to generate the rendered audio content.
16. An apparatus configured for implementing the method of any one of claims 1-15.
17. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of claims 1-15.
18. An audio capture device, comprising: one or more processors (641 ; and one or more processor-readable media ( 661 ) storing instructions which, when executed by the one or more processors, cause performance of: capturing audio content from a plurality of microphones (502): encoding a subset of audio content channels associated with a subset of the plurality of microphones to form a first audio stream and encoding the remaining subset of audio content channels associated with the remaining microphones of the plurality of microphones to form a second audio stream (504): determining sensor data (110) associated with the at least one audio capture device and device information (108) associated with the at least one audio capture device (506): storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream (508)-, and generating an audio package (116) comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non-interactive audio content (510).
19. The audio capture device of claim 18, wherein the instructions further cause performance of transmitting the audio package to a user device configured to render and playback audio content based on the first audio stream and the second audio stream.
20. The audio capture device of any one of claims 18 or 19, wherein the sensor data comprises data indicative of an orientation of the audio capture device during capture of the audio content.
21. An audio playback device comprising: one or more processors (641 ); and one or more processor-readable media ( 661 ) storing instructions which, when executed by the one or more processors, cause performance of: receiving an audio package (116) comprising a first audio stream and a metadata stream, wherein metadata of the metadata stream is usable to decode and reconstruct a second audio stream and perform spatially interactive rendering and playback of the first audio stream and the second audio stream (552); decoding the first audio stream to generate a first set of audio channels and generating a second set of audio channels associated with the second audio stream (554); recovering a full set of audio channels based on the first set of audio channels and the second set of audio channels using the metadata (556); obtaining headtracking data (406) based on sensor data associated with a set of earbuds or headphones paired with the audio playback device (558); rendering the full set of audio channels based on the headtracking data and using sensor data encoded in the metadata to generate rendered audio content (560); and presenting the rendered audio content using the set of earbuds or headphones (562).
22. The audio playback device of claim 21, wherein the metadata comprises sensor data and device information associated with an audio capture device that recorded the first set of audio channels and / or the second set of audio channels, and wherein the sensor data and / or the device information are used to render the full set of audio channels.
23. The audio playback device of any one of claims 21 or 22, wherein rendering the full set of audio channels comprises performing beamforming using the second set of audio channels and the headtracking data to generate a beamformed second set of audio channels.
Citation Information
Patent Citations
Audio Recording and Playback Apparatus
US20160219392A1
System and method for capturing, encoding, distributing, and decoding immersive audio
US20160227337A1
Headtracking for parametric binaural output system and method
US20180359596A1
Two stage audio focus for spatial audio processing
US20190394606A1
Spatial audio file format for storing capture metadata
US20200409995A1