Spatial Audio Monoization via Data Exchange
By exchanging data between portable devices and adopting a simplified 3D sound field communication solution, the computing complexity and resource limitation problems of portable devices when generating high-quality head tracking immersive audio, achieving longer 3D sound field reproduction duration and better user experience.
Patent Information
- Application Number
- CN202280036257.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-27
- Filing Date
- 2022-05-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-25
AI Technical Summary
Existing portable devices face computational complexity and resource limitations when generating high-quality head tracking immersive audio, especially in maintaining the duration of 3D sound field reproduction.
By exchanging data between the first audio output device and the second audio output device, convolution and stereo decoding operations performed on the personal audio device are reduced, and a simplified 3D sound field communication scheme is adopted to generate mono audio output to reduce computational complexity and balance resource requirements of both devices.
Effectively reduces the computational complexity performed on personal audio devices, extends the duration of 3D sound field reproduction, and improves the user experience.
Smart Images

Figure CN117378220B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority of co - owned U.S. Non - Provisional Patent Application No. 17 / 332,798, filed on May 27, 2021, the entire content of which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to using data exchange to facilitate the generation of monaural audio output based on spatial audio data. Background Art
[0004] Advances in technology have led to smaller and more powerful computing devices. For example, there are currently various portable personal computing devices, including wireless telephones (e.g., mobile phones and smart phones), small, lightweight, and easily portable tablet computers and laptop computers. These devices can transmit voice and data packets over a wireless network. In addition, many such devices incorporate additional functions, such as digital cameras, digital video cameras, digital recorders, and audio file players. Further, such devices can process executable instructions, including software applications, such as a web browser application, which can be used to access the Internet. Accordingly, these devices can include key computing capabilities.
[0005] The proliferation of such devices has facilitated a change in media consumption. For example, personal video games have increased, where an individual uses a handheld or portable video game system to play video games. As another example, personal media consumption has increased, where a handheld or portable media player outputs media (e.g., audio, video, augmented reality media, virtual reality media, etc.) to an individual. Such personalized or individualized media consumption typically involves relatively small portable (e.g., battery - powered) devices for generating the output. Due to the size, weight constraints, power constraints, or other reasons of portable devices, the processing resources available for such portable devices may be limited. Therefore, it may be challenging to provide a high - quality user experience using these resource - constrained devices. Summary of the Invention
[0006] According to a particular aspect of the present disclosure, a device includes: a memory configured to store instructions; and one or more processors configured to execute the instructions to obtain spatial audio data at a first audio output device. The one or more processors are further configured to perform a data exchange that exchanges data between the first audio output device and a second audio output device based on the spatial audio data. The one or more processors are further configured to generate a first monaural audio output at the first audio output device based on the spatial audio data.
[0007] According to certain aspects of the present disclosure, a method includes: obtaining spatial audio data at a first audio output device. The method further includes: performing a data exchange for exchanging data between the first audio output device and a second audio output device based on the spatial audio data. The method further includes: generating a first mono audio output at the first audio output device based on the spatial audio data.
[0008] According to another implementation of the present disclosure, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to obtain spatial audio data at a first audio output device. The instructions, when executed, further cause the one or more processors to perform a data exchange for exchanging data between the first audio output device and a second audio output device based on the spatial audio data. The instructions, when executed, further cause the one or more processors to generate a first mono audio output at the first audio output device based on the spatial audio data.
[0009] According to another implementation of the present disclosure, an apparatus includes: a unit for obtaining spatial audio data at a first audio output device. The apparatus further includes: a unit for performing a data exchange for exchanging data between the first audio output device and a second audio output device based on the spatial audio data. The apparatus further includes: a unit for generating a first mono audio output at the first audio output device based on the spatial audio data.
[0010] Other aspects, advantages, and features of the present disclosure will become apparent after reviewing the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and Claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a block diagram of a particular illustrative aspect of a system according to some examples of the present disclosure, the system including a plurality of audio output devices configured to exchange data to enable generation of a mono audio output from spatial audio data.
[0012] Figure 2 is according to some examples of the present disclosure Figure 1 of a block diagram of a particular illustrative example of a system.
[0013] Figure 3 is according to some examples of the present disclosure Figure 1 of a block diagram of another particular illustrative example of a system.
[0014] Figure 4 is according to some examples of the present disclosure Figure 1 of a block diagram of another particular illustrative example of a system.
[0015] Figure 5 A diagram of a head-mounted device (such as headphones) operable to perform data exchange to enable generation of a mono audio output from spatial audio data, according to some examples of the present disclosure.
[0016] Figure 6 A diagram of an earbud operable to perform data exchange to enable generation of a mono audio output from spatial audio data, according to some examples of the present disclosure.
[0017] Figure 7 A diagram of a head-mounted device (such as a virtual reality or augmented reality head-mounted device) operable to perform data exchange to enable generation of a mono audio output from spatial audio data, according to some examples of the present disclosure.
[0018] Figure 8 A diagram of a specific illustrative implementation of a method for generating a mono audio output from spatial audio data performed by one or more of the Figure 1 audio output devices, according to some examples of the present disclosure.
[0019] Figure 9 A diagram of another specific illustrative implementation of a method for generating a mono audio output from spatial audio data performed by one or more of the Figure 1 audio output devices, according to some examples of the present disclosure. DETAILED DESCRIPTION
[0020] Audio information can be captured or generated in a manner that enables rendering of an audio output to represent a three-dimensional (3D) sound field. For example, ambisonics (e.g., first-order ambisonics (FOA) or higher-order ambisonics (HOA)) can be used to represent the 3D sound field for later playback. During playback, the 3D sound field can be reconstructed in a manner that enables a listener to distinguish the position and / or distance between the listener and one or more audio sources in the 3D sound field.
[0021] In certain aspects of the present disclosure, a personal audio device (such as a head-mounted device, headphones, earbuds, or another audio playback device configured to generate different audio outputs for each ear of a user (e.g., two mono audio output streams)) can be used to render a 3D sound field. One challenge in rendering 3D audio using a personal audio device is the computational complexity of such rendering. By way of illustration, personal audio devices are typically configured to be worn by a user such that movement of the user's head changes the relative positions of the user's ears and audio sources in the 3D sound field to generate head-tracked immersive audio. Such personal audio devices are typically battery-powered and have limited on-board computational resources. Generating head-tracked immersive audio with such resource constraints is challenging. One way to avoid some of the power and processing constraints of a personal audio device is to perform most of the processing at a host device (such as a laptop computer or a mobile computing device). However, the more processing that is performed on the host device, the greater the latency between head movement and sound output, which results in a less satisfactory user experience.
[0022] Additionally, many personal audio devices include a pair of different audio output devices, such as a pair of earbuds including one earbud for each ear. In such a configuration, it is useful to balance the power requirements imposed on each audio output device such that one audio output device does not run out of power before the other. Since simulating a 3D sound field requires providing sound to both ears of a user, a failure in one of the audio output devices (e.g., due to battery depletion) will prematurely stop the generation of 3D audio output.
[0023] Aspects disclosed herein facilitate a reduction in the computational complexity for generating head-tracked immersive audio by using a simplified 3D sound field communication scheme to reduce the number of convolution operations performed on the personal audio device, exchanging data between audio output devices to reduce the number of stereo decoding operations performed on the personal audio device, and generating mono audio outputs. Aspects disclosed herein also facilitate balancing the resource requirements between a pair of audio output devices to extend the duration of 3D sound field reproduction that can be provided by the audio output devices.
[0024] Certain aspects of the present disclosure are described below with reference to the accompanying drawings. In the specification, common features are designated by common reference numerals. As used herein, various terms are for the purpose of describing particular implementations only and are not intended to limit the implementations. For example, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Additionally, some features described herein are singular in some implementations and plural in other implementations. By way of illustration, Figure 1 depicts including one or more processors ( Figure 1the first audio output device 110 of the "processor" 112), indicating that in some implementations the first audio output device 110 includes a single processor 112, and in other implementations the first audio output device 110 includes multiple processors 112.
[0025] As used herein, the terms "comprise", "comprises" and "comprising" may be used interchangeably with "include", "includes" or "including". In addition, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" indicates examples, implementations and / or aspects and should not be construed as restrictive or indicating a preference or preferred implementation. As used herein, ordinal terms (e.g., "first", "second", "third", etc.) used to modify elements (e.g., structures, components, operations, etc.) do not themselves indicate any priority or order of the element relative to another element, but merely distinguish the element from another element having the same name (but using an ordinal term). As used herein, the term "set" refers to one or more specific elements, and the term "plurality" refers to a plurality (e.g., two or more) of specific elements.
[0026] As used herein, "coupled" may include "communicatively coupled", "electrically coupled" or "physically coupled", and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired network, wireless network or a combination thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or different devices and may be connected via electronics, one or more connectors or inductive coupling, as shown in the non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled (such as electrically communicating) may directly or indirectly send and receive signals (e.g., digital signals or analog signals) via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include coupling (e.g., communicatively coupled, electrically coupled or physically coupled) of two devices without an intermediate component.
[0027] In the present disclosure, terms such as "determine", "calculate", "estimate", "shift", "adjust", etc. may be used to describe how to perform one or more operations. It should be noted that such terms should not be construed as restrictive, and other techniques may be utilized to perform similar operations. Additionally, as mentioned herein, "generate", "calculate", "estimate", "use", "select", "access", and "determine" may be used interchangeably. For example, "generate", "calculate", "estimate", or "determine" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing (e.g., by another component or device) a parameter (or signal) that has already been generated.
[0028] Reference Figure 1 , certain illustrative aspects of system 100 include two or more audio output devices, such as first audio output device 110 and second audio output device 140, which are configured to perform data exchange to generate a mono audio output based on spatial audio data 106. In Figure 1 the specific implementation illustrated, the spatial audio data 106 is received from a host device 102 or accessed from a memory 114 of one of the audio output devices 110, 140.
[0029] The spatial audio data 106 represents sound from one or more sources in three dimensions (3D), which may include real or virtual sources, such that the audio output representing the spatial audio data 106 can simulate the distance and direction between a listener and the one or more sources. The spatial audio data 106 may be encoded using various coding schemes, such as first-order ambisonics (FOA), higher-order ambisonics (HOA), or equivalent spatial domain (ESD) representation (described further below). As an example, a total of four channels (such as two stereo channels) may be used to encode the FOA coefficients or ESD data for representing the spatial audio data 106.
[0030] Each of the audio output devices 110, 140 is configured to generate a mono audio output based on the spatial audio data 106. In a specific example, the first audio output device 110 is configured to generate a first mono audio output 152 for a user's first ear, and the second audio output device 140 is configured to generate a second mono audio output 154 for a user's second ear. In this example, the first mono audio output 152 and the second mono audio output 154 together simulate the spatial relationship of the sound source relative to the user's ears, such that the user perceives the mono audio output as spatial audio.
[0031] Figure 1FIG. 160 in shows the conversion of spatial audio data 106 to a mono audio output 188, which corresponds to or includes one or more of a first mono audio output 152 and / or a second mono audio output 154. In some implementations, each of the audio output devices 110, 140 performs one or more of the operations shown in FIG. 160. In other implementations, one of the audio output devices 110, 140 performs one or more of the operations illustrated in FIG. 160 and shares the results of one or more of the operations with the other of the audio output devices 110, 140 as exchange data, as further described below.
[0032] In FIG. 160, the spatial audio data 106 is in an ESD representation 164 when received. In the ESD representation 164, the spatial audio data 106 includes four channels, which represent virtual speakers 168, 170, 172, 174 placed around the user 166. For example, the ESD representation 164 of the spatial audio data can be encoded as first audio data corresponding to a first plurality of virtual speakers or other sound sources of a 3D sound field and second audio data corresponding to a second plurality of virtual speakers or other sound sources of the 3D sound field, where the first plurality of sound sources is different from the second plurality of sound sources. For illustration, in FIG. 160, the sounds corresponding to speakers 168 and 170 can be encoded in the first audio data (e.g., as a first stereo channel), and the sounds corresponding to speakers 172 and 174 can be encoded in the second audio data (e.g., as a second stereo channel). By controlling the amplitude, frequency, and / or phase of the sounds assigned to each of the virtual speakers 168, 170, 172, 174, the ESD representation 164 can simulate sounds from one or more virtual sound sources at various distances from the user 166 and in various directions relative to the user 166.
[0033] In the example illustrated in FIG. 160, the ESD representation 164 is converted to a stereophonic reverberation representation (e.g., stereophonic reverberation coefficients) of the spatial audio data 106 at block 176. For example, block 176 may receive a set of audio input signals of the ESD representation 164 and convert it to a set of audio output signals in a stereophonic reverberation domain (e.g., FOA or HOA domain). In some implementations, the stereophonic reverberation data is in a stereophonic audio channel number (ACN) or semi-normalized 3D (SN3D) data format. The audio input signals of the ESD representation 164 correspond to the spatial audio data 106 (e.g., immersive audio content) rendered at a pre-determined set of virtual speaker positions. In the ESD domain of the ESD representation 164, the virtual speakers may be located at different positions around a sphere (e.g., flye points) to preserve the stereophonic reverberation rendering energy per area on the sphere or the stereophonic reverberation rendering energy per volume within the sphere. In a particular implementation, the operations of block 176 may be performed using a stereophonic reverberation decoder module (e.g., the "AmbiX decoder" available from https: / / github.com / kronihias / ambix with a transformation matrix that takes into account the ESD representation 164).
[0034] In this example, the stereophonic reverberation representation of the spatial audio data 106 is used to perform a rotation operation 180 based on motion data 178 from one or more motion sensors (e.g., the motion sensor 116 of the first audio output device 110). In a particular implementation, the technique described at https: / Ambonics.iem.at / xchange / fileformat / docs / spherical-harmonics-rotation is used to perform the sound field rotation. For illustration, a sound field rotator module (such as the "AmbiX sound field rotator" plugin available from https: / github.com / kronihias / ambix) may be used to perform the sound field rotation operation 180. The rotation operation 180 takes into account the change in the relative positions of the user 166 and the virtual speakers 168, 170, 172, 174 due to the movement indicated by the motion data 178.
[0035] Continuing with the above example, at block 182, the rotated stereophonic reverberation representation of the spatial audio data 106 is converted back to the ESD domain as the rotated ESD representation 184. In a particular implementation, the ESD domain of the ESD representation 184 is different from the ESD domain of the ESD representation 164. For illustration, the virtual speakers of the ESD domain including the ESD representation 184 may be located at different positions around the sphere than the virtual speakers of the ESD domain of the ESD representation 164.
[0036] In certain aspects, block 182 receives (N + 1) 2 signals of the stereophonic reverberation representation, where N is an integer representing the order of the stereophonic reverberation (e.g., for first-order stereophonic reverberation, N = 1, for second-order stereophonic reverberation, N = 2, etc.). In this particular aspect, block 182 outputs (N + 1) 2 signals, where each signal corresponds to a virtual loudspeaker in the ESD domain. When the arrangement of the virtual loudspeakers is appropriately selected (e.g., based on a t-design grid of points on a sphere), the following property holds: H * E = I, where I is an (N + 1) 2 x (N + 1) 2 identity matrix, H is a matrix representing the pressure signal entering the stereophonic reverberation domain, and E is the ESD transformation matrix of the virtual loudspeakers that converts the stereophonic reverberation signal to the ESD domain. In this particular aspect, the conversion between the stereophonic reverberation domain and the ESD domain is lossless.
[0037] In the example shown in FIG. 160, one or more head-related transfer (HRT) functions 130 are applied to the rotated ESD representation 184 to generate a monaural audio output 188. The application of the HRT function 130 determines the sound level, frequency, phase, other audio information, or a combination of one or more of these to generate a monaural audio output 188 that is provided to an audio output device (e.g., an audio output device that provides the audio output to one ear of the user 166) to simulate the perception of one or more virtual sound sources at various distances from the user 166 and in various directions relative to the user 166.
[0038] Although Figure 1 FIG. 160 of shows that the spatial audio data 106 is received in the ESD representation 164, converted to a rotated stereophonic reverberation representation, and then converted to the rotated ESD representation 184, in other examples, the spatial audio data 106 is received in the stereophonic reverberation representation, and the conversion operation of block 176 is omitted. In other examples, the rotation operation 180 is performed using the ESD representation 164, and the conversion operations of blocks 176 and 182 are omitted.
[0039] In certain aspects, each of the audio output devices 110, 140 is configured to perform at least a subset of the operations illustrated in FIG. 160 such that a pair of monaural audio outputs (e.g., a first monaural audio output 152 and a second monaural audio output 154) are generated during the operation to simulate the entire 3D sound field. In a particular example, a personal audio device includes the audio output devices 110, 140, which generate separate sound outputs for each ear of the user, such as a head-mounted device (an example of which is shown in Figure 5 ), earbuds (an example of which is shown in Figure 6shown) or a multimedia headset device (an example of which is shown in Figure 7 shown).
[0040] In Figure 1 FIG. 11A, the first audio output device 110 includes one or more processors 112, a memory 114, a receiver 126, and a transceiver 124. The processor 112 is configured to execute instructions 132 from the memory 114 to perform one or more of the operations shown in FIG. 160. In Figure 1 the example shown in FIG. 11A, the receiver 126 and the transceiver 124 are separate components of the first audio output device 110; however, in other implementations, the transceiver 124 includes the receiver 126, or the receiver 126 and the transceiver 124 are combined in a radio chipset. In Figure 1 FIG. 11A, the first audio output device 110 also includes a modem 122 coupled to one or more processors and coupled to the receiver 126, the transceiver 124, or both. The first audio output device 110 also includes an audio codec 120, which is coupled to the modem 122, one or more processors 112, or both, and is coupled to one or more audio transducers 118. In Figure 1 FIG. 11A, the first audio output device 110 includes one or more motion sensors 116 coupled to the processor 112.
[0041] In Figure 1 the example shown in FIG. 11B, the second audio output device 140 includes a transceiver 148, a modem 146, an audio codec 144, and one or more audio transducers 142. In other examples, the second audio output device 140 also includes one or more processors, a memory, one or more motion sensors, other components, or a combination thereof. For illustration, in some implementations, the second audio output device 140 includes the same features and components as the first audio output device 110.
[0042] In a particular implementation, the receiver 126 is configured to receive a wireless transmission 104 from the host device 102, and the transceiver 124 is configured to support data exchange with the second audio output device 140. For example, the transceiver 124 of the first audio output device 110 and the transceiver 148 of the second audio output device 140 can be configured to establish a wireless peer-to-peer ad-hoc link 134 to support data exchange. In this example, the wireless peer-to-peer ad-hoc link 134 can include a connection that complies with the protocol specification (BLUETOOTH is a registered trademark of Bluetooth SIG, Inc., of Kirkland, Wash., USA), complies with A connection that complies with a protocol specification (IEEE is a registered trademark of the Institute of Electrical and Electronics Engineers, Inc. in Piscataway, New Jersey, USA), a connection that complies with a proprietary protocol, or another wireless peer-to-peer ad-hoc connection. The wireless peer-to-peer ad-hoc link 134 between the first audio output device 110 and the second audio output device 140 may use the same protocol as the wireless transmission 104 from the host device 102, or may use one or more different protocols. For illustration, the host device 102 may transmit spatial audio data 106 via a BLUETOOTH connection, and the first audio output device 110 and the second audio output device 140 may exchange data via a proprietary connection.
[0043] During operation, the first audio output device 110 obtains the spatial audio data 106 by reading the spatial audio data 106 from the memory 114 or via the wireless transmission 104 from the host device 102. In certain aspects, the first audio output device 110 and the second audio output device 140 each receive a part or all of the spatial audio data 106 from the host device 102, and the audio output devices 110, 140 exchange data to generate their respective mono audio outputs 152, 154.
[0044] In some implementations, the spatial audio data 106 includes four channels of audio data encoded as two stereo channels corresponding to the ESD representation 164 of the spatial audio data 106. In some such implementations, two of the four channels are encoded (e.g., as the first stereo channel) and sent to the first audio output device 110, and the other two of the four channels are encoded (e.g., as the second stereo channel) and sent to the second audio output device 140. Figure 2 Examples illustrating such implementations are described further below. In other implementations, the spatial audio data 106 is encoded as two stereo channels or stereo reverb coefficients and sent to only one of the audio output devices 110, 140, such as sent to the first audio output device 110. Figure 3 Examples illustrating such implementations. In other implementations, the spatial audio data 106 is encoded as two stereo channels or stereo reverb coefficients and sent to the two audio output devices 110, 140. Figure 4 Examples illustrating such implementations.
[0045] In an implementation where a first part of the spatial audio data 106 is sent to the first audio output device 110 and a second part of the spatial audio data 106 is sent to the second audio output device 140 (such as Figure 2In the example of [description], the first audio output device 110 decodes the first part of the spatial audio data 106 and sends the first exchange data 136 to the second audio output device 140. In this implementation, the first exchange data 136 may include data representing the decoded first part of the spatial audio data 106, such as stereo reverberation coefficients corresponding to the first part of the spatial audio data, audio waveform data (e.g., pulse code modulation (PCM) data), and / or other data representing the decoded first part of the spatial audio data 106. Similarly, in this implementation, the second audio output device 140 decodes the second part of the spatial audio data 106 and sends the second exchange data 150 to the first audio output device 110. The data exchange may also include synchronization data 138 (generated by one or both of the audio output devices 110, 140) to facilitate synchronization of the playback of the audio output devices 110, 140. In this implementation, after the data exchange, each of the audio output devices 110, 140 has both the first part and the second part of the spatial audio data 106. In some exemplary implementations, each audio output device may assemble the first part and the second part of the spatial audio data 106 for mono audio output.
[0046] In an implementation where the spatial audio data 106 is only sent to the first audio output device 110 (such as Figure 3 In the example of [description], the first audio output device 110 decodes the spatial audio data 106 and sends the first exchange data 136 to the second audio output device 140. In this implementation, the first exchange data 136 includes data representing the decoded spatial audio data 106, such as stereo reverberation coefficients corresponding to the spatial audio data, audio waveform data (e.g., pulse code modulation (PCM) data), and / or other data representing the decoded spatial audio data 106. The first audio output device 110 may send the decoded spatial audio data 106 to the second audio output device 140 before performing the operations described with reference to FIG. 160 or after any of the operations described with reference to FIG. 160. For illustration, the first exchange data 136 may include the ESD representation 164 of the spatial audio data 106, the stereo reverberation coefficients output by block 176, the rotated stereo reverberation coefficients generated by the rotation operation 180 based on the motion data 178, or the rotated ESD representation 184.
[0047] The data exchange may also include synchronization data 138 to facilitate synchronization of the playback of the audio output devices 110, 140. In this implementation, after the data exchange, each of the audio output devices 110, 140 has the entire content of the spatial audio data 106.
[0048] In an implementation where the spatial audio data 106 is sent to both the first audio output device 110 and the second audio output device 140 (such as Figure 4 example), each of the audio output devices 110, 140 decodes the spatial audio data 106, and one or both of the audio output devices 110, 140 send the exchange data 136, 150 to the other audio output devices 110, 140. In this implementation, the exchange data 136, 150 includes or corresponds to the synchronization data 138 to facilitate the synchronization of the playback of the audio output devices 110, 140. In this implementation, each of the audio output devices 110, 140 has the entire content of the spatial audio data 106 before the data exchange, and the data exchange is used for synchronized playback.
[0049] In some implementations, each of the audio output devices 110, 140 includes one or more motion sensors, such as the motion sensor 116. In such an implementation, each of the audio output devices 110, 140 performs a rotation operation 180 based on the motion data 178 from the corresponding motion sensor. In other implementations, only one of the audio output devices 110, 140 includes a motion sensor, and the exchange data includes the motion data 178, or the exchange data sent from one audio output device to the other (e.g., from the first audio output device 110 to the second audio output device 140) includes the rotated ESD representation 184 of the spatial audio data 106, the rotated stereo reverberation coefficient, or other data representing the rotated 3D sound field.
[0050] In various implementations, the audio output devices 110, 140 may have more or fewer components than shown in Figure 1 . In a particular implementation, the processor 112 includes one or more central processing units (CPUs), one or more digital signal processors (DSPs), one or more other single-core or multi-core processing devices, or a combination thereof (e.g., CPU and DSP). The processor 112 may include a voice and music codec, which includes a voice codec (“vocoder”) encoder, a vocoder decoder, or a combination thereof.
[0051] In a particular implementation, a portion of the first audio output device 110, a portion of the second audio output device 140, or both may be included in a system-in-package or system-on-chip device. In a particular implementation, the memory 114, the processor 112, the audio codec 120, the modem 122, the transceiver 124, the receiver 126, the motion sensor 116, or a subset or combination thereof is included in a system-in-package or system-on-chip device.
[0052] In certain aspects, system 100 facilitates the generation of head-tracked immersive audio (e.g., first mono audio output 152 and second mono audio output 154) by using a simplified 3D sound field communication scheme (e.g., ESD representation 164) to reduce the number of convolution operations performed on first audio output device 110 and second audio output device 140, by exchanging data between first audio output device 110 and second audio output device 140 to reduce the number of stereo decoding operations performed to generate first mono audio output 152 and second mono audio output 154, by generating mono audio outputs (e.g., first mono audio output 152 and second mono audio output 154), or by a combination thereof. In certain aspects, system 100 facilitates balancing resource requirements between a pair of audio output devices (e.g., first audio output device 110 and second audio output device 140) to extend the duration of 3D sound field reproduction that can be provided by the audio output devices.
[0053] Figure 2 , Figure 3 and Figure 4 are block diagrams of specific illustrative examples of Figure 1 system 100 in accordance with some aspects of the present disclosure. Figure 2 , Figure 3 and Figure 4 each illustrate additional aspects of Figure 1 host device 102, first audio output device 110, and second audio output device 140 of Figure 2 illustrates an example in which a first portion of spatial audio data 106 is sent by host device 102 to first audio output device 110 and a second portion of spatial audio data 106 is sent to second audio output device 140. Figure 3 illustrates an example in which spatial audio data 106 is sent by host device 102 only to first audio output device 110. Figure 4 illustrates an example in which spatial audio data 106 is sent by host device 102 to both first audio output device 110 and second audio output device 140. Although Figures 2 - 4 illustrates different hardware configurations of host device 102 and / or one or more of audio output devices 110, 140, in some implementations, Figures 2 - 4 represents different operating modes of the same hardware. For example, each of audio output devices 110, 140 may include two stereo decoders, as shown in Figure 4 ; however, when host device 102 or audio output devices 110, 140 operate in a particular operating mode corresponding to Figure 2 , each audio output device 110, 140 uses only one of its stereo decoders.
[0054] In the examples shown in each of Figures 2 - 4 , host device 102 includes a receiver 204 and a modem 206 to receive and decode media that includes or represents spatial audio data 106. In one particular example, the media includes a game that generates an audio output corresponding to the spatial audio data 106. In another particular example, the media includes virtual reality and / or augmented reality media that generates an audio output corresponding to the spatial audio data 106. In other examples, the media includes and / or corresponds to audio, video (e.g., 2D or 3D video), mixed reality media, other media content, or a combination of one or more of these. Host device 102 also includes a memory 202 for storing the downloaded media for subsequent processing.
[0055] In Figures 2 - 4 , host device 102 includes a 3D audio converter 208. The 3D audio converter 208 is configured to generate data representing a 3D sound field based on the media. For example, an ESD representation can be used to represent the 3D sound field, as described in reference to Figure 1 . Using the ESD representation enables encoding the 3D sound field in four channels, one channel per virtual speaker, which conveniently enables transmitting the entire 3D sound field using two stereo audio channels.
[0056] In Figures 2 - 4 , host device 102 includes a pair of stereo encoders, including a first stereo encoder 210 and a second stereo encoder 212. In a particular aspect, the first stereo encoder 210 is configured to encode a first portion 215 of the spatial audio data 106, and the second stereo encoder 212 is configured to encode a second portion 217 of the spatial audio data 106. In an example where the 3D audio converter 208 uses the Figure 1 ESD representation 164 to generate data representing the 3D sound field, the first portion 215 represents stereo audio (e.g., differential audio) for a first pair of virtual speakers (e.g., virtual speakers 168 and 170), and the second portion 217 represents stereo audio for a second pair of virtual speakers (e.g., virtual speakers 172 and 174).
[0057] In Figures 2 - 4 , host device 102 includes one or more modems 214 and one or more transmitters 216, which are coupled to the stereo encoders 210, 212 and are configured to encode and transmit the first portion 215 and the second portion 217 of the spatial audio data 106. In the example shown in Figure 2 , the first portion 215 is sent to a first audio output device 110, and the second portion 217 is sent to a second audio output device 140. In Figure 3In the example shown, the first portion 215 and the second portion 217 are combined, interleaved, or otherwise sent together to the first audio output device 110. In Figure 4 the example shown, the first portion 215 and the second portion 217 are combined, interleaved, or otherwise sent together to both the first audio output device 110 and the second audio output device 140.
[0058] Referring Figure 2 to the example illustrated in Figure 1 the first audio output device 110 includes the receiver 126, the modem 122, the audio codec 120, the memory 114, the transceiver 124, the motion sensor 116, and the processor 112 described in Figure 2 In addition, in Figure 2 the second audio output device 140 includes components similar to those of the first audio output device 110. For example, in Figure 1 the second audio output device 140 includes the modem 122, the audio codec 144, the transceiver 148, one or more motion sensors 246, and one or more processors 244 described in Figure 2 In the example of Figure 1 the receiver 238, the memory 242, the motion sensor 246, and the processor 244 are generally similar to the receiver 126, the memory 114, the motion sensor 116, and the processor 112 of the first audio output device 140, respectively, and operate in the same or generally similar manner as described in
[0059] In Figure 2 the audio codec 120 includes the stereo decoder 220 configured to decode the first portion 215 of the spatial audio data 106. The audio codec 120 is configured to provide the decoded first portion 215 of the spatial audio data 106 to the transceiver 124 for transmission to the second audio output device 140 together with the exchange data 250. The audio codec 120 is also configured to store the decoded first portion 215 of the spatial audio data 106 at the buffer 222 of the memory 114.
[0060] In Figure 2In [the above], the audio codec 144 further includes a stereo decoder 240 configured to decode a second portion 217 of the spatial audio data 106. The audio codec 120 is configured to provide the decoded second portion 217 of the spatial audio data 106 to the transceiver 148 for transmission to the first audio output device 110 together with the exchange data 250. The audio codec 120 is further configured to store the decoded second portion 217 of the spatial audio data 106 at a buffer 254 of the memory 242.
[0061] The decoded first portion 215 of the spatial audio data 106 stored in the buffers 222 and 252 includes data frames (e.g., time window segments of audio data) representing two virtual audio sources of a 3D sound field, and the decoded second portion 217 of the spatial audio data 106 stored in the buffers 224 and 254 includes data frames representing two other virtual audio sources of the 3D sound field. Each data frame in the data frames may include synchronization data (such as a frame sequence identifier, a play timestamp, or other synchronization data) or be associated with the synchronization data. In certain aspects, the synchronization data is transmitted between the audio output devices 110, 140 via the exchange data 250.
[0062] In Figure 2 the example of Figure 1 the instructions 132 of Figure 1 are configured to perform various operations, such as one or more operations described with reference to Figure 2 the figure 160 of Figure 1 For example, in Figure 2 the (one or more) processors 112 include an aligner 226, a 3D sound converter 228, a sound field rotator 230, a 3D sound converter 232, and a monoizer 234. Similarly, the processor 244 is configured to execute instructions to perform various operations, such as one or more operations described with reference to
[0063] the figure 160 of Figure 1 Figure 1to align the data frames with the synchronized data 138. Similarly, the aligner 256 is configured to obtain the decoded first portion 215 of the data frame from the buffer 252 and combine or align the decoded first portion 215 of the data frame with the corresponding (e.g., time-aligned) data frame of the decoded second portion 217 from the buffer 254. In each case, the decoded first portion 215 of the data frame and the corresponding data frame of the decoded second portion 217 together form a data frame (e.g., Figure 1 a data frame or time window segment of the ESD representation 164) of the spatial audio data representing the 3D sound field.
[0064] The 3D sound converter 228 is configured to convert the spatial audio data representing the 3D sound field into a computationally efficient format to perform a sound field rotation operation. For example, the 3D sound converter 228 can perform an ESD-to-stereo reverberation conversion as described in reference to Figure 1 block 176. In this example, the 3D sound converter 228 converts the data frame of the ESD representation of the spatial audio data into the corresponding stereo reverberation coefficients. The 3D sound converter 258 is configured to perform a similar conversion of the spatial audio data, such as by generating stereo reverberation coefficients based on the data frame of the ESD representation of the spatial audio data.
[0065] The sound field rotator 230 is configured to modify the 3D sound field based on the motion data (e.g., motion data 178) from the motion sensor 116. For example, the sound field rotator 230 can perform a rotation operation 180 by determining a transformation matrix based on the motion data and applying the transformation matrix to the stereo reverberation coefficients to generate a rotated 3D sound field. Similarly, the sound field rotator 260 is configured to modify the 3D sound field based on the motion data (e.g., motion data 178) from the motion sensor 246 to generate a rotated 3D sound field.
[0066] The 3D sound converter 232 is configured to convert the rotated 3D sound field into a computationally efficient format to perform a mono operation. For example, the 3D sound converter 232 can perform a stereo reverberation-to-ESD conversion as described in reference to Figure 1 block 182. In this example, the 3D sound converter 228 converts the data frame of the corresponding rotated stereo reverberation coefficients into an ESD representation. The 3D sound converter 262 is configured to perform a similar conversion of the spatial audio data, such as by generating an ESD representation based on the rotated stereo reverberation coefficients.
[0067] The mono - mizer 234 is configured to generate first mono - audio data 152 by applying a head - related transfer (HRT) function 236 to the rotated 3D sound field to generate audio data (e.g., mono - audio data) for a single ear of the user, simulating how that ear would receive sound in the 3D sound field. The mono - mizer 264 is configured to generate second mono - audio data 154 by applying an HRT function 266 to the rotated 3D sound field to generate audio data for the other ear of the user, simulating how the other ear would receive sound in the 3D sound field.
[0068] In certain aspects, Figure 2 system 100 facilitates the generation of head - tracked immersive audio (e.g., first mono - audio output 152 and second mono - audio output 154) by using a simplified 3D sound - field communication scheme (e.g., ESD representation 164) to reduce the number of convolution operations performed on the first audio output device 110 and the second audio output device 140, by exchanging data between the first audio output device 110 and the second audio output device 140 to reduce the number of stereo - decoding operations performed to generate the first mono - audio output 152 and the second mono - audio output 154, by generating mono - audio outputs (e.g., first mono - audio output 152 and second mono - audio output 154), or by a combination thereof. In certain aspects, system 100 facilitates balancing the resource requirements between a pair of audio output devices (e.g., first audio output device 110 and second audio output device 140) to extend the duration of 3D sound - field reproduction that can be provided by the audio output devices.
[0069] Referring Figure 3 to the example illustrated in Figure 1 and Figure 2 the first audio output device 110 includes the receiver 126, modem 122, audio codec 120, memory 114, transceiver 124, motion sensor 116, and processor 112 described in Figure 3 In addition, in Figure 1 and Figure 2 the second audio output device 140 includes the transceiver 148, memory 242, motion sensor 246, and processor 244 described in
[0070] In Figure 3In, the audio codec 120 includes a stereo decoder 220 configured to decode a first portion 215 of the spatial audio data 106 and a stereo decoder 320 configured to decode a second portion 217 of the spatial audio data 106. The audio codec 120 is configured to provide the decoded spatial audio data 106 (e.g., both the first portion 215 and the second portion 217) to the transceiver 124 for transmission to the second audio output device 140 together with the exchange data 350. The audio codec 120 is further configured to store the decoded spatial audio data 106 at a buffer 322 of the memory 114. In Figure 3 In, the transceiver 148 of the second audio output device is configured to store the decoded spatial audio data received via the exchange data 350 at a buffer 352 of the memory 242.
[0071] The decoded spatial audio data 106 stored in the buffers 322 and 352 includes data frames representing four virtual audio sources of a 3D sound field or stereo reverberation coefficients (e.g., a time window segment of the audio data). Each data frame in the data frames may include or be associated with synchronization data such as a frame sequence identifier, a play timestamp, or other synchronization data. In a particular aspect, the synchronization data is transmitted between the audio output devices 110, 140 via the exchange data 350.
[0072] In Figure 3 the example of, the processors 112 and 244 are configured to execute instructions (e.g., Figure 1 the instructions 132) to perform various operations such as one or more operations described with reference to Figure 1 Figure 160 of. For example, in Figure 3 the, the processor 112 includes an aligner 226, a 3D sound converter 228, a sound field rotator 230, a 3D sound converter 232, and a monoizer 234 described with reference to Figure 2 . Similarly, the processor 244 includes an aligner 256, a 3D sound converter 258, a sound field rotator 260, a 3D sound converter 262, and a monoizer 264 described with reference to Figure 2 .
[0073] In a particular aspect, Figure 3System 100 facilitates the generation of head-tracked immersive audio (e.g., first mono audio output 152 and second mono audio output 154) by using a simplified 3D sound field communication scheme (e.g., ESD representation 164) to reduce the number of convolution operations performed on first audio output device 110 and second audio output device 140, by exchanging data between first audio output device 110 and second audio output device 140 to reduce the number of stereo decoding operations performed to generate first mono audio output 152 and second mono audio output 154, by generating mono audio outputs (e.g., first mono audio output 152 and second mono audio output 154), or by a combination thereof.
[0074] Referring Figure 4 to the example illustrated in Figure 1 and Figure 2 described, first audio output device 110 includes Figure 4 and Figure 1 the Figure 2 receiver 126, modem 122, audio codec 120, memory 114, transceiver 124, motion sensor 116, and processor 112 described. Additionally, in
[0075] In Figure 4 , audio codec 120 includes stereo decoder 220 configured to decode a first portion 215 of spatial audio data 106 and stereo decoder 320 configured to decode a second portion 217 of spatial audio data 106. Audio codec 120 is configured to store the decoded spatial audio data 106 at buffer 322 of memory 114. Additionally, audio codec 144 includes stereo decoder 440 configured to decode a first portion 215 of spatial audio data 106 and stereo decoder 240 configured to decode a second portion 217 of spatial audio data 106. Audio codec 144 is configured to store the decoded spatial audio data 106 at buffer 352 of memory 242.
[0076] In Figure 4In [the system], transceivers 124 and 148 exchange exchange data 450, which includes synchronization data to facilitate the synchronization of first mono audio output data 152 and second mono audio output data 154. For example, when reading data frames from respective buffers 322, 352, one or both of aligners 226, 256 may initiate the transmission of synchronization data. As another example, another component of either processor 112, processor 244, or audio output devices 110, 140 may generate a synchronization signal (e.g., a clock signal) communicated via exchange data 450.
[0077] In Figure 4 the example of [the system], processors 112 and 244 are configured to execute instructions (e.g., Figure 1 instructions 132 of [the system]) to perform various operations, such as one or more operations described with reference to Figure 1 FIG. 160 of [the system]. For example, in Figure 4 the [system], processor 112 includes aligner 226, 3D sound converter 228, sound field rotator 230, 3D sound converter 232, and monoizer 234 described with reference to Figure 2 [the system]. Similarly, processor 244 includes aligner 256, 3D sound converter 258, sound field rotator 260, 3D sound converter 262, and monoizer 264 described with reference to Figure 2 [the system].
[0078] In certain aspects, system 100 facilitates the generation of head-tracked immersive audio (e.g., first mono audio output 152 and second mono audio output 154) by using a simplified 3D sound field communication scheme (e.g., ESD representation 164) to reduce the number of convolution operations performed on first audio output device 110 and second audio output device 140, by exchanging data between first audio output device 110 and second audio output device 140 to reduce the number of stereo decoding operations performed to generate first mono audio output 152 and second mono audio output 154, by generating mono audio outputs (e.g., first mono audio output 152 and second mono audio output 154), or by a combination thereof. In certain aspects, system 100 facilitates the balancing of resource requirements between a pair of audio output devices (e.g., first audio output device 110 and second audio output device 140) to extend the duration of 3D sound field reproduction that can be provided by the audio output devices.
[0079] Figure 5 is a diagram of a head-mounted device 500 (e.g., a specific example of a personal audio device) (such as headphones) operable to perform data exchange to enable the generation of mono audio outputs from spatial audio data according to some examples of the present disclosure. In Figure 5Among them, the first earcup of the head-mounted device 500 includes a first audio output device 110, and the second earcup of the head-mounted device 500 includes a second audio output device 140. The head-mounted device 500 may further include one or more microphones 502. In a specific example, the audio output devices 110, 140 operate as described in any one of the references Figures 1 - 4 For example, Figure 5 The first audio output device 110, the second audio output device 140, or both are configured to obtain spatial audio data; perform data exchange for exchanging data with another audio output device based on the spatial audio data; and generate a mono audio output based on the spatial audio data.
[0080] Figure 6 FIG. is a diagram of an earbud 600 (e.g., another specific example of a personal audio device) operable to perform data exchange to enable generation of a mono audio output from spatial audio data according to some examples of the present disclosure. In Figure 6 Among them, the first earbud 602 includes or corresponds to the first audio output device 110, and the second earbud 604 includes or corresponds to the second audio output device 140. One or both of the earbuds 600 may further include one or more microphones. In a specific example, the audio output devices 110, 140 operate as described in any one of the references Figures 1 - 4 For example, Figure 6 The first audio output device 110, the second audio output device 140, or both are configured to obtain spatial audio data; perform data exchange for exchanging data with another audio output device based on the spatial audio data; and generate a mono audio output based on the spatial audio data.
[0081] Figure 7 FIG. is a diagram of a head-mounted device 700 (e.g., another specific example of a personal audio device) (such as a virtual reality head-mounted device, an augmented reality head-mounted device, or a mixed reality head-mounted device) operable to perform data exchange to enable generation of a mono audio output from spatial audio data according to some examples of the present disclosure. In Figure 7 Among them, the first audio output device 110 is included in the head-mounted device 700 or coupled to the head-mounted device 700 at a position close to the user's first ear, and the second audio output device 140 is included in the head-mounted device 700 or coupled to the head-mounted device 700 at a position close to the user's second ear. The head-mounted device 700 may further include one or more microphones 710 and one or more display devices 712. In a specific instance, the audio output devices 110, 140 operate as described in any one of the references Figures 1 to 4 For example, Figure 7The first audio output device 110, the second audio output device 140, or both are configured to obtain spatial audio data; perform data exchange of exchanging data with another audio output device based on the spatial audio data; and generate a mono audio output based on the spatial audio data.
[0082] Figure 8 is a diagram of a specific implementation of a method 800 for generating a mono audio output from spatial audio data performed by one or more of the audio output devices according to some examples of the present disclosure. In a particular aspect, one or more operations of the method 800 are performed by Figures 1 - 7 at least one of the first audio output device 110, the processor 112, the receiver 126, the transceiver 124, the audio transducer 118, the second audio output device 140, the transceiver 148, the audio transducer 142, or a combination of one or more of its components of Figures 1 - 4 In another particular aspect, one or more operations of the method 800 are performed by Figures 2 - 4 at least one of the receiver 238 or the processor 244 of
[0083] The method 800 includes, at block 802, obtaining spatial audio data at a first audio output device. For example, Figures 1 - 7 the first audio output device 110 of
[0084] can obtain the spatial audio data 106 from the host device 102 or from the memory 114 via a wireless transmission 104. Figures 1 - 7 The method 800 includes: at block 804, performing data exchange of exchanging data between the first audio output device and the second audio output device based on the spatial audio data. For example,
[0085] the first audio output device 110 and the second audio output device 140 of Figures 1 to 7 can exchange data via a wireless transmission (such as via a wireless peer-to-peer ad-hoc link between the audio output devices 110, 140).
[0086] Figure 8 The method 800 of Figure 8 can be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. For example, Figure 8 the method 800 ofFigure 1 as described.
[0087] Although method 800 has been generally described from the perspective of the first audio output device 110, Figure 8 the second audio output device 140 may perform the operations of method 800 in addition to or instead of the first audio output device 110.
[0088] In certain aspects, method 800 enables the generation of a mono audio output (such as head-tracked immersive audio (e.g., Figures 1 - 4 the first mono audio output 152 and the second mono audio output 154)) based on spatial audio data. In certain aspects, method 800 enables the use of a simplified 3D sound field communication scheme (e.g., ESD representation 164) to reduce the number of convolution operations performed on an audio output device (e.g., the first audio output device 110 and the second audio output device 140), exchange data between audio output devices (e.g., the first audio output device 110 and the second audio output device 140) to reduce the number of stereo decoding operations performed to generate a mono audio output (e.g., the first mono audio output 152 and the second mono audio output 154) at each of the audio output devices, generate a mono audio output (e.g., the first mono audio output 152 and the second mono audio output 154), or a combination thereof. In certain aspects, method 800 is performed in a manner that balances the resource requirements between a pair of audio output devices (e.g., the first audio output device 110 and the second audio output device 140) to extend the duration of 3D sound field reproduction that can be provided by the audio output devices.
[0089] Figure 9 is a diagram of another specific implementation of method 900 for generating a mono audio output from spatial audio data performed by one or more of the Figures 1 - 7 audio output devices according to some examples of the present disclosure. In certain aspects, one or more operations of method 900 are performed by Figures 1 - 4 at least one of the first audio output device 110, the processor 112, the receiver 126, the transceiver 124, the audio transducer 118, the second audio output device 140, the transceiver 148, the audio transducer 142, or a combination of one or more of its components. In another specific aspect, one or more operations of method 800 are performed by Figures 2 - 4 at least one of the receiver 238 or the processor 244 or a combination thereof.
[0090] Method 900 includes, at block 902, obtaining spatial audio data at a first audio output device. In certain aspects, obtaining spatial audio data in method 900 includes: receiving spatial audio data from a host device via a wireless peer-to-peer ad-hoc link at block 904. For example, Figures 1 - 7 the first audio output device 110 may obtain spatial audio data 106 from the host device 102 via a wireless transmission 104.
[0091] Method 900 includes: at block 906, performing a data exchange that exchanges data between the first audio output device and the second audio output device based on the spatial audio data. In Figure 9 the specific example illustrated, performing the data exchange that exchanges data includes: sending first exchange data to the second audio output device at block 908, and receiving second exchange data from the second audio output device at block 910. In various implementations, as illustrative examples, the exchanged data includes audio waveform data, stereo reverberation coefficients, decoded stereo data, and / or ESD representations.
[0092] Method 900 includes: at block 912, determining a stereo reverberation coefficient based on the spatial audio data. For example, the 3D sound converter 228 may perform operations (e.g., the ESD to stereo reverberation operation at block 176) to determine the stereo reverberation coefficient based on the ESD representation of the spatial audio data. In some implementations, the 3D sound field representation representing the spatial audio data is generated partially based on the exchanged data.
[0093] Method 900 includes: at block 914, modifying the stereo reverberation coefficient based on motion data from one or more motion sensors to generate a modified stereo reverberation coefficient representing a rotated 3D sound field. For example, the sound field rotator 230 may perform operations (e.g., the rotation operation 180 based on the motion data 178) to determine the rotated 3D sound field.
[0094] Method 900 includes: at block 916, applying a head-related transfer function to the modified stereo reverberation coefficient to generate first mono audio data corresponding to the first audio output device. For example, the monoizer 234 may apply the HRT function 236 to the stereo reverberation coefficient to generate mono audio data 152.
[0095] Method 900 includes: at block 918, generating a first mono audio output at the first audio output device based on the spatial audio data. For example, the audio transducer 118 may generate the first mono audio output 152.
[0096] Although described generally from the perspective of the first audio output device 110 Figure 9Method 900, except that in addition to or instead of the first audio output device 110, the second audio output device 140 may perform the operations of method 900.
[0097] In certain aspects, method 900 enables the generation of a mono audio output (such as head-tracked immersive audio (e.g., Figures 1 - 4 the first mono audio output 152 and the second mono audio output 154)) based on spatial audio data. In certain aspects, method 800 enables the use of a simplified 3D sound field communication scheme (e.g., ESD representation 164) to reduce the number of convolution operations performed on the audio output devices (e.g., the first audio output device 110 and the second audio output device 140), exchange data between the audio output devices (e.g., the first audio output device 110 and the second audio output device 140) to reduce the number of stereo decoding operations performed to generate a mono audio output (e.g., the first mono audio output 152 and the second mono audio output 154) at each of the audio output devices, generate a mono audio output (e.g., the first mono audio output 152 and the second mono audio output 154), or a combination thereof. In certain aspects, method 900 is performed in a manner that balances the resource requirements between a pair of audio output devices (e.g., the first audio output device 110 and the second audio output device 140) to extend the duration of 3D sound field reproduction that can be provided by the audio output devices.
[0098] In combination with the described implementations, an apparatus includes a unit for obtaining spatial audio data at a first audio output device. For example, the unit for obtaining spatial audio data may correspond to the first audio output device 110, the processor 112, the receiver 126, one or more other circuits or components configured to obtain spatial audio data, or any combination thereof.
[0099] The apparatus further includes a unit for performing data exchange for exchanging data based on the spatial audio data. For example, the unit for performing data exchange may correspond to the first audio output device 110, the processor 112, the transceiver 124, one or more other circuits or components configured to perform data exchange, or any combination thereof.
[0100] The apparatus further includes: a unit for generating a first mono audio output at the first audio output device based on the spatial audio data. For example, the unit for generating the first mono audio output may correspond to the first audio output device 110, the processor 112, the audio transducer 118, one or more other circuits or components configured to generate a mono audio output, or any combination thereof.
[0101] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 114) includes instructions (e.g., instructions 132) that, when executed by one or more processors (e.g., processor 112), cause the one or more processors to obtain spatial audio data (e.g., spatial audio data 106) at a first audio output device (e.g., first audio output device 110). The instructions, when executed by the one or more processors, also cause the one or more processors to perform a data exchange of exchange data (e.g., first exchange data 136, second exchange data 150, synchronization data 138, or a combination thereof) based on the spatial audio data. The instructions, when executed by the one or more processors, also cause the one or more processors to generate a first mono audio output (e.g., first mono audio output 152) at the first audio output device based on the spatial audio data.
[0102] Specific aspects of the present disclosure are described in the following first set of related clauses:
[0103] According to Clause 1, a device includes: a memory configured to store instructions; and one or more processors configured to execute the instructions to perform the following operations: obtain spatial audio data at a first audio output device; perform a data exchange of exchange data between the first audio output device and a second audio output device based on the spatial audio data; and generate a first mono audio output at the first audio output device based on the spatial audio data.
[0104] Clause 2 includes the device according to Clause 1, and further includes a receiver coupled to the one or more processors and configured to receive the spatial audio data from a host device via a wireless peer-to-peer ad-hoc link.
[0105] Clause 3 includes the device according to Clause 1 or Clause 2, and further includes a transceiver coupled to the one or more processors, the transceiver being configured to send first exchange data to the second audio output device via a wireless link between the first audio output device and the second audio output device.
[0106] Clause 4 includes the device according to Clause 3, wherein the transceiver is further configured to receive second exchange data from the second audio output device via a wireless link.
[0107] Clause 5 includes the apparatus according to any one of Clauses 1 to 4, and further includes a transceiver coupled to the one or more processors, the transceiver being configured to send first exchange data to the second audio output device and receive second exchange data from the second audio output device, wherein generating the first mono audio output includes determining a 3D sound field representation based on the spatial audio data and the second exchange data.
[0108] Clause 6 includes the apparatus according to any one of Clauses 1 to 5, and further includes: a modem coupled to the one or more processors and configured to obtain spatial audio data via wireless transmission; and an audio codec coupled to the modem and coupled to the one or more processors, wherein the audio codec is configured to generate audio waveform data based on the spatial audio data.
[0109] Clause 7 includes the apparatus according to Clause 6, wherein the exchange data includes the audio waveform data.
[0110] Clause 8 includes the apparatus according to Clause 6, wherein the audio waveform data includes pulse code modulation (PCM) data.
[0111] Clause 9 includes the apparatus according to any one of Clauses 1 to 8, wherein the exchange data includes synchronization data to facilitate synchronization of the first mono audio output with a second mono audio output at the second audio output device.
[0112] Clause 10 includes the apparatus according to any one of Clauses 1 to 9, wherein the exchange data includes a stereo reverberation coefficient.
[0113] Clause 11 includes the apparatus according to any one of Clauses 1 to 10, wherein the first audio output device corresponds to a first earbud, a first speaker, or a first earcup of a headset device, and wherein the second audio output device corresponds to a second earbud, a second speaker, or a second earcup of the headset device.
[0114] Clause 12 includes the apparatus according to any one of Clauses 1 to 11, wherein the spatial audio data includes a stereo reverberation coefficient representing a 3D sound field.
[0115] Clause 13 includes the apparatus according to any one of Clauses 1 to 12, wherein the spatial audio data includes first audio data corresponding to a first plurality of sound sources of a 3D sound field and second audio data corresponding to a second plurality of sound sources of the 3D sound field, wherein the first plurality of sound sources is different from the second plurality of sound sources.
[0116] Clause 14 includes the apparatus according to any one of Clauses 1 to 13, wherein the spatial audio data includes first audio data corresponding to a first plurality of sound sources of a 3D sound field, and wherein performing the data exchange includes receiving second audio data from the second audio output device, wherein the second audio data corresponds to a second plurality of sound sources of the 3D sound field, and wherein the first plurality of sound sources is different from the second plurality of sound sources.
[0117] Clause 15 includes the apparatus according to any one of Clauses 1 to 14, and further includes one or more motion sensors coupled to the one or more processors, wherein the one or more processors are further configured to: determine a stereo reverberation coefficient based on the spatial audio data; modify the stereo reverberation coefficient based on motion data from the one or more motion sensors to generate a modified stereo reverberation coefficient representing a rotated 3D sound field; and apply a head-related transfer function to the modified stereo reverberation coefficient to generate first audio data corresponding to the first mono audio output.
[0118] Clause 16 includes the apparatus according to Clause 15, wherein the exchanged data includes data representing a rotated 3D sound field.
[0119] Clause 17 includes the apparatus according to Clause 15, wherein the stereo reverberation coefficient is further determined based on second exchanged data received from the second audio output device.
[0120] Clause 18 includes the apparatus according to any one of Clauses 1 to 17, and further includes a memory coupled to the one or more processors, wherein the spatial audio data is obtained from the memory.
[0121] According to Clause 19, a method includes: obtaining spatial audio data at a first audio output device; performing a data exchange of exchanging data between the first audio output device and the second audio output device based on the spatial audio data; and generating a first mono audio output at the first audio output device based on the spatial audio data.
[0122] Clause 20 includes the method according to Clause 19, and further includes receiving the spatial audio data from a host device via a wireless peer-to-peer ad-hoc link.
[0123] Clause 21 includes the method according to Clause 19 or Clause 20, wherein performing the data exchange includes transmitting first exchanged data to the second audio output device via a wireless link between the first audio output device and the second audio output device.
[0124] Clause 22 includes the method according to Clause 21, wherein performing data exchange includes receiving second exchange data from the second audio output device via a wireless link.
[0125] Clause 23 includes the method according to any one of Clauses 19 to 22, wherein performing data exchange includes sending first exchange data to the second audio output device and receiving second exchange data from the second audio output device, and wherein generating the first mono audio output includes determining a 3D sound field representation based on the spatial audio data and the second exchange data.
[0126] Clause 24 includes the method according to any one of Clauses 19 to 23, and further includes generating audio waveform data based on the spatial audio data.
[0127] Clause 25 includes the method according to Clause 24, wherein the exchange data includes the audio waveform data.
[0128] Clause 26 includes the method according to Clause 24, wherein the audio waveform data includes pulse code modulation (PCM) data.
[0129] Clause 27 includes the method according to any one of Clauses 19 to 26, wherein the exchange data includes synchronization data to facilitate synchronization of the first mono audio output with a second mono audio output at the second audio output device.
[0130] Clause 28 includes the method according to any one of Clauses 19 to 27, wherein the exchange data includes stereo reverberation coefficients.
[0131] Clause 29 includes the method according to any one of Clauses 19 to 28, wherein the first audio output device corresponds to a first earbud, a first speaker, or a first earcup of a headset device, and wherein the second audio output device corresponds to a second earbud, a second speaker, or a second earcup of the headset device.
[0132] Clause 30 includes the method according to any one of Clauses 19 to 29, wherein the spatial audio data includes stereo reverberation coefficients representing a 3D sound field.
[0133] Clause 31 includes the method according to any one of Clauses 19 to 30, wherein the spatial audio data includes first audio data corresponding to a first plurality of sound sources of a 3D sound field and second audio data corresponding to a second plurality of sound sources of the 3D sound field, wherein the first plurality of sound sources is different from the second plurality of sound sources.
[0134] Clause 32 includes the method according to any one of Clauses 19 to 31, wherein the spatial audio data includes first audio data corresponding to a first plurality of sound sources of a 3D sound field, and wherein performing the data exchange includes receiving second audio data from the second audio output device, wherein the second audio data corresponds to a second plurality of sound sources of the 3D sound field, and wherein the first plurality of sound sources is different from the second plurality of sound sources.
[0135] Clause 33 includes the method according to any one of Clauses 19 to 32, and further includes: determining a stereophonic reverberation coefficient based on the spatial audio data; modifying the stereophonic reverberation coefficient based on motion data from one or more motion sensors to generate a modified stereophonic reverberation coefficient representing a rotated 3D sound field; and applying a head-related transfer function to the modified stereophonic reverberation coefficient to generate first mono audio data corresponding to the first audio output device.
[0136] Clause 34 includes the method according to Clause 33, wherein the exchanged data includes data representing a rotated 3D sound field.
[0137] Clause 35 includes the method according to Clause 33, wherein the stereophonic reverberation coefficient is further determined based on second exchanged data received from the second audio output device.
[0138] Clause 36 includes the method according to any one of Clauses 19 to 35, wherein the spatial audio data is obtained from a memory.
[0139] According to Clause 37, a device includes: a unit for obtaining spatial audio data at a first audio output device; a unit for performing data exchange for exchanging data between the first audio output device and a second audio output device based on the spatial audio data; and a unit for generating a first mono audio output at the first audio output device based on the spatial audio data.
[0140] Clause 38 includes the device according to Clause 37, wherein the unit for performing data exchange includes a unit for receiving spatial audio data from a host device via a wireless peer-to-peer ad-hoc link.
[0141] Clause 39 includes the device according to Clause 37 or Clause 38, wherein the unit for performing data exchange includes a unit for sending first exchanged data to the second audio output device via a wireless link between the first audio output device and the second audio output device.
[0142] Clause 40 includes the apparatus according to Clause 39, wherein the unit for performing data exchange further includes a unit for receiving second exchange data from the second audio output device via the wireless link.
[0143] Clause 41 includes the apparatus according to any one of Clauses 37 to 40, wherein the unit for performing data exchange includes a unit for sending first exchange data to the second audio output device and receiving second exchange data from the second audio output device, and wherein generating the first mono audio output includes determining a 3D sound field representation based on the spatial audio data and the second exchange data.
[0144] Clause 42 includes the apparatus according to any one of Clauses 37 to 41, and further includes a unit for generating audio waveform data based on the spatial audio data.
[0145] Clause 43 includes the apparatus according to any one of Clauses 37 to 42, wherein the exchange data includes synchronization data to facilitate synchronization of the first mono audio output with the second mono audio output at the second audio output device.
[0146] Clause 44 includes the apparatus according to any one of Clauses 37 to 43, wherein the unit for obtaining spatial audio data, the unit for performing data exchange, and the unit for generating the first mono audio output are integrated within a first earbud, a first speaker, or a first earcup of a headset device, and wherein the second audio output device corresponds to a second earbud, a second speaker, or a second earcup of a headset device.
[0147] Clause 45 includes the apparatus according to any one of Clauses 37 to 44, wherein the spatial audio data includes first audio data corresponding to a first plurality of sound sources of a 3D sound field and second audio data corresponding to a second plurality of sound sources of the 3D sound field, and wherein the first plurality of sound sources is different from the second plurality of sound sources.
[0148] Clause 46 includes the apparatus according to any one of Clauses 37 to 45, wherein the spatial audio data includes first audio data corresponding to a first plurality of sound sources of a 3D sound field, and wherein the unit for performing data exchange includes a unit for receiving second audio data from the second audio output device, and wherein the second audio data corresponds to a second plurality of sound sources of the 3D sound field, and wherein the first plurality of sound sources is different from the second plurality of sound sources.
[0149] Clause 47 includes the apparatus according to any one of Clauses 37 to 46, and further includes a unit for storing the spatial audio data.
[0150] According to Clause 48, a non - transitory computer - readable storage device stores instructions that, when executable by one or more processors, cause the processors to perform the following operations: obtain spatial audio data at a first audio output device; perform a data exchange for exchanging data between the first audio output device and the second audio output device based on the spatial audio data; and generate a first mono audio output at the first audio output device based on the spatial audio data.
[0151] Clause 49 includes the non - transitory computer - readable storage device according to Clause 48, wherein the instructions further cause the one or more processors to perform the following operations: send first exchange data to the second audio output device and receive second exchange data from the second audio output device, wherein generating the first mono audio output includes determining a 3D sound field representation based on the spatial audio data and the second exchange data.
[0152] Clause 50 includes the non - transitory computer - readable storage device according to Clause 48 or Clause 49, wherein the instructions further cause the one or more processors to perform the following operations: generate audio waveform data based on the spatial audio data.
[0153] Clause 51 includes the non - transitory computer - readable storage device according to any one of Clauses 48 to 50, wherein the exchange data includes synchronization data to facilitate synchronization of the first mono audio output with a second mono audio output at the second audio output device.
[0154] Clause 52 includes the non - transitory computer - readable storage device according to any one of Clauses 48 to 52, wherein the spatial audio data includes first audio data corresponding to a first plurality of sound sources of a 3D sound field and second audio data corresponding to a second plurality of sound sources of the 3D sound field, wherein the first plurality of sound sources is different from the second plurality of sound sources.
[0155] Clause 53 includes the non - transitory computer - readable storage device according to any one of Clauses 48 to 53, wherein the instructions further cause the one or more processors to perform the following operations: determine a stereo reverberation coefficient based on the spatial audio data; modify the stereo reverberation coefficient based on motion data from one or more motion sensors to generate a modified stereo reverberation coefficient representing a rotated 3D sound field; and apply a head - related transfer function to the modified stereo reverberation coefficient to generate first mono audio data corresponding to the first audio output device.
[0156] Those skilled in the art will also understand that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of the two. The various illustrative components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the particular application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, and such implementation decisions will not be construed as departing from the scope of the present disclosure.
[0157] The steps of a method or algorithm described in connection with the implementations disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module can reside in random access memory (RAM), flash memory, read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, a hard disk, a removable disk, a compact disc read only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.
[0158] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features defined by the following claims.
Claims
1. A device for generating an audio output, comprising: a memory configured to store a head-related transfer function; and one or more processors configured to perform the following operations: obtain a first portion of spatial audio data representing a three-dimensional (3D) sound field at a first audio output device; decode the first portion of the spatial audio data at the first audio output device to obtain first exchange data, the first exchange data including data representing the decoded first portion of the spatial audio data; perform data exchange between the first audio output device and a second audio output device based on the spatial audio data to send the first exchange data from the first audio output device to the second audio output device and receive second exchange data from the second audio output device by the first audio output device, wherein the second exchange data includes data representing the decoded second portion of the spatial audio data; determine a stereo reverberation coefficient based on the spatial audio data; modify the stereo reverberation coefficient based on motion data from one or more motion sensors to generate a modified stereo reverberation coefficient representing a rotated 3D sound field; and apply the head-related transfer function to the modified stereo reverberation coefficient to generate a first mono audio output at the first audio output device based on a combination of data representing the decoded first portion of the spatial audio data and data representing the decoded second portion of the spatial audio data.
2. The device according to claim 1, further comprising: a receiver coupled to the one or more processors and configured to receive the first portion of the spatial audio data from a host device via a wireless peer-to-peer ad-hoc link.
3. The device according to claim 1, further comprising: a transceiver coupled to the one or more processors, the transceiver being configured to send the first exchange data to the second audio output device via a wireless link between the first audio output device and the second audio output device.
4. The device according to claim 3, wherein the transceiver is further configured to receive the second exchange data from the second audio output device via the wireless link.
5. The device according to claim 1, further comprising: a transceiver coupled to the one or more processors, the transceiver being configured to send the first exchange data to the second audio output device and receive the second exchange data from the second audio output device, wherein generating the first mono audio output includes: determining a 3D sound field representation based on a combination of data representing the decoded first portion of the spatial audio data and data representing the decoded second portion of the spatial audio data.
6. The device according to claim 1, further comprising: a modem coupled to the one or more processors and configured to obtain the first portion of the spatial audio data via wireless transmission; and An audio codec, coupled to the modem and coupled to the one or more processors, wherein the audio codec is configured to generate audio waveform data based on a first portion of the spatial audio data.
7. The apparatus according to claim 6, wherein, the first exchanged data includes the audio waveform data.
8. The apparatus according to claim 6, wherein, the audio waveform data includes pulse code modulation (PCM) data.
9. The apparatus according to claim 1, wherein, the first exchanged data or the second exchanged data includes synchronization data to facilitate synchronization of the first mono audio output with the second mono audio output at the second audio output device.
10. The apparatus according to claim 1, wherein, the first exchanged data or the second exchanged data includes the stereo reverberation coefficient.
11. The apparatus according to claim 1, wherein, the first audio output device corresponds to a first earbud, a first speaker, or a first earcup of a headset device, and wherein the second audio output device corresponds to a second earbud, a second speaker, or a second earcup of the headset device.
12. The apparatus according to claim 1, wherein, the spatial audio data includes the stereo reverberation coefficient.
13. The apparatus according to claim 1, wherein, the spatial audio data includes first audio data corresponding to a first plurality of sound sources of the 3D sound field and second audio data corresponding to a second plurality of sound sources of the 3D sound field, and wherein the first plurality of sound sources is different from the second plurality of sound sources.
14. The apparatus according to claim 1, wherein, a first portion of the spatial audio data includes first audio data corresponding to a first plurality of sound sources of the 3D sound field, wherein second exchanged data received from the second audio output device includes second audio data, wherein the second audio data corresponds to a second plurality of sound sources of the 3D sound field, and wherein the first plurality of sound sources is different from the second plurality of sound sources.
15. The apparatus according to claim 1, further comprising: one or more motion sensors coupled to the one or more processors, wherein the stereo reverberation coefficient is determined based on a first portion of the spatial audio data.
16. The apparatus according to claim 15, wherein, the stereo reverberation coefficient is further determined based on second exchanged data received from the second audio output device.
17. The apparatus according to claim 1, wherein, a first portion of the spatial audio data is obtained from the memory.
18. A method for generating an audio output, comprising: obtaining, at a first audio output device, a first portion of spatial audio data representing a three-dimensional (3D) sound field; decoding, at the first audio output device, the first portion of the spatial audio data to obtain first exchanged data, the first exchanged data including data representing the decoded first portion of the spatial audio data; Perform data exchange between the first audio output device and the second audio output device based on the spatial audio data, to send the first exchange data from the first audio output device to the second audio output device, and receive second exchange data from the second audio output device at the first audio output device, where the second exchange data includes data representing a decoded second part of the spatial audio data; Determine a stereo reverberation coefficient based on the spatial audio data; and Modify the stereo reverberation coefficient based on motion data from one or more motion sensors to generate a modified stereo reverberation coefficient representing a rotated 3D sound field; and Apply a head-related transfer function to the modified stereo reverberation coefficient to generate a first mono audio output at the first audio output device based on a combination of data representing a decoded first part of the spatial audio data and data representing a decoded second part of the spatial audio data.
19. The method according to claim 18, further comprising: Receiving a first part of the spatial audio data from a host device via a wireless peer-to-peer ad-hoc link.
20. The method according to claim 18, wherein generating the first mono audio output comprises: Determining a 3D sound field representation based on a combination of data representing a decoded first part of the spatial audio data and data representing a decoded second part of the spatial audio data.
21. The method according to claim 18, wherein the first exchange data or the second exchange data includes synchronization data to facilitate synchronization of the first mono audio output with a second mono audio output at the second audio output device.
22. The method according to claim 18, wherein the first exchange data or the second exchange data includes the stereo reverberation coefficient.
23. The method according to claim 18, wherein the first part of the spatial audio data includes first audio data corresponding to a first plurality of sound sources of the 3D sound field, where second exchange data received from the second audio output device includes second audio data, where the second audio data corresponds to a second plurality of sound sources of the 3D sound field, and where the first plurality of sound sources is different from the second plurality of sound sources.
24. The method according to claim 18, wherein the stereo reverberation coefficient is determined based on the first part of the spatial audio data.
25. An apparatus for generating an audio output, comprising: A unit for obtaining a first part of spatial audio data representing a three-dimensional (3D) sound field at a first audio output device; A unit for decoding the first part of the spatial audio data at the first audio output device to obtain first exchange data, the first exchange data including data representing a decoded first part of the spatial audio data; A unit for performing data exchange between the first audio output device and the second audio output device based on the spatial audio data, to send the first exchange data from the first audio output device to the second audio output device, and for the first audio output device to receive second exchange data from the second audio output device, wherein the second exchange data includes data representing a decoded second part of the spatial audio data; and A unit for determining a stereo reverberation coefficient based on the spatial audio data; and A unit for modifying the stereo reverberation coefficient based on motion data from one or more motion sensors to generate a modified stereo reverberation coefficient representing a rotated 3D sound field; and A unit for applying a head-related transfer function to the modified stereo reverberation coefficient to generate a first mono audio output at the first audio output device based on a combination of data representing a decoded first part of the spatial audio data and data representing a decoded second part of the spatial audio data.
26. The apparatus according to claim 25, wherein, The unit for performing data exchange includes: a unit for receiving a first part of the spatial audio data from a host device via a wireless peer-to-peer ad-hoc link.
27. The apparatus according to claim 25, wherein, The unit for performing data exchange includes: a unit for sending the first exchange data to the second audio output device via a wireless link between the first audio output device and the second audio output device.
28. The apparatus according to claim 27, wherein, The unit for performing data exchange further includes: a unit for receiving the second exchange data from the second audio output device via the wireless link.
29. A non-transitory computer-readable storage device storing instructions executable by one or more processors to cause the one or more processors to perform the following operations: Obtain a first part of spatial audio data representing a three-dimensional (3D) sound field at a first audio output device; Decode the first part of the spatial audio data at the first audio output device to obtain first exchange data, the first exchange data including data representing a decoded first part of the spatial audio data; Perform data exchange between the first audio output device and the second audio output device based on the spatial audio data, to send the first exchange data from the first audio output device to the second audio output device, and for the first audio output device to receive second exchange data from the second audio output device, wherein the second exchange data includes data representing a decoded second part of the spatial audio data; Determine a stereo reverberation coefficient based on the spatial audio data; and Modify the stereo reverberation coefficient based on motion data from one or more motion sensors to generate a modified stereo reverberation coefficient representing a rotated 3D sound field; and Apply a head-related transfer function to the modified stereo reverberation coefficients to generate a first mono audio output at the first audio output device based on a combination of data representing a first decoded portion of the spatial audio data and data representing a second decoded portion of the spatial audio data.
30. The non-transitory computer-readable storage device according to claim 29, wherein, generating the first mono audio output includes: determining a 3D sound field representation based on a combination of data representing a first decoded portion of the spatial audio data and data representing a second decoded portion of the spatial audio data.
Citation Information
Patent Citations
An audio data processing device for and a method of synchronized audio data processing
CN101263735A