Directional Audio Generation with Multiple Sound Source Arrangements
By processing audio data on the host device and selecting appropriate directional audio data in the portable device, the computing resource limitation and audio delay problems of portable devices when generating interactive audio content is solved, improving the quality and latency performance of audio output.
Patent Information
- Application Number
- CN202280037735.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-27
- Filing Date
- 2022-05-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-25
AI Technical Summary
Existing portable devices face problems of computing resource limitations and audio output delays when generating high-quality interactive audio content.
By performing most of the processing on the host device, multiple directional audio data sets are generated, and directional audio data corresponding to the detected position data is selected in the personal audio device to reduce audio delay and save device resources.
It effectively reduces the audio delay related to interactive audio content rendering, improves user experience quality, and indirectly solves the problem of device resource constraints.
Smart Images

Figure CN117378222B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority of co - owned U.S. Non - Provisional Patent Application No. 17 / 332,813, filed on May 27, 2021, the content of which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to generating directional audio using multiple arrangements of sound sources. Background Art
[0004] Advances in technology have led to smaller and more powerful computing devices. For example, there are currently a variety of portable personal computing devices, including wireless telephones such as mobile and smart phones, tablets, and laptop computers, which are small in size, light in weight, and easy for users to carry. These devices can communicate voice and data packets over a wireless network. In addition, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Further, such devices can process executable instructions, including software applications that can be used to access the Internet, such as a web browser application. Thus, these devices can include significant computing power.
[0005] The proliferation of such devices has facilitated a change in media consumption. There has been an increase in interactive audio content, such as in personal electronic games, where a handheld or portable electronic game system is used to play an electronic game and the audio content is based on the user's interaction with the game. This personalized or individualized media consumption typically involves relatively small portable (e.g., battery - powered) devices for generating the output. Due to the size, weight constraints, power constraints, or other reasons of portable devices, the processing resources available to such portable devices may be limited. In some cases, waiting for user interaction to initiate the rendering of interactive audio content may result in a delay in the audio output. Thus, providing a high - quality user experience can be challenging. Summary of the Invention
[0006] According to one embodiment of the present disclosure, a device includes a memory and a processor. The memory is configured to store instructions. The processor is configured to execute the instructions to obtain spatial audio data representing audio from one or more sound sources. The processor is further configured to execute the instructions to generate first directional audio data based on the spatial audio data. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to the audio output device. The processor is further configured to execute the instructions to generate second directional audio data based on the spatial audio data. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The processor is further configured to execute the instructions to generate an output stream based on the first directional audio data and the second directional audio data.
[0007] According to another embodiment of the present disclosure, a device includes a memory and a processor. The memory is configured to store instructions. The processor is configured to execute the instructions to receive first directional audio data representing audio from one or more sound sources from a host device. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to the audio output device. The processor is further configured to execute the instructions to receive second directional audio data representing audio from one or more sound sources from the host device. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The processor is further configured to receive position data indicating the position of the audio output device. The processor is further configured to generate an output stream based on the first directional audio data, the second directional audio data, and the position data. The processor is further configured to provide the output stream to the audio output device.
[0008] According to another embodiment of the present disclosure, a method includes obtaining, at a device, spatial audio data representing audio from one or more sound sources. The method further includes generating, at the device, first directional audio data based on the spatial audio data. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to the audio output device. The method further includes generating, at the device, second directional audio data based on the spatial audio data. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The method further includes generating, at the device, an output stream based on the first directional audio data and the second directional audio data. The method further includes providing the output stream from the device to an audio output device.
[0009] According to another embodiment of the present disclosure, a method includes receiving, at a device, first directional audio data representing audio from one or more sound sources from a host device. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to an audio output device. The method further includes receiving, at the device, second directional audio data representing audio from the one or more sound sources from the host device. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The method further includes receiving, at the device, location data indicating a location of the audio output device. The method further includes generating, at the device, an output stream based on the first directional audio data, the second directional audio data, and the location data. The method further includes providing the output stream from the device to the audio output device.
[0010] According to another embodiment of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain spatial audio data representing audio from one or more sound sources. The instructions, when executed by the one or more processors, further cause the one or more processors to generate first directional audio data based on the spatial audio data. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to an audio output device. The instructions, when executed by the one or more processors, further cause the one or more processors to generate second directional audio data based on the spatial audio data. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The instructions, when executed by the one or more processors, further cause the one or more processors to generate an output stream based on the first directional audio data and the second directional audio data. The instructions, when executed by the one or more processors, further cause the one or more processors to provide the output stream to the audio output device.
[0011] According to another embodiment of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to receive from a host device first directional audio data representing audio from one or more sound sources. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to an audio output device. The instructions, when executed by the one or more processors, further cause the one or more processors to receive from the host device second directional audio data representing audio from the one or more sound sources. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The instructions, when executed by the one or more processors, further cause the one or more processors to receive location data indicating the location of the audio output device. The instructions, when executed by the one or more processors, further cause the one or more processors to generate an output stream based on the first directional audio data, the second directional audio data, and the location data. The instructions, when executed by the one or more processors, further cause the one or more processors to provide the output stream to the audio output device.
[0012] According to another embodiment of the present disclosure, a device includes means for obtaining spatial audio data representing audio from one or more sound sources. The device further includes means for generating first directional audio data based on the spatial audio data. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to an audio output device. The device further includes means for generating second directional audio data based on the spatial audio data. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The device further includes means for generating an output stream based on the first directional audio data and the second directional audio data. The device further includes means for providing the output stream to the audio output device.
[0013] According to another embodiment of the present disclosure, a device includes means for receiving from a host device first directional audio data representing audio from one or more sound sources. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to an audio output device. The device further includes means for receiving from the host device second directional audio data representing audio from the one or more sound sources. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The device further includes means for receiving location data indicating the location of the audio output device. The device further includes means for generating an output stream based on the first directional audio data, the second directional audio data, and the location data. The device further includes means for providing the output stream to the audio output device.
[0014] After reading the entire application, other aspects, advantages, and features of the present disclosure will become apparent, including the following sections: Brief Description of the Drawings, Detailed Description, and Claims. Brief Description of the Drawings
[0015] Figure 1 is a block diagram of specific illustrative aspects of a system operable to generate directional audio having a plurality of sound source arrangements in accordance with some examples of the present disclosure.
[0016] Figure 2A is in accordance with some examples of the present disclosure Figure 1 of the operation of a stream generator.
[0017] Figure 2B is in accordance with some examples of the present disclosure by Figure 1 the data generated by the stream generator.
[0018] Figure 2C is in accordance with some examples of the present disclosure by Figure 1 the data generated by the stream generator.
[0019] Figure 3 is in accordance with some examples of the present disclosure Figure 2A of the operation of a parameter generator of a stream generator.
[0020] Figure 4 is in accordance with some examples of the present disclosure Figure 1 of the operation of a stream selector.
[0021] Figure 5 is another illustrative aspect of a system operable to generate directional audio having a plurality of sound source arrangements in accordance with some examples of the present disclosure.
[0022] Figure 6 is another illustrative aspect of a system operable to generate directional audio having a plurality of sound source arrangements in accordance with some examples of the present disclosure.
[0023] Figure 7 is in accordance with some examples of the present disclosure Figure 1 , 5 or any one of 6 of the operation of a stream generator and a stream selector.
[0024] Figure 8 shows an example of an integrated circuit operable to generate directional audio using a plurality of sound source arrangements in accordance with some examples of the present disclosure.
[0025] Figure 9Diagram of a wearable electronic device operable to generate directional audio using a plurality of sound source arrangements, according to some examples of the present disclosure.
[0026] Figure 10 Diagram of a voice-controlled speaker system operable to generate directional audio using a plurality of sound source arrangements, according to some examples of the present disclosure.
[0027] Figure 11 Diagram of a headset (such as a virtual reality or augmented reality headset) operable to generate directional audio using a plurality of sound source arrangements, according to some examples of the present disclosure.
[0028] Figure 12 Diagram of a first example of a vehicle operable to generate directional audio using a plurality of sound source arrangements, according to some examples of the present disclosure.
[0029] Figure 13 Diagram of a second example of a vehicle operable to generate directional audio using a plurality of sound source arrangements, according to some examples of the present disclosure.
[0030] Figure 14 Specific embodiments of a method for generating directional audio using a plurality of sound source arrangements that can be performed by a device of any one of Figure 1 , Figure 5 , Figure 6 , Figures 8 to 13 and Figure 16 according to some examples of the present disclosure.
[0031] Figure 15 Specific embodiments of a method for generating directional audio using a plurality of sound source arrangements that can be performed by a device of any one of Figure 1 , 5 or 6 according to some examples of the present disclosure.
[0032] Figure 16 Block diagram of a specific illustrative example of a device operable to generate directional audio having a plurality of sound source arrangements, according to some instances of the present disclosure. Detailed Description
[0033] Audio information can be captured or generated in a manner that enables an audio output to be presented to represent a three-dimensional (3D) sound field. For example, high-fidelity stereophonic reproduction (ambisonics) (e.g., first-order high-fidelity stereophonic reproduction (FOA) or higher-order high-fidelity stereophonic reproduction (HOA)) can be used to represent the 3D sound field for later playback. During playback, the 3D sound field can be reconstructed in a manner that enables a listener to distinguish the position and / or distance between the listener and one or more sound sources of the 3D sound field.
[0034] In accordance with certain aspects of the present disclosure, a personal audio device (such as a headset, headphones, earbuds, or another audio playback device configured to generate a directional audio output for a binaural user experience) can be used to present a 3D sound field. One challenge in rendering 3D audio using a personal audio device is the computational complexity of such rendering. By way of illustration, personal audio devices are typically configured to be worn by a user such that movement of the user's head changes the relative positions of the user's ears and sound sources in the 3D sound field to generate head-tracked immersive audio. Such personal audio devices are typically battery-powered and have limited on-board computational resources. Generating head-tracked immersive audio with such resource constraints is challenging. Another challenge associated with rendering interactive audio content is that waiting for user interaction to initiate rendering of the corresponding audio content can increase audio latency.
[0035] Some aspects disclosed herein facilitate a shift of certain power and processing constraints of a personal audio device by performing most of the processing at a host device (such as a laptop computer or a mobile computing device). Additionally, multiple directional audio data sets are generated, where each directional audio data set corresponds to a user location of the user, a reference location of a reference point, or both. In a particular example, the reference point includes the host device, a virtual reference point, a display screen, or a combination thereof. Some aspects disclosed herein facilitate a reduction in audio output latency by generating the directional audio data sets based on predicted user interactions. The multiple directional audio data sets are provided to the personal audio device, and the personal audio device selects the directional audio data corresponding to the detected location data for output. In some examples, the host device pre-generates multiple directional audio data sets (e.g., based on predicted location data) and provides the selected directional audio data set to the personal audio device corresponding to the detected location data to further offload processing from the personal audio device. In some examples, a single audio device (e.g., having certain power and processing capabilities) pre-generates multiple directional audio data sets (e.g., based on predicted location data), selects the directional audio data set corresponding to the detected location data, and outputs the selected directional audio data to reduce the audio latency associated with rendering interactive audio content.
[0036] Certain aspects of the present disclosure are described below with reference to the accompanying drawings. In the specification, common features are denoted by common reference numerals. As used herein, various terms are used for the purpose of describing particular embodiments only and are not intended to limit the embodiments. For example, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Additionally, some features described herein are singular in some embodiments and plural in other embodiments. By way of illustration, Figure 1 depicts including one or more selection parameters (Figure 1 a stream generator 140 of "selection parameter" 156), which indicates that in some embodiments, the stream generator 140 generates a single selection parameter 156, and in other embodiments, the stream generator 140 generates multiple selection parameters 156.
[0037] As used herein, the terms "comprise", "comprises" and "comprising" may be used interchangeably with "include", "includes" or "including". Additionally, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" indicates examples, embodiments and / or aspects, and should not be construed as limiting or indicating a preference or preferred embodiment. As used herein, ordinal terms (e.g., "first", "second", "third", etc.) used to modify elements (such as structures, components, operations, etc.) do not themselves indicate any priority or order of the element relative to another element, but merely distinguish the element from another element having the same name (but using ordinal terms). As used herein, the term "set" refers to one or more of a particular element, and the term "plurality" refers to a plurality of a particular element (e.g., two or more).
[0038] As used herein, "coupled" may include "communicatively coupled", "electrically coupled" or "physically coupled", and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired network, wireless network or a combination thereof), etc. As an illustrative non-limiting example, two electrically coupled devices (or components) may be included in the same device or different devices, and may be connected via electronics, one or more connectors or inductive coupling. In some embodiments, two devices (or components) that are communicatively coupled (e.g., electrically communicatively) may directly or indirectly send and receive signals (e.g., digital signals or analog signals) via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled or physically coupled) without an intermediate component.
[0039] In the present invention, terms such as "determine", "calculate", "estimate", "shift", "adjust", etc. may be used to describe how to perform one or more operations. It should be noted that these terms should not be construed as restrictive, and other techniques may be utilized to perform similar operations. Additionally, as mentioned herein, "generate", "calculate", "estimate", "use", "select", "access", and "determine" may be used interchangeably. For example, "generate", "calculate", "estimate", or "determine" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal) or may refer to using, selecting, or accessing, for example, a parameter (or signal) that has been generated by another component or device.
[0040] Refer to Figure 1 , specific illustrative aspects of a system configured to generate directional audio having a plurality of sound source arrangements are disclosed and are generally designated as 100. System 100 includes a device 102 (e.g., a host device) configured to communicate with a device 104 (e.g., an audio output device).
[0041] Spatial audio data 170 represents sound from one or more sound sources 184 in three dimensions (3D), which may include real or virtual sources, such that an audio output representing the spatial audio data 170 can simulate the distance and direction between a listener and the one or more sound sources 184. The spatial audio data 170 may be encoded using various coding schemes, such as first-order high-fidelity stereophonic sound reproduction (FOA), higher-order high-fidelity stereophonic sound reproduction (HOA), or equivalent spatial domain (ESD) representation (described further below). As an example, the FOA coefficients or ESD data representing the spatial audio data 170 may be encoded using a total of four channels (e.g., two stereo channels).
[0042] Device 102 is configured to process the spatial audio data 170 using a stream generator 140 to generate a directional audio data set corresponding to a plurality of sound source arrangements, as further described with reference to Figure 2A In a particular aspect, the stream generator 140 is configured to obtain user interactivity data 111, spatial audio data 170, or both from an application (e.g., a video player, a video game, an online meeting, etc.) of the device 102. In a particular aspect, the user interactivity data 111 indicates the position of a virtual object in a virtual space, a mixed reality space, or an augmented reality space.
[0043] In certain aspects, the spatial audio data 170 represents sound from a sound source 184 that, when played, will be perceived as coming from a position 192 (e.g., to the left and at a particular distance) relative to a reference point 143 (e.g., the device 102, a display screen, another physical reference point, a virtual reference point, or a combination thereof). In certain aspects, the reference point 143 can have a fixed position (e.g., the driver's seat) in a reference frame (e.g., a vehicle). For example, regardless of whether the user wearing the device 104 looks out of a side window or straight ahead, the sound from the sound source 184 will be perceived as coming from the driver's seat of the vehicle. In another aspect, the reference point 143 (e.g., a non-player character (NPC)) can move within a reference frame (e.g., a virtual world). For example, the sound from the sound source 184 will be perceived as coming from the NPC that the user wearing the device 104 is following in the virtual world, regardless of whether the user wearing the device 104 is looking at the NPC or turning their head to look in another direction.
[0044] In certain aspects, the position sensor 186 is configured to generate user position data 115 indicative of the position of the user of the device 104. In certain aspects, the position sensor 188 is configured to generate device position data 109 indicative of the position of the reference point 143 (e.g., the device 102, the display screen of the device 102, another physical reference point, or a combination thereof). In certain aspects, the user interactivity data 111 includes virtual reference position data 107 indicative of the position of the reference point 143 (e.g., a virtual reference point, such as a virtual building in a game) at a first virtual reference position time.
[0045] In certain embodiments, the position sensor 188 is external to the device 102. For example, the position sensor 188 includes a camera configured to capture an image (e.g., device position data 109) indicative of the position of the device 102. In certain embodiments, the position sensor 188 is integrated in the device 102. For example, the position sensor 188 includes an accelerometer configured to generate sensor data (e.g., device position data 109) indicative of the position of the device 102. In certain aspects, the position sensor 188 is configured to generate device position data 109 indicative of the relative position (e.g., rotation, displacement, or both), absolute position (e.g., orientation, position, or both), or a combination thereof of the device 102.
[0046] In certain embodiments, the position sensor 186 is external to the device 104. For example, the position sensor 186 includes a camera configured to capture an image (e.g., user location data 115) indicative of the position of the user, the device 104, or both. In certain embodiments, the position sensor 186 is integrated in the device 104. For example, the position sensor 186 includes an accelerometer configured to generate sensor data (e.g., user location data 115) indicative of the position of the device 104, the user, or both. In certain aspects, the position sensor 186 is configured to generate user location data 115 indicative of the relative position (e.g., rotation, displacement, or both), absolute position (e.g., orientation, location, or both), or a combination thereof of the device 104.
[0047] In certain aspects, the stream generator 140 is configured to determine the reference position data 113 based on the device location data 109, the virtual reference position data 107, or both. The reference position data 113 indicates the position of the reference point 143. For example, the reference position data 113 is based on the device location data 109 indicative of the position of a physical reference point, the virtual reference position data 107 indicative of the position of a virtual reference point, or both.
[0048] In certain embodiments, the stream generator 140 is configured to generate one or more directional audio data sets based at least in part on the reference position data 113, the user location data 115, or both, as further described in the reference Figure 2A In certain embodiments, the stream selector 142 is configured to select one of the directional audio data sets from the directional audio data sets based at least in part on the reference position data 157 received from the device 102, the user location data 185 received from the position sensor 186, or both, as further described in the reference Figure 4 further described.
[0049] The device 104 includes the speaker 120, the speaker 122, or both. The stream generator 140 is configured to provide the directional audio data set to the device 104. The device 104 is configured to select a directional audio data set from the directional audio data sets using the stream selector 142, generate acoustic data 172 based on the directional audio data set, and output the acoustic data 172 via the speaker 120, the speaker 122, or both, as further described in the reference Figure 4 further described.
[0050] In some embodiments, device 102, device 104, or both correspond to or are included in various types of devices. In certain aspects, device 102 includes at least one of a mobile device, a gaming console, a communication device, a computer, a display device, a vehicle, a camera, or a combination thereof. In certain aspects, device 104 includes at least one of a headset, an extended reality (XR) headset, a gaming device, headphones, a speaker, or a combination thereof. In an illustrative example, stream generator 140, stream selector 142, or both are integrated in a headset device including speakers 120 and 122, such as described with reference to Figure 1 and Figure 6 In some examples, stream generator 140, stream selector 142, or both are integrated in a mobile phone or tablet computer device as described with reference to Figure 1 , 5 and 6, a wearable electronic device as described with reference to Figure 9 , a voice-controlled speaker system as described with reference to Figure 10 , or a virtual reality headset or an augmented reality headset as described with reference to Figure 11 In another illustrative example, stream generator 140, stream selector 142, or both are integrated into a vehicle that also includes speakers 120 and 122, such as further described with reference to Figure 12 and Figure 13
[0051] During operation, stream generator 140 obtains spatial audio data 170 representing audio from one or more sound sources 184. In certain aspects, stream generator 140 retrieves spatial audio data 170, user interactivity data 111, or a combination thereof from a memory. In another aspect, stream generator 140 receives spatial audio data 170, user interactivity data 111, or a combination thereof from an audio data source (e.g., a server). In a particular instance, a user of device 104 (e.g., headphones) initiates an application (e.g., a game, a video player, an online meeting, or a music player) on device 102, and the application outputs spatial audio data 170, user interactivity data 111, or a combination thereof. In certain aspects, stream generator 140 obtains user interactivity data 111 while obtaining spatial audio data 170.
[0052] Stream generator 140 processes spatial audio data 170 based on one or more selection parameters 156 to generate a plurality of directional audio data sets. For example, stream generator 140 processes spatial audio data 170 based on location data 174 (e.g., default location data, detected location data, or both) to generate directional audio data 152, as described with reference to Figure 2A is further described. In a particular example, the location data 174 includes default location data indicating a default location of the device 104, a default head location of a user of the device 104, a default location of the reference point 143, a default relative location of the device 102 and the reference point 143, a default relative movement of the device 102 and the reference point 143, or a combination thereof. In a particular aspect, the default relative location of the reference point 143 and the device 104 corresponds to a user of the device 104 facing the reference point 143.
[0053] In a particular aspect, the location data 174 includes detected location data (which indicates a detected location of the device 104), a detected movement of the device 104, a detected head location of a user of the device 104, a detected head movement of a user of the device 104, a detected location of the reference point 143, a detected movement of the reference point 143, a detected relative location of the device 104 and the reference point 143, a detected relative movement of the device 104 and the reference point 143, or a combination thereof. By way of illustration, the location data 174 includes reference location data 103 indicating a first location (e.g., location, orientation, or both) of the reference point 143, user location data 105 indicating a first location (e.g., location, orientation, or both) of a user of the device 104, or both.
[0054] In a particular instance, the device 102 receives user location data 115 indicating a first location, a first movement, or both detected by the location sensor 186 at a first user location time. The stream generator 140 generates (e.g., updates) the user location data 105 based on the user location data 115. For example, the user location data 105 indicates a first absolute location of a user of the device 104, the user location data 115 indicates a change in location of a user of the device 104, and the stream generator 140 updates the user location data 105 by applying the change in location to the first absolute location to indicate a second absolute location of a user of the device 104.
[0055] In a particular example, the stream generator 140 receives device position data 109 indicating a first position, a first movement, or both of a reference point 143 (e.g., device 102, display screen, or another physical reference point) detected by the position sensor 188 at a first device position time. In a particular example, the stream generator 140 receives virtual reference position data 107 indicating a first position, a first movement, or both of a reference point 143 (e.g., virtual reference point) detected (e.g., occurred) at a first virtual reference position time. The stream generator 140 determines reference position data 113 based on the device position data 109, the virtual reference position data 107, or both. The stream generator 140 generates (e.g., updates) reference position data 103 based on the reference position data 113. For example, the reference position data 103 indicates a first absolute position of the reference point 143, the reference position data 113 indicates a change in position of the reference point 143, and the stream generator 140 updates the reference point 143 by applying the change in position to the first absolute position to indicate a second absolute position of the reference point 143.
[0056] The directional audio data 152 corresponds to an arrangement 162 of one or more sound sources 184 relative to a listener (e.g., device 104). In a particular aspect, the spatial audio data 170 represents sound from the sound sources 184 that will be perceived as coming from a position 192 relative to the reference point 143 when the spatial audio data 170 is broadcast. As an illustrative example, the user position data 105 and the reference position data 103 indicate a first position (e.g., 0 degrees (deg.)) In some embodiments, the position of the user of the wearable device 104 is determined relative to the reference point 143. In a particular aspect, the user defaults to a first position relative to the reference point 143. In another aspect, it is detected that the user (e.g., as indicated by the user position data 115) has a first position relative to the reference point 143.
[0057] The stream generator 140 generates the directional audio data 152 to have an arrangement 162 such that the sound from the sound sources 184 is perceived as coming from a second direction (e.g., the right side) of the listener (e.g., device 104). When the directional audio data 152 is broadcast, such that when the user has a user position indicated by the user position data 105 and the reference point 143 has a reference position indicated by the reference position data 103, the sound will be perceived as coming from a position 192 relative to the reference point 143.
[0058] In a particular aspect, the stream generator 140 processes the spatial audio data 170 based on one or more position data sets (e.g., predetermined position data, predicted position data, or both) to generate one or more directional audio data sets, as referenced Figure 2AFurther description. For example, the stream generator 140 processes the spatial audio data 170 based on the location data 176 to generate the directional audio data 154.
[0059] In certain aspects, the location data 176 includes the reference location data 123 indicating a second location (e.g., location, orientation, or both) of the reference location data 123, the user location data 125 indicating a second location (e.g., location, orientation, or both) of the user of the device 104, or both.
[0060] In certain examples, the location data 176 includes predetermined location data indicating a predetermined location of the device 104, a predetermined head location of the user of the device 104, a predetermined location of the reference point 143, a predetermined relative location between the device 102 and the reference point 143, a predetermined relative movement between the device 102 and the reference point 143, or a combination thereof. In certain aspects, the predetermined relative location between the reference point 143 and the device 104 corresponds to the user of the device 104 facing the reference point 143.
[0061] In certain aspects, the location data 176 includes predicted location data indicating a predicted location of the device 104, a predicted movement of the device 104, a predicted head location of the user of the device 104, a predicted head movement of the user of the device 104, a predicted location of the reference point 143, a predicted movement of the reference point 143, a predicted relative location between the device 104 and the reference point 143, a predicted relative movement between the device 104 and the reference point 143, or a combination thereof. By way of illustration, the location data 176 includes the reference location data 103 indicating a first location (e.g., location, orientation, or both) of the reference point 143, the user location data 105 indicating a first location (e.g., location, orientation, or both) of the user of the device 104, or both.
[0062] In certain aspects, the reference location data 123, the user location data 125, or both correspond to a predetermined location of the user of the device 104 relative to the reference point 143. For example, a predetermined location (e.g., 90 degrees) corresponds to the user of the device 104 rotating relative to the reference point 143 in a particular direction.
[0063] In certain aspects, the stream generator 140 generates a directional audio data set based on a predetermined range of positions of a user of the device 104 relative to a reference point 143 (e.g., 0 degrees, 45 degrees, 90 degrees, 135 degrees, and 180 degrees). In certain aspects, the range of predetermined positions is based on the user position detected at a first user position time (e.g., as indicated by the user position data 115), the reference position detected at a first reference position time (e.g., as indicated by the reference position data 113), or both. For example, in response to determining that the reference position data 113 and the user position data 115 indicate a relative position of the device 104 relative to the reference point 143 (e.g., 90 degrees), the stream generator 140 determines the range of predetermined positions based on the relative position (e.g., from 80 degrees to 100 degrees) (e.g., starting at the relative position, ending at the relative position, ending around the relative position, or centered on the relative position). The stream generator 140 determines first directional audio data corresponding to a first predetermined position (e.g., 80 degrees), directional audio data 154 corresponding to a second predetermined position (e.g., 90 degrees), third directional audio data corresponding to a third predetermined position (e.g., 100 degrees), or a combination thereof.
[0064] In certain aspects, the reference position data 123 corresponds to a predicted reference position of the reference point 143, the user position data 125 corresponds to a predicted user position of a user of the device 104, or both. In a particular example, the stream generator 140 determines the predicted reference position based on the reference position data 113 (e.g., detected position, detected movement, or both), predicted device position data, predicted user interactivity data, or a combination thereof, as further described in the reference Figure 3 as further described. In a particular example, the stream generator 140 determines the predicted user position data based on the user position data 115 (e.g., detected position, detected movement, or both), user interactivity data 111 (e.g., detected user interactivity data), predicted user interactivity data, or a combination thereof, as further described in the reference Figure 3 as further described.
[0065] In certain aspects, the stream generator 140 generates a directional audio data set based on a plurality of predicted positions of a user of the device 104 relative to a reference point 143. In certain aspects, each of the predicted positions is based on reference position data 113 (e.g., detected position, detected movement, or both), predicted device position data, predicted user interactivity data, or a combination thereof. For example, in response to determining that a first predicted position of a user of the device 104 relative to the reference point 143 has a first predicted probability greater than a threshold probability, the stream generator 140 determines first directional audio data corresponding to the first predicted position. As another example, in response to determining that a second predicted position of a user of the device 104 relative to the reference point 143 has a second predicted probability greater than a threshold probability, the stream generator 140 determines second directional audio data corresponding to the second predicted position.
[0066] The directional audio data 154 corresponds to an arrangement 164 of one or more sound sources 184 relative to a listener (e.g., the device 104). In certain aspects, the arrangement 164 is different from the arrangement 162. As an illustrative example, the user position data 125 and the reference position data 123 indicate a second position (e.g., 90 degrees) of the user of the device 104 relative to the reference point 143. In the illustrative example, the user is facing (e.g., as predetermined or predicted) the position 192. The stream generator 140 generates the directional audio data 154 to have the arrangement 164 such that sound from the sound source 184 is perceived as coming from a particular direction (e.g., the front) of the listener (e.g., the device 104). When the directional audio data 154 is broadcast, such that when the user has the user position indicated by the user position data 125 and the reference point 143 has the reference position indicated by the reference position data 123, the sound will be perceived as coming from the position 192 relative to the reference point 143.
[0067] In a particular embodiment, the stream generator 140 is configured to initiate the transmission of an output stream 150 that includes a directional audio data set (e.g., directional audio data 152, directional audio data 154, one or more additional directional audio data sets, or a combination thereof) to the device 104. In a particular aspect, the stream generator 140 also concurrently transmits to the device 104 one or more selection parameters 156 in conjunction with the transmission of the output stream 150 to the device 104. The one or more selection parameters 156 indicate a user location, a reference location, or both, associated with a particular set of directional audio data. For example, the one or more selection parameters 156 indicate that the directional audio data 152 is based on the reference location data 103, the user location data 105, or both, of the location data 174. As another example, the one or more selection parameters 156 indicate that the directional audio data 154 is based on the reference location data 123, the user location data 125, or both, of the location data 176. In a particular instance, the one or more selection parameters 156 indicate that the additional directional audio data set is based on particular location data (e.g., corresponding to a predetermined location or a predicted location).
[0068] The stream selector 142 receives the output stream 150 and the one or more selection parameters 156 from the device 102. The stream selector 142 renders (e.g., generates) acoustic data 172 based on the output stream 150, the reference location data 157, the user location data 185, or both. In a particular aspect, the position sensor 188 generates second device location data indicating the device location of a reference point 143 (e.g., the device 102, a display screen, or another physical reference point) detected at a second device location time. In a particular aspect, the second device location time is after the first device location time associated with the device location data 109. In a particular aspect, the user interactivity data 111 includes second virtual reference location data indicating the reference location of a reference point 143 (e.g., a virtual reference point) detected at a second virtual reference location time. In a particular aspect, the second virtual reference location time is after the first virtual reference location time associated with the virtual reference location data 107. The stream selector 142 determines the reference location data 157 based on the second device location data, the second virtual location data, or both.
[0069] In a particular embodiment, the device 102 sends the reference location data 157 to the device 104 while sending the output stream 150 to the device 104. In an alternative embodiment, the second device location time, the second virtual reference location time, or both, are after the transmission time of the output stream 150 from the device 102 to the device 104. In this embodiment, the device 102 sends the reference location data 157 to the device 104 after sending the output stream 150 to the device 104.
[0070] User location data 185 indicates the location of the user of device 104. For example, location sensor 186 generates user location data 185 indicating the location of the user of device 104 detected at a second user location time. In certain aspects, the second user location time is after the first user location time associated with user location data 115. In example 160, user location data 185 and reference location data 157 indicate that the user of device 104 has a detected location (e.g., 60 degrees) relative to reference point 143.
[0071] In certain aspects, arrangement 162 corresponds to a first position of sound source 184 relative to the listener (e.g., device 104) (e.g., from the right side of the listener (e.g., device 104)). When device 104 has a detected location (e.g., 60 degrees) relative to reference point 143, arrangement 162 corresponds to position 196 of sound source 184 relative to reference point 143. In certain aspects, arrangement 164 corresponds to a second position of sound source 184 relative to the listener (e.g., device 104) (e.g., from in front of the listener (e.g., device 104)). When device 104 has a detected location (e.g., 60 degrees) relative to reference point 143, arrangement 164 corresponds to position 194 of sound source 184 relative to reference point 143.
[0072] In certain embodiments, stream selector 142 selects one or a combination of directional audio data 152, directional audio data 154, and one or more additional sets of directional audio data based on the detected location of device 104 relative to reference point 143 (e.g., 60 degrees), as further described in reference Figure 4 Spatial audio data 170 represents sound from sound source 184 that, when played, will be perceived as coming from position 192 relative to reference point 143. Stream selector 142 selects directional audio data 154 in response to determining that the match of position 194 to position 192 is closer than the match of position 196 to position 192. For example, stream selector 142 selects directional audio data 154 in response to determining that the difference between position 194 (corresponding to arrangement 164) and position 192 is less than or equal to the difference between position 196 (corresponding to arrangement 162) and position 192. Stream selector 142 decodes directional audio data 154 (e.g., the selected set of directional audio data) to generate acoustic data 172.
[0073] In certain embodiments, stream selector 142 generates acoustic data 172 (e.g., the output stream) by combining directional audio data 152 and directional audio data 154 based on the detected position of device 104 relative to reference point 143, as further described in reference Figure 4is further described. In certain aspects, the stream generator 140 generates the acoustic data 172 to have an arrangement 166 such that when the acoustic data 172 is played, the sound from the sound source 184 is perceived as coming from a particular direction (e.g., partially to the right) from the listener (e.g., the device 104), such that the sound will be perceived as coming from a particular location (e.g., when the user has a user location indicated by the user location data 185 and the reference point 143 has a reference location indicated by the reference location data 157, the position 192 of the sound source 184 relative to the reference point 143). The particular location (e.g., the location 192) is between the location 194 and the location 196. For example, when a greater weight is applied to the directional audio data 152 to generate the acoustic data 172, the particular location is closer to the location 196. As another example, when a greater weight is applied to the directional audio data 154 to generate the acoustic data 172, the particular location is closer to the location 194.
[0074] In certain aspects, the stream selector 142 outputs the acoustic data 172 via the speaker 120 (e.g., an audio output device). For example, the stream selector 142 outputs the acoustic data 172 via the speaker 120 (e.g., the right speaker) corresponding to a particular channel in response to determining that the acoustic data 172 corresponds to a particular channel (e.g., the right channel).
[0075] Thus, the system 100 enables the generation of the acoustic data 172 such that as the position of the listener (e.g., the orientation, the position, or both) changes relative to the reference point 143, the acoustic arrangement of one or more sound sources 184 relative to the listener (e.g., the user of the device 104) is updated. Most of the processing for generating the acoustic data 172 (such as generating the directional audio data set) is performed at the device 102 to conserve resources (e.g., power and computational cycles) at the device 104. In a particular example, pre-generating at least some of the directional audio data sets based on predicted location data and selecting one of the directional audio data sets based on the detected location data to generate the acoustic data 172 reduces the latency between detecting the location data and outputting the acoustic data 172 based on the corresponding directional audio data.
[0076] Although the device 104 is shown as including the speaker 120 and the speaker 122, in other embodiments, fewer than two or more than two speakers are integrated in or coupled to the device 104. Although the stream generator 140 and the stream selector 142 are shown as being included in separate devices, in other implementations, the stream generator 140 and the stream selector 142 may be included in a single device, as further described in reference Figures 5 - 6 is further described.
[0077] In certain embodiments, the stream generator 140 is configured to generate multiple directional audio data sets corresponding to various bitrates. For example, the stream generator 140 generates a first copy of the directional audio data 152 corresponding to a first bitrate (e.g., a higher bitrate), a second copy of the directional audio data 152 corresponding to a second bitrate (e.g., a lower bitrate), a first copy of the directional audio data 154 corresponding to the first bitrate, a second copy of the directional audio data 154 corresponding to the second bitrate, or a combination thereof.
[0078] The stream generator 140 selects a bitrate (e.g., the first bitrate, the second bitrate, or both) based on the ability, condition, or both of detecting the communication link with the stream selector 142. For example, the stream generator 140 selects the first bitrate in response to determining that the first bandwidth of the communication link is greater than a threshold bandwidth. As another example, the stream generator 140 selects the second bitrate in response to determining that the first bandwidth of the communication link is less than or equal to the threshold bandwidth.
[0079] The stream generator 140 provides the directional audio data associated with the selected bitrate as the output stream 150 to the stream selector 142. For example, the stream generator 140 provides a first copy of the directional audio data 152, a first copy of the directional audio data 154, or both as the output stream 150 to the stream selector 142 in response to determining that the first bandwidth of the communication link is greater than the threshold bandwidth. As another example, the stream generator 140 provides a second copy of the directional audio data 152, a second copy of the directional audio data 154, or both as the output stream 150 to the stream selector 142 in response to determining that the first bandwidth of the communication link is less than or equal to the threshold bandwidth.
[0080] In certain embodiments, the stream generator 140 provides one or more of the directional audio data 152, the directional audio data 154, one or more additional directional audio data sets, or a combination thereof as the output stream 150 based on the ability, condition, or both of the communication link with the stream selector 142. For example, the stream generator 140 provides one of the directional audio data 152, the directional audio data 154, one or more additional directional audio data sets, or a combination thereof as the output stream 150 to the stream selector 142 in response to determining that the first bandwidth of the communication link is less than or equal to the threshold bandwidth. As another example, the stream generator 140 provides more than one of the directional audio data 152, the directional audio data 154, one or more additional directional audio data sets, or a combination thereof as the output stream 150 to the stream selector 142 in response to determining that the first bandwidth of the communication link is greater than the threshold bandwidth.
[0081] In certain embodiments, the stream generator 140 provides, as the output stream 150, one of the directional audio data 152, the directional audio data 154, one or more additional directional audio data sets, or a combination thereof, based on the capabilities, conditions, or both of the communication link with the stream selector 142. For example, in response to determining that the first bandwidth of the communication link is less than or equal to a threshold bandwidth, the stream generator 140 provides, as the output stream 150, one of the directional audio data 152, the directional audio data 154, one or more additional directional audio data sets, or a combination thereof, to the stream selector 142. As another example, in response to determining that the first bandwidth of the communication link is greater than the threshold bandwidth, the stream generator 140 provides, as the output stream 150, another one of the directional audio data 152, the directional audio data 154, one or more additional directional audio data sets, or a combination thereof, to the stream selector 142.
[0082] Reference Figure 2A , FIG. 200 shows illustrative aspects of the operation of the stream generator 140. In certain aspects, the stream generator 140 is coupled to an audio data source 202 (e.g., a memory, a server, a storage device, or another audio data source). In certain aspects, the audio data source 202 is external to the Figure 1 device 102. For example, the device 102 includes a modem configured to receive audio data from the audio data source 202. In an alternative aspect, the audio data source 202 is integrated within the device 102.
[0083] The stream generator 140 includes an audio decoder 204 coupled to a reference position adjuster 208 via a user position adjuster 206. The reference position adjuster 208 is coupled to one or more renderers, such as renderer 212, renderer 214, one or more additional renderers, or a combination thereof. The stream generator 140 also includes a parameter generator 210 coupled to at least one renderer (such as renderer 214), one or more additional renderers, or a combination thereof.
[0084] In certain aspects, the audio decoder 204 receives encoded audio data 203 from the audio data source 202. The audio decoder 204 decodes the encoded audio data 203 to generate spatial audio data 205. In Figure 2B FIG. 260 shows an example of data generated by the stream generator 140. For example, the previous spatial audio data has an arrangement 262. A first value 264 of the user position data 105 indicates the previous position of the user of the device 104 corresponding to the arrangement 262. For example, the first value 264 indicates the position 272 (e.g., the first position coordinates) and the orientation 276 (e.g., north) of the user of the device 104. The spatial audio data 205 corresponds to the first position of the sound source 184 relative to the listener (e.g., on the right side of the listener).
[0085] The stream generator 140 receives user location data 115 from the position sensor 186. The user location data 115 indicates a change in the position of the user of the device 104. In a particular embodiment, the user location data 115 indicates that the user of the device 104 has changed the orientation (e.g., rotated counterclockwise) by a particular amount (e.g., 90 degrees) while staying in the same position (e.g., no displacement). The user location adjuster 206 determines that the user has moved from the orientation 276 (e.g., facing north) to the orientation 278 (e.g., facing west) based on the orientation 276 (e.g., facing north) and the orientation change (e.g., 90 degrees counterclockwise) indicated by the user location data 115. The user location adjuster 206 determines that the user remains in the same position (e.g., position 272) based on the position 272 and the displacement (e.g., none) indicated by the user location data 115. In another embodiment, the user location data 115 indicates that the user of the device 104 has the orientation 278 (e.g., facing west) at the position 272. The user location adjuster 206 determines that the user has changed the orientation (e.g., rotated counterclockwise 90 degrees) while staying in the same position (e.g., no displacement) based on a comparison of the first value 264 of the user location data 105 with the user location data 115.
[0086] The user location adjuster 206 generates the spatial audio data 207 by adjusting the spatial audio data 205 based on a change in the user location (e.g., an orientation change, a displacement, or both) indicated by the user location data 115, the first value 264 of the user location data 105, or both. For example, the user location adjuster 206 generates the spatial audio data 207 by adjusting the spatial audio data 205 based on the change in the user location such that the sound source 184 has a second position relative to the listener (e.g., behind the listener).
[0087] The user location adjuster 206 determines (e.g., updates) the user location data 105 based on the user location data 115. For example, the user location adjuster 206 updates the user location data 105 to a second value 266 indicating the position 272, the orientation 278, or both. In a particular aspect, the user location adjuster 206 provides the user location data 105 (e.g., the second value 266) to the parameter generator 210.
[0088] The user location adjuster 206 provides the spatial audio data 207 to the reference location adjuster 208. In Figure 2C FIG. 280 shows additional examples of data generated by the stream generator 140. For example, the first value 284 of the reference location data 103 indicates the previous position of the reference point 143 corresponding to the arrangement 262 (e.g., associated with the previous spatial audio data). For illustration, the first value 284 indicates the position 292 (e.g., the second position coordinate) and the orientation 294 (e.g., facing south) of the reference point 143.
[0089] The reference position adjuster 208 obtains reference position data 113 (e.g., device position data 109, virtual reference position data 107 indicated by user interactivity data 111, or both). The reference position data 113 indicates a change in the position of the reference point 143. In a particular embodiment, the reference position data 113 indicates that the reference point 143 changes its orientation (e.g., rotates counterclockwise by 90 degrees) and has a first displacement (e.g., moves a first distance westward and a second distance southward). The reference position adjuster 208 determines that the reference point 143 has moved from the orientation 294 (e.g., facing south) to the orientation 298 (e.g., facing east) based on the orientation 294 (e.g., facing south) and the change in orientation (e.g., counterclockwise 90 degrees) indicated by the reference position data 113. The reference position adjuster 208 determines that the reference point 143 has moved from the position 292 to the position 296 (e.g., a third position coordinate) based on the position 292 and the displacement (e.g., a first distance westward and a second distance southward) indicated by the reference position data 113. In another embodiment, the reference position data 113 indicates that the reference point 143 has the orientation 298 (e.g., facing east) at the position 296. The reference position adjuster 208 determines that the reference point 143 has changed its orientation (e.g., rotates counterclockwise by 90 degrees) and has a first displacement (e.g., moves a first distance westward and a second distance southward) based on a comparison of the first value 284 of the reference position data 103 with the reference position data 113.
[0090] The reference position adjuster 208 generates the spatial audio data 170 by adjusting the spatial audio data 207 based on the change in the position of the reference point 143 (e.g., change in orientation, displacement, or both) indicated by the reference position data 113, the first value 284 of the reference position data 103, or both. For example, the reference position adjuster 208 generates the spatial audio data 170 by adjusting the spatial audio data 207 based on the change in the reference point position such that the sound source 184 has a position 192 relative to the reference point 143 (e.g., to the left of the reference point 143).
[0091] The reference position adjuster 208 determines (e.g., updates) the reference position data 103 based on the reference position data 113. For example, the reference position adjuster 208 updates the reference position data 103 to a second value 286 indicating the position 296, the orientation 298, or both. In a particular aspect, the reference position adjuster 208 provides the reference position data 103 (e.g., the second value 286) to the parameter generator 210.
[0092] Return to Figure 2A, the parameter generator 210 generates one or more selection parameters 156 indicating that the spatial audio data 170 is associated with the position data 174 (e.g., the second value 286 of the reference position data 103, the second value 266 of the user position data 105, or both). The parameter generator 210 generates one or more position data sets (e.g., predicted position data, predetermined position data, or both). For example, the parameter generator 210 generates the position data 176 indicating the reference position data 123, the user position data 125, or both, such as the reference position data 123. Figure 3 Further described. In some examples, parameter generator 210 generates one or more additional location data sets. Parameter generator 210 provides each location data set to a specific renderer. For example, parameter generator 210 provides location data 176 to renderer 214, provides an additional set of location data to an additional renderer, or both.
[0093] The reference position adjuster 208 provides the spatial audio data 170 to one or more renderers (e.g., a renderer 212, a renderer 214, one or more additional renderers, or a combination thereof). The renderer 212 generates one or more directional audio data sets based on the spatial audio data 170. For example, the renderer 212 performs binaural processing on the spatial audio data 170 to generate directional audio data 152 corresponding to a first channel (e.g., a right channel) and directional audio data 252 corresponding to a second channel (e.g., a left channel). The spatial audio data 170 is associated with position data 174 (e.g., detected position data, default position data, or both).
[0094] The renderer 214 generates the spatial audio data 270 by adjusting the spatial audio data 170 based on the position data 174 and the position data 176. In certain aspects, the spatial audio data 170 represents the sound from the sound source 184 that will be perceived as coming from the position 192 relative to the reference point 143 (e.g., to the left and from a certain distance). The spatial audio data 170 corresponds to the arrangement 162 of the sound source 184 relative to the listener (e.g., the user of the device 104), as shown in the reference point 143. Figure 1 and 2C The renderer 214 generates the spatial audio data 270 to have Figure 1 arrangement 164 such that sound from sound source 184 is perceived as coming from a particular direction (e.g., in front) of a listener (e.g., a user of device 104), and when spatial audio data 270 is played out, such that when the user has a user position indicated by user position data 125 and reference point 143 has a reference position indicated by reference position data 123, the sound will be perceived as coming from position 192 relative to reference point 143.
[0095] The renderer 214 generates one or more directional audio data sets based on the spatial audio data 270. For example, the renderer 214 performs binaural processing on the spatial audio data 270 to generate directional audio data 154 corresponding to a first channel (e.g., the right channel) and directional audio data 254 corresponding to a second channel (e.g., the left channel). The spatial audio data 270 is associated with location data 176 (e.g., predicted location data, predetermined location data, or both).
[0096] In some examples, one or more additional renderers generate additional directional audio data sets. For example, the additional renderer generates specific spatial audio data by adjusting the spatial audio data 170 based on the location data 174 and specific location data. The specific spatial audio data corresponds to a specific sound arrangement. The additional renderer 214 generates one or more additional directional audio data sets based on the specific spatial audio data. For example, the additional renderer performs binaural processing on the specific spatial audio data to generate first directional audio data corresponding to a first channel (e.g., the right channel) and second directional audio data corresponding to a second channel (e.g., the left channel).
[0097] The stream generator 140 provides the directional audio data 152, the directional audio data 252, the directional audio data 154, the directional audio data 254, one or more additional directional audio data sets, or a combination thereof as an output stream 150 to the stream selector 142. In certain aspects, the stream generator 140 provides one or more selection parameters 156 to the stream selector 142 while providing the output stream 150 to the stream selector 142. The one or more selection parameters 156 indicate that the directional audio data 152, the directional audio data 252, or both are associated with the location data 174. The one or more selection parameters 156 indicate that the directional audio data 154, the directional audio data 254, or both are associated with the location data 176. In some examples, the one or more selection parameters 156 indicate that one or more additional directional audio data sets are associated with additional location data.
[0098] Reference Figure 3 , FIG. 300 showing illustrative aspects of the operation of the parameter generator 210. In certain aspects, the parameter generator 210 includes a user interactivity predictor 374 coupled to a reference location predictor 376, a user location predictor 378, or both. In certain aspects, the parameter generator 210 includes a predetermined location data generator 380.
[0099] The user interactivity predictor 374 is configured to generate predicted user interactivity data 375 by processing user interactivity data 111. In a particular embodiment, the user interactivity predictor 374 determines predicted interaction data 393 based on user interactivity data 111 including application data indicating future events, application data history, or a combination thereof. By way of illustration, the predicted interaction data 393 indicates a predicted occurrence event (e.g., an explosion at a particular virtual location in a video game). In a particular aspect, the user interactivity predictor 374 (e.g., a neural network) generates predicted virtual reference position data 391 based on virtual reference position data 107 indicated by the user interactivity data 111, the predicted interaction data 393, or both. The predicted virtual reference position data 391 indicates a predicted position of a reference point 143 (e.g., a virtual reference point). In a particular aspect, the user interactivity predictor 374 provides the predicted user interactivity data 375 to the reference position predictor 376, the user position predictor 378, or both.
[0100] The reference position predictor 376 determines predicted reference position data 377 based on reference position data 113, the predicted virtual reference position data 391, the predicted interaction data 393, or a combination thereof. The predicted reference position data 377 indicates a predicted position of the reference point 143 (e.g., an absolute position or a change in position). In a particular aspect, the reference point 143 includes a virtual reference point, and the predicted reference position data 377 indicates the predicted virtual reference position data 391. In a particular aspect, the reference point 143 corresponds to a fixed reference point (e.g., a television), and the predicted reference position data 377 indicates that the predicted reference point 143 has the same position as indicated by the reference position data 113. In a particular aspect, the reference point 143 is movable, and the reference position predictor 376 tracks the movement of the reference point 143 based on the reference position data 113, previous reference position data, or a combination thereof to generate the predicted reference position data 377.
[0101] The user location predictor 378 determines predicted user location data 379 based on user location data 115, predicted reference location data 377, predicted interaction data 393, or a combination thereof. The predicted user location data 379 indicates the predicted location of the user of the device 104 (e.g., an absolute location or a location change). In certain aspects, the user location predictor 378 determines the predicted user location data 379 based on an event predicted by the predicted interaction data 393, the predicted location of a reference point 143 indicated by the predicted reference location data 377, or both. For example, the predicted user location data 379 generates a user location predictor 378 to indicate that the predicted user moves away from a predicted event (e.g., an explosion in a video game), that the predicted user follows the reference point 143 (e.g., an NPC), or both. In a particular aspect, the user location predictor 378 tracks the movement of the user of the device 104 based on the user location data 115, previous user location data, or a combination thereof to generate the predicted user location data 379.
[0102] The predetermined location data generator 380 is configured to generate predetermined location data (e.g., predetermined reference location data 381, predetermined user location data 383, or both). In certain aspects, the predetermined location data generator 380 generates the predetermined reference location data 381 based on the reference location data 113 and a set of predetermined values. For example, the predetermined location data generator 380 generates a predetermined reference orientation of the predetermined reference location data 381 by incrementing (or decrementing) a reference orientation indicated by the reference location data 113 by a predetermined orientation (e.g., 10 degrees) indicated by a predetermined set of values. As another example, the predetermined location data generator 380 generates a predetermined reference location of the predetermined reference location data 381 by incrementing (or decrementing) a reference location indicated by the reference location data 113 by a predetermined displacement (e.g., a specific distance in a specific direction) indicated by a predetermined set of values.
[0103] In certain aspects, the predetermined location data generator 380 generates the predetermined user location data 383 based on the user location data 115 and a set of predetermined values. For example, the predetermined location data generator 380 generates a predetermined reference orientation of the predetermined reference location data 381 by incrementing (or decrementing) a reference orientation indicated by the reference location data 113 by a predetermined orientation (e.g., 10 degrees) indicated by a predetermined set of values. As another example, the predetermined location data generator 380 generates a predetermined reference location of the predetermined reference location data 381 by incrementing (or decrementing) a reference location indicated by the reference location data 113 by a predetermined displacement (e.g., a specific distance in a specific direction) indicated by a predetermined set of values.
[0104] In certain aspects, the parameter generator 210 generates location data 176 based on predicted reference location data 377, predicted user location data 379, predetermined reference location data 381, predetermined user location data 383, or a combination thereof. For example, the reference location data 123 is based on the predicted reference location data 377, the predetermined reference location data 381, or both. In a particular example, the user location data 125 is based on the predicted user location data 379, the predetermined user location data 383, or both.
[0105] In certain aspects, the parameter generator 210 generates one or more additional location data sets, and the selection parameter 156 includes the one or more additional location data sets. In some examples, the reference location predictor 376 generates multiple predicted reference location data sets corresponding to multiple predicted reference locations, the user location predictor 378 generates multiple predicted user location data sets corresponding to multiple predicted user locations, or both. The parameter generator 210 generates multiple location data sets based on the multiple predicted reference locations, the multiple predicted user locations, or a combination thereof. In some examples, the predetermined location data generator 380 generates multiple predetermined reference location data sets corresponding to multiple predetermined reference locations and multiple predetermined user location data sets corresponding to multiple predetermined user locations. The parameter generator 210 generates multiple location data sets based on the multiple predetermined reference locations, the multiple predetermined user locations, or a combination thereof.
[0106] Reference Figure 4 , FIG. 400 shows an illustrative aspect of the operation of the stream selector 142. The stream selector 142 includes a combination factor (CF) generator 404 and one or more audio decoders (e.g., audio decoder 406A, audio decoder 406B, one or more additional audio decoders, or a combination thereof). The combination factor generator 404 is coupled to each of one or more acoustic stream generators (e.g., acoustic stream generator 408A, acoustic stream generator 408B, one or more additional acoustic stream generators, or a combination thereof). The one or more audio decoders are coupled to the one or more acoustic stream generators. For example, the audio decoder 406A is coupled to the acoustic stream generator 408A. As another example, the audio decoder 406B is coupled to the acoustic stream generator 408B.
[0107] The stream selector 142 receives user location data 115 from the location sensor 186 indicating the location of the device 104, the user of the device 104, or both detected at a first user location time. The stream selector 142 provides the user location data 115 to the stream generator 140 at a first time. The stream selector 142 receives the output stream 150, one or more selection parameters 156, or a combination thereof at a second time after the first time.
[0108] In certain aspects, output stream 150 includes directional audio data 152 (e.g., right channel data) and directional audio data 252 (e.g., left channel data) based on location data 174 (e.g., detected location data, default location data, or both). In certain aspects, output stream 150 includes directional audio data 154 (e.g., right channel data) and directional audio data 254 (e.g., left channel data) based on location data 176 (e.g., predetermined location data, predicted location data, or both). In some examples, output stream 150 includes additional directional audio data sets based on additional location data sets.
[0109] In certain aspects, audio decoder 406A decodes the directional audio data of the first audio channel (e.g., right channel), and audio decoder 406B decodes the directional audio data of the second audio channel (e.g., left channel). For example, audio decoder 406A decodes directional audio data 152 to generate acoustic data 452, decodes directional audio data 154 to generate acoustic data 454, decodes additional directional audio data to generate additional acoustic data, or a combination thereof. Audio decoder 406B decodes directional audio data 252 to generate acoustic data 456, decodes directional audio data 254 to generate acoustic data 458, decodes additional directional audio data to generate additional acoustic data, or a combination thereof. In some examples, an additional audio decoder decodes the directional audio data of an additional audio channel.
[0110] Combination factor generator 404 receives user location data 185 from location sensor 186 indicating the location of device 104, a user of device 104, or both, detected at a second user location time after a first user location time associated with user location data 115. In certain aspects, combination factor generator 404 receives reference location data 157 from stream generator 140. For example, reference location data 157 corresponds to an updated location (e.g., detected location) of reference point 143 relative to the location of reference point 143 indicated by reference location data 103.
[0111] Combination factor generator 404 generates combination factor 405 based on location data 476 (e.g., user location data 185, reference location data 157, or both), one or more selection parameters 156, or a combination thereof. In certain aspects, location data 174 corresponds to previously detected location data or default location data, location data 176 corresponds to predetermined location data or predicted location data, and location data 476 corresponds to most recently detected location data. In certain aspects, one or more selection parameters 156 include additional location data sets (e.g., corresponding to additional predetermined locations, additional predicted locations, or a combination thereof).
[0112] The combined factor generator 404 generates a combined factor 405 based on a comparison of the location data 476 with the location data 174, the location data 176, the location data of one or more additional groups, or a combination thereof. In certain aspects, the combined factor generator 404 determines a first reference difference based on a comparison of a reference location indicated by the reference location data 103 (e.g., a default reference location or a previously detected reference location) with a reference location indicated by the reference location data 157 (e.g., the most recently detected reference location). The combined factor generator 404 determines a second reference difference based on a comparison of a reference location indicated by the reference location data 123 (e.g., a predetermined reference location or a predicted reference location) with a reference location indicated by the reference location data 157 (e.g., the most recently detected reference location). The combined factor generator 404 determines a first user difference based on a comparison of a user location indicated by the user location data 105 (e.g., a default user location or a previously detected user location) with a user location indicated by the user location data 185 (e.g., the most recently detected user location). The combined factor generator 404 determines a second user difference based on a comparison of a user location indicated by the user location data 125 (e.g., a predetermined user location or a predicted user location) with a user location indicated by the user location data 185 (e.g., the most recently detected user location).
[0113] The combined factor generator 404 generates a first difference indicator based on the first reference difference, the first user difference, or both. The combined factor generator 404 generates a second difference indicator based on the second reference difference, the second user difference, or both. The first difference indicator indicates the level of difference between the location data 174 and the location data 476. The second difference indicator indicates the level of difference between the location data 176 and the location data 476. In certain aspects, the combined factor generator 404 generates one or more additional difference indicators based on one or more additional location data sets.
[0114] In a particular embodiment, the combination factor generator 404 generates a combination factor 405 to have a first value (e.g., 0) based on determining that the location data 476 is closer to or an equal match with the location data 174 than with the location data 176. For example, the combination factor generator 404 generates the combination factor 405 to have the first value (e.g., 0) in response to determining that a first difference indicator indicates a difference level that is less than or equal to the difference level indicated by a second difference indicator (e.g., first difference indicator ≤ second difference indicator). Alternatively, the combination factor generator 404 generates the combination factor 405 to have a second value (e.g., 1) based on determining that the match of the location data 476 with the location data 176 is closer than the match with the location data 174. For example, the combination factor generator 404 generates the combination factor 405 to have the second value (e.g., 1) in response to determining that the first difference indicator indicates a difference level that is greater than the difference level indicated by the second difference indicator (e.g., first difference indicator > second difference indicator).
[0115] In an alternative embodiment, the combination factor generator 404 generates a combination factor 405 that is greater than or equal to a first value (e.g., 0) and less than or equal to a second value (e.g., 1) based on the relative difference between the location data 476 and the location data 174 and the location data 176. For example, the combination factor generator 404 generates the combination factor 405 to have a value based on the ratio of the first difference indicator and the second difference indicator (e.g., combination factor 405 = first difference indicator / (first difference indicator + second difference indicator)). In a particular aspect, the combination factor generator 404 generates the combination factor 405 to have a specific value corresponding to an additional location data set that is closer to or an equal match with the location data 476 compared to other location data sets.
[0116] The combination factor generator 404 provides the combination factor 405 to each of the acoustic streaming generators 408A and 408B. In a particular aspect, the acoustic streaming generator 408 selects acoustic data corresponding to position data associated with a particular value of the combination factor 405 in response to determining that the combination factor 405 has the particular value. In a particular implementation, the acoustic streaming generator 408 selects audio data associated with the position data 174 in response to determining that the combination factor 405 has a first value (e.g., 0). For example, the acoustic streaming generator 408A selects the acoustic data 452 associated with the position data 174 as the acoustic data 172 in response to determining that the combination factor 405 has the first value (e.g., 0). In response to determining that the combination factor 405 has the first value (e.g., 0), the acoustic streaming generator 408B selects the acoustic data 456 associated with the position data 174 as the acoustic data 472. Alternatively, the acoustic streaming generator 408 selects audio data associated with the position data 176 in response to determining that the combination factor 405 has a second value (e.g., 1). For example, the acoustic streaming generator 408A selects the acoustic data 454 associated with the position data 176 as the acoustic data 172 in response to determining that the combination factor 405 has the second value (e.g., 1). In response to determining that the combination factor 405 has the second value (e.g., 1), the acoustic streaming generator 408B selects the acoustic data 458 associated with the position data 176 as the acoustic data 472.
[0117] In a particular implementation, the acoustic streaming generator 408 combines audio data associated with a set of position data (e.g., audio data associated with the position data 174, audio data associated with the position data 176, audio data associated with one or more additional sets of position data, or a combination thereof) based on the combination factor 405. In a particular example, the acoustic streaming generator 408A generates a first weight (e.g., first weight = 1 - combination factor 405) based on the combination factor 405 and generates a second weight (e.g., second weight = combination factor 405) based on the combination factor 405. The acoustic streaming generator 408A generates the acoustic data 172 based on a weighted sum of the acoustic data 452 and the acoustic data 454. For example, the acoustic data 172 corresponds to a combination of the first weight applied to the acoustic data 452 and the second weight applied to the acoustic data 454 (e.g., acoustic data 172 = first weight (acoustic data 452) + second weight (acoustic data 454)).
[0118] In a particular example, the acoustic streaming generator 408B generates a first weight based on the combination factor 405 (e.g., first weight = 1 - combination factor 405) and generates a second weight based on the combination factor 405 (e.g., second weight = combination factor 405). The acoustic streaming generator 408B generates the acoustic data 472 based on a weighted sum of the acoustic data 456 and the acoustic data 458. For example, the acoustic data 472 corresponds to a combination of the first weight applied to the acoustic data 456 and the second weight applied to the acoustic data 458 (e.g., acoustic data 472 = first weight (acoustic data 456)+second weight (acoustic data 458)).
[0119] In a particular aspect, the stream selector 142 enables the generation of the acoustic data 172 such that the difference between the acoustic data 172 and the acoustic data 452 (corresponding to the directional audio data 152) and the acoustic data 454 (corresponding to the directional audio data 154) corresponds to the difference between the location data 476 and the location data 174 and the location data 176. For example, when the location data 476 (e.g., most recently detected location data) is closer to the location data 174 (e.g., previously detected location data or default location data), the acoustic data 172 is closer to the acoustic data 452 (e.g., based on the location data 174). Alternatively, when the location data 476 (e.g., most recently detected location data) is closer to the location data 176 (e.g., predetermined location data or predicted location data), the acoustic data 172 is closer to the acoustic data 454 (e.g., based on the location data 176).
[0120] The stream selector 142 outputs the acoustic data 172 and the acoustic data 472 as the output stream 450 to one or more speakers. For example, in response to determining that the acoustic data 172 is associated with a first channel (e.g., right channel), the stream selector 142 outputs the acoustic data 172 to the speaker 120 associated with the first channel. As another example, in response to determining that the acoustic data 472 is associated with a second channel (e.g., left channel), the stream selector 142 outputs the acoustic data 472 to the speaker 122 associated with the second channel.
[0121] In certain aspects, the stream selector 142 receives the output stream 150 from the stream generator 140 before receiving the user location data 185, the reference location data 157, or both. Thus, the stream selector 142 can generate the output stream 450 when the location data 476 is received, without the latency associated with generating the directional audio data 152, the directional audio data 154, or both. In certain aspects, generating the acoustic data 172 based on the acoustic data 452 and the acoustic data 454 uses fewer resources than generating either the directional audio data 152 or the directional audio data 154 based on the spatial audio data 170 and the location data 476. Thus, having the stream generator 140 on the device 102 offloads some processing from the device 104.
[0122] Reference Figure 5 , shows a system 500 operable to generate directional audio having a plurality of sound source arrangements. The device 102 (e.g., a host device) includes a stream generator 140 coupled via a stream selector 142 to one or more audio encoders (e.g., audio encoders 542A, audio encoders 542B, one or more additional audio encoders, or combinations thereof). The device 104 includes one or more audio decoders, e.g., audio decoders 506A, audio decoders 506B, one or more additional audio decoders, or combinations thereof.
[0123] The device 104 provides the user location data 115 to the device 102 at a first time. The stream generator 140 generates the output stream 150, one or more selection parameters 156, or combinations thereof based on the spatial audio data 170, the reference location data 113, the user location data 115, or combinations thereof, as described in reference Figure 2A . The stream generator 140 provides the output stream 150, one or more selection parameters 156, or combinations thereof to the stream selector 142.
[0124] The stream selector 142 receives the output stream 150, one or more selection parameters 156, or combinations thereof from the stream generator 140. The device 104 provides the user location data 185 to the device 102 at a second time after the first time. In certain aspects, the stream selector 142 receives the reference location data 157 from the stream generator 140. In an alternative aspect, the stream selector 142 determines the reference location data 157. For example, the stream selector 142 receives user interactivity data 111 indicating a second virtual reference location data of a reference point 143 (e.g., a virtual reference point), and determines the reference location data 157 at least in part based on the second virtual reference location data. In a particular example, the stream selector 142 receives second device location data from the location sensor 188, and determines the reference location data 157 at least in part based on the second device location data.
[0125] The stream selector 142 generates acoustic data 172, acoustic data 472, or both, based on the output stream 150, one or more selection parameters 156, location data 476 (e.g., reference location data 157, user location data 185, or both), or a combination thereof, as referenced Figure 4 as described. In a particular embodiment, the stream selector 142 does not include the audio decoder 406A or the audio decoder 406B. In this embodiment, the stream selector 142 provides the directional audio data 152 as acoustic data 452 and the directional audio data 154 as acoustic data 454 to the acoustic stream generator 408A. The stream selector 142 provides the directional audio data 252 as acoustic data 456 and the directional audio data 254 as acoustic data 458 to the acoustic stream generator 408B. The acoustic stream generator 408A combines the directional audio data 152 (e.g., acoustic data 452) and the directional audio data 154 (e.g., acoustic data 454) based on the combination factor 405 to generate acoustic data 172. In a particular aspect, the acoustic stream generator 408A selects one of the directional audio data 152 (e.g., acoustic data 452) or the directional audio data 154 (e.g., acoustic data 454) as acoustic data 172 based on the combination factor 405. Similarly, the acoustic stream generator 408B generates acoustic data 472 based on the directional audio data 252 and the directional audio data 254.
[0126] The stream selector 142 provides the acoustic data 172 to the audio encoder 542A, the acoustic data 472 to the audio encoder 542B, or both. The audio encoder 542A generates the directional audio data 552 by encoding the acoustic data 172. The audio encoder 542B generates the directional audio data 554 by encoding the acoustic data 472. The device 102 initiates the transmission of the directional audio data 552, the directional audio data 554, or both, as the output stream 550 to the device 104.
[0127] The device 104 receives the output stream 550 from the device 102. The audio decoder 506A generates the acoustic data 172 by decoding the directional audio data 552. The audio decoder 506B generates the acoustic data 472 by decoding the directional audio data 554. In response to determining that the acoustic data 172 is associated with the first channel (e.g., the right channel), the audio decoder 506A provides the acoustic data 172 to the speaker 120 associated with the first channel. In response to determining that the acoustic data 472 is associated with the second channel (e.g., the left channel), the audio decoder 506B provides the acoustic data 472 to the speaker 122 associated with the second channel.
[0128] Accordingly, system 500 enables most of the processing to be offloaded from device 104 to device 102. System 500 also enables the stream generator 140 and the stream selector 142 to operate with traditional audio output devices such as device 104.
[0129] Reference Figure 6 , a system 600 is shown that is operable to generate directional audio with a plurality of sound source arrangements. System 600 includes a device 604, and device 604 includes a stream generator 140 and a stream selector
[0130] 142. Device 604 is coupled to one or more speakers (e.g., speaker 120, speaker 122, one or more additional speakers, or a combination thereof). In certain aspects, device 604 includes or is coupled to one or more position sensors (e.g., position sensor 186, position sensor 188, or both). In example 620, device 102 includes device 604. In example 640, device 104 includes device 604.
[0131] The stream generator 140 receives user location data 115 from the position sensor 186 at a first time. The stream generator 140 generates an output stream 150, one or more selection parameters 156, or a combination thereof based on the spatial audio data 170, the reference location data 113, the user location data 115, or a combination thereof, as described in the reference Figure 2A . The stream generator 140 provides the output stream 150, one or more selection parameters 156, or a combination thereof to the stream selector 142.
[0132] The stream selector 142 receives the output stream 150, one or more selection parameters 156, or a combination thereof from the stream generator 140. The stream selector 142 receives user location data 185 from the position sensor 186 at a second time after the first time. In certain aspects, the stream selector 142 receives reference location data 157 from the stream generator 140. In an alternative aspect, the stream selector 142 determines the reference location data 157 based on second virtual reference location data indicated by user interactivity data 111, second device location data from the position sensor 188, or both.
[0133] The stream selector 142 generates acoustic data 172, acoustic data 472, or both based on the output stream 150, one or more selection parameters 156, location data 476 (e.g., reference location data 157, user location data 185, or both), or a combination thereof, as described in the reference Figure 4As described. In a particular embodiment, the stream selector 142 does not include the audio decoder 406A or the audio decoder 406B. In this embodiment, the stream selector 142 provides the directional audio data 152 as the acoustic data 452 and the directional audio data 154 as the acoustic data 454 to the acoustic stream generator 408A. The stream selector 142 provides the directional audio data 252 as the acoustic data 456 and the directional audio data 254 as the acoustic data 458 to the acoustic stream generator 408B.
[0134] The stream selector 142 provides the acoustic data 172, the acoustic data 472, or both as the output stream 650 to one or more speakers. For example, in response to determining that the acoustic data 172 is associated with a first channel (e.g., the right channel), the stream selector 142 renders an acoustic output based on the acoustic data 172 and provides the acoustic output to the speaker 120 associated with the first channel. In response to determining that the acoustic data 472 is associated with a second channel (e.g., the left channel), the stream selector 142 renders an acoustic output based on the acoustic data 472 and provides the acoustic output to the speaker 122 associated with the second channel.
[0135] Accordingly, the system 600 enables the stream generator 140 to reduce audio latency by generating the output stream 150 before receiving the position data 476 (the reference position data 157, the user position data 185, or both). In certain aspects, generating the acoustic data 172 and the acoustic data 472 from the output stream 150 when the position data 476 is available is faster than adjusting the spatial audio data 170 based on the position data 476 to generate the acoustic data.
[0136] Figure 7 FIG. 700 is a diagram of illustrative aspects of the operation of the stream generator 140 and the stream selector 142. The stream generator 140 is configured to receive spatial audio data 170 corresponding to a sequence of audio data samples, such as a sequence of continuously captured frames, which are shown as a first frame (F1) 712, a second frame (F2) 714, and one or more additional frames including an Nth frame (FN) 716 (where N is an integer greater than 2). The stream generator 140 is configured to output directional audio data 152 corresponding to the sequence of audio data samples, such as a sequence of frames, which is shown as a first frame (F1) 722, a second frame (F2) 724, and one or more additional sets including an nth frame (FN) 726. The stream generator 140 is configured to output the directional audio data 154 while outputting the directional audio data 152. For example, the stream generator 140 is configured to output directional audio data 154 corresponding to a sequence of audio samples (e.g., a sequence of frames), the frame sequence being illustrated as a first frame (F1) 732, a second frame (F2) 734, and one or more additional sets including an nth frame (FN) 736.
[0137] The stream selector 142 is configured to receive the directional audio data 152 and the directional audio data 154 and generate acoustic data 172. For example, the stream selector 142 is configured to output acoustic data 172 corresponding to a sequence of audio samples, such as a sequence of frames, which is shown as a first frame (F1) 742, a second frame (F2) 744, and one or more additional sets including an nth frame (FN) 746.
[0138] During operation, the stream generator 140 processes the first frame 712 to generate a first frame 722 and a first frame 732. The stream selector 142 generates the first frame 742 based on the first frame 722 and the first frame 732. For example, the stream selector 142 selects one of the first frame 722 or the first frame 732 as the first frame 742. As another example, the stream selector 142 combines the first frame 722 and the first frame 732 to generate the first frame 742. This processing continues, including the stream generator 140 processing the nth frame 716 to generate an nth frame 726 and an nth frame 736, and the stream selector 142 generating the nth frame 746 based on the nth frame 726 and the nth frame 736. In certain aspects, the stream generator 140 generates the directional audio data 154 at least in part based on position data associated with a previous frame. For example, as audio is processed across multiple frames, the accuracy of position prediction can be improved.
[0139] Figure 8 An embodiment 800 of an integrated circuit 802 including one or more processors 890 is depicted. The one or more processors 890 include the stream generator 140, the stream selector 142, the position sensor 186, the position sensor 188, or a combination thereof. In certain aspects, the integrated circuit 802 includes any one of or is included in the following: the device 102, Figure 1 , 5 , the device 104 of 6, Figure 6 the device 604 of, or a combination thereof.
[0140] The integrated circuit 802 includes an audio input 804, such as one or more bus interfaces, to enable the receipt of audio data 850 for processing. The integrated circuit 802 also includes an audio output 806, such as a bus interface, to enable the transmission of an output stream 870. In certain aspects, the audio data 850 contains user position data 115, spatial audio data 170, reference position data 113, user interactivity data 111, device position data 109, or a combination thereof, and the output stream 870 contains an output stream 150, one or more selection parameters 156, reference position data 157, or a combination thereof.
[0141] In certain aspects, the audio data 850 includes the output stream 150, one or more selection parameters 156, reference location data 157, user location data 185, or a combination thereof, and the output stream 870 includes acoustic data 172, acoustic data 472, output stream 450, or a combination thereof. In certain aspects, the audio data 850 includes user location data 115, spatial audio data 170, reference location data 113, user interactivity data 111, device location data 109, reference location data 157, user location data 185, or a combination thereof, and the output stream 870 includes directional audio data 552, directional audio data 554, output stream 550, or a combination thereof.
[0142] In certain aspects, the audio data 850 includes user location data 115, spatial audio data 170, reference location data 113, user interactivity data 111, device location data 109, reference location data 157, user location data 185, or a combination thereof, and the output stream 870 includes acoustic data 172, acoustic data 472, output stream 650, or a combination thereof.
[0143] Integrated circuit 802 implements an embodiment of directional audio generation, where multiple sound sources are arranged as components in a system including speakers, such as Figure 9 the wearable electronic device shown, such as Figure 10 the voice-controlled speaker system shown, such as Figure 11 the virtual reality headset or augmented reality headset shown, or such as Figure 12 or Figure 13 the vehicle shown.
[0144] Figure 9 An embodiment 900 of the wearable electronic device 902 shown as a "smartwatch" is depicted. In certain aspects, the wearable electronic device 902 includes the device 102, Figure 1 , 5 , the device 104 of 6, Figure 6 the device 604 of, or a combination thereof.
[0145] The stream generator 140, the stream selector 142, or both are integrated into the wearable electronic device 902. In certain aspects, the wearable electronic device 902 is coupled to or includes the position sensor 186, the position sensor 188, the speaker 120, the speaker 122, or a combination thereof. In a particular example, the stream generator 140 and the stream selector 142 operate to detect user speech activity in the acoustic data 172 and then process the acoustic data 172 to perform one or more operations at the wearable electronic device 902, such as launching a graphical user interface or otherwise displaying other information associated with the user's speech at the display screen 904 of the wearable electronic device 902. By way of illustration, the wearable electronic device 902 may include a display screen configured to display notifications based on user speech detected by the wearable electronic device 902. In a particular example, the wearable electronic device 902 includes a haptic device that provides a haptic notification (e.g., vibration) in response to detecting user speech activity. For example, the haptic notification may cause the user to look at the wearable electronic device 902 to see the displayed notification indicating the detected keyword spoken by the user. Thus, the wearable electronic device 902 can alert a user with hearing impairment or a user wearing headphones of the detected user speech activity.
[0146] Figure 10 is an embodiment 1000 of a wireless speaker and voice-activated device 1002. In certain aspects, the wireless speaker and voice-activated device 1002 includes the device 102, Figure 1 , 5 , the device 104 of 6, Figure 6 the device 604 thereof, or a combination thereof.
[0147] The wireless speaker and voice-activated device 1002 may have a wireless network connection and be configured to perform auxiliary operations. One or more processors 890 including the stream generator 140, the stream selector 142, or both are included in the wireless speaker and voice-activated device 1002. In certain aspects, the wireless speaker and voice-activated device 1002 includes or is coupled to the position sensor 186, the position sensor 188, the speaker 120, the speaker 122, or a combination thereof. During operation, in response to receiving an oral command recognized as user speech via the operation of the stream generator 140, the stream selector 142, or both, the wireless speaker and voice-activated device 1002 may perform auxiliary operations, such as via the execution of a voice-activated system (e.g., an integrated assistant application). Auxiliary operations may include adjusting the temperature, playing music, turning on the lights, etc. For example, in response to receiving a command to perform an assistant operation after a keyword or key phrase (e.g., "Hello Assistant").
[0148] Figure 11Illustrates an embodiment 1100 of a portable electronic device corresponding to a virtual reality, augmented reality, or mixed reality headset 1102. In certain aspects, the headset 1102 includes device 102, Figure 1 , 5 , device 104 of 6, Figure 6 , device 604 of
[0149] Figure 12 or a combination thereof. A stream generator 140, a stream selector 142, a position sensor 186, a position sensor 188, speakers 120, speakers 122, or a combination thereof are integrated into the headset 1102. In certain aspects, the acoustic data 172 is output by the stream selector 142 via the speaker 120. A visual interface device is located in front of the user's eyes to enable the display of augmented reality or virtual reality images or scenes to the user when wearing the headset 1102. Figure 1 , 5 , device 104 of 6, Figure 6 , device 604 of
[0150] or a combination thereof. A stream generator 140, a stream selector 142, a position sensor 186, a position sensor 188, speakers 120, speakers 122, or a combination thereof are integrated into the vehicle 1202. In certain aspects, the acoustic data 172 is output by the stream selector 142 via the speaker 120, for example, for delivery instructions from an authorized user of the vehicle 1202.
[0151] Figure 13 Illustrates another embodiment 1300 of a vehicle 1302 shown as an automobile. In certain aspects, the vehicle 1202 includes device 102, Figure 1 , 5 , device 104 of 6, Figure 6 , device 604 of
[0152] The vehicle 1302 includes a stream generator 140, a stream selector 142, a position sensor 186, a position sensor 188, speakers 120, speakers 122, or a combination thereof. In some examples, the stream generator 140 of the vehicle 1302 generates Figure 1 , an output stream 150, and provides the output stream 150 to the device 104 of a passenger of the vehicle 1302. In some examples, the stream selector 142 Figure 6The output stream 650 is provided to speaker 120, speaker 122, or both. In a particular embodiment, the voice activation system initiates one or more operations of vehicle 1302 based on one or more keywords detected in output stream 150 (e.g., "unlock", "start engine", "play music", "display weather forecast", or another voice command), such as by providing feedback or information via display 1320 or one or more speakers (e.g., speaker 120, speaker 122, or both).
[0153] Reference Figure 14 , a particular embodiment of method 1400 for generating directional audio using a multi-source arrangement is shown. In a particular aspect, one or more operations of method 1400 are performed by Figure 1 stream generator 140, device 102, device 104, system 100, Figure 6 device 604, or a combination of at least one thereof.
[0154] Method 1400 includes obtaining spatial audio data representing audio from one or more sound sources at 1402. For example, Figure 1 stream generator 140 obtains spatial audio data 170 representing audio from one or more sound sources 184, as described in reference Figure 1 .
[0155] Method 1400 further includes generating first directional audio data at 1404 based on the spatial audio data, the first directional audio data corresponding to a first arrangement of one or more sound sources relative to the audio output device. For example, Figure 1 stream generator 140 generates directional audio data 152 based on spatial audio data 170. Directional audio data 152 corresponds to the arrangement 162 of one or more sound sources 184 relative to device 104, speaker 120, or both, as described in reference Figure 1 .
[0156] Method 1400 further includes generating second directional audio data at 1406 based on the spatial audio data, the second directional audio data corresponding to a second arrangement of one or more sound sources relative to the audio output device, where the second arrangement is different from the first arrangement. For example, Figure 1 stream generator 140 generates directional audio data 154 based on spatial audio data 170. Directional audio data 154 corresponds to the arrangement 164 of one or more sound sources 184 relative to device 104, speaker 120, or both, as described in reference Figure 1 .
[0157] Method 1400 further includes generating an output stream at 1408 based on the first directional audio data and the second directional audio data. For example, Figure 1The stream generator 140 generates an output stream 150 based on the directional audio data 152 and the directional audio data 154, as referenced in Figure 1 as described. In another example, the stream selector 142 generates an output stream 550 based on the directional audio data 152 and the directional audio data 154, as referenced in Figure 5 as described. In certain aspects, the stream selector 142, the device 604, or both generate an output stream 650 based on the directional audio data 152 and the directional audio data 154, as described in Figure 6 reference.
[0158] Method 1400 further includes providing the output stream to an audio output device at 1410. For example, Figure 1 the stream generator 140 provides the output stream 150 to the device 104, the stream selector 142, or both, as referenced in Figure 1 as described. In another example, the stream selector 142 provides the output stream 550 to the device 104, the stream selector 142, or both, as referenced in Figure 5 as described. In certain aspects, the stream selector 142, the device 604, or both provide the output stream 650 to the speaker 120, the speaker 122, or both, as described in Figure 6 reference.
[0159] Method 1400 may reduce audio latency by generating the directional audio data 152, the directional audio data 154, or both before receiving the location data 476. In some examples, method 1400 offloads some processing from the audio output device to the host device.
[0160] Figure 14 Method 1400 may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (e.g., a central processing unit (CPU)), a digital signal processor (DSP), a graphics processing unit (GPU), a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 14 method 1400 may be executed by a processor that executes instructions, such as those described in Figure 16 reference.
[0161] Reference Figure 15 illustrates a particular implementation of method 1500 for generating directional audio using multiple sound source arrangements. In certain aspects, one or more operations of method 1500 are performed by Figure 1 the stream generator 140, the device 102, the device 104, the system 100, Figure 6 the device 604 of
[0162] Method 1500 includes receiving, at 1502, first directional audio data representing audio from one or more sound sources from a host device, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device. For example, the stream selector 142 of device 104, Figure 1 or both receive directional audio data 152 representing audio from one or more sound sources 184. The directional audio data 152 corresponds to an arrangement 162 of the one or more sound sources 184 relative to a listener (e.g., device 104, speaker 120, or both), as described with reference to Figure 1 this.
[0163] Method 1500 further includes receiving, at 1504, second directional audio data representing audio from one or more sound sources from the host device, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement. For example, the stream selector 142 of device 104, Figure 1 or both receive directional audio data 154 representing audio from one or more sound sources 184. The directional audio data 154 corresponds to an arrangement 164 of the one or more sound sources 184 relative to a listener (e.g., device 104, speaker 120, or both), as described with reference to Figure 1 this.
[0164] Method 1500 further includes receiving, at 1506, position data indicating the location of the audio output device. For example, the stream selector 142 of device 104, Figure 1 or both receive user position data 185 indicating the location of device 104, speaker 120, or both, as described with reference to Figure 1 this.
[0165] Method 1500 further includes generating, at 1508, an output stream based on the first directional audio data, the second directional audio data, and the position data. For example, Figure 1 device 104, the stream selector 142, or both generate an output stream 450 based on the directional audio data 152, the directional audio data 154, and the user position data 185, as described with reference to Figure 4 this. In another example, device 604, the stream selector 142, or both generate an output stream 650 based on the directional audio data 152, the directional audio data 154, and the user position data 185, as described with reference to Figure 6 this.
[0166] Method 1500 further includes providing, at 1510, the output stream to the audio output device. For example, Figure 1 device 104, the stream selector 142, or both provide the output stream 450 to speaker 120, speaker 122, or both, as described with reference toFigure 4 as described. In another example, device 604, stream selector 142, or both provide output stream 650 to speaker 120, speaker 122, or both, as referenced Figure 6 as described.
[0167] Method 1500 can reduce audio latency by receiving directional audio data 152, directional audio data 154, or both before receiving location data 476, and generating acoustic data 172 based on directional audio data 152, directional audio data 154, location data 476, or a combination thereof. In some examples, method 1500 offloads some processing from the audio output device to the host device.
[0168] Figure 15 Method 1500 can be implemented by an FPGA device, an ASIC, a processing unit (e.g., CPU, DSP, GPU), a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 15 method 1500 can be executed by a processor that executes instructions, such as referenced Figure 16 as described.
[0169] Referenced Figure 16 , depicts a block diagram of a particular illustrative implementation of a device and is generally designated as 1600. In various implementations, device 1600 can have more or fewer components than Figure 16 shown. In the illustrative implementation, device 1600 can correspond to device 102, Figure 1 device 104 of Figure 6 device 604 of Figures 1 - 15 or a combination thereof. In the illustrative implementation, device 1600 can perform one or more operations described with reference to
[0170] In a particular implementation, device 1600 includes a processor 1606 (e.g., a CPU). Device 1600 can include one or more additional processors 1610 (e.g., one or more DSPs, one or more GPUs, or a combination thereof). In a particular aspect, Figure 8 one or more processors 890 of
[0171] Device 1600 may include memory 1686 and CODEC 1634. Memory 1686 may include instructions 1656, which may be executed by one or more additional processors 1610 (or processor 1606) to implement the functions described for reference stream generator 140, stream selector 142, or both. Device 1600 may include a modem 1640 coupled to antenna 1652 via transceiver 1650. In certain aspects, modem 1640 is configured to receive Figure 2A encoded audio data 203 from audio data source 202. In certain aspects, modem 1640 is configured to exchange data (e.g., Figure 1 user location data 115, output stream 150, one or more selection parameters 156, user location data 185, reference location data 157, Figure 2A encoded audio data 203, Figure 5 output stream 550, or a combination thereof) with device 102, device 104, audio data source 202, device 604, or a combination thereof.
[0172] Device 1600 may include a display 1628 coupled to display controller 1626. One or more speakers 1692, one or more microphones 1690, or a combination thereof may be coupled to CODEC 1634. In certain aspects, one or more speakers 1692 include speaker 120, speaker 122, or both. CODEC 1634 may include a digital-to-analog converter (DAC) 1602, an analog-to-digital converter (ADC) 1604, or both. In certain embodiments, CODEC 1634 may receive an analog signal from one or more microphones 1690, convert the analog signal to a digital signal using analog-to-digital converter 1604, and provide the digital signal to voice and music codec 1608. Voice and music codec 1608 may process the digital signal, and the digital signal may be further processed by stream generator 140, stream selector 142, or both. In certain embodiments, voice and music codec 1608 may provide the digital signal to codec 1634. CODEC 1634 may convert the digital signal to an analog signal using digital-to-analog converter 1602 and may provide the analog signal to one or more speakers 1692.
[0173] In certain embodiments, the device 1600 may be included in a system-in-package or system-on-chip device 1622. In certain embodiments, the memory 1686, the processor 1606, the processor 1610, the display controller 1626, the CODEC 1634, and the modem 1640 are included in the system-in-package or system-on-chip device 1622. In certain embodiments, the input device 1630 and the power supply 1644 are coupled to the system-on-chip device 1622. Additionally, in certain embodiments, as Figure 16 shown, the display 1628, the input device 1630, one or more speakers 1692, one or more microphones 1690, the antenna 1652, and the power supply 1644 are external to the system-on-chip device 1622. In certain embodiments, each of the display 1628, the input device 1630, one or more speakers 1692, one or more microphones 1690, the antenna 1652, and the power supply 1644 may be coupled to a component of the system-on-chip device 1622, such as an interface or a controller.
[0174] The device 1600 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a gaming device, headphones, a headset, an augmented reality headset, a virtual reality headset, an extended reality headset, an aircraft, a home automation system, a voice-activated device, a speaker, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a host device, an audio output device, a virtual reality (VR) device, a mixed reality (MR) device, an augmented reality (AR) device, an extended reality (XR) device, a base station, a mobile device, or any combination thereof.
[0175] In connection with the described embodiments, the apparatus includes components for obtaining spatial audio data representing audio from one or more sound sources. For example, the components for obtaining spatial audio data may correspond to Figure 1 the stream generator 140, the device 102, the device 104, the system 100, Figure 2A the audio decoder 204, the renderer 212, the renderer 214, Figure 6 the device 604, the antenna 1652, the transceiver 1650, the modem 1640, the voice and music codec 1608, the processor 1606, one or more additional processors 1610, one or more other circuits or components configured to obtain spatial audio data, or any combination thereof.
[0176] The apparatus further includes components for generating first directional audio data based on the spatial audio data. The first directional audio data corresponds to a first arrangement of one or more sound sources relative to the audio output device. For example, the device for generating the first directional audio data may correspond to Figure 1 the stream generator 140, device 102, device 104, system 100, Figure 2A the renderer 212, renderer 214, Figure 6 device 604, voice and music codec 1608, processor 1606, one or more additional processors 1610, one or more other circuits or components configured to generate directional audio data, or any combination thereof.
[0177] The apparatus further includes components for generating second directional audio data based on the spatial audio data. The second directional audio data corresponds to a second arrangement of one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. For example, the device for generating the second directional audio data may correspond to Figure 1 the stream generator 140, device 102, device 104, system 100, Figure 2A the renderer 212, renderer 214, Figure 6 device 604, voice and music codec 1608, processor 1606, one or more additional processors 1610, one or more other circuits or components configured to generate directional audio data, or any combination thereof.
[0178] The apparatus further includes components for generating an output stream based on the first directional audio data and the second directional audio data. For example, the components for generating the output stream may correspond to Figure 1 the stream generator 140, stream selector 142, device 102, device 104, system 100, Figure 2A the renderer 212, renderer 214, Figure 6 device 604, voice and music codec 1608, codec 1634, processor 1606, one or more additional processors 1610, voice and music codec 1608, one or more other circuits or components configured to generate the output stream, or any combination thereof.
[0179] The apparatus further includes components for providing the output stream to the audio output device. For example, the components for providing the output stream may correspond to Figure 1 the stream generator 140, stream selector 142, device 102, device 104, system 100, Figure 2A the renderer 212, renderer 214, Figure 6Device 604, antenna 1652, transceiver 1650, modem 1640, voice and music codec 1608, codec 1634, processor 1606, one or more additional processors 1610, one or more other circuits or components configured to provide an output stream, or any combination thereof.
[0180] Also in connection with the described embodiments, the device includes means for receiving from a host device first directional audio data representative of audio from one or more sound sources. The first directional audio data corresponds to a first arrangement of the one or more sound sources relative to the audio output device. For example, the means for receiving may correspond to Figure 1 stream selector 142, device 104, system 100, Figure 4 audio decoder 406A, audio decoder 406B, sound stream generator 408A, sound stream generator 408B, antenna 1652, transceiver 1650, modem 1640, voice and music codec 1608, codec 1634, processor 1606, one or more additional processors 1610, one or more other circuits or components configured to receive directional audio data from a host device, or any combination thereof.
[0181] The device also includes means for receiving from the host device second directional audio data representative of audio from one or more sound sources. The second directional audio data corresponds to a second arrangement of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. For example, the means for receiving may correspond to Figure 1 stream selector 142, device 104, system 100, Figure 4 audio decoder 406A, audio decoder 406B, sound stream generator 408A, sound stream generator 408B, antenna 1652, transceiver 1650, modem 1640, voice and music codec 1608, codec 1634, processor 1606, one or more additional processors 1610, one or more other circuits or components configured to receive directional audio data from a host device, or any combination thereof.
[0182] The apparatus also includes means for receiving position data indicative of the location of the audio output device. For example, the means for receiving may correspond to Figure 1 stream selector 142, device 104, system 100, Figure 4audio decoder 406A, combination factor generator 404, antenna 1652, transceiver 1650, modem 1640, voice and music codec 1608, codec 1634, processor 1606, one or more additional processors 1610, one or more other circuits or components configured to receive location data, or any combination thereof.
[0183] The apparatus further includes means for generating an output stream based on the first directional audio data, the second directional audio data, and the location data. For example, the means for generating the output stream may correspond to Figure 1 stream selector 142, device 104, system 100, Figure 2A renderer 212, renderer 214, voice and music codec 1608, codec 1634, processor 1606, one or more additional processors 1610, voice and music codec 1608, one or more other circuits or components configured to generate the output stream, or any combination thereof.
[0184] The apparatus further includes means for providing the output stream to an audio output device. For example, the means for providing the output stream may correspond to Figure 1 stream selector 142, device 104, system 100, Figure 2A renderer 212, renderer 214, antenna 1652, transceiver 1650, modem 1640, voice and music codec 1608, codec 1634, processor 1606, one or more additional processors 1610, one or more other circuits or components configured to provide the output stream, or any combination thereof.
[0185] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 1686) includes instructions (e.g., instructions 1656) that, when executed by one or more processors (e.g., one or more processors 1610, processor 1606, or one or more processors 890), cause the one or more processors to obtain spatial audio data (e.g., spatial audio data 170) representative of audio from one or more sound sources (e.g., one or more sound sources 184). The instructions, when executed by the one or more processors, also cause the one or more processors to generate first directional audio data (e.g., directional audio data 152) based on the spatial audio data. The first directional audio data corresponds to a first arrangement (e.g., arrangement 162) of the one or more sound sources relative to an audio output device (e.g., device 104, speaker 120, or both). The instructions, when executed by the one or more processors, also cause the one or more processors to generate second directional audio data (e.g., directional audio data 154) based on the spatial audio data. The second directional audio data corresponds to a second arrangement (e.g., arrangement 164) of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The instructions, when executed by the one or more processors, also cause the one or more processors to generate an output stream (e.g., output stream 150, output stream 450, output stream 550, output stream 650, or a combination thereof) based on the first directional audio data and the second directional audio data. The instructions, when executed by the one or more processors, also cause the one or more processors to provide the output stream to the audio output device.
[0186] In some embodiments, a non - transitory computer - readable medium (e.g., a computer - readable storage device such as memory 1686) includes instructions (e.g., instructions 1656) that, when executed by one or more processors (e.g., one or more processors 1610, processor 1606, or one or more processors 890), cause the one or more processors to receive, from a host device (e.g., device 104), first directional audio data (e.g., directional audio data 152) representing audio from one or more sound sources (e.g., one or more sound sources 184). The first directional audio data corresponds to a first arrangement (e.g., arrangement 162) of the one or more sound sources relative to an audio output device (e.g., device 104, speaker 120, or both). The instructions, when executed by the one or more processors, also cause the one or more processors to receive, from the host device, second directional audio data (e.g., directional audio data 154) representing audio from the one or more sound sources. The second directional audio data corresponds to a second arrangement (e.g., arrangement 164) of the one or more sound sources relative to the audio output device. The second arrangement is different from the first arrangement. The instructions, when executed by the one or more processors, also cause the one or more processors to receive position data (e.g., user position data 185) indicating the location of the audio output device. The instructions, when executed by the one or more processors, also cause the one or more processors to generate an output stream (e.g., output stream 450, output stream 650, or both) based on the first directional audio data, the second directional audio data, and the position data. The instructions, when executed by the one or more processors, also cause the one or more processors to provide the output stream to the audio output device.
[0187] The following describes specific aspects of the present disclosure in a set of related clauses:
[0188] According to Clause 1, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to: obtain spatial audio data representing audio from one or more sound sources; generate first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; generate second directional audio data based on the spatial audio data, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; and generate an output stream based on the first directional audio data and the second directional audio data.
[0189] Clause 2 includes the apparatus of Clause 1, wherein the first arrangement is based on default location data that indicates a default location of the audio output device, a default head position, a default location of the host device, a default relative position of the audio output device and the host device, or a combination thereof.
[0190] Clause 3 includes the apparatus of Clause 1 or Clause 2, wherein the first arrangement is based on detected location data that indicates a detected location of the audio output device, a detected movement of the audio output device, a detected head position, a detected head movement, a detected location of the host device, a detected movement of the host device, a detected relative position of the audio output device and the host device, a detected relative movement of the audio output device and the host device, or a combination thereof.
[0191] Clause 4 includes the apparatus of any one of Clauses 1 to 3, wherein the first arrangement is based on user interaction data.
[0192] Clause 5 includes the apparatus of any one of Clauses 1 to 4, wherein the second arrangement is based on predetermined location data that indicates a predetermined location of the audio output device, a predetermined head position, a predetermined location of the host device, a predetermined relative position of the audio output device and the host device, or a combination thereof.
[0193] Clause 6 includes the apparatus of any one of Clauses 1 to 5, wherein the second arrangement is based on predicted location data that indicates a predicted location of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted location of the host device, a predicted movement of the host device, a predicted relative position of the audio output device and the host device, a predicted relative movement of the audio output device and the host device, or a combination thereof.
[0194] Clause 7 includes the apparatus of any one of Clauses 1 to 6, wherein the second arrangement is based on predicted user interaction data.
[0195] Clause 8 includes the apparatus of any one of Clauses 1 to 7, wherein the processor is configured to execute instructions to: receive first location data indicating a first location of the audio output device; select, at least in part based on the first location data, one of the first directional audio data or the second directional audio data as the output stream; and initiate transmission of the output stream to the audio output device.
[0196] Clause 9 includes the apparatus of any one of Clauses 1 to 8, wherein the processor is configured to execute instructions to: receive first position data indicating a first position of an audio output device; at least partially based on the first position data, combine the first directional audio data and the second directional audio data to generate the output stream; and initiate transmission of the output stream to the audio output device.
[0197] Clause 10 includes the apparatus of any one of Clauses 1 to 9, wherein the processor is configured to execute instructions to: receive first position data indicating a first position of an audio output device; determine a combining factor at least partially based on the first position data; combine the first directional audio data and the second directional audio data based on the combining factor to generate the output stream; and initiate transmission of the output stream to the audio output device.
[0198] Clause 11 includes the apparatus of any one of Clauses 1 to 7, wherein the processor is configured to execute instructions to initiate transmission of the first directional audio data and the second directional audio data as an output stream to the audio output device.
[0199] Clause 12 includes the apparatus of any one of Clauses 1 to 7 or Clause 11, wherein the processor is configured to execute instructions to: generate second directional audio data based on one or more parameters; and initiate transmission of the one or more parameters to the audio output device concurrently with transmission of the output stream to the audio output device.
[0200] Clause 13 includes the apparatus of Clause 12, wherein the one or more parameters are based on predetermined position data, predicted position data, predicted user interaction data, or a combination thereof.
[0201] Clause 14 includes the apparatus of any one of Clauses 1 to 13, wherein the audio output device includes a speaker, and wherein the processor is configured to execute instructions to: render an acoustic output based on the output stream; and provide the acoustic output to the speaker.
[0202] Clause 15 includes the apparatus of any one of Clauses 1 to 14, wherein the audio output device includes a headset, an extended reality (XR) headset, a gaming device, headphones, a speaker, or a combination thereof.
[0203] Clause 16 includes the apparatus of any one of Clauses 1 to 15, wherein the processor is integrated in the audio output device.
[0204] Clause 17 includes the apparatus of any one of Clauses 1 to 16, wherein the processor is integrated in a mobile device, a gaming console, a communication device, a computer, a display device, a vehicle, a camera, or a combination thereof.
[0205] Clause 18 includes the device of any one of Clauses 1 to 17, and further includes a modem configured to receive audio data from an audio data source, and the spatial audio data is based on the audio data.
[0206] Clause 19 includes the device of any one of Clauses 1 to 18, wherein the processor is further configured to execute instructions to generate one or more additional directional audio data sets based on the spatial audio data, and the output stream is based on the one or more additional directional audio data sets.
[0207] According to Clause 20, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to: receive first directional audio data representing audio from one or more sound sources from a host device, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to the audio output device; receive second directional audio data representing audio from the one or more sound sources from the host device, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; receive position data indicating the position of the audio output device; generate an output stream based on the first directional audio data, the second directional audio data, and the position data; and provide the output stream to the audio output device.
[0208] Clause 21 includes the device of Clause 20, wherein the processor is configured to execute instructions to select, at least in part based on the position data, one of the first audio data corresponding to the first directional audio data or the second audio data corresponding to the second directional audio data as the output stream.
[0209] Clause 22 includes the device as described in Clause 20 or Clause 21, wherein the first directional audio data is based on a first position of the audio output device, the second directional audio data is based on a second position of the audio output device, and the processor is configured to execute the instructions to select, based on a comparison of the position with the first position and the second position, one of the first audio data or the second audio data as the output stream.
[0210] Clause 23 includes the device of any one of Clauses 20 to 22, wherein the processor is configured to execute instructions to combine the first audio data corresponding to the first directional audio data and the second audio data corresponding to the second directional audio data, at least in part based on the position data, to generate the output stream.
[0211] Clause 24 includes a device according to any one of Clauses 20 to 23, wherein the processor is configured to execute instructions to: determine a combination factor based at least in part on location data; and combine a first audio data corresponding to first directional audio data and a second audio data corresponding to second directional audio data based on the combination factor to generate an output stream.
[0212] Clause 25 includes the device of Clause 24, wherein the first directional audio data is based on a first location of the audio output device, the second directional audio data is based on a second location of the audio output device, and the combination factor is based on a comparison of the location with the first location and the second location.
[0213] Clause 26 includes the device of any one of Clauses 20 to 25, wherein the processor is configured to execute instructions to provide to a host device first location data indicating a first location of the audio output device detected at a first time, wherein the first directional audio data is based on the first location data.
[0214] Clause 27 includes the device of any one of Clauses 20 to 26, wherein the processor is configured to execute instructions to receive from the host device one or more parameters indicating that the first directional audio data is based on a first location of the audio output device, the second directional audio data is based on a second location of the audio output device, or both.
[0215] Clause 28 includes the device of Clause 27, wherein the first location is based on a default location of the audio output device, a detected location of the audio output device, a detected movement of the audio output device, or a combination thereof.
[0216] Clause 29 includes the device of Clause 27 or Clause 28, wherein the second location is based on a predetermined location of the audio output device, a predicted location of the audio output device, a predicted movement of the audio output device, or a combination thereof.
[0217] Clause 30 includes the device of any one of Clauses 20 to 29, wherein the processor is configured to execute instructions to receive from the host device one or more additional directional audio data sets representing audio from one or more sound sources, and an output stream is generated based on the one or more additional directional audio data sets.
[0218] According to clause 31, a method includes: obtaining, at a device, spatial audio data representing audio from one or more sound sources; generating, at the device, first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; generating, at the device, second directional audio data based on the spatial audio data, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; generating, at the device, an output stream based on the first directional audio data and the second directional audio data; and providing the output stream from the device to the audio output device.
[0219] Clause 32 includes the method of clause 31, wherein the first arrangement is based on default location data that indicates a default location of the audio output device, a default head position, a default location of the host device, a default relative position of the audio output device and the host device, or a combination thereof.
[0220] Clause 33 includes the method of clause 31 or clause 32, wherein the first arrangement is based on detected location data that indicates a detected location of the audio output device, a detected movement of the audio output device, a detected head position, a detected head movement, a detected location of the host device, a detected movement of the host device, a detected relative position of the audio output device and the host device, a detected relative movement of the audio output device and the host device, or a combination thereof.
[0221] Clause 34 includes the method of any one of clauses 31 to 33, wherein the first arrangement is based on user interaction data.
[0222] Clause 35 includes the method of any one of clauses 31 to 34, wherein the second arrangement is based on predetermined location data that indicates a predetermined location of the audio output device, a predetermined head position, a predetermined location of the host device, a predetermined relative position of the audio output device and the host device, or a combination thereof.
[0223] Clause 36 includes the method of any one of clauses 31 to 35, wherein the second arrangement is based on predicted location data that indicates a predicted location of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted location of the host device, a predicted movement of the host device, a predicted relative position of the audio output device and the host device, a predicted relative movement of the audio output device and the host device, or a combination thereof.
[0224] Clause 37 includes the method of any one of clauses 31 to 36, wherein the second arrangement is based on predicted user interaction data.
[0225] Clause 38 includes the method of any one of Clauses 31 to 37, and further includes: receiving first position data indicating a first position of an audio output device; selecting, at least in part based on the first position data, one of the first directional audio data or the second directional audio data as the output stream; and initiating transmission of the output stream to the audio output device.
[0226] Clause 39 includes the method of any one of Clauses 31 to 38, and further includes: receiving first position data indicating a first position of an audio output device; combining the first directional audio data and the second directional audio data, at least in part based on the first position data, to generate the output stream; and initiating transmission of the output stream to the audio output device.
[0227] Clause 40 includes the method of any one of Clauses 31 to 39, and further includes: receiving first position data indicating a first position of an audio output device; determining a combination factor at least in part based on the first position data; combining the first directional audio data and the second directional audio data based on the combination factor to generate the output stream; and initiating transmission of the output stream to the audio output device.
[0228] Clause 41 includes the method of any one of Clauses 31 to 37, and further includes: initiating transmission of the first directional audio data and the second directional audio data as an output stream to an audio output device.
[0229] Clause 42 includes the method of any one of Clauses 31 to 37 or Clause 41, and further includes: generating second directional audio data based on one or more parameters; and initiating transmission of the one or more parameters to the audio output device concurrently with transmission of the output stream to the audio output device.
[0230] Clause 43 includes the method of Clause 42, wherein the one or more parameters are based on predetermined position data, predicted position data, predicted user interaction data, or a combination thereof.
[0231] Clause 44 includes the method of any one of Clauses 31 to 43, wherein the audio output device includes a speaker, and further includes: presenting an acoustic output based on the output stream; and providing the acoustic output to the speaker.
[0232] Clause 45 includes the method of any one of Clauses 31 to 44, wherein the audio output device includes a headset, an extended reality (XR) headset, a gaming device, headphones, a speaker, or a combination thereof.
[0233] Clause 46 includes the method of any one of Clauses 31 to 45, wherein the audio output device includes a speaker, a second device, or both.
[0234] Clause 47 includes the method of any one of Clauses 31 to 46, wherein the device includes a mobile device, a game console, a communication device, a computer, a display device, a vehicle, a camera, or a combination thereof.
[0235] Clause 48 includes the method of any one of Clauses 31 to 47, further including receiving audio data from an audio data source via a modem, and the spatial audio data is based on the audio data.
[0236] Clause 49 includes the method of any one of Clauses 31 to 48, further including generating one or more additional directional audio data sets based on the spatial audio data, wherein the output stream is based on the one or more additional directional audio data sets.
[0237] According to Clause 50, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method of any one of Clauses 31 to 49.
[0238] According to Clause 51, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform the method of any one of Clauses 31 to 49.
[0239] According to Clause 52, a device includes components for performing the method of any one of Clauses 31 to 49.
[0240] According to Clause 53, a method includes: receiving, at a device, first directional audio data from a host device representing audio from one or more sound sources, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; receiving, at the device, second directional audio data from the host device representing audio from the one or more sound sources, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; receiving, at the device, position data indicating the position of the audio output device; generating, at the device, an output stream based on the first directional audio data, the second directional audio data, and the position data; and providing the output stream from the device to the audio output device.
[0241] Clause 54 includes the method of Clause 53, further including selecting, at least in part based on the position data, one of first audio data corresponding to the first directional audio data or second audio data corresponding to the second directional audio data as the output stream.
[0242] Clause 55 includes the method of Clause 53 or Clause 54, wherein the first directional audio data is based on a first position of the audio output device, wherein the second directional audio data is based on a second position of the audio output device, and further includes selecting one of the first audio data or the second audio data as an output stream based on a comparison of the position with the first position and the second position.
[0243] Clause 56 includes the method of any one of Clauses 53 to 55, and further includes combining the first audio data corresponding to the first directional audio data and the second audio data corresponding to the second directional audio data at least partially based on position data to generate an output stream.
[0244] Clause 57 includes the method of any one of Clauses 53 to 56, and further includes: determining a combination factor at least partially based on position data; and combining the first audio data corresponding to the first directional audio data and the second audio data corresponding to the second directional audio data based on the combination factor to generate an output stream.
[0245] Clause 58 includes the method of Clause 57, wherein the first directional audio data is based on a first position of the audio output device, wherein the second directional audio data is based on a second position of the audio output device, and wherein the combination factor is based on a comparison of the position with the first position and the second position.
[0246] Clause 59 includes the method of any one of Clauses 53 to 58, and further includes providing first position data indicating a first position of the audio output device detected at a first time to a host device, wherein the first directional audio data is based on the first position data.
[0247] Clause 60 includes the method of any one of Clauses 53 to 59, and further includes receiving from a host device one or more parameters indicating that the first directional audio data is based on a first position of the audio output device, the second directional audio data is based on a second position of the audio output device, or both.
[0248] Clause 61 includes the method of Clause 60, wherein the first position is based on a default position of the audio output device, a detected position of the audio output device, a detected movement of the audio output device, or a combination thereof.
[0249] Clause 62 includes the method of Clause 60 or Clause 61, wherein the second position is based on a predetermined position of the audio output device, a predicted position of the audio output device, a predicted movement of the audio output device, or a combination thereof.
[0250] Clause 63 includes the method of any one of Clauses 53 to 62, and further includes receiving, from a host device, one or more additional directional audio data sets representing audio from one or more sound sources, wherein an output stream is generated based on the one or more additional directional audio data sets.
[0251] According to Clause 64, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method of any one of Clauses 53 to 63.
[0252] According to Clause 65, a non - transitory computer - readable medium stores instructions that, when executed by a processor, cause the processor to perform the method of any one of Clauses 53 to 63.
[0253] According to Clause 66, an apparatus includes components for performing the method of any one of Clauses 53 to 63.
[0254] According to Clause 67, a non - transitory computer - readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to: obtain spatial audio data representing audio from one or more sound sources; generate first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; generate second directional audio data based on the spatial audio data, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; generate an output stream based on the first directional audio data and the second directional audio data; and provide the output stream to the audio output device.
[0255] According to Clause 68, a non - transitory computer - readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to: receive, from a host device, first directional audio data representing audio from one or more sound sources, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; receive, from the host device, second directional audio data representing audio from the one or more sound sources, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; receive position data indicating the position of the audio output device; generate an output stream based on the first directional audio data, the second directional audio data, and the position data; and provide the output stream to the audio output device.
[0256] According to Clause 69, a device includes: components for obtaining spatial audio data representing audio from one or more sound sources; components for generating first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; components for generating second directional audio data based on the spatial audio data; the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; components for generating an output stream based on the first directional audio data and the second directional audio data; and components for providing the output stream to the audio output device.
[0257] According to Clause 70, a device includes: components for receiving first directional audio data representing audio from one or more sound sources from a host device, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; components for receiving second directional audio data representing audio from the one or more sound sources from the host device, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement; components for receiving position data indicating the position of the audio output device; components for generating an output stream based on the first directional audio data, the second directional audio data, and the position data; and components for providing the output stream to the audio output device.
[0258] Those skilled in the art will further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or combinations of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been described generally above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application, and such implementation decisions should not be construed as causing a departure from the scope of the present invention.
[0259] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. A software module may reside in random access memory (RAM), flash memory, read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, a hard disk, a removable disk, a compact disc read only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integrated into the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.
[0260] The previous description of the disclosed aspects is provided to enable a person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the invention is not intended to be limited to the aspects shown herein but is to be accorded the widest possible scope consistent with the principles and novel features as defined by the appended claims.
Claims
1. A device for generating directional audio, comprising: a processor configured to: obtain spatial audio data representing audio from one or more sound sources; generate first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; generate second directional audio data based on the spatial audio data, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data that indicates a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of a host device, a predicted movement of the host device, a predicted relative position between the audio output device and the host device, a predicted relative movement between the audio output device and the host device, or a combination thereof; and generate an output stream based on the first directional audio data and the second directional audio data.
2. The device according to claim 1, wherein the first arrangement is based on default position data that indicates a default position of the audio output device, a default head position, a default position of the host device, a default relative position between the audio output device and the host device, or a combination thereof.
3. The device according to claim 1, wherein the first arrangement is based on detected position data that indicates a detected position of the audio output device, a detected movement of the audio output device, a detected head position, a detected head movement, a detected position of the host device, a detected movement of the host device, a detected relative position between the audio output device and the host device, a detected relative movement between the audio output device and the host device, or a combination thereof.
4. The device according to claim 1, wherein the first arrangement is based on user interaction data.
5. The device according to claim 1, wherein the second arrangement is based on predetermined position data that indicates a predetermined position of the audio output device, a predetermined head position, a predetermined position of the host device, a predetermined relative position between the audio output device and the host device, or a combination thereof.
6. The device according to claim 1, wherein the second arrangement is based on predicted user interaction data.
7. The device according to claim 1, wherein the processor is configured to: receive first position data indicating a first position of the audio output device; select, at least in part based on the first position data, one of the first directional audio data or the second directional audio data as the output stream; and initiate transmission of the output stream to the audio output device.
8. The device according to claim 1, wherein the processor is configured to: receive first position data indicating a first position of the audio output device; Combining the first directional audio data and the second directional audio data to generate the output stream, at least in part based on the first position data; and initiating transmission of the output stream to the audio output device.
9. The apparatus according to claim 1, wherein the processor is configured to: Receive first position data indicating a first position of the audio output device; Determine a combination factor at least in part based on the first position data; Combine the first directional audio data and the second directional audio data based on the combination factor to generate the output stream; and Initiate transmission of the output stream to the audio output device.
10. The apparatus according to claim 1, wherein the processor is configured to initiate transmission of the first directional audio data and the second directional audio data as the output stream to the audio output device.
11. The apparatus according to claim 1, wherein the processor is configured to: Generate the second directional audio data based on one or more parameters; and Initiate transmission of the one or more parameters to the audio output device concurrently with transmission of the output stream to the audio output device.
12. The apparatus according to claim 11, wherein the one or more parameters are based on predetermined position data, predicted position data, predicted user interaction data, or a combination thereof.
13. The apparatus according to claim 1, wherein, the audio output device includes a speaker, and wherein the processor is configured to: Render an acoustic output based on the output stream; and Provide the acoustic output to the speaker.
14. The apparatus according to claim 1, wherein the audio output device comprises a headset, an extended reality (XR) headset, a gaming device, headphones, a speaker, or a combination thereof.
15. The apparatus according to claim 1, wherein, the processor is integrated in the audio output device.
16. The apparatus according to claim 1, wherein the processor is integrated in a mobile device, a gaming console, a communication device, a computer, a display device, a vehicle, a camera, or a combination thereof.
17. The apparatus according to claim 1, further comprising a modem configured to receive audio data from an audio data source, the spatial audio data being based on the audio data.
18. The apparatus according to claim 1, wherein the processor is configured to: Generate a first copy and a second copy of the first directional audio data, the first copy and the second copy of the first directional audio data corresponding to different bitrates; and Generate a first copy and a second copy of the second directional audio data, the first copy and the second copy of the second directional audio data corresponding to different bitrates.
19. An apparatus for generating directional audio, comprising: A processor configured to: Receive first directional audio data from a host device representing audio from one or more sound sources, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; Receive second directional audio data representing audio from the one or more sound sources from the host device, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data indicating a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of the host device, a predicted movement of the host device, a predicted relative position between the audio output device and the host device, a predicted relative movement between the audio output device and the host device, or a combination thereof; Receive position data indicating the position of the audio output device; Generate an output stream based on the first directional audio data, the second directional audio data, and the position data; And Provide the output stream to the audio output device.
20. The apparatus according to claim 19, wherein the processor is configured to select, at least in part based on the position data, one of first audio data corresponding to the first directional audio data or second audio data corresponding to the second directional audio data as the output stream.
21. The apparatus according to claim 20, wherein the first directional audio data is based on a first position of the audio output device, wherein the second directional audio data is based on a second position of the audio output device, and wherein the processor is configured to select one of the first audio data or the second audio data as the output stream based on a comparison of the position with the first position and the second position.
22. The apparatus according to claim 19, wherein the processor is configured to combine, at least in part based on the position data, first audio data corresponding to the first directional audio data and second audio data corresponding to the second directional audio data to generate the output stream.
23. The apparatus according to claim 19, wherein the processor is configured to: Determine a combination factor at least in part based on the position data; and Combine first audio data corresponding to the first directional audio data and second audio data corresponding to the second directional audio data based on the combination factor to generate the output stream.
24. The apparatus according to claim 23, wherein the first directional audio data is based on a first position of the audio output device, wherein the second directional audio data is based on a second position of the audio output device, and wherein the combination factor is based on a comparison of the position with the first position and the second position.
25. The apparatus according to claim 19, wherein the processor is configured to provide the host device with first position data indicating a first position of the audio output device detected at a first time, wherein the first directional audio data is based on the first position data.
26. The apparatus according to claim 19, wherein the processor is configured to receive from the host device one or more parameters indicating that the first directional audio data is based on a first position of the audio output device, the second directional audio data is based on a second position of the audio output device, or both, wherein the first position is based on a default position of the audio output device, a detected position of the audio output device, a detected movement of the audio output device, or a combination thereof, and wherein the second position is based on a predetermined position of the audio output device, a predicted position of the audio output device, a predicted movement of the audio output device, or a combination thereof.
27. The apparatus according to claim 19, wherein the processor is configured to receive from the host device one or more additional directional audio data sets representing the audio from the one or more sound sources, wherein the output stream is generated based on the one or more additional directional audio data sets.
28. A method for generating directional audio, comprising: obtaining, at a device, spatial audio data representing audio from one or more sound sources; generating, at the device, first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; generating, at the device, second directional audio data based on the spatial audio data, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data indicating a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of the host device, a predicted movement of the host device, a predicted relative position of the audio output device and the host device, a predicted relative movement of the audio output device and the host device, or a combination thereof; generating, at the device, an output stream based on the first directional audio data and the second directional audio data; and providing the output stream from the device to the audio output device.
29. A method for generating directional audio, comprising: receiving, at a device, from a host device first directional audio data representing audio from one or more sound sources, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; Receive, at the device, second directional audio data representing audio from the one or more sound sources from the host device, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data indicating a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of the host device, a predicted movement of the host device, a predicted relative position between the audio output device and the host device, a predicted relative movement between the audio output device and the host device, or a combination thereof; Receive, at the device, position data indicating the position of the audio output device; Generate, at the device, an output stream based on the first directional audio data, the second directional audio data, and the position data; And Provide the output stream from the device to the audio output device.
30. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: Obtain spatial audio data representing audio from one or more sound sources; Generate first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; Generate second directional audio data based on the spatial audio data, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data indicating a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of a host device, a predicted movement of the host device, a predicted relative position between the audio output device and the host device, a predicted relative movement between the audio output device and the host device, or a combination thereof; And Generate an output stream based on the first directional audio data and the second directional audio data.
31. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: Receive, from a host device, first directional audio data representing audio from one or more sound sources, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; Receiving second directional audio data representing audio from the one or more sound sources from the host device, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data indicating a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of the host device, a predicted movement of the host device, a predicted relative position of the audio output device and the host device, a predicted relative movement of the audio output device and the host device, or a combination thereof; Receiving position data indicating the position of the audio output device; Generating an output stream based on the first directional audio data, the second directional audio data, and the position data; And Providing the output stream to the audio output device.
32. An apparatus for generating directional audio, Comprising: Means for obtaining spatial audio data representing audio from one or more sound sources; Means for generating first directional audio data based on the spatial audio data, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; Means for generating second directional audio data based on the spatial audio data, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data indicating a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of the host device, a predicted movement of the host device, a predicted relative position of the audio output device and the host device, a predicted relative movement of the audio output device and the host device, or a combination thereof; And Means for generating an output stream based on the first directional audio data and the second directional audio data.
33. An apparatus for generating directional audio, Comprising: Means for receiving first directional audio data representing audio from one or more sound sources from a host device, the first directional audio data corresponding to a first arrangement of the one or more sound sources relative to an audio output device; A component for receiving second directional audio data representing audio from the one or more sound sources from the host device, the second directional audio data corresponding to a second arrangement of the one or more sound sources relative to the audio output device, wherein the second arrangement is different from the first arrangement and is based on predicted position data indicating a predicted position of the audio output device, a predicted movement of the audio output device, a predicted head position, a predicted head movement, a predicted position of the host device, a predicted movement of the host device, a predicted relative position of the audio output device and the host device, a predicted relative movement of the audio output device and the host device, or a combination thereof; A component for receiving position data indicating the position of the audio output device; A component for generating an output stream based on the first directional audio data, the second directional audio data, and the position data; And A component for providing the output stream to the audio output device.
Citation Information
Patent Citations
Concept for generating an enhanced sound field description or a modified sound field description using a multi-point sound field description
CN111149155A