Audio processing

By detecting and mitigating listener pose anomalies, and utilizing pose sensors and local pre-rendering technology, the problem of resource waste and latency caused by inaccurate listener poses is solved, thereby improving the accuracy and efficiency of immersive audio.

CN121533039APending Publication Date: 2026-02-13QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480047073.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2024-07-25
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Conventional immersive audio systems suffer from resource waste and delays in the immersive audio experience due to inaccurate listener pose data, especially the latency and resource waste issues when head tracking information is updated remotely.

Method used

By detecting and mitigating abnormal values ​​in listener pose, accurate listener pose data is generated using pose sensors, and pose is adjusted based on pose constraints to reduce unnecessary audio rendering. Local pre-rendering and pose prediction techniques are used to optimize the audio rendering process.

Benefits of technology

It improves the accuracy and efficiency of immersive audio environments, reduces resource waste, enhances user experience, and lowers audio rendering latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121533039A_ABST
    Figure CN121533039A_ABST
Patent Text Reader

Abstract

A device includes a memory configured to store data associated with an immersive audio environment and one or more processors configured to obtain pose data for a listener in the immersive audio environment. The processor is configured to determine a current listener pose based on the pose data and the one or more pose constraints. The processor is configured to obtain pose data based on the pose update parameter. The processor is configured to obtain rendered assets associated with the immersive audio environment based on the current listener pose. The processor is configured to generate an output audio signal based on the rendered asset.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims priority to commonly owned U.S. Provisional Patent Application No. 63 / 515,648, filed July 26, 2023, and U.S. Non-Provisional Patent Application No. 18 / 783,186, filed July 24, 2024, the contents of each of which are expressly incorporated herein by reference in their entirety. TECHNICAL FIELD

[0002] The present disclosure relates generally to audio processing, and more particularly to processing immersive audio. BACKGROUND

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there currently exist a variety of portable personal computing devices, including wireless telephones such as mobile and smart phones, tablets, and laptop computers, that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Further, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Still further, such devices can process executable instructions, including software applications, such as web browser applications that can be used to access the Internet. As such, these devices can include significant computing, data storage, and communication capabilities.

[0004] One application of such devices includes providing immersive audio to a user. As an example, a headphone device worn by a user can receive streaming audio data from a remote server for playback to the user. Conventional multi-source spatial audio systems are typically designed to use a relatively high complexity rendering of audio streams from multiple audio sources, with the goal of ensuring that the worst case performance of the headphone device still provides an acceptable quality of immersive audio to the user. However, real-time local rendering of immersive audio is resource intensive (e.g., in terms of processor cycles, time, power, and memory utilization).

[0005] Another conventional approach is to offload local rendering of immersive audio to a streaming device. For example, a headphone device can detect rotation of a user’s head and send head tracking information to a remote server. The remote server updates an audio scene based on the head tracking information, generates binaural audio data based on the updated audio scene, and sends the binaural audio data to the headphone device for playback to the user.

[0006] Performing audio scene updates and binauralization at a remote server allows users to experience immersive audio through headphones with relatively limited processing resources. However, such systems can result in unnaturally high motion-to-sound latency due to the latency associated with sending head tracking information to the remote server, updating audio data based on head rotation, and sending the updated binaural audio data to the headphones. In other words, the time delay between the user's head rotation and the corresponding modified spatial audio played at the user's ears can be unnaturally long, potentially diminishing the user's immersive audio experience.

[0007] Typically, an immersive audio environment is generated based on streaming audio data corresponding to one or more audio sources in the audio environment, rendered according to the listener's pose. The listener's pose is based on pose data generated by one or more sensors on the listener's playback device. Inaccuracies in the pose data lead to inaccurate listener poses. Audio playback systems that use inaccurate listener poses to initiate updates to the immersive audio environment can result in wasted resources, such as requesting, sending, and initiating the rendering of unwanted audio streams based on inaccurate estimates of the listener's position. Summary of the Invention

[0008] According to one or more aspects of this disclosure, an apparatus includes a memory configured to store data associated with an immersive audio environment; and one or more processors configured to acquire pose data of a listener in the immersive audio environment. The one or more processors are configured to determine a current listener pose based on the pose data and one or more pose constraints. The one or more processors are configured to acquire the pose data based on pose update parameters. The one or more processors are configured to acquire rendered assets associated with the immersive audio environment based on the current listener pose. The one or more processors are configured to generate an output audio signal based on the rendered assets.

[0009] According to one or more aspects of this disclosure, a method includes obtaining pose data of a listener in an immersive audio environment at one or more processors. The method includes determining a current listener pose at the one or more processors based on the pose data and one or more pose constraints. The method includes obtaining a rendered asset associated with the immersive audio environment at the one or more processors and based on the current listener pose. The method includes generating an output audio signal at the one or more processors based on the rendered asset.

[0010] According to one or more aspects of this disclosure, a non-transitory computer-readable device stores instructions executable by one or more processors to cause the one or more processors to obtain pose data of a listener in an immersive audio environment. The instructions cause the one or more processors to determine the current listener pose based on the pose data and one or more pose constraints. The instructions cause the one or more processors to obtain rendering assets associated with the immersive audio environment based on the current listener pose. The instructions cause the one or more processors to generate an output audio signal based on the rendering assets.

[0011] According to one or more aspects of this disclosure, an apparatus includes components for acquiring pose data of a listener in an immersive audio environment. The apparatus includes components for determining a current listener pose based on the pose data and one or more pose constraints. The apparatus includes components for acquiring a rendered asset associated with the immersive audio environment based on the current listener pose. The apparatus includes components for generating an output audio signal based on the rendered asset.

[0012] Other aspects, advantages, and features of this disclosure will become apparent upon reading the entire application, which comprises the following sections: description of the drawings, detailed description, and claims. Attached Figure Description

[0013] FIG. 1 This is a block diagram of various aspects of a system, based on some examples of this disclosure, capable of operating to process data associated with an immersive audio environment.

[0014] FIG. 2 Based on some examples of this disclosure, it is possible to... FIG. 1 A diagram illustrating the operations performed by the system.

[0015] FIG. 3A Based on some examples of this disclosure, it is possible to... FIG. 1 A diagram illustrating the operations performed by the system.

[0016] FIG. 3B Based on some examples of this disclosure, it is possible to... FIG. 1 A diagram illustrating the operations performed by the system.

[0017] FIG. 4 Based on some examples of this disclosure FIG. 1 A block diagram of various aspects of the system.

[0018] FIG. 5 Based on some examples of this disclosure FIG. 1 A block diagram of various aspects of the system.

[0019] FIG. 6 Based on some examples of this disclosure FIG. 1A block diagram of various aspects of the system.

[0020] FIG. 7 Based on some examples of this disclosure FIG. 1 A block diagram of various aspects of the system.

[0021] FIG. 8 Based on some examples of this disclosure FIG. 1 The diagram illustrates the illustrative aspects of the operation of the system's components.

[0022] FIG. 9 Examples of integrated circuits capable of operating to process data associated with an immersive audio environment are illustrated according to some examples of this disclosure.

[0023] FIG. 10 This is a block diagram illustrating an exemplary implementation of a system for processing data associated with an immersive audio environment and including external speakers.

[0024] FIG. 11 These are illustrations of mobile devices capable of operating to process data associated with an immersive audio environment, based on some examples of this disclosure.

[0025] FIG. 12 This is an illustration of a headset capable of operating to process data associated with an immersive audio environment, based on some examples of this disclosure.

[0026] FIG. 13 This is an illustration of an in-ear headphone, according to some examples of this disclosure, capable of operating to process data associated with an immersive audio environment.

[0027] FIG. 14 These are illustrations of mixed reality or augmented reality glasses devices, based on some examples of this disclosure, capable of operating to process data associated with an immersive audio environment.

[0028] FIG. 15 This is an illustration of a headset (such as a virtual reality, mixed reality, or augmented reality headset) that is capable of operating to process data associated with an immersive audio environment, according to some examples of this disclosure.

[0029] FIG. 16 This is a diagram illustrating a first example of a vehicle capable of operating to process data associated with an immersive audio environment, based on some examples of this disclosure.

[0030] FIG. 17 This is a diagram illustrating a second example of a vehicle capable of operating to process data associated with an immersive audio environment, based on some examples of this disclosure.

[0031] FIG. 18 Based on some examples of this disclosure, it is possible to...FIG. 1 A diagram illustrating a specific implementation of a method for processing data associated with an immersive audio environment performed by a device.

[0032] FIG. 19 This is a block diagram of a particular exemplary example of a device capable of operating to process data associated with an immersive audio environment, based on some examples of this disclosure. Detailed Implementation

[0033] Systems and methods for providing an immersive audio environment based on the listener's pose are described. Typically, a conventional immersive audio environment is generated based on rendering streaming audio data corresponding to one or more audio sources in the audio environment according to the listener's pose, and the listener's pose is based on pose data generated by one or more sensors of the listener's playback device. Inaccuracies in the pose data lead to inaccurate listener poses. Audio playback systems that use inaccurate listener poses to initiate updates to the immersive audio environment can lead to wasted resources in the audio playback system, such as requesting, sending, and / or initiating the rendering of unwanted audio streams based on inaccurate estimates of the listener's position.

[0034] The described systems and methods improve the accuracy and efficiency of immersive audio environments by identifying and mitigating outliers in listener pose based on one or more pose constraints. For example, one or more pose sensors can generate listener pose data indicating the listener's pose and determine whether the pose violates one or more constraints based on human movement, or one or more spatial constraints, or a combination thereof. According to one aspect, constraints based on human movement include one or more body pose constraints, such as constraints on the pose of the listener's head relative to the listener's hands and / or torso, velocity constraints, acceleration constraints, or combinations thereof. Spatial constraints may include one or more spatial boundaries associated with the immersive audio environment, such as positional constraints corresponding to 6-DOF rendering operations.

[0035] When an outlier listener pose that violates one or more of the human movement constraints or spatial constraints is detected, the disclosed techniques include determining a listener pose value that does not violate any constraints, rather than using the outlier. For example, the listener's pose can be set to the listener's most recent (non-outlier) previous pose. As another example, the listener's pose can be determined by adjusting the outlier pose to not violate any constraints. Detection and mitigation of such pose outliers reduce the inefficiencies experienced by conventional systems due to handling inaccurate listener poses, including wasted resources due to requests, transmissions, and initiation of rendering of unwanted audio streams caused by inaccurate estimates of the listener's position. Furthermore, detection and mitigation of such pose outliers improve the listener experience by preventing audio rendering based on estimates of listener movement that may be erroneous and / or exceed the spatial boundaries associated with the immersive audio environment.

[0036] According to some aspects, in addition to detecting and mitigating outliers in the listener's current pose, the disclosed techniques also include detecting and mitigating outliers in the predicted listener pose. For example, the predicted listener pose can be determined based on the listener's current pose and used to pre-fetch assets, such as audio data associated with one or more audio sources, based on the listener's predicted future position in the immersive audio environment. Detecting and mitigating outliers in the predicted listener pose improves the efficiency of the audio rendering system by reducing erroneous predictions associated with pre-fetching assets used to render the immersive audio environment, such as by reducing the consumption of processing resources associated with acquiring and processing assets based on incorrect predictions, and reducing transmission bandwidth usage associated with pre-fetching assets based on incorrect predictions.

[0037] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In this description, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to limit the scope of the embodiments. For example, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some embodiments and plural in others. For example, FIG. 4 It describes a system that includes one or more processors ( FIG. 4 The term "processor 410" in the system 400 indicates that in some embodiments the system 400 includes a single processor 410, while in other embodiments the system 400 includes multiple processors 410. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular or optional plural form (as indicated by "(multiple)"), unless the aspect described relates to multiples of features.

[0038] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference numerals are used for each feature, and these different instances are distinguished by adding letters to the reference numerals. Reference numerals are used without distinguishing letters when a feature is referenced herein as a group or a type of feature (e.g., when a specific feature among these features is not referenced). However, reference numerals are used with distinguishing letters when a specific feature among multiple features of the same type is mentioned herein. For example, see reference... FIG. 7 Multiple pose sensors are illustrated and associated with reference numerals 108A and 108B. When referring to a specific pose sensor among these pose sensors (such as pose sensor 108A), the distinguishing letter "A" is used. However, when referring to any one of these pose sensors or referring to these pose sensors as a group, reference numeral 108 is used without the distinguishing letter.

[0039] As used herein, the term "comprising" is used interchangeably with "including". Additionally, the term "in which" is used interchangeably with "wherein". As used herein, "exemplary" indicates an example, specific implementation, and / or aspect, and should not be construed as restrictive or indicating a preference or preferred implementation. As used herein, ordinal terms used to modify elements (such as structures, components, operations, etc.) (e.g., "first", "second", "third", etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term "set" refers to one or more specific elements among specific elements, while the term "multiple" refers to multiple (e.g., two or more) specific elements.

[0040] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling. In some specific implementations, two communicationally coupled (such as electrical communication) devices (or components) may directly or indirectly transmit and receive signals (e.g., digital or analog signals) via one or more wires, buses, networks, etc. As used herein, “direct coupling” can include two devices coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) without intermediate components.

[0041] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” can be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as restrictive, and similar operations can be performed using other techniques. Additionally, as mentioned herein, “obtain,” “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “obtain,” “generate,” “calculate,” “estimate,” or “determine” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining a parameter (or signal), or it can refer to (e.g., using, obtaining, selecting, reading, receiving, retrieving, or accessing, such as a parameter (or signal) already generated by another component or device, from a memory, buffer, container, data structure, lookup table, transmission channel, etc.)

[0042] FIG. 1 This is a block diagram of various aspects of a system 100 capable of operating to process data associated with an immersive audio environment, according to some examples of this disclosure. System 100 includes one or more media output devices 102 coupled to or including an immersive audio renderer 122. Each media output device 102 is configured to output media content to a user. For example, each media output device includes one or more speakers 104, one or more displays 106, or both. The media content may include sound based on an output audio signal 180 (e.g., binaural or multi-channel audio content). Optionally, the media content may also include video content, game content, or other visual content.

[0043] System 100 also includes one or more pose sensors 108. The pose sensors 108 are configured to generate pose data 110 associated with the pose of a user of at least one media output device in media output device 102. As used herein, “pose” indicates the position and orientation of media output device 102, the position and orientation of a user of media output device 102, or both. In some embodiments, at least one of the pose sensors 108 is integrated within a wearable device such that the pose data 110 indicates the user’s pose when the wearable device is worn by a user of media output device 102. In some such embodiments, the wearable device may include pose sensor 108 and at least one media output device in media output device 102. For illustration, pose sensor 108 and at least one media output device in media output device 102 may be combined in a head-mounted wearable device including speaker 104, display 106, or both. Examples of sensors that can be used as wearable pose sensors include, but are not limited to, inertial sensors (e.g., accelerometers or gyroscopes), compasses, positioning sensors (e.g., Global Positioning System (GPS) receivers), magnetometers, inclinometers, optical sensors, and one or more other sensors for detecting acceleration, position, velocity, angular orientation, angular velocity, angular acceleration, or any combination thereof. For illustration, pose sensor 108 may include GPS, electronic maps, and electronic compasses that use inertial and magnetic sensor technologies to determine orientation (such as a 3-axis magnetometer for measuring the Earth's geomagnetic field and a 3-axis accelerometer for providing a horizontal reference to the Earth's magnetic field vector based on the direction of gravity).

[0044] In some implementations, at least one pose sensor in pose sensor 108 is not configured to be worn by a user. For example, at least one pose sensor in pose sensor 108 may include one or more optical sensors (e.g., cameras) for tracking movement of a user or media output device 102. In some implementations, pose sensor 108 may include a combination of user-worn and non-user-worn sensors, wherein the sensor combination is configured to cooperate to generate pose data 110.

[0045] Pose data 110 indicates the pose of the user or media output device 102, or indicates the movement (e.g., change of pose) of the user or media output device 102. In this context, "movement" includes rotation (e.g., a change of orientation without a change in position, such as a change of rolling, tilting, or yaw), translation (e.g., non-rotational movement), or a combination thereof.

[0046] exist FIG. 1In this context, the immersive audio renderer 122 is configured to process immersive audio data based on pose data 110 to generate an output audio signal 180. The immersive audio data corresponds to multiple immersive audio assets (…). FIG. 1 The term "asset" in this context refers to a data structure (such as a file) that stores data representing at least a portion of an immersive audio environment. Generating an output audio signal 180 based on pose data 110 includes generating a sound field representation of the immersive audio data in a manner that takes into account the current or predicted listener pose within the immersive audio environment. For example, an immersive audio renderer 122 is configured to perform rendering operations on assets (e.g., remote asset 144, local asset 142, or both) to generate a rendered asset 126. The rendered asset (whether pre-rendered or rendered as needed, for example, in real-time) may include data describing the sound from multiple sound sources within the immersive audio environment, as such sound sources would be perceived by a listener at a specific location within the immersive audio environment or at a specific location within the immersive audio environment and with a specific orientation. For example, for a particular listener pose, the rendered asset may include data representing sound field characteristics such as the azimuth angle of the direction of the average intensity vector associated with a set of sources in the immersive audio environment. θ ) and elevation angle ( φ ); the signal energy associated with this set of sources in the immersive audio environment ( e The direct and total energy ratio associated with this group of sources in an immersive audio environment ( ); r ); and the interpolated audio signals from this set of sources in an immersive audio environment ( In this example, each of these sound field characteristics can be calculated for each of the frame (f), subframe (k), and frequency range (b).

[0047] The immersive audio renderer 122 includes a binauralizer 128 configured to binauralize the output of a rendering operation (e.g., render asset 126) to generate an output audio signal 180. According to one aspect, the output audio signal 180 includes an output binaural signal provided to a speaker 104 for playback. The rendering operation and binauralization may include sound field rotation (e.g., three degrees of freedom (3DOF)), rotation and finite translation (e.g., 3DOF+), or rotation and translation (e.g., 6DOF) based on the listener's pose.

[0048] exist FIG. 1In this immersive audio renderer 122, an audio asset selector 124 is included or coupled to an audio asset selector configured to select one or more assets based on pose data 110. In some specific implementations, the audio asset selector 124 selects one or more assets for rendering to generate one of the rendered assets 126 based on the current listener pose indicated by the pose data 110. "Current listener pose" refers to the listener's position, orientation, or both in the immersive audio environment as indicated by the pose data 110. In another example, the audio asset selector 124 may select one or more previously rendered assets 126 for output based on the current listener pose indicated by the pose data 110. For illustration, the audio asset selector 124 selects one of the rendered assets 126 for binauralization and output via an output audio signal 180 based on the current listener pose indicated by the pose data 110.

[0049] In the same or different implementations, the audio asset selector 124 is configured to select one or more assets for rendering based on the predicted listener pose. As explained further below, the pose predictor may determine the predicted listener pose based on pose data 110, etc. One benefit of selecting assets based on the predicted listener pose is that the immersive audio renderer 122 can retrieve and / or process (e.g., render) the assets before they are needed, thereby avoiding delays due to asset retrieval and processing.

[0050] After selecting a target asset, the audio asset selector 124 generates an asset retrieval request 138. The asset retrieval request 138 identifies at least one target asset to be retrieved for processing by the immersive audio renderer 122. In a specific implementation where assets are stored in two or more locations (such as remote memory 112 and local memory 170), system 100 includes an asset location selector 130 configured to receive the target asset retrieval request 138 and determine which available memory to retrieve the asset from. In some cases, a particular asset may only be available from one of the available memories. For example, assets 172 stored in local memory 170 may include a subset of assets 114 stored in remote memory 112. For illustration, as further described below, some assets in assets 114 may be retrieved from remote memory 112 (e.g., pre-fetched) and stored in assets 172 in local memory 170 before being processed by the immersive audio renderer 122.

[0051] In some implementations, the asset location selector 130 is configured to retrieve the target asset from local memory 170 if the target asset is among assets 172 stored in local memory 170. In such implementations, based on the determination that the target asset is not stored in local memory 170, the asset location selector 130 chooses to obtain the target asset from remote memory 112. For example, the asset location selector 130 may send an asset retrieval request 138 to client 120, and client 120 may initiate a retrieval of the target asset from remote memory 112 via asset request 136. Otherwise, based on the determination that the target asset is stored in local memory 170, the asset location selector 130 chooses to obtain the target asset from local memory 170. For example, the asset location selector 130 may send an asset retrieval request 138 to local memory 170 to initiate a retrieval of the target asset as local asset 142 in the immersive audio renderer 122.

[0052] exist FIG. 1 In the illustrated example, remote device 116 includes a remote memory that stores multiple assets 114 corresponding to representations of audio content associated with the immersive audio environment. For example, assets 114 stored at remote memory 112 may include one or more scene-based assets 114A, one or more object-based assets 114B, one or more channel-based assets 114C, one or more pre-rendered assets 114D, or combinations thereof. Remote memory 112 is configured to provide client 120 with a list of assets 134 available at remote memory 112, such as a streaming list. Remote memory 112 is configured to receive requests from client 120 for one or more specific assets, such as asset requests 136, and in response to such requests, provide the target asset, such as audio asset 132, to client 120.

[0053] FIG. 1 The pre-rendered assets 114D may include those that have already undergone rendering operations (e.g., as referenced). FIG. 6 (Further description) to generate sound field representations for specific listener locations or specific listener locations and orientations. FIG. 1The scene-based asset 114A may include various versions, such as a first ambisonics representation 114AA, a second ambisonics representation 114AB, a third ambisonics representation 114AC, and one or more additional ambisonics representations including an Nth ambisonics representation 114AN. One or more of the ambisonics representations 114AA to 114AN may correspond to a complete set of ambisonics coefficients corresponding to a specific ambisonic order (such as first-order ambisonics, second-order ambisonics, third-order ambisonics, etc.). Alternatively or additionally, one or more of the ambisonics representations 114AA to 114AN may correspond to a mixed-order ambisonics coefficient set that provides enhanced resolution for a specific listener orientation (e.g., higher resolution in the listener's line of sight direction compared to directions outside the listener's line of sight), while using less bandwidth compared to a complete set of ambisonics coefficients corresponding to the enhanced resolution.

[0054] In some implementations, asset 172 may include assets of the same type as asset 114. For example, asset 172 may include scene-based assets, object-based assets, channel-based assets, pre-rendered assets, or combinations thereof. For example, as noted above, in some implementations, one or more assets in asset 114 may be retrieved from remote memory 112 and stored in asset 172 at local memory 170 before being processed by immersive audio renderer 122. When remote memory 112 provides assets to client 120, the assets may be encoded and / or compressed for transmission (e.g., via one or more networks). In some implementations, client 120 includes or is coupled to decoder 121, which is configured to decode and / or decompress assets for storage at local memory 170, for transmission as remote asset 144 to immersive audio renderer 122, or both. In some such implementations, one or more assets in asset 172 are stored in local memory 170 in an encoded and / or compressed format, and decoder 121 is operable to decode and / or decompress selected assets in asset 172, which are then passed to immersive audio renderer 122 as local asset 142. For example, when a target asset identified in asset retrieval request 138 is in asset 172 stored in local memory 170, asset location selector 130 may determine whether the asset is stored in an encoded and / or compressed format. Based on this determination, asset location selector 130 may selectively cause decoder 121 to decode and / or decompress the asset.

[0055] exist FIG. 1In this system, system 100 includes a pose outlier detector and a mitigator 150 configured to determine the current listener pose based on pose data 110 and one or more pose constraints 158. For example, pose data 110 may indicate the current pose 154 of a listener in an immersive audio environment, and the pose outlier detector and mitigator 150 determines whether to adjust the value of the current pose 154 based on whether the current pose 154 violates one or more pose constraints 158. Additionally or alternatively, the pose outlier detector and mitigator 150 is configured to determine whether to adjust the value of one or more predicted poses 156 based on the current listener pose and pose constraints 158.

[0056] For example, as referenced FIG. 2 In further detail, the pose outlier detector and mitigator 150 is configured to obtain a pose based on pose data 110 and determine whether the pose violates at least one of the pose constraints 158. Based on the determination that the pose does not violate the pose constraint 158, the pose outlier detector and mitigator 150 is configured to use the pose as the current listener pose 154.

[0057] According to some aspects, pose constraint 158 ​​includes human movement constraints. For example, human movement constraints may correspond to velocity constraints or acceleration constraints, and pose outlier detector and mitigator 150 may determine the listener's velocity and / or acceleration based on the listener's current pose indicated by pose data 110 and based on one or more of the listener's previous poses 152. Pose outlier detector and mitigator 150 may determine whether a human movement constraint is violated based on comparing the determined velocity with a velocity constraint and / or by comparing the determined acceleration with an acceleration constraint. In another example, human movement constraints correspond to constraints on the pose of the listener's hands or torso relative to the listener's head pose, and outlier detector and mitigator 150 may determine whether a human movement constraint is violated based on determining the relationship between the listener's head and the listener's hands and / or the listener's torso (e.g., rotational offset, positional difference, etc.) and comparing the determined head / body relationship with constraints.

[0058] According to some aspects, pose constraint 158 ​​includes boundary constraints that indicate the boundaries associated with the immersive audio environment. Pose outlier detector and mitigator 150 can compare the listener position indicated by pose data 110 with the boundaries to determine whether the boundary constraints have been violated.

[0059] If the pose outlier detector and mitigator 150 determines that one or more pose constraints in pose constraints 158 have been violated, the pose outlier detector and mitigator 150 may generate or update the value of the current listener pose such that the current listener pose satisfies pose constraints 158. For example, the most recent previous pose 152 that does not violate any pose constraints 158 may be used as the current pose 154, as referenced. FIG. 3A Further described. As another example, the pose indicated by the pose data 110 can be adjusted based on pose constraints 158 to obtain a current pose 154 that does not violate any pose constraints in pose constraints 158, as described in reference. FIG. 3B Further description.

[0060] Similarly, the pose anomaly detector and mitigator 150 can obtain a predicted pose 156 corresponding to the predicted listener pose from the pose predictor, such as a reference. FIG. 4 Further described. The pose outlier detector and mitigator 150 can determine the acceleration and / or velocity associated with a predicted movement of the listener from the current pose 154 to the predicted pose 156 to determine whether the predicted movement of the listener to the predicted pose 156 will violate the human movement constraints in pose constraints 158. The pose outlier detector and mitigator 150 can also determine whether the predicted pose 156 violates constraints on the listener's hand or torso pose relative to the listener's head pose, boundary constraints, or one or more other pose constraints in pose constraints 158. If it is determined that the predicted pose 156 violates any pose constraint in pose constraints 158, the pose outlier detector and mitigator 150 can generate or update the value of the predicted pose 156 such that the generated or updated value of the predicted pose 156 satisfies pose constraints 158.

[0061] The technical advantage of detecting and mitigating listener pose anomalies lies in reducing or eliminating audio rendering based on listener movements that may be erroneous and / or exceed the spatial boundaries associated with the immersive audio environment. This saves processing resources for system 100 and improves the listener experience. Similarly, in addition to detecting and mitigating anomalies in the listener's current pose, detecting and mitigating anomalies in the predicted listener pose improves the efficiency of system 100 by reducing erroneous predictions associated with pre-fetching assets used to render the immersive audio environment. This includes reducing the consumption of processing resources associated with acquiring and processing assets based on incorrect predictions, and reducing bandwidth usage associated with pre-fetching assets based on incorrect predictions.

[0062] The technical advantages described above can be obtained even when the pose sensor 108 performs filtering of the pose sensor data to remove outliers during the generation of the pose data 110. For example, one or more pose sensors 108 may implement filtering (e.g., Kalman filtering) to remove outliers in the pose sensor data and / or the pose data 110 itself. However, such filtering is typically performed without accessing specific pose constraints 158 (such as human motion constraints and audio scene boundaries) associated with the rendering of the audio scene at system 100. Therefore, the pose data 110 may still include listener poses that are identified as outliers by the pose outlier detector and mitigator 150.

[0063] although FIG. 1 Examples include both local storage 170 storing asset 172 and remote storage 112 storing asset 114; however, in other implementations, only one or the other of storage devices 112 and 170 stores assets for playback. For example, in some implementations of system 100 or some operating modes, asset 172 is downloaded to local storage 170 for use, and asset location selector 130 always retrieves asset 172 from local storage 170. For illustration, system 100 can operate in local-only mode when a network connection to remote storage 112 is unavailable (e.g., when the device is in "airplane mode"). As another example, in some implementations of system 100 or some operating modes, asset 114 is downloaded from remote storage 112 for use, and asset location selector 130 always retrieves asset 114 from remote storage 112 via client 120. For illustration, system 100 can operate in remote-only mode when streaming content from a streaming service associated with remote storage 112. In some implementations, system 100 may be configured for remote operation only, and local storage 170, asset location selector 130, or both may be omitted.

[0064] FIG. 1 This illustrates a specific, non-limiting arrangement of the components of system 100. In other specific embodiments, the components may be arranged in a different manner. FIG. 1 The illustrated different arrangements and interconnections are as follows. For example, decoder 121 may be different from client 120 and located outside of that client. As another example, audio asset selector 124 may be different from immersive audio renderer 122 and located outside of it. For illustration, audio asset selector 124 and asset location selector 130 may be combined. As another example, pose anomaly detector and mitigator 150 may be combined with audio asset selector 124.

[0065] In some implementations, many components of system 100 are integrated within media output device 102. For example, media output device 102 may include a head-mounted wearable device such as a headset, helmet, or earbuds, comprising client 120, local memory 170, asset location selector 130, immersive audio renderer 122, motion estimator 460, pose sensor 108, or any combination thereof. As another example, media output device 102 may include a head-mounted wearable device and a separate player device such as a game console, computer, or smartphone. In this example, at least a pair of speakers 104 and at least one pose sensor 108 may be integrated within the head-mounted wearable device, and other components of system 100 may be integrated into the player device or partitioned between the player device and the head-mounted wearable device.

[0066] FIG. 2 An example of a pose outlier detection operation 200 that can be performed by the pose outlier detector and mitigator 150 is depicted. Operation 200 includes determining at box 210 whether one or more pose constraints are violated.

[0067] For example, the pose anomaly detector and mitigator 150 processes pose information 204 including a previous pose 152, a current pose 154, a predicted pose 156, and one or more hand / torso poses 214. For illustration, one or more pose sensors in pose sensors 108 may be configured to track movement of the listener's hands, such as pose sensors 108 included in (or coupled to) a handheld controller device, virtual reality and / or haptic gloves, a smartwatch, or other hand-based wearable devices. Additionally or alternatively, one or more pose sensors in pose sensors 108 may be configured to track movement of the listener's torso, such as pose sensors 108 included in (or coupled to) portable electronic devices such as smartphones or tablets, virtual reality and / or haptic vests.

[0068] The determination of whether one or more pose constraints are violated at box 210 is based on human movement constraint 206 and spatial physical constraint 208. For example, human movement constraint 206 may be included in pose constraint 158 ​​and may include velocity constraint 216, acceleration constraint 218, and constraints on the pose of the listener's hands or torso relative to the listener's head pose (which is exemplified as relative head / body constraint 220).

[0069] Spatial physical constraints 208 include one or more boundary constraints 222, such as 6DOF boundaries associated with the rendering of the immersive audio environment. For example, spatial physical constraints 208 may include scene boundary distances along six directions, such as boundary distances in the +x, -x, +y, -y, +z, and -z directions, where (x, y, z) correspond to the coordinate system used by the immersive audio renderer 122 to represent the listener's position and the audio source associated with the sound scene.

[0070] Determining whether one or more constraints are violated at box 210 may include determining a velocity associated with the current pose 154 and comparing that velocity to velocity constraint 216. For example, the velocity may correspond to a rotational velocity associated with the difference in listener head orientation between the selected previous pose 152 and the current pose 154 within a time period between a first timestamp associated with the selected previous pose 152 (e.g., the most recent previous pose 152) and a second timestamp associated with the current pose 154. Velocity constraint 216 may include a rotational velocity threshold, and the determined rotational velocity may be compared to the rotational velocity threshold to determine whether velocity constraint 216 is violated. As another example, the velocity may correspond to a translational velocity associated with the difference in listener position between the selected previous pose 152 and the current pose 154 within a time period between a first timestamp associated with the selected previous pose 152 (e.g., the most recent previous pose 152) and a second timestamp associated with the current pose 154. Velocity constraint 216 may include a translational velocity threshold, and a determined translational velocity may be compared with the translational velocity threshold to determine whether velocity constraint 216 is violated. Similar comparisons may be made to determine whether the predicted pose 156 violates the rotational and / or translational velocity thresholds of velocity constraint 216 by determining rotational velocity based on changes in head orientation between the current pose 154 and the predicted pose 156 within a time period between a second timestamp associated with the current pose 154 and a predicted third timestamp associated with the predicted pose 156, and / or by determining translational velocity based on changes in listener position between the current pose 154 and the predicted pose 156 within a time period between a second timestamp associated with the current pose 154 and a predicted third timestamp associated with the predicted pose 156.

[0071] Determining whether one or more constraints are violated at box 210 may include determining an acceleration associated with the current pose 154 and comparing the acceleration to acceleration constraint 218. For example, the acceleration may correspond to rotational acceleration associated with the difference in listener head rotation speed between the selected previous pose 152 and the current pose 154 over a time period between a first timestamp associated with the selected previous pose 152 (e.g., the most recent previous pose 152) and a second timestamp associated with the current pose 154. Acceleration constraint 218 may include a rotational acceleration threshold, and the determined rotational acceleration may be compared to the rotational acceleration threshold to determine whether acceleration constraint 218 is violated. As another example, the acceleration may correspond to translational acceleration associated with the difference in listener translation speed between the selected previous pose 152 and the current pose 154 over a time period between a first timestamp associated with the selected previous pose 152 (e.g., the most recent previous pose 152) and a second timestamp associated with the current pose 154. Acceleration constraint 218 may include a translational acceleration threshold, and the determined translational acceleration may be compared with the translational acceleration threshold to determine whether acceleration constraint 218 is violated. Similar comparisons may be made to determine rotational acceleration based on changes in head rotational velocity between the current pose 154 and the predicted pose 156 within a time period between a second timestamp associated with the current pose 154 and a third prediction timestamp associated with the predicted pose 156, and / or to determine translational acceleration based on changes in listener translational velocity between the current pose 154 and the predicted pose 156 within a time period between a second timestamp associated with the current pose 154 and a third prediction timestamp associated with the predicted pose 156, to determine whether the predicted pose 156 violates the rotational and / or translational acceleration thresholds of acceleration constraint 218.

[0072] Determining whether one or more constraints are violated at box 210 may include determining the current hand and / or torso position of the listener relative to the listener's head pose indicated by the current pose 154, and comparing the current hand and / or torso position relative to the listener's head pose with the relative head / body constraint 220. For example, hand / torso pose 214 may include a hand pose indicating the position of the listener's hand, and the position of the listener's hand relative to the listener's head (indicated by the current pose 154) may be determined and compared with the hand-head relative position constraint in the relative head / body constraint 220. As another example, hand / torso pose 214 may include a body pose indicating the position of the listener's torso, and the position of the listener's torso relative to the listener's position may be determined and compared with the torso-head relative position constraint in the relative head / body constraint 220. Similar operations can be performed to compare the hand-head relative rotation and / or torso-head relative rotation with the hand-head relative rotation constraints and / or torso-head relative rotation constraints in the relative head / body constraint 220, respectively. In some implementations, a similar comparison can be made between the predicted hand-head relative position and / or rotation, the predicted torso-head relative position and / or rotation, or combinations thereof, with the corresponding constraints in the relative head / body constraint 220 to determine whether the predicted pose 156 and the predicted head / torso pose violate the relative head / body constraint 220.

[0073] Determining whether one or more constraints are violated at box 210 may include determining the listener's position in the audio scene as indicated by the current pose 154, and comparing the listener's position to spatial physical constraints 208. For example, the listener's position may be compared to boundary constraints 222 to determine whether the listener is inside or outside the spatial boundary defined by boundary constraints 222. Determining whether one or more constraints are violated at box 210 may include determining the predicted listener position in the audio scene as indicated by the predicted pose 156, and comparing the predicted listener position to spatial physical constraints 208. For example, the listener's predicted position may be compared to boundary constraints 222 to determine whether the listener is predicted to be inside or outside the spatial boundary defined by boundary constraints 222.

[0074] In response to a determination based on pose information 204 that one or more human movement constraints 206 and / or one or more spatial physical constraints 208 have been violated, at block 212, the pose outlier detector and mitigator 150 sets an outlier detection indicator 224. The outlier detection indicator 224 can be used to trigger the execution of pose outlier mitigation operations, as shown in reference... FIG. 3A to FIG. 3B Further description.

[0075] In some implementations, the outlier detection indicator 224 includes an indication of which constraint(s) were violated, an indication of whether a violation was detected for the current pose 154, for the predicted pose 156, or any combination thereof. Information associated with determining that one or more constraints have been violated, such as calculated velocity, acceleration, position, relative movement, etc., may also be saved for reuse during outlier mitigation.

[0076] FIG. 3A A first example of a pose outlier mitigation operation 300, which can be performed by the pose outlier detector and mitigator 150, is depicted. Operation 200 includes determining whether one or more pose constraints have been violated, such as by determining at box 340 whether an outlier detection indicator 224 has been set.

[0077] Based on the determination that the current pose 154 does not violate one or more pose constraints (e.g., outlier detection indicator 224 has not yet been set), the pose outlier detector and mitigator 150 are configured to use the current pose 154 as the current listener pose for the purpose of asset retrieval (e.g., stream selection) at the audio asset selector 124 and / or rendering at the immersive audio renderer 122.

[0078] Otherwise, based on the determination that the current pose 154 violates at least one of one or more pose constraints, the pose outlier detector and mitigator 150 is configured to determine the current listener pose based on a previous listener pose that did not violate the pose constraints. For illustration, if the outlier detection indicator 224 has been set to indicate that the current pose 154 is an outlier, then at operation 342, the current pose 154 is set to the previous pose. For example, the current pose 154 may be set to the value of the most recent previous pose 152 that was not determined to be an outlier. Setting the current pose 154 to the value of the most recent non-outlier previous pose 152 may include changing one or more values ​​of the current pose 154 to an equal corresponding value of the previous pose 152, replacing the current pose 154 with the previous pose 152, or adjusting the current pose 154 to match the previous pose 152, as illustrative and non-limiting examples.

[0079] In some implementations, the pose outlier detector and mitigator 150 are further configured to determine a predicted pose 156 based on a previously predicted listener pose associated with a previous listener pose, based on the determination that the current pose 154 violates at least one of one or more pose constraints. For example, a previously predicted listener pose 152, selected as the most recent non-outlier previously predicted pose 152 for adjusting the current pose 154, may be associated with the previously predicted listener pose, and at operation 344, the predicted pose 156 is set to the value of that previously predicted listener pose.

[0080] One effect of operations 342 and 344 at box 340 is that the poses used for asset selection and audio rendering can effectively repeat the most recent previous non-outlier poses and predicted poses, as if the listener had not moved since the pose data 110 indicated a previous non-outlier pose. In some cases, when a violation of one or more of the human movement constraints 206 is detected, the violation can indicate an unusually large or physiologically unlikely (or improbable) movement by the listener. For example, human movement constraints 206 can be associated with corresponding thresholds 306, including a velocity threshold 316 corresponding to velocity constraint 216, an acceleration threshold 318 corresponding to acceleration constraint 218, and a body pose threshold 320 corresponding to body pose threshold 320. Each threshold in 306 can correspond to the maximum limit above which the listener's movement would prevent the listener from tracking changes in the audio scene, thus saving processing resources consumed by selecting, acquiring, and rendering assets when the listener's movement exceeds one or more of the thresholds in 306 without adversely affecting the listener's experience. In other cases, when a violation of one or more of the spatial physical constraints 208 is detected, the violation may instruct the listener to move outside the boundaries of the audio scene or to an area of ​​the audio scene without an audio source, and repeat the most recent previous non-outlier pose and predicted pose to prevent the generation of invalid target asset retrieval requests 138.

[0081] FIG. 3B A second example of a pose outlier mitigation operation 350, which can be performed by the pose outlier detector and mitigator 150, is depicted. Operation 200 includes determining whether one or more pose constraints have been violated, such as by determining at box 370 whether an outlier detection indicator 224 has been set.

[0082] Based on the determination that the current pose 154 does not violate one or more pose constraints (e.g., outlier detection indicator 224 has not yet been set), the pose outlier detector and mitigator 150 are configured to use the current pose 154 as the current listener pose for the purpose of asset retrieval (e.g., stream selection) at the audio asset selector 124 and / or rendering at the immersive audio renderer 122.

[0083] Otherwise, based on the determination that the current pose 154 violates at least one of one or more pose constraints, the pose outlier detector and mitigator 150 is configured to determine the current listener pose based on adjusting the current pose 154 to satisfy one or more pose constraints. For example, if the outlier detection indicator 224 has been set to indicate that the current pose 154 is an outlier, then at operation 372, the value of the current pose 154 is adjusted to match a threshold 306 or spatial boundary associated with the violated pose constraint. For instance, when the listener velocity determined based on the current pose 154 exceeds a velocity threshold 316, at operation 372, the velocity associated with the current pose 154 may be adjusted, for example via a clipping operation, to match the velocity threshold 316, such that one or more values ​​associated with the current pose are clipped so that none of the thresholds 306 are exceeded. As another example, when the listener's position based on the current pose 154 is outside the boundary indicated by the velocity constraint 216, the position associated with the current pose 154 can be clipped so that it does not cross the boundary (e.g., allowing the listener to move to the boundary but not beyond it).

[0084] In some implementations, the pose outlier detector and mitigator 150 are also configured to determine a predicted pose 156 based on adjusting a previously predicted listener pose associated with a previous listener pose, based on the determination that the current pose 154 violates at least one of one or more pose constraints. For example, the most recent non-outlier previous pose 152 may be selected, and the previously predicted pose 156 associated with the selected previous pose 152 may alternatively be used as the predicted pose 156. If using the selected previously predicted pose 156 results in exceeding one or more thresholds or spatial boundaries in thresholds 306, then at operation 374, the previously predicted pose 156 may be adjusted (e.g., clipped) in a manner similar to that described for operation 372 to match the threshold or boundary associated with the violated pose constraint.

[0085] One effect of operations 372 and 374 at box 370 is that the pose used for asset selection and audio rendering can track the listener pose indicated by pose data 110 as closely as possible, but the listener pose is not allowed to violate any threshold or boundary constraint 222 in threshold 306.

[0086] Although they have been described separately FIG. 2 , FIG. 3A and FIG. 3B The specific implementation of each of operations 200, 300, and 350, but in alternative implementations... FIG. 2 , FIG. 3A and FIG. 3BOperations 200, 300, and 350 may omit one or more of the human movement constraint 206 and associated threshold 306, respectively, and incorporate one or more additional human movement constraints 206 and associated thresholds 306, or combinations thereof. In some implementations, the human movement constraint 206 and associated threshold 306 may be omitted, and outlier detection and mitigation may be performed solely based on the spatial physical constraint 208. Similarly, FIG. 2 , FIG. 3A and FIG. 3B Operations 200, 300, and 350 may omit one or more of the spatial physical constraints 208 and associated boundary constraints 222, respectively, and incorporate one or more spatial physical constraints 208 and associated spatial conditions, or a combination thereof. In some implementations, spatial physical constraints 208 may be omitted, and outlier detection and mitigation may be performed solely based on human movement constraints 206.

[0087] FIG. 4 It includes FIG. 1 A block diagram of system 400 showing various aspects of system 100. For example, system 400 includes... FIG. 1 The media output device 102, immersive audio renderer 122, asset location selector 130, client 120, local storage 170, and remote storage 112, each of which is as described in the reference. FIG. 1 It operates as described. In system 400, an immersive audio renderer 122, an asset location selector 130, a client 120, and local memory 170 are included in an immersive audio player 402, which is configured (e.g., via modem 420) to communicate with a pose sensor 108 and a media output device 102. In other examples, the media output device 102 and the immersive audio player 402 are integrated into a single device (such as a wearable device), which may include a speaker 104, a display 106, a pose sensor 108, or a combination thereof.

[0088] exist FIG. 4 In the immersive audio player 402, one or more processors 410 are configured to execute instructions (e.g., instructions 174 from local memory 170) to perform operations of an immersive audio renderer 122, an asset location selector 130, a client 120, a pose anomaly detector and mitigator 150, an audio asset selector 124, a motion estimator 460, a pose predictor 450, or a combination thereof.

[0089] FIG. 4 Examples FIG. 1An example of system 100, wherein the pose outlier detector and mitigator 150 is one aspect of the immersive audio renderer 122. For example, the pose outlier detector and mitigator 150 may be integrated within or coupled to the pose predictor 450. The pose outlier detector and mitigator 150 may be configured to process the current listener pose (e.g., current pose 154) indicated by pose data 110 and the predicted listener pose (e.g., predicted pose 156) generated by the pose predictor 450 before the current listener pose and / or predicted listener pose are used by the audio asset selector 124, motion estimator 460, pose predictor 450, immersive audio renderer 122, or any combination thereof, to detect and mitigate pose outliers. In some specific implementations, the pose outlier detector and mitigator 150 may also be configured to process the contextual motion estimation data 462 in a manner similar to that described above for detecting and mitigating outliers in the predicted listener pose, which indicates the expected amount or type of listener pose movement at a particular time.

[0090] In some implementations, the motion estimator 460 is configured to determine contextual motion estimation data 462 based on a previous pose 152, a current pose 154, a predicted pose 156, or a combination thereof. For example, the motion estimator 460 may determine contextual motion estimation data 462 based on historical motion rates, wherein historical motion rates are determined based on differences between previous poses 152, between previous poses 152 and current pose 154, between previous poses 152 or current pose 154 and predicted pose 156, or a combination thereof. In this context, the previous pose 152 may include historical pose data 110; while the current pose 154 refers to the pose indicated by the most recent set of samples of pose data 110.

[0091] The motion estimator 460 can make the contextual motion estimation data 462 based on various types of information. For example, the motion estimator 460 can generate the motion estimation data 462 based on pose data 110. For illustration, pose data 110 may indicate the current listener pose, and the motion estimator 460 may generate the motion estimation data 462 based on the current listener pose or a set of recent changes in the current listener pose over time. As an example, the motion estimator 460 may generate the contextual motion estimation data 462 based on the rate and / or type of recent changes in the listener pose, based on a set of recent listener pose data, where "recent" is determined based on a specified time limit (e.g., the last minute, the last five minutes, etc.) or based on a specified number of pose data 110 samples (e.g., the most recent ten samples, the most recent one hundred samples, etc.).

[0092] As another example, motion estimator 460 may generate motion estimation data 462 based on predicted pose. For example, a pose predictor may generate a predicted listener pose based at least in part on pose data 110. The predicted listener pose may indicate the listener's position and / or orientation in the immersive audio environment at some future time. In this example, motion estimator 460 may generate motion estimation data 462 based on movement that will occur (e.g., is predicted to occur) to change the listener pose from the current listener pose to the predicted listener pose.

[0093] As another example, motion estimator 460 may generate motion estimation data 462 based on historical interaction data 458 associated with an asset, an immersive audio environment, a scene within the immersive audio environment, or a combination thereof. Historical interaction data 458 may indicate the interactions of the current user of media output device 102, the interactions of other users who have consumed a particular asset or have interacted with the immersive audio environment, or a combination thereof. For example, historical interaction data 458 may include motion trajectory data describing the movements of a group of users (which may include the current user) who have interacted with the immersive audio environment. In this example, motion estimator 460 may use historical interaction data 458 to estimate how much the current user may move in the near future (e.g., during consumption of an asset or part of a scene that the user is currently consuming). To illustrate, when the immersive audio environment is associated with game content, the scene of the game content may depict a fright event (e.g., an explosion, a collision, a jump scare, etc.) that has historically caused the user to quickly look in a particular direction or look around the environment, as indicated by historical interaction data 458. In this exemplary example, contextual movement estimation data 462 may be based on historical interaction data 458 to indicate that the rate and / or type of movement of the listener's pose may increase when a startling event occurs.

[0094] As another example, motion estimator 460 may generate motion estimation data 162 based on one or more contextual cues 454 (also referred to herein as “motion cues”) associated with the immersive audio environment. One or more contextual cues 454 may be explicitly provided in the metadata representing the asset of the immersive audio environment. For example, the metadata associated with the asset may include fields indicating the contextual motion estimation data 462. To illustrate, a game creator or publisher may indicate in the metadata associated with a particular asset that the asset or a portion of the asset is expected to cause a change in the listener’s movement rate. As an example, if the game’s scenario includes events that may cause the user to move more (or less), the game’s metadata may indicate when the event occurs during the asset’s playback, where the event occurs (e.g., the location of a sound source in the immersive audio environment), the type of event, the expected outcome of the event (e.g., increased or decreased translation in a particular direction, increased or decreased head rotation, etc.), the duration of the event, etc.

[0095] In some implementations, one or more contextual cues in context cue 454 are implicit rather than explicit. For example, metadata associated with an asset may indicate the genre of the asset, and the movement estimator 460 may generate contextual movement estimation data 462 based on the genre of the asset. For illustration, the movement estimator 460 may anticipate that head movement during playback of an immersive audio environment representing a classical music genre is not as fast as anticipated during playback of an immersive audio environment representing a first-person shooter game.

[0096] The motion estimator 460 is configured to set one or more pose update parameters 456 based on contextual motion estimation data 462. In a particular aspect, the pose update parameter 456 indicates the pose data update rate of the pose data 110. For example, the motion estimator 460 can set the pose update parameter 456 by transmitting it to the pose sensor 108 so that the pose sensor 108 provides the pose data 110 at a rate associated with the pose update parameter 456. In some embodiments, the system 100 includes two or more pose sensors 108. In such embodiments, the motion estimator 460 can transmit the same pose update parameter 456 to each of the two or more pose sensors 108, or the motion estimator 460 can transmit different pose update parameters 456 to different pose sensors 108. For illustration, system 100 may include a first pose sensor 108 configured to generate pose data 110 indicating translational positioning of a listener in an immersive audio environment and a second pose sensor 108 configured to generate pose data 110 indicating rotational orientation of a listener in the immersive audio environment. In this example, motion estimator 460 may transmit different pose update parameters 456 to the first and second pose sensors 108. For example, contextual motion estimation data 462 may indicate that the rate of head rotation is expected to increase, while the rate of translation is expected to remain unchanged. In this example, motion estimator 460 may transmit the first pose update parameter 456 to cause the second pose sensor to increase the generation rate of pose data 110 indicating rotational orientation of the listener, and may avoid transmitting pose update parameter 456 to the first pose sensor (or may transmit the second pose update parameter 456) to cause the first pose sensor to continue generating pose data 110 indicating translational orientation at the same rate as before.

[0097] One technical advantage of using contextual motion estimation data 462 to set pose update parameters 456 is that the pose data 110 update rate can be set based on the user's movement rate, which enables resource savings and an improved user experience. For example, when a relatively high movement rate is expected (as indicated by contextual motion estimation data 462), pose update parameters 456 can be set to increase the rate at which pose data 110 is updated. The increased update rate of pose data 110 reduces motion / sound latency in the output audio signal 180. For illustration, in this example, user movement (e.g., head rotation) is reflected more quickly in the output audio signal 180 because pose data 110 reflecting user movement is available more quickly to the immersive audio renderer 122. Conversely, when a relatively low movement rate is expected (as indicated by contextual motion estimation data 462), pose update parameters 456 can be set to decrease the rate at which pose data 110 is updated. The reduced update rate of pose data 110 saves resources associated with rendering and binauralization of the immersive audio renderer 122 (e.g., computation cycle, power, memory), resources associated with transmitting pose data 110 (e.g., bandwidth, power), or a combination thereof.

[0098] In a specific aspect, the pose predictor 450 is configured to use prediction techniques to determine the predicted pose 156, such as extrapolation based on the previous pose 152 and / or the current pose 154; inference using one or more artificial intelligence models; probability-based estimation based on the previous pose 152 and / or the current pose 154; based on... FIG. 1 Probability-based estimation of 458 historical interaction data; FIG. 1 Contextual hints 454; or combinations thereof. Predicted pose 156 can be used to reduce motion-to-sound delay of output audio signal 180. For example, immersive audio renderer 122 can generate asset retrieval request 138 for one or more assets associated with predicted pose 156. In this example, immersive audio renderer 122 can process assets associated with a specific predicted pose 156 to generate rendered assets. Rendered assets represent the sound field of the immersive audio environment as the sound field will be perceived by a listener with a specific predicted pose 156. In this example, rendered assets are used to generate output audio signal 180 when (or if) pose data 110 indicates that the specific predicted pose 156 used for rendering assets is the current pose 154. By using the predicted pose 156 to select and / or render assets, the immersive audio renderer 122 is able to perform many complex rendering operations in advance, thereby reducing the latency of providing the output audio signal 180 representing a specific asset and pose compared to selecting, requesting, receiving, and rendering assets on demand (e.g., rendering assets specifically based on the current pose 154).

[0099] In some implementations, the immersive audio renderer 122 may render two or more assets based on predicted poses 156. For example, in some cases, there may be significant uncertainty about which possible pose the user will move to in the set of possible poses in the future. For illustration, in a game environment, the user may face several choices, and a particular choice made by the user may change the asset to be rendered, the future listener pose, or both. In this example, predicted poses 156 may include multiple poses at a specific future time, and the immersive audio renderer 122 may render one asset based on two or more predicted poses 156, or render two or more different assets based on two or more predicted poses 156, or both. In this example, when the current pose 154 is aligned with one of the predicted poses 156, the corresponding rendered asset is used to generate the output audio signal 180.

[0100] In some specific implementations, the immersive audio renderer 122 can render assets in stages, as shown in the reference. FIG. 8 As described. For example, an immersive audio renderer 122 may perform a first set of operations to localize the sound field representation of the immersive audio environment to the listener's position, and perform a second set of operations to rotate the sound field representation of the immersive audio environment to the listener's orientation. In some such embodiments, the immersive audio renderer 122 may perform only the first set of operations (e.g., localization operations) or only the second set of operations (e.g., rotation operations) based on the predicted pose 156. In such embodiments, when (or if) one of the predicted poses 156 becomes the current pose 154, the remaining operations (e.g., localization or rotation operations) are performed to generate an output audio signal 180.

[0101] In certain aspects, the mode or rate of pose prediction by pose predictor 450 may be related to pose update parameter 456. For example, pose predictor 450 may be turned off when pose update parameter 456 has a specific value. For illustration, when contextual motion estimation data 462 indicates that little or no user movement is expected within a specific time period, pose update parameter 456 may be set such that pose sensor 108 is turned off or provides pose data 110 at a low rate, and pose predictor 450 is turned off. Conversely, when contextual motion estimation data 462 indicates that rapid user movement is expected within a specific time period, pose update parameter 456 may be set such that pose sensor 108 provides pose data 110 at a high rate, and pose predictor 450 generates predicted pose 156. Additionally or alternatively, pose predictor 450 may generate predicted pose 156 in a first mode for a future first distance time, and in a second mode for a future second distance time, wherein the mode is selected based on pose update parameter 456.

[0102] The technical advantage of the pose predictor 450 adjusting the mode or rate of pose prediction based on the pose update parameter 456 is that the pose predictor 450 can generate more predicted poses 156 for periods of expected more movement and fewer predicted poses 156 for periods of expected less movement. Generating more predicted poses 156 for periods of expected more movement allows the immersive audio renderer 122 to have a higher probability of rendering assets that will be used to generate the output audio signal 180 in advance. For example, the immersive audio renderer 122 can render assets associated with predicted poses 156, and when the current pose 154 corresponds to a predicted pose for a specific asset among these assets, use that specific rendered asset to generate the output audio signal 180. In this example, having more predicted poses 156 and corresponding rendered assets means that there is a higher probability that the current pose 154 at some point in the future will correspond to one of the predicted poses 156, thus enabling the use of the corresponding rendered asset to generate the output audio signal 180 instead of performing real-time rendering. On the other hand, pose prediction and rendering assets based on the predicted pose 156 are resource-intensive and may be wasteful if assets are not rendered based on the predicted pose 156. Therefore, generating fewer predicted poses 156 for periods of less expected movement allows the immersive audio renderer 122 to save resources.

[0103] FIG. 5 It includes FIG. 1 A block diagram of system 500 showing various aspects of system 100. For example, system 500 includes an immersive audio player 402, a media output device 102, and a remote memory 112. Furthermore, the immersive audio player 402 includes a processor 410, a modem 420, and local memory 170, and the processor 410 is configured to execute instructions 174 to perform operations of the immersive audio renderer 122, the asset location selector 130, and the client 120. In addition to what is described below, FIG. 5 The immersive audio player 402, media output device 102, remote storage 112, immersive audio player 402, processor 410, modem 420, local storage 170, immersive audio renderer 122, asset location selector 130, and client 120 are each as described in the reference. FIG. 1 to FIG. 4 Operate as described.

[0104] In system 500, pose sensor 108, pose outlier detector and mitigator 150, pose predictor 450, and motion estimator 460 are on media output device 102 (e.g., integrated within the media output device). This is to enable immersive audio renderer 122 to render certain assets before they are needed (e.g., based on predicted pose 156). FIG. 5The pose data 110 includes a predicted pose 156 and a current pose 154. The current pose 154 and the predicted pose 156 can be processed by a pose outlier detector and mitigator 150 to detect and mitigate pose outliers before being transmitted in the pose data 110 to the immersive audio renderer 122. (See reference...) FIG. 4 As described, the motion estimator 460 can determine contextual motion estimation data 462, which can be used to set pose update parameters 456. These pose update parameters affect the rate at which the pose sensor 108 transmits updated pose data 110 to the immersive audio player 402 and may optionally affect the operation of the pose predictor 450. (See reference...) FIG. 1 and FIG. 4 As described, setting pose update parameters 456 based on contextual motion estimation data 462 can achieve resource savings and an improved user experience.

[0105] FIG. 6 It includes FIG. 1 A block diagram of various aspects of system 600 is provided for system 100. For example, system 600 includes an immersive audio player 402, a media output device 102, and a remote memory 112. Furthermore, the immersive audio player 402 includes a processor 410, a modem 420, and local memory 170, and the processor 410 is configured to execute instructions 174 to perform operations of the immersive audio renderer 122, the asset location selector 130, and the client 120. In addition to the descriptions below, FIG. 6 The immersive audio player 402, media output device 102, remote storage 112, immersive audio player 402, processor 410, modem 420, local storage 170, immersive audio renderer 122, asset location selector 130, and client 120 are each as described in the reference. FIG. 1 and FIG. 4 Operate as described.

[0106] FIG. 6 Examples of the pose outlier detector and mitigator 150, motion estimator 460, and pose predictor 450 are aspects of the client 120. FIG. 1 Example of system 100. FIG. 6 An example of a system 100 in which historical interaction data 458 is based on movement trajectory data 602, movement trajectory data 606, or both is also shown.

[0107] For reference FIG. 4 As described, the motion estimator 460 is configured to determine contextual motion estimation data 462 and set pose update parameters 456 based on the contextual motion estimation data 462. FIG. 6In the example shown, pose update parameter 456 is provided to pose predictor 450, immersive audio renderer 122, pose sensor 108, or a combination thereof. (See reference...) FIG. 1 As described, the motion estimator 460 is configured to set the pose update parameters 456 based on contextual motion estimation data 462. In certain aspects, the pattern or rate of pose prediction by the pose predictor 450 may be correlated with the pose update parameters 456.

[0108] In certain aspects, the pose predictor 450 is configured to use a reference FIG. 4 The described prediction technique determines a predicted pose 156. Client 120 provides the predicted pose 156 (e.g., after processing by the pose outlier detector and mitigator 150) to immersive audio renderer 122. Immersive audio renderer 122 issues an asset retrieval request 138 for assets associated with one or more predicted poses 156 and processes the retrieved assets associated with the predicted poses 156 to generate rendered assets 126. By using the predicted poses 156 to select and / or render assets, immersive audio renderer 122 is able to perform many complex rendering operations in advance, thereby reducing the latency of providing a representation of a specific asset and pose via the output audio signal compared to rendering assets on demand (e.g., rendering assets specifically based on the current pose 154).

[0109] exist FIG. 6 In this context, the motion estimator 460 may use historical interaction data 458 (optionally, along with other information) to determine contextual motion estimation data 462. Additionally or alternatively, the pose predictor 450 may use historical interaction data 458 (optionally, along with other information) to determine predicted pose 156. Historical interaction data 458 may include or correspond to motion trajectory data associated with the immersive audio environment. For example, local memory 170 may store motion trajectory data 606, which may indicate how one or more users of the immersive audio player 402 move during playback in an immersive audio environment, during playback in another immersive audio environment, or both. In this example, motion trajectory data 606 may include information describing the immersive audio environment (e.g., by title, genre, etc.), specific movements or listener poses detected during playback and the time index when such movements or poses were detected, other user interactions detected during playback (e.g., game inputs) and their associated time indexes, etc.

[0110] In some implementations, the movement trajectory data 602 stored at remote memory 112 is a copy (e.g., identical) of the movement trajectory data 606 stored at local memory 170. In some implementations, the movement trajectory data 602 stored at remote memory 112 includes the same type of information (e.g., data fields) as the movement trajectory data 606 stored at local memory 170, but includes information describing how users of other immersive audio player devices have interacted with the immersive audio environment. For example, the movement trajectory data 602 may aggregate historical user interactions associated with the immersive audio environment across multiple users of immersive audio player 402 and other immersive audio players.

[0111] In one implementation where the motion estimator 460 determines contextual motion estimation data 462 based on historical interaction data 458, the historical interaction data 458 may indicate or be used to determine motion probability information associated with a specific scene or specific asset in the immersive audio environment. For example, motion probability information may indicate the likelihood of a specific rate of movement during a specific segment of the playback, based on how the user or other users have moved during that segment. As another example, motion probability information may indicate the likelihood of a specific type of movement (e.g., translation in a specific direction, rotation in a specific direction, etc.) during that segment, based on how the user or other users have moved during that segment. Therefore, the motion estimator 460 may set pose update parameters 456 to prepare for anticipated movements associated with the playback of the immersive audio environment. For example, when historical interaction data 458 indicates that an upcoming portion of the immersive audio environment has historically been associated with rapid rotation of the listener's pose, the motion estimator 460 may set the pose update parameter 456 to increase the rate at which rotation-related pose data 110 is provided by the pose sensor 108, thereby reducing motion-to-sound latency associated with the playback of that upcoming portion. Conversely, when historical interaction data 458 indicates that an upcoming portion of the immersive audio environment has historically been associated with little or no change in the listener's pose, the motion estimator 460 may set the pose update parameter 456 to decrease the rate at which pose data 110 is provided by the pose sensor 108, thereby saving power and computational resources.

[0112] In one implementation where the pose predictor 450 determines the predicted pose 156 based on historical interaction data 458, the historical interaction data 458 may indicate or be used to determine pose probability information associated with a specific scene or specific asset in the immersive audio environment. For example, the pose probability information may indicate the probability of a specific listener position, a specific listener orientation, or a specific listener pose during the playback of a specific part of the immersive audio environment based on historical listener poses during the playback of that specific part.

[0113] The technical benefit of determining historical interaction data 458 based on motion trajectory data 602, 606 is that motion trajectory data 602, 606 provides an accurate estimate of how a real user interacts with the immersive audio environment, thereby enabling more accurate pose prediction, more accurate contextual motion estimation, or both. Furthermore, motion trajectory data 602, 606 can be easily captured. For example, during the playback of content associated with a specific immersive audio environment using immersive audio player 402, immersive audio player 402 can store motion trajectory data 606 at local memory 170. Immersive audio player 402 can transmit motion trajectory data 606 to remote memory 112 to update motion trajectory data 606 at any convenient time, such as after playback of content associated with a specific immersive audio environment has finished, or when immersive audio player 402 is connected to remote memory 112 and the connection to remote memory 112 has available bandwidth. Motion trajectory data 602 may include an aggregation of historical interaction data from the user of immersive audio player 402 and other users.

[0114] FIG. 7 It includes FIG. 1 A block diagram of various aspects of system 700 is provided for system 100. For example, system 700 includes an immersive audio player 402, a media output device 102, and a remote memory 112. Furthermore, the immersive audio player 402 includes a processor 410, a modem 420, and local memory 170, and the processor 410 is configured to execute instructions 174 to perform operations of the immersive audio renderer 122, the asset location selector 130, and the client 120. In addition to the descriptions below, FIG. 7 The immersive audio player 402, media output device 102, remote storage 112, immersive audio player 402, processor 410, modem 420, local storage 170, immersive audio renderer 122, asset location selector 130, and client 120 are each as described in the reference. FIG. 1 and FIG. 4 to FIG. 6 Operate as described by either of them.

[0115] FIG. 7 Examples FIG. 1 An example of system 100, which includes at least two pose sensors 108, such as pose sensor 108A and pose sensor 108B. FIG. 7In the example shown, pose sensor 108A is integrated within one of the media output devices 102, and pose sensor 108B is shown external to media output device 102; however, in other embodiments, pose sensor 108A is integrated within a first media output device of media output device 102, and pose sensor 108B is integrated within a second media output device of media output device 102. For example, pose sensor 108A may be included in a head-mounted media output device such as headphones or earphones, and pose sensor 108B may be included in a non-head-mounted media output device such as a game console, computer, or smartphone.

[0116] In certain aspects, pose sensors 108A and 108B work together to determine the listener's pose. For example, in some implementations, pose sensor 108A provides pose data 110A representing rotation (e.g., rotation of the user's head), and pose sensor 108B provides pose data 110B indicating translation (e.g., movement of the user's body). As another example, pose data 110A may include first translation data, and pose data 110B may include second translation data. In this example, the first and second translation data may be combined (e.g., subtracted) to determine changes in the listener's pose in an immersive audio environment. Additionally or alternatively, pose data 110A may include first rotation data, and pose data 110B may include second rotation data. In this example, the first and second rotation data may be combined (e.g., subtracted) to determine changes in the listener's pose in an immersive audio environment.

[0117] exist FIG. 7 In the example shown, the motion estimator 460 can set the pose update parameters 456A for the pose sensor 108A separately from the pose update parameters 456B for the pose sensor 108B. For example, a user's perception of the sound field may change faster due to head rotation than due to translation while sitting or walking. Therefore, it may be desirable to set a higher update rate for the pose data 110 indicating rotation than for the pose data 110 indicating translation.

[0118] FIG. 8 Depicting what can be FIG. 1 and FIG. 4 to FIG. 7 An example of operation 800 implemented in any of the immersive audio renderers 122. FIG. 8 In this context, the operations are divided between rendering operation 820 and blending and binarization operation 822.

[0119] In a particular aspect, the mixing and binauralizing operation 822 can be performed by a mixer and a binauralizer 814, which includes, corresponding to FIG. 1 andFIG. 4 to FIG. 7 The binaural unit 128 of any of them may be included therein. FIG. 8 In this process, rendering operation 820 can be performed by one or more of the following modules: preprocessing module 802, positioning preprocessing module 804, spatial analysis module 806, spatial metadata interpolation module 808, and signal interpolation module 810. In a specific implementation, operation 800 generates an output audio signal 180 based on the processing of assets representing the immersive audio environment using an environmental stereo representation. FIG. 8 The middle corresponds to the binaural output signal s out ( j ).

[0120] When the assets for rendering are received, the preprocessing module 802 is configured to receive Head-Related Impulse Response Information (HRIR) and audio source location information. i (where bold text indicates vectors, and where...) i This refers to the audio source index, such as the (x, y, z) coordinates of the location of each audio source in the audio scene. The preprocessing module 802 is configured to generate a representation of the audio source locations and an HRTF, as a set of triangles with an audio source at each triangle vertex. T 1..NT (in N T (Indicates the number of triangles).

[0121] The positioning preprocessing module 804 is configured to receive a representation of the location of the audio source. T 1..NT Audio source location information p i Frames indicating the audio data to be rendered j Listener location information p L ( j (e.g., x, y, z coordinates). The positioning preprocessing module 804 is configured to generate an indication of the listener's position relative to the audio source, such as an active triangle in a triangle set including the listener's position. T A ( j Audio source selection indicator m C ( j (e.g., an index of the selected source used for signal interpolation (e.g., a high-order ambient stereo (HOA) source)); and spatial metadata interpolation weights. (For example, frames) j subframe k (Selected spatial metadata interpolation weights).

[0122] The spatial analysis module 806 receives the audio signal from the audio stream, such as s ESD ( i, j As illustrated (e.g., each source) i and frame j The equivalent spatial domain representation of the signal), and also receives the active triangle including the listener. T A ( j The spatial analysis module 806 can convert the input audio signal into HOA format and generate orientation information of the HOA source (e.g., Representing a frame j subframe k HOA source i and frequency range b The azimuth parameters, and (representing altitude parameters) and energy information (e.g., This represents the parameter directly related to the overall energy ratio, and (Represents energy values). The spatial analysis module 806 also generates a frequency domain representation of the input audio, such as representing the HOA source. i time-frequency domain signal .

[0123] Spatial metadata interpolation module 808 is based on source orientation information. i Listener orientation information L ( j The spatial metadata interpolation module 808 performs spatial metadata interpolation using HOA source orientation and energy information from the spatial analysis module 806, and spatial metadata interpolation weights from the positioning preprocessing module 804. The spatial metadata interpolation module 808 generates energy and orientation information, including frequency band representations. b HOA source i and audio frames j Average (on a subframe) energy , indicates a frame j and frequency range b HOA source i The azimuth parameters , indicates a frame j and frequency range b HOA source i altitude parameters and representing frames j and frequency range b HOA source i The direct relationship between the total energy ratio parameter and the total energy ratio parameter .

[0124] Signal interpolation module 810 receives energy information (e.g., from spatial metadata interpolation module 808) Energy information from the Space Analysis Module 806 (e.g.) ) and the frequency domain representation of the input audio (e.g. ) and audio source selection indication from positioning preprocessing module 804. m C ( j The signal interpolation module 810 generates the interpolated audio signal. The completion of rendering operation 820 generates source orientation information. i Interpolated audio signals from signal interpolation module 810 and spatial metadata interpolation module 808 Rendered assets corresponding to interpolation orientation and energy parameters (e.g., FIG. 1 and FIG. 4 to FIG. 7 Rendered assets of any one of them (126).

[0125] The mixer and binaural amplifier 814 receive source orientation information. i Listener orientation information L ( j ), HRTF, and interpolated audio signals from signal interpolation module 810 and spatial metadata interpolation module 808, respectively. And interpolation orientation and energy parameters. When the asset is a pre-rendered asset 824, the mixer and binauralizer 814 receive source orientation information. i HRTF and interpolated audio signals Interpolation orientation and energy parameters are included as part of the pre-rendered asset 824. Optionally, if the listener pose associated with the pre-rendered asset 824 is specified in advance, the pre-rendered asset 824 also includes listener orientation information. L ( j Alternatively, if the listener pose associated with the pre-rendered asset 824 is not specified in advance, the pre-rendered asset 824 receives listener orientation information based on the listener pose. L ( j ).

[0126] The mixer and binauralizer 814 are configured to: apply one or more rotation operations based on the orientation of each interpolated signal and the listener orientation; binauralize the signal using HRTF; combine the signals (e.g., after binauralization) if multiple interpolated signals are received; perform one or more other operations; or any combination thereof, to generate an output audio signal 180.

[0127] FIG. 9 This is a block diagram illustrating a specific implementation 900 of integrated circuit 902. Integrated circuit 902 includes one or more processors 920, such as one or more processors 410. The one or more processors 920 include an immersive audio component 922. FIG. 9In this embodiment, the immersive audio component 922 includes an immersive audio renderer 122 and a pose anomaly detector and mitigator 150. Optionally, the immersive audio component 922 may include a pose predictor 450, a motion estimator 460, a client 120, a decoder 121, an asset position selector 130, or a combination thereof. Furthermore, the immersive audio renderer 122 or the pose anomaly detector and mitigator 150 may include, be included in, or be coupled to the audio asset selector 124, the pose predictor 450, the motion estimator 460, or a combination thereof. In some specific embodiments, the integrated circuit 902 also includes one or more pose sensors 108.

[0128] Integrated circuit 902 also includes signal input 904, such as a bus interface and / or modem 420, to enable processor 920 to receive input data 906, such as target assets (e.g., local asset 142 or remote asset 144), pose data 110, historical interaction data 458, contextual cues 454, contextual motion estimation data 462, pose update parameters 456, asset list 134, audio assets 132, or combinations thereof. Integrated circuit 902 also includes signal output 912, such as one or more bus interfaces and / or modems 420, to enable processor 920 to provide output data 914 to one or more other devices. For example, output data 914 may include output audio signal 180, pose update parameters 456, asset retrieval request 138, asset request 136, or combinations thereof.

[0129] Integrated circuit 902 enables immersive audio processing to be implemented as a component in one of a variety of devices, such as... FIG. 10 The described loudspeaker array, such as FIG. 11 The described mobile devices, such as FIG. 12 The described headphone device, such as FIG. 13 The described in-ear headphones, such as FIG. 14 The described augmented reality glasses, such as FIG. 15 The depicted extended reality headset, or as FIG. 16 or FIG. 17 The vehicles depicted.

[0130] FIG. 10This is a block diagram illustrating a specific implementation of a system 1000 for immersive audio processing, wherein an immersive audio component 922 is integrated within a speaker array (such as a soundbar device 1002). The soundbar device 1002 is configured to perform beamforming operations to direct binaural signals to a location associated with a user. The soundbar device 1002 may receive audio assets 132 (e.g., an ambient stereo representation of an immersive audio environment) from a remote streaming server via a wireless network 1006. The soundbar device 1002 may include... FIG. 9 One or more processors 920 (e.g., the one or more processors include an immersive audio renderer 122, a pose anomaly detector, and a mitigator 150, or both). Optionally, in FIG. 10 In the soundbar device 1002, pose sensor 108 is included or coupled to generate a sound field for rendering and binauralizing one or more assets to generate an immersive audio environment and pose data 110 for outputting binaural audio using beam manipulation operations.

[0131] The soundbar device 1002 includes or is coupled to a pose sensor 108 (e.g., a camera, structured light sensor, ultrasound, lidar, etc.) to enable the detection of the pose of the listener 1020 and the generation of head tracker data for the listener 1020. For example, the soundbar device 1002 may detect the pose of the listener 1020 at a first position 1022 (e.g., at a first angle to a reference 1024), adjust the sound field based on the listener 1020's pose, and perform beam manipulation operations to make the emitted sound 1004 perceived by the listener 1020 as a pose-adjusted binaural signal. In this example, the beam manipulation operation is based on the listener 1020's first position 1022 and a first orientation (e.g., facing the soundbar device 1002). In response to a change in the pose of the listener 1020, such as the listener 1020 moving to the second position 1032, the soundbar device 1002 adjusts the sound field (e.g., according to 3DOF / 3DOF+ or 6DOF operation) and performs beam manipulation operations so that the resulting emitted sound 1004 is perceived by the listener 1020 as a binaural signal of pose adjustment at the second position 1032.

[0132] FIG. 11 A specific implementation 1100 in which a mobile device 1102 is configured to perform immersive audio processing is depicted. FIG. 11 In this context, as a non-limiting example, mobile device 1102 may include a telephone or a tablet. FIG. 11In the example shown, mobile device 1102 includes a microphone 1104, multiple speakers 104, and a display 106. An immersive audio component 922 and optionally one or more pose sensors 108 are integrated into mobile device 1102 and are illustrated using dashed lines to indicate internal components of mobile device 1102 that are typically not visible to the user. In a particular example, immersive audio component 922 includes an immersive audio renderer 122, a pose anomaly detector and mitigator 150, and optionally a reference... FIG. 1 to FIG. 10 Other components described. For example, in some specific implementations, the mobile device 1102 is configured to perform the reference. FIG. 4 to FIG. 7 The operation described in the immersive audio player 402 of any of the following is illustrated. For example, the mobile device 1102 may acquire pose data of a listener in an immersive audio environment, determine the current listener pose based on the pose data and one or more pose constraints, acquire a rendering asset associated with the immersive audio environment based on the current listener pose, and generate an output audio signal based on the rendering asset. In such a specific implementation, the output audio signal may be provided to another device, such as... FIG. 10 1002 soundbar equipment FIG. 12 Headphones FIG. 13 Earbuds FIG. 14 Extended reality glasses FIG. 15 Extended reality headsets, or FIG. 16 or FIG. 17 A speaker in a vehicle. The mobile device 1102 obtains pose data for rendering assets from the pose sensor 108 of the mobile device 1102, from the pose sensor of another device, or from a combination of pose data from the pose sensor 108 of the mobile device 1102 and pose data from another device.

[0133] FIG. 12 A specific implementation 1200 in which a headset device 1202 is configured to perform immersive audio processing is depicted. FIG. 12 In the example shown, the headset device 1202 includes a speaker 104 and optionally a microphone 1204. An immersive audio component 922 and optionally one or more pose sensors 108 are integrated into the headset device 1202. In a particular example, the immersive audio component 922 includes an immersive audio renderer 122, a pose anomaly detector and mitigator 150, and optionally a reference... FIG. 1 to FIG. 10 Other components described. For example, in some specific implementations, the headset device 1202 is configured to perform the reference. FIG. 4 to FIG. 7The operation described in the immersive audio player 402 of any of the above. For example, the headphone device 1202 may acquire pose data of a listener in an immersive audio environment, determine the current listener pose based on the pose data and one or more pose constraints, acquire a rendering asset associated with the immersive audio environment based on the current listener pose, and generate an output audio signal based on the rendering asset.

[0134] In some specific implementations, the headset device 1202 is configured to perform reference FIG. 1 and FIG. 4 to FIG. 7 The operation described for any of the media output devices 102. For example, a headset device 1202 may generate pose data 110 and receive an output audio signal 180 representing immersive audio content rendered based on the pose data 110 after pose anomaly detection and mitigation have been performed to ensure that one or more human movement constraints and / or one or more spatial constraints are not violated. In this example, the headset device 1202 may output sound based on the output audio signal 180.

[0135] FIG. 13 A specific embodiment 1300 is depicted in which a pair of earbud headphones 1306 (including a first earbud headphone 1302 and a second earbud headphone 1304) are configured to perform immersive audio processing. Although earbud headphones are described, it should be understood that the technology disclosed herein can be applied to other in-ear or over-ear playback devices.

[0136] exist FIG. 13 In the illustrated example, the first earbud 1302 includes a first microphone 1320 (such as a high signal-to-noise ratio microphone positioned to capture the speech of the wearer of the first earbud 1302), an array of one or more other microphones configured to detect ambient sound and spatially distributed to support beamforming (illustrated as microphones 1322A, 1322B, and 1322C), an "inner" microphone 1324 located near the wearer's ear canal (e.g., to assist in active noise cancellation), and a self-speech microphone 1326 (such as a bone conduction microphone configured to convert sound vibrations from the wearer's ear bones or skull into audio signals). The second earbud 1304 may be configured in a substantially similar manner to the first earbud 1302.

[0137] An immersive audio component 922 and optionally one or more pose sensors 108 are integrated into at least one of the earbuds 1306 (e.g., in a first earbud 1302, a second earbud 1304, or both). In a particular example, the immersive audio component 922 includes an immersive audio renderer 122, a pose anomaly detector and mitigator 150, and optionally a reference... FIG. 1 to FIG. 10Other components described. For example, in some specific implementations, the earbud headphone 1306 is configured to perform the reference. FIG. 4 to FIG. 7 The operation described in the immersive audio player 402 of any of the above. For example, the earphone 1306 can acquire pose data of a listener in an immersive audio environment, determine the current listener pose based on the pose data and one or more pose constraints, acquire a rendering asset associated with the immersive audio environment based on the current listener pose, and generate an output audio signal based on the rendering asset.

[0138] In some specific implementations, the earbud-type headphone 1306 is configured as the execution reference. FIG. 1 or FIG. 4 to FIG. 7 The operation described for any of the media output devices 102. For example, an earphone device 1306 may generate pose data 110 and receive an output audio signal 180 representing immersive audio content rendered based on the pose data 110 after pose anomaly detection and mitigation have been performed to ensure that one or more human movement constraints and / or one or more spatial constraints are not violated. In this example, the earphone 1306 may output sound via a speaker 104 based on the output audio signal 180.

[0139] FIG. 14 A specific implementation 1400 is depicted in which extended reality (e.g., augmented reality or mixed reality) glasses 1402 are configured to perform immersive audio processing. Glasses 1402 includes a holographic projection unit 1404 configured to project visual data onto the surface of a lens 1406, or to reflect the visual data from the surface of the lens 1406 onto the wearer's retina. An immersive audio component 922 and optionally one or more pose sensors 108 are integrated into glasses 1402. In a particular example, the immersive audio component 922 includes an immersive audio renderer 122, a pose anomaly detector and mitigator 150, and optionally a reference... FIG. 1 to FIG. 10 Other components described. For example, in some specific implementations, glasses 1402 are configured to perform reference... FIG. 4 to FIG. 7 The operation described in the immersive audio player 402 of any of the following is illustrated: For example, glasses 1402 acquires pose data of a listener in the immersive audio environment, determines the current listener pose based on the pose data and one or more pose constraints, acquires a rendering asset associated with the immersive audio environment based on the current listener pose, and generates an output audio signal based on the rendering asset.

[0140] In some specific implementations, glasses 1402 are configured to perform reference. FIG. 1 or FIG. 4 to FIG. 7The operation described for any of the media output devices 102. For example, glasses 1402 may generate pose data 110 and receive an output audio signal 180 representing immersive audio content rendered based on the pose data 110 after pose anomaly detection and mitigation have been performed to ensure that one or more human movement constraints and / or one or more spatial constraints are not violated. In this example, glasses 1402 may output sound via speaker 104 based on the output audio signal 180.

[0141] FIG. 15 A specific implementation 1500 is depicted in which an extended reality (e.g., virtual reality, mixed reality, or augmented reality) headset 1502 is configured to perform immersive audio processing. FIG. 15 In the example shown, the headset 1502 includes a speaker 104 and a display 106. An immersive audio component 922 and optionally one or more pose sensors 108 are integrated into the headset 1502. In a particular example, the immersive audio component 922 includes an immersive audio renderer 122, a pose anomaly detector and mitigator 150, and optionally a reference... FIG. 1 to FIG. 10 Other components described. For example, in some specific implementations, the headset 1502 is configured to perform the reference. FIG. 4 to FIG. 7 The operation described in the immersive audio player 402 of any of the above. For example, the headset 1502 may acquire pose data of a listener in an immersive audio environment, determine the current listener pose based on the pose data and one or more pose constraints, acquire a rendering asset associated with the immersive audio environment based on the current listener pose, and generate an output audio signal based on the rendering asset.

[0142] In some specific implementations, the headset 1502 is configured to perform the reference. FIG. 1 or FIG. 4 to FIG. 7 The operation described for any of the media output devices 102. For example, a headset 1502 may generate pose data 110 and receive an output audio signal 180 representing immersive audio content rendered based on the pose data 110 after pose anomaly detection and mitigation have been performed to ensure that one or more human movement constraints and / or one or more spatial constraints are not violated. In this example, the headset 1502 may output sound based on the output audio signal 180.

[0143] FIG. 16 Another specific implementation 1600, in which vehicle 1602 is configured to perform immersive audio processing, is depicted. FIG. 16In this example, vehicle 1602 is exemplified as a car. An immersive audio component 922 is integrated into vehicle 1602. In a particular example, immersive audio component 922 includes an immersive audio renderer 122, a pose anomaly detector and mitigator 150, and optionally a reference... FIG. 1 to FIG. 10 Other components described. For example, in some specific implementations, vehicle 1602 is configured to perform the reference. FIG. 4 to FIG. 7 The operation described in the immersive audio player 402 of any of the following is illustrated. For example, vehicle 1602 may acquire pose data of a listener in an immersive audio environment, determine the current listener pose based on the pose data and one or more pose constraints, acquire a rendering asset associated with the immersive audio environment based on the current listener pose, and generate an output audio signal based on the rendering asset.

[0144] In some specific implementations, vehicle 1602 is configured as a reference for execution. FIG. 1 or FIG. 4 to FIG. 7 The operation described by media output device 102 of any of the above. For example, vehicle 1602 may generate pose data 110 and receive an output audio signal 180 representing immersive audio content rendered based on pose data 110 after pose anomaly detection and mitigation have been performed to ensure that one or more human movement constraints and / or one or more spatial constraints are not violated. In this example, vehicle 1602 may output sound based on the output audio signal 180 via a set of speakers.

[0145] FIG. 17 The illustration depicts a specific implementation 1700 in which a vehicle 1702 is configured to perform immersive audio processing. FIG. 17 In this context, vehicle 1702 is exemplified as an unmanned aerial vehicle, such as a personal drone or a package delivery drone. An immersive audio component 922 is integrated into vehicle 1702. In a specific example, the immersive audio component 922 includes an immersive audio renderer 122, a pose anomaly detector and mitigator 150, and optionally a reference... FIG. 1 to FIG. 10 Other components described. For example, in some specific implementations, vehicle 1702 is configured to perform the reference. FIG. 4 to FIG. 7 The operation described in the immersive audio player 402 of any of the following is illustrated. For example, vehicle 1702 may acquire pose data of a listener in an immersive audio environment, determine the current listener pose based on the pose data and one or more pose constraints, acquire a rendering asset associated with the immersive audio environment based on the current listener pose, and generate an output audio signal based on the rendering asset.

[0146] In some specific implementations, vehicle 1702 is configured as a reference for execution. FIG. 1 or FIG. 4 to FIG. 7The operation described by media output device 102 of any of the above. For example, vehicle 1702 may generate pose data 110 and receive an output audio signal 180 representing immersive audio content rendered based on pose data 110 after pose anomaly detection and mitigation have been performed to ensure that one or more human movement constraints and / or one or more spatial constraints are not violated. In this example, vehicle 1702 may output sound via speaker 104 based on output audio signal 180.

[0147] In some specific implementations, FIG. 16 , FIG. 17 One or both of the vehicles are implemented as or correspond to robotic devices, such as radio-controlled (RC) or autonomous flying devices (e.g., amateur quadcopters), land devices (e.g., RC cars or robotic vacuum cleaners), or water devices (e.g., RC boats or submarines). Such robotic devices may include one or more cameras for capturing a visual scene corresponding to the device's environment, a microphone for capturing an audio scene corresponding to the device's environment, or both. The device may also include one or more pose sensors 108 for tracking the device's pose (position and orientation), and one or more wireless transceivers for transmitting captured audio content, video content, or both to a remote user and optionally enabling control signals to be received from the remote user. Thus, a remote user can experience an immersive spatial audio environment and optionally a visual environment of the device, such as via a VR headset, as if the remote user were at the device's location and oriented according to the device's pose. In some implementations, the camera, microphone, or both are mounted to an adjustable component of the device with one or more degrees of freedom to accommodate changes in the camera and / or microphone's rotational orientation, tilt angle, or both relative to the device body. The adjustable component thus enables the orientation associated with audio / video capture to be adjusted by a remote user via remote control of the adjustable component or via autonomous control of the device, without compromising the body orientation associated with the device's movement.

[0148] refer to FIG. 18 This illustrates a specific implementation of a method 1800 for processing immersive audio data. In a particular aspect, one or more operations of method 1800 are performed by... FIG. 1 One or more components of system 100 or FIG. 4 to FIG. 7 It can be executed on any of the systems from 400 to 700.

[0149] Method 1800 includes, at block 1802, obtaining pose data of a listener in an immersive audio environment at one or more processors. The pose data may be received from one or more pose sensors, such as pose sensor 108. In some embodiments, the pose data includes first pose data associated with the listener's head and second pose data associated with at least one of the listener's torso or the listener's hands.

[0150] Method 1800 includes, at block 1804, determining the current listener pose at one or more processors based on pose data and one or more pose constraints. For example, pose outlier detector and mitigator 150 determines the current pose 154 based on pose data 110 and pose constraints 158.

[0151] Method 1800 includes, at block 1806, at one or more processors and based on the current listener pose, obtaining a rendered asset associated with an immersive audio environment. For example, an immersive audio renderer 122 may perform rendering operations to generate rendered assets based on local assets 142 or remote assets 144. For illustration, the rendering operation may include referencing FIG. 8 One or more rendering operations in the described rendering operation 820. In some specific implementations, local asset 142 or remote asset 144 may include pre-rendered assets (e.g., one of the pre-rendered assets in pre-rendered asset 114D). As an example, obtaining a rendered asset may include determining a target asset based on pose data (e.g., predicted pose or current pose) and generating an asset retrieval request to retrieve the target asset from a storage location. The target asset may include a pre-rendered asset associated with a specific listener pose or an asset that has not yet been pre-rendered. For example, when the target asset is a pre-rendered asset, generating an output audio signal may include applying a head-related transfer function to the target asset to generate a binaural output signal. When the target asset has not yet been pre-rendered, obtaining a rendered asset may include rendering the target asset based on pose data to generate a rendered asset, and applying a head-related transfer function to the rendered asset to generate a binaural output signal.

[0152] Method 1800 also includes, at block 1808, generating an output audio signal based on a rendered asset at one or more processors. For example, an immersive audio renderer 122 may generate the output audio signal 180 based on a rendered asset (e.g., rendered asset 126). For illustration, generating the output audio signal 180 may include performing a binauralization operation (e.g., via binauralizer 128) referenced to... FIG. 8 One or more of the described mixing and binauralizing operations 822.

[0153] Depending on some aspects, one or more pose constraints include human movement constraints. For example, human movement constraints may correspond to velocity constraints (such as...) FIG. 2Velocity constraints 216), acceleration constraints (such as acceleration constraint 218), and constraints on the listener's hand or torso posture relative to the listener's head position (such as... FIG. 2 The relative head / body constraints 220 or any combination thereof. According to some aspects, one or more pose constraints include boundary constraints indicating the boundaries associated with the immersive audio environment, such as one or more boundary constraints in boundary constraints 222, and the current listener pose is determined such that the current listener pose is constrained by the boundaries.

[0154] Method 1800 optionally includes obtaining a pose based on pose data and determining whether the pose violates at least one of one or more pose constraints, such as a reference. FIG. 2 The pose anomaly detection operation 200 is described. Method 1800 may include: based on determining that the pose does not violate one or more pose constraints, using the pose as the current listener pose, such as a reference pose. FIG. 3A Pose anomaly reduction operation 300 and FIG. 3B The pose anomaly mitigation operation 350 is described. In some specific implementations, method 1800 includes: determining the current listener pose based on a previous listener pose that does not violate one or more pose constraints, such as a reference, based on determining that the pose violates at least one of one or more pose constraints. FIG. 3A As described in operation 342. Additionally, based on the determination that the pose violates at least one of one or more pose constraints, the predicted listener pose can be determined based on a previously predicted listener pose associated with the previous listener pose, such as a reference. FIG. 3A Operation 344 is described.

[0155] In some specific implementations, method 1800 includes: determining the current listener pose based on determining that the pose violates at least one of one or more pose constraints, such as a reference pose, by adjusting the pose to satisfy one or more pose constraints. FIG. 3B As described in operation 372. For illustration, determining the current listener pose may include adjusting the pose value to match a threshold associated with a violated pose constraint, such as one or more thresholds in threshold 306. Based on determining that the pose violates at least one of one or more pose constraints, a predicted listener pose may be determined based on a previously predicted listener pose associated with a previous listener pose, and determining the predicted listener pose based on determining that the previously predicted listener pose violates at least one of one or more pose constraints may include adjusting the previously predicted listener pose to match a threshold associated with the violated pose constraint, as described with reference to operation 374.

[0156] In some implementations, the pose data includes first data indicating the listener's translational positioning within the immersive audio environment and second data indicating the listener's rotational orientation within the immersive audio environment. In such implementations, the first and second data may be received from the same device, from different devices, or a combination thereof. For example, the first data may be received from a first device, and the second data may be received from a second device different from the first device. As another example, the first translational data may be obtained from the first device, the second translational data may be obtained from a second device different from the first device, and the first data indicating the listener's translational positioning within the immersive audio environment may be determined based on the first and second translational data.

[0157] In some implementations, the pose data obtained at block 1802 is associated with a first time, and method 1800 includes determining a predicted listener pose associated with a second time following the first time. In such implementations, at block 1806, the render asset may include at least one render asset among those associated with the predicted listener pose. In some implementations, more than one predicted listener pose may be predicted for a particular time, and each predicted listener pose may be used to obtain a render asset. For example, the pose data obtained at block 1801 is associated with a first time, and method 1800 may include determining two or more predicted listener poses associated with a second time following the first time, obtaining a first render asset associated with a first predicted listener pose, and obtaining a second render asset associated with a second predicted listener pose. In this example, method 1800 may also include selectively generating an output audio signal based on either the first or second render asset. For illustration, selectively generating an output audio signal based on a first rendered asset or a second rendered asset may include: obtaining a first target asset associated with a first predicted listener pose; rendering the first target asset to generate a first rendered asset; obtaining a second target asset associated with a second predicted listener pose; rendering the second target asset to generate a second rendered asset; obtaining pose data associated with a second time; and selecting either the first rendered asset or the second rendered asset based on the pose data associated with the second time for further processing.

[0158] FIG. 18 Method 1800 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, FIG. 18 Method 1800 can be executed by a processor that executes instructions, such as references FIG. 19 As described.

[0159] refer toFIG. 19 A block diagram depicting a specific, exemplary embodiment of the device is provided, and is generally designated as 1900. In various embodiments, device 1900 may have the same... FIG. 19 The illustrated components may be more or fewer than the number of components. In an illustrative embodiment, device 1900 may correspond to one or more media output devices in media output device 102, to immersive audio player 402, or a combination thereof. In an illustrative embodiment, device 1900 may perform reference... FIG. 1 to FIG. 18 One or more operations as described.

[0160] In a particular implementation, device 1900 includes a processor 1906 (e.g., a central processing unit (CPU)). Device 1900 may include one or more additional processors 1910 (e.g., one or more DSPs). In a particular aspect, FIG. 4 to FIG. 7 Processor 410 of any of these corresponds to processor 1906, processor 1910, or a combination thereof. Processor 1910 may include a speech and music encoder-decoder (CODEC) 1908, which includes a speech decoder (“phonetic coder”) encoder 1936, a phonetic coder decoder 1938, an immersive audio component 922, or a combination thereof. Immersive audio component 922 may include, for example, an immersive audio renderer 122 and a pose anomaly detector and mitigator 150. Optionally, immersive audio component 922 may include a pose predictor 450, a motion estimator 460, a client 120, a decoder 121, an asset location selector 130, or a combination thereof. Immersive audio renderer 122, pose anomaly detector and mitigator 150, or motion estimator 460 may include an audio asset selector 124, a pose predictor 450, or both. Optionally, pose sensor 108 may be included within or coupled to device 1900.

[0161] Device 1900 may include memory 1986 and CODEC 1934. Memory 1986 may include instructions 1956 that can be executed by one or more additional processors 1910 (or processor 1906) to implement references. FIG. 1 to FIG. 18 The function described by either of them. In FIG. 19 In addition, device 1900 also includes modem 420 coupled to antenna 1952 via transceiver 1950.

[0162] Device 1900 may include a display 106 coupled to display controller 1926. Speaker 104 and microphone 1994 may be coupled to CODEC 1934. CODEC 1934 may include digital-to-analog converter (DAC) 1902, analog-to-digital converter (ADC) 1904, or both. In a particular embodiment, CODEC 1934 may receive analog signals from microphone 1994, use ADC 1904 to convert the analog signals to digital signals, and provide the digital signals to speech and music codec 1908. Speech and music codec 1908 may process digital signals, and digital signals or other digital signals (e.g., one or more assets associated with an immersive audio environment) may be further processed by immersive audio component 922. In a particular embodiment, speech and music codec 1908 may provide digital signals to CODEC 1934. CODEC 1934 may use ADC 1902 to convert digital signals to analog signals and may provide analog signals to speaker 104.

[0163] In a particular embodiment, device 1900 may be included in a system-in-package (SoC) or a system-on-a-chip (SoC) 1922. In a particular embodiment, memory 1986, processor 1906, processor 1910, display controller 1926, CODEC 1934, and modem 420 are included in a SoC or SoC 1922. In a particular embodiment, pose sensor 108, input device 1930, and power supply 1944 are coupled to the SoC or SoC 1922. Furthermore, in a particular embodiment, such as... FIG. 19 As illustrated, display 106, input device 1930, speaker 104, microphone 1994, pose sensor 108, antenna 1952, and power supply 1944 are external to system-in-package or system-on-chip device 1922. In a particular implementation, each of display 106, input device 1930, speaker 104, microphone 1994, pose sensor 108, antenna 1952, and power supply 1944 may be coupled to components of system-in-package or system-on-chip device 1922, such as interfaces (e.g., signal input 904 or signal output 912) or controllers.

[0164] Device 1900 may include smart speakers, speaker bars, mobile communication devices, smartphones, cellular phones, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, headsets, augmented reality headsets, mixed reality headsets, virtual reality headsets, air vehicles, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.

[0165] In conjunction with the described specific embodiments, an apparatus includes components for acquiring pose data of a listener in an immersive audio environment. For example, the components for acquiring pose data may correspond to a pose sensor 108, a pose anomaly detector and mitigator 150, a motion estimator 460, an immersive audio renderer 122, an audio asset selector 124, a client 120, an immersive audio player 402, a processor 410, a modem 420, a pose predictor 450, a binauralizer 128, a media output device 102, a processor 1906, one or more processors 1910, one or more other circuits or components configured to acquire pose data, or any combination thereof.

[0166] The apparatus includes components for determining the current listener pose based on pose data and one or more pose constraints. For example, the components for obtaining the current listener pose may correspond to a pose outlier detector and mitigator 150, an immersive audio renderer 122, an audio asset selector 124, an asset location selector 130, a client 120, an immersive audio player 402, a processor 410, a modem 420, a motion estimator 460, a media output device 102, a processor 1906, one or more processors 1910, one or more other circuits or components configured to obtain rendered assets, or any combination thereof.

[0167] The apparatus includes components for obtaining rendered assets associated with an immersive audio environment based on the current listener's pose. For example, the components for obtaining rendered assets may correspond to an immersive audio renderer 122, an audio asset selector 124, an asset location selector 130, a client 120, a decoder 121, an immersive audio player 402, a processor 410, a modem 420, a binauralizer 128, a media output device 102, a processor 1906, one or more processors 1910, one or more other circuits or components configured to obtain rendered assets, or any combination thereof.

[0168] The apparatus includes components for generating an output audio signal based on rendered assets. For example, the components for generating the output audio signal may correspond to an immersive audio renderer 122, an immersive audio player 402, a processor 410, a binauralizer 128, a media output device 102, a processor 1906, one or more processors 1910, one or more other circuits or components configured to generate an output audio signal, or any combination thereof.

[0169] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 1986 or local memory 170) includes instructions (e.g., instruction 1956 or instruction 174) that, when executed by one or more processors (e.g., one or more processors 410, one or more processors 1910, or processor 1906), cause the one or more processors to: obtain pose data of a listener in an immersive audio environment; determine the current listener pose based on the pose data and one or more pose constraints; obtain a rendering asset associated with the immersive audio environment based on the current listener pose; and generate an output audio signal based on the rendering asset.

[0170] Specific aspects of this disclosure are described below in a collection of related embodiments:

[0171] According to Embodiment 1, an apparatus includes: a memory configured to store audio data associated with an immersive audio environment; and one or more processors configured to: obtain pose data of a listener in the immersive audio environment; determine a current listener pose based on the pose data and one or more pose constraints; obtain a rendering asset associated with the immersive audio environment based on the current listener pose; and generate an output audio signal based on the rendering asset.

[0172] Example 2 includes the device according to Example 1, wherein the one or more pose constraints include human movement constraints.

[0173] Example 3 includes the device according to Example 2, wherein the human movement constraint corresponds to a speed constraint.

[0174] Example 4 includes the device according to Example 2 or Example 3, wherein the human movement constraint corresponds to the acceleration constraint.

[0175] Example 5 includes the device according to any one of Examples 2 to 4, wherein the human body movement constraint corresponds to the constraint on the hand or torso posture of the listener relative to the head posture of the listener.

[0176] Example 6 includes a device according to any one of Examples 1 to 5, wherein the one or more pose constraints include boundary constraints indicating a boundary associated with the immersive audio environment, and wherein the one or more processors are configured to determine the current listener pose such that the current listener pose is constrained by the boundary.

[0177] Example 7 includes a device according to any one of Examples 1 to 6, wherein the one or more processors are configured to: obtain a pose based on the pose data; and determine whether the pose violates at least one of the one or more pose constraints.

[0178] Example 8 includes the device according to Example 7, wherein the one or more processors are configured to use the pose as the current listener pose based on determining that the pose does not violate the one or more pose constraints.

[0179] Example 9 includes the device according to Example 7 or Example 8, wherein the one or more processors are configured to determine the current listener pose based on a previous listener pose that does not violate the one or more pose constraints, based on determining that the pose violates at least one of the one or more pose constraints.

[0180] Example 10 includes the device according to Example 9, wherein the one or more processors are configured to determine a predicted listener pose based on a previously predicted listener pose associated with the previous listener pose, based on the determination that the pose violates at least one of the one or more pose constraints.

[0181] Example 11 includes the device according to Example 7, wherein the one or more processors are configured to determine the current listener pose based on adjusting the pose to satisfy the one or more pose constraints based on determining that the pose violates at least one of the one or more pose constraints.

[0182] Example 12 includes the device according to Example 11, wherein one or more processors are configured to adjust the value of the pose to match a threshold associated with a violated pose constraint.

[0183] Example 13 includes the device according to Example 12, wherein the one or more processors are configured to determine a predicted listener pose based on a previously predicted listener pose associated with the previous listener pose, based on the determination that the pose violates at least one of the one or more pose constraints.

[0184] Example 14 includes the device according to Example 13, wherein the one or more processors are configured to determine the predicted listener pose based on adjusting the previously predicted listener pose to match a threshold associated with the violated pose constraint, based on determining that the previously predicted listener pose violates at least one of the one or more pose constraints.

[0185] Example 15 includes the device according to any one of Examples 1 to 14, wherein the pose data is received from one or more pose sensors.

[0186] Example 16 includes the device according to any one of Examples 1 to 15, wherein the pose data includes first pose data associated with the listener’s head and second pose data associated with at least one of the listener’s torso or the listener’s hands.

[0187] Example 17 includes the device according to Example 16, wherein the first pose data includes head translation data, head rotation data, or both, and wherein the second pose data includes body translation data, body rotation data, or both.

[0188] Example 18 includes the device according to Example 16 or Example 17, wherein the first pose data is obtained from the first device, and wherein the second pose data is received from the second device, which is different from the first device.

[0189] Example 19 includes a device according to any one of Examples 1 to 18, wherein, in order to obtain the rendered asset, the one or more processors are configured to: determine a target asset based on the pose data; and generate an asset retrieval request to retrieve the target asset from a storage location.

[0190] Example 20 includes the device according to Example 19, wherein the memory includes the storage location.

[0191] Example 21 includes the device according to Example 19, wherein the storage location is at a remote device.

[0192] Example 22 includes a device according to any one of Examples 19 to 21, wherein the target asset is a pre-rendered asset, and wherein, in order to generate the output audio signal, the one or more processors are configured to apply a head-related transfer function to the target asset to generate a binaural output signal.

[0193] Example 23 includes the device according to any one of Examples 19 to 21, wherein, in order to obtain the rendered asset, the one or more processors are configured to render the target asset based on the current listener pose, and wherein, in order to generate the output audio signal, the one or more processors are configured to apply a head-related transfer function to the rendered asset to generate a binaural output signal.

[0194] Example 24 includes the device according to any one of Examples 1 to 23, and the device further includes a pose sensor coupled to the one or more processors, and wherein the pose sensor is configured to provide at least a portion of the pose data.

[0195] Example 25 includes the device according to Example 24, wherein the pose sensor and the one or more processors are integrated within a head-mounted wearable device.

[0196] Example 26 includes a device according to any one of Examples 1 to 25, wherein one or more processors are integrated within an immersive audio player device.

[0197] Example 27 includes a device according to any one of Examples 1 to 26, and the device further includes a modem coupled to the one or more processors and configured to receive the pose data from a device including a pose sensor.

[0198] According to embodiment 28, a method includes: obtaining pose data of a listener in an immersive audio environment at one or more processors; determining a current listener pose at the one or more processors based on the pose data and one or more pose constraints; obtaining a rendering asset associated with the immersive audio environment at the one or more processors and based on the current listener pose; and generating an output audio signal at the one or more processors based on the rendering asset.

[0199] Example 29 includes the method according to Example 28, wherein the one or more pose constraints include human movement constraints.

[0200] Example 30 includes the method according to Example 29, wherein the human movement constraint corresponds to a velocity constraint.

[0201] Example 31 includes the method according to Example 29 or Example 30, wherein the human movement constraint corresponds to an acceleration constraint.

[0202] Example 32 includes the method according to any one of Examples 29 to 31, wherein the human body movement constraint corresponds to the constraint on the hand or torso pose of the listener relative to the head pose of the listener.

[0203] Example 33 includes the method according to any one of Examples 28 to 32, wherein the one or more pose constraints include boundary constraints indicating a boundary associated with the immersive audio environment, and wherein the current listener pose is determined such that the current listener pose is constrained by the boundary.

[0204] Example 34 includes the method according to any one of Examples 28 to 33, and the method further includes: obtaining a pose based on the pose data; and determining whether the pose violates at least one of the one or more pose constraints.

[0205] Example 35 includes the method according to Example 34, and the method further includes: using the pose as the current listener pose based on determining that the pose does not violate one or more pose constraints.

[0206] Example 36 includes the method according to Example 34, and the method further includes: determining the current listener pose based on a previous listener pose that does not violate the one or more pose constraints, based on determining that the pose violates at least one of the one or more pose constraints.

[0207] Example 37 includes the method according to Example 36, and the method further includes: determining a predicted listener pose based on a previously predicted listener pose associated with the previous listener pose, based on the determination that the pose violates at least one of the one or more pose constraints.

[0208] Example 38 includes the method according to Example 34, and the method further includes: determining the current listener pose based on adjusting the pose to satisfy the one or more pose constraints, based on determining that the pose violates at least one of the one or more pose constraints.

[0209] Example 39 includes the method according to Example 38, wherein determining the current listener pose includes adjusting the value of the pose to match a threshold associated with a violated pose constraint.

[0210] Example 40 includes the method according to Example 38 or Example 39, and the method further includes: determining a predicted listener pose based on a previously predicted listener pose associated with a previous listener pose, based on the determination that the pose violates at least one of the one or more pose constraints.

[0211] Example 41 includes the method according to Example 40, wherein determining the predicted listener pose based on determining that the previously predicted listener pose violates at least one of the one or more pose constraints includes adjusting the previously predicted listener pose to match a threshold associated with the violated pose constraint.

[0212] Example 42 includes the method according to any one of Examples 28 to 41, wherein the pose data is received from one or more pose sensors.

[0213] Example 43 includes the method according to any one of Examples 28 to 42, wherein the pose data includes first pose data associated with the listener’s head and second pose data associated with at least one of the listener’s torso or the listener’s hands.

[0214] Example 44 includes the method according to Example 43, wherein the first pose data includes head translation data, head rotation data, or both, and wherein the second pose data includes body translation data, body rotation data, or both.

[0215] Example 45 includes the method according to Example 43 or Example 44, wherein the first pose data is obtained from a first device, and wherein the second pose data is received from a second device different from the first device.

[0216] Example 46 includes the method according to any one of Examples 28 to 41, wherein obtaining the rendered asset includes: determining a target asset based on the pose data; and generating an asset retrieval request to retrieve the target asset from a storage location.

[0217] Example 47 includes the method according to Example 46, wherein the storage location is at local memory.

[0218] Example 48 includes the device according to Example 46, wherein the storage location is at a remote device.

[0219] Example 49 includes the method according to any one of Examples 46 to 48, wherein the target asset is a pre-rendered asset, and wherein generating the output audio signal includes applying a head-related transfer function to the target asset to generate a binaural output signal.

[0220] Example 50 includes the method according to any one of Examples 46 to 48, wherein obtaining the rendered asset further includes rendering the target asset based on the current listener pose, and wherein generating the output audio signal includes applying a head-related transfer function to the rendered asset to generate a binaural output signal.

[0221] According to Embodiment 51, a non-transitory computer-readable device stores instructions that can be executed by one or more processors to cause the one or more processors to: obtain pose data of a listener in an immersive audio environment; determine a current listener pose based on the pose data and one or more pose constraints; obtain a rendering asset associated with the immersive audio environment based on the current listener pose; and generate an output audio signal based on the rendering asset.

[0222] Example 52 includes a non-transitory computer-readable device according to Example 51, wherein the one or more pose constraints include human movement constraints.

[0223] Example 53 includes a non-transitory computer-readable device according to Example 52, wherein the human movement constraint corresponds to a velocity constraint.

[0224] Example 54 includes a non-transitory computer-readable device according to Example 52 or Example 53, wherein the human movement constraint corresponds to an acceleration constraint.

[0225] Example 55 includes a non-transitory computer-readable device according to any one of Examples 52 to 54, wherein the human body movement constraint corresponds to a constraint on the hand or torso pose of the listener relative to the head pose of the listener.

[0226] Example 56 includes a non-transitory computer-readable device according to any one of Examples 51 to 55, wherein the one or more pose constraints include boundary constraints indicating a boundary associated with the immersive audio environment, and wherein the current listener pose is determined such that the current listener pose is constrained by the boundary.

[0227] Example 57 includes a non-transitory computer-readable device according to any one of Examples 51 to 56, wherein the instructions cause the one or more processors to: obtain a pose based on the pose data; and determine whether the pose violates at least one of the one or more pose constraints.

[0228] Example 58 includes a non-transitory computer-readable device according to Example 57, wherein, based on determining that the pose does not violate one or more pose constraints, the instructions cause the one or more processors to use the pose as the current listener pose.

[0229] Example 59 includes a non-transitory computer-readable device according to Example 57 or Example 58, wherein, based on determining that the pose violates at least one of the one or more pose constraints, the instructions cause the one or more processors to determine the current listener pose based on a previous listener pose that does not violate the one or more pose constraints.

[0230] Example 60 includes a non-transitory computer-readable device according to Example 59, wherein, based on the determination that the pose violates at least one of the one or more pose constraints, the instructions cause the one or more processors to determine a predicted listener pose based on a previously predicted listener pose associated with the previous listener pose.

[0231] Example 61 includes a non-transitory computer-readable device according to Example 57, wherein, based on determining that the pose violates at least one of the one or more pose constraints, the instructions cause the one or more processors to determine the current listener pose based on adjusting the pose to satisfy the one or more pose constraints.

[0232] Example 62 includes a non-transitory computer-readable device according to Example 61, wherein, in order to determine the current listener pose, the instructions cause the one or more processors to adjust the value of the pose to match a threshold associated with a violated pose constraint.

[0233] Example 63 includes a non-transitory computer-readable device according to Example 61 or Example 62, wherein, based on the determination that the pose violates at least one of the one or more pose constraints, the instructions cause the one or more processors to determine a predicted listener pose based on a previously predicted listener pose associated with a previous listener pose.

[0234] Example 64 includes a non-transitory computer-readable device according to Example 63, wherein, based on determining that the previously predicted listener pose violates at least one of the one or more pose constraints, the instructions cause the one or more processors to determine the predicted listener pose based on adjusting the previously predicted listener pose to match a threshold associated with the violated pose constraint.

[0235] Example 65 includes a non-transitory computer-readable device according to any one of Examples 51 to 64, wherein the pose data is received from one or more pose sensors.

[0236] Example 66 includes a non-transitory computer-readable device according to any one of Examples 51 to 65, wherein the pose data includes first pose data associated with the head of the listener and second pose data associated with at least one of the torso of the listener or the hand of the listener.

[0237] Example 67 includes a non-transitory computer-readable device according to Example 66, wherein the first pose data includes head translation data, head rotation data, or both, and wherein the second pose data includes body translation data, body rotation data, or both.

[0238] Example 68 includes a non-transitory computer-readable device according to Example 66 or Example 67, wherein the first pose data is obtained from a first device, and wherein the second pose data is received from a second device different from the first device.

[0239] Example 69 includes a non-transitory computer-readable device according to any one of Examples 51 to 68, wherein, in order to obtain the rendered asset, the instructions cause the one or more processors to: determine a target asset based on the pose data; and generate an asset retrieval request to retrieve the target asset from a storage location.

[0240] Example 70 includes a non-transitory computer-readable device according to Example 69, wherein the storage location is at local memory.

[0241] Example 71 includes a non-transitory computer-readable device according to Example 69, wherein the storage location is at a remote device.

[0242] Example 72 includes a non-transitory computer-readable device according to any one of Examples 69 to 71, wherein the target asset is a pre-rendered asset, and wherein, in order to generate the output audio signal, the instructions cause the one or more processors to apply a head-related transfer function to the target asset to generate a binaural output signal.

[0243] Example 73 includes a non-transitory computer-readable device according to any one of Examples 69 to 71, wherein the instructions cause the one or more processors to: render the target asset based on the current listener pose to generate a rendered asset; and apply a head-related transfer function to the rendered asset to generate a binaural output signal.

[0244] According to embodiment 74, an apparatus includes: components for obtaining pose data of a listener in an immersive audio environment; components for determining a current listener pose based on the pose data and one or more pose constraints; components for obtaining a rendering asset associated with the immersive audio environment based on the current listener pose; and components for generating an output audio signal based on the rendering asset.

[0245] Example 75 includes the apparatus according to Example 74, wherein the one or more pose constraints include human movement constraints.

[0246] Example 76 includes the apparatus according to Example 75, wherein the human movement constraint corresponds to a speed constraint.

[0247] Example 77 includes the apparatus according to Example 75 or Example 76, wherein the human movement constraint corresponds to an acceleration constraint.

[0248] Example 78 includes the apparatus according to any one of Examples 75 to 77, wherein the human body movement constraint corresponds to the constraint on the position of the listener's hands or torso relative to the head position of the listener.

[0249] Example 79 includes an apparatus according to any one of Examples 74 to 78, wherein the one or more pose constraints include boundary constraints indicating a boundary associated with the immersive audio environment, and wherein the current listener pose is determined such that the current listener pose is constrained by the boundary.

[0250] Example 80 includes the apparatus according to any one of Examples 74 to 78, and the apparatus further includes: a component for obtaining a pose based on the pose data; and a component for determining whether the pose violates at least one of the one or more pose constraints.

[0251] Example 81 includes the apparatus according to Example 80, and the apparatus further includes a component for using the pose as the current listener pose based on determining that the pose does not violate the one or more pose constraints.

[0252] Example 82 includes the apparatus according to Example 80 or Example 81, and the apparatus further includes a component for determining the current listener pose based on a previous listener pose that does not violate the one or more pose constraints.

[0253] Example 83 includes the apparatus according to Example 82, and the apparatus further includes a component for determining a predicted listener pose based on a previously predicted listener pose associated with the previous listener pose.

[0254] Example 84 includes the apparatus according to Example 80, and the apparatus further includes components for determining the current listener pose based on adjusting the pose to satisfy one or more pose constraints.

[0255] Example 85 includes the apparatus according to Example 84, wherein the current listener pose is adjusted based on the value of the pose to match a threshold associated with a violated pose constraint.

[0256] Example 86 includes the apparatus according to Example 84 or Example 85, and the apparatus further includes a component for determining a predicted listener pose based on a previously predicted listener pose associated with a previous listener pose.

[0257] Example 87 includes the apparatus according to Example 86, wherein the component for determining the predicted listener pose includes a component for adjusting the previously predicted listener pose to match a threshold associated with a violated pose constraint.

[0258] Example 88 includes the apparatus according to any one of Examples 74 to 87, wherein the pose data is received from one or more pose sensors.

[0259] Example 89 includes the apparatus according to any one of Examples 74 to 88, wherein the pose data includes first pose data associated with the listener’s head and second pose data associated with at least one of the listener’s torso or the listener’s hands.

[0260] Example 90 includes the apparatus according to Example 89, wherein the first pose data includes head translation data, head rotation data, or both, and wherein the second pose data includes body translation data, body rotation data, or both.

[0261] Example 91 includes the apparatus according to Example 89 or Example 90, wherein the first pose data is obtained from a first device, and wherein the second pose data is received from a second device different from the first device.

[0262] Example 92 includes an apparatus according to any one of Examples 74 to 91, wherein the component for obtaining the rendered asset associated with the immersive audio environment includes: a component for determining a target asset based on the pose data; and a component for generating an asset retrieval request to retrieve the target asset from a storage location.

[0263] Example 93 includes the apparatus according to Example 92, wherein the storage location is in local memory.

[0264] Example 94 includes the apparatus according to Example 92, wherein the storage location is at a remote device.

[0265] Example 95 includes an apparatus according to any one of Examples 92 to 94, wherein the target asset is a pre-rendered asset, and wherein the component for generating the output audio signal includes a component for applying a head-related transfer function to the target asset to generate a binaural output signal.

[0266] Example 96 includes an apparatus according to any one of Examples 92 to 94, wherein the component for obtaining the rendered asset associated with the immersive audio environment further includes a component for rendering the target asset based on the current listener pose, and wherein the component for generating the output audio signal includes a component for applying a head-related transfer function to the rendered asset to generate a binaural output signal.

[0267] Those skilled in the art will also understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the specific embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various exemplary components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, and such implementation decisions shall not be construed as departing from the scope of this disclosure.

[0268] The steps of the methods or algorithms described in conjunction with the specific embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compressed optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and storage medium may reside as discrete components in a computing device or a user terminal.

[0269] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but should be granted the broadest scope that may be consistent with the principles and novel features as defined by the following claims.

Claims

1. An apparatus, said apparatus comprising: A memory configured to store audio data associated with an immersive audio environment; and One or more processors, said one or more processors being configured to: Obtain the pose data of the listener in the immersive audio environment; The current listener's pose is determined based on the pose data and one or more pose constraints. Rendering assets associated with the immersive audio environment are obtained based on the current listener pose. as well as The output audio signal is generated based on the rendered assets.

2. The device according to claim 1, wherein the one or more pose constraints include human movement constraints.

3. The device according to claim 2, wherein the human movement constraint corresponds to a speed constraint.

4. The device according to claim 2, wherein the human movement constraint corresponds to an acceleration constraint.

5. The device of claim 2, wherein the human body movement constraint corresponds to the constraint on the position of the listener's hands or torso relative to the head position of the listener.

6. The device of claim 1, wherein the one or more pose constraints include boundary constraints indicating a boundary associated with the immersive audio environment, and wherein the one or more processors are configured to determine the current listener pose such that the current listener pose is constrained by the boundary.

7. The device of claim 1, wherein the one or more processors are configured to: The pose is obtained based on the pose data; and Determine whether the pose violates at least one of the one or more pose constraints.

8. The device of claim 7, wherein the one or more processors are configured to use the pose as the current listener pose based on determining that the pose does not violate the one or more pose constraints.

9. The device of claim 7, wherein the one or more processors are configured to determine the current listener pose based on a previous listener pose that does not violate the one or more pose constraints, based on determining that the pose violates at least one of the one or more pose constraints.

10. The device of claim 9, wherein the one or more processors are configured to determine a predicted listener pose based on a previously predicted listener pose associated with the previous listener pose, based on the determination that the pose violates at least one of the one or more pose constraints.

11. The device of claim 7, wherein the one or more processors are configured to determine the current listener pose based on adjusting the pose to satisfy the one or more pose constraints, based on determining that the pose violates at least one of the one or more pose constraints.

12. The device of claim 1, wherein the pose data includes first pose data associated with the listener's head and second pose data associated with at least one of the listener's torso or the listener's hand.

13. The device of claim 12, wherein the first pose data is obtained from the first device, and wherein the second pose data is received from a second device different from the first device.

14. The device according to claim 1, wherein, In order to obtain the rendered assets, the one or more processors are configured to: The target asset is determined based on the pose data. as well as Generate an asset retrieval request to retrieve the target asset from the storage location.

15. The device of claim 1, further comprising a pose sensor coupled to the one or more processors, wherein the pose sensor is configured to provide at least a portion of the pose data.

16. The device of claim 15, wherein the pose sensor and the one or more processors are integrated within a head-mounted wearable device.

17. The device of claim 1, wherein the one or more processors are integrated within the immersive audio player device.

18. The device of claim 1, further comprising a modem coupled to the one or more processors and configured to receive the pose data from a device including a pose sensor.

19. A method comprising: Acquire listener pose data in an immersive audio environment at one or more processors; The current listener pose is determined at one or more processors based on the pose data and one or more pose constraints. At one or more processors and based on the current listener pose, render assets associated with the immersive audio environment are obtained; as well as The output audio signal is generated based on the rendered asset at one or more processors.

20. A non-transitory computer-readable device storing instructions executable by one or more processors to cause the one or more processors to: Obtain the pose data of the listener in an immersive audio environment; The current listener's pose is determined based on the pose data and one or more pose constraints. Rendering assets associated with the immersive audio environment are obtained based on the current listener pose. as well as The output audio signal is generated based on the rendered assets.