Head-tracked split rendering and head-related transfer function personalization
By implementing split rendering techniques with DOA-based head-tracking and HRTF personalization, the method addresses computational complexity and latency issues in XR devices, improving immersive audio quality and extending battery life.
Patent Information
- Application Number
- JP2025514700
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-03
- Filing Date
- 2023-09-11
- Publication Date
- 2025-09-25
AI Technical Summary
Extended reality (XR) devices, particularly AR glasses, face challenges in providing immersive audio due to high computational complexity and latency issues when rendering is performed on network entities, leading to motion-sound latency and quality degradation.
A method for split rendering where head pose-specific processing is performed on a second device, utilizing techniques like DOA-based head-tracking and HRTF personalization, distributing decoding and rendering operations between a first and second device to reduce computational load on the wearable device.
This approach reduces processing demands on wearable devices, extends battery life, and minimizes motion-to-sound latency by using lightweight rendering based on current head pose information, thereby enhancing the immersive audio experience.
Smart Images

Figure 2025531871000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Application No. 63 / 405,538, filed September 12, 2022, and U.S. Provisional Application No. 63 / 422,331, filed November 3, 2022, each of which is incorporated by reference herein in its entirety.
[0002] [Technical field] FIELD OF THE DISCLOSURE The present disclosure relates to audio processing, and in particular to audio rendering. [Background technology]
[0003] Extended reality (XR) (AR / MR / VR) increasingly relies on very power-constrained end devices. Augmented reality (AR) glasses are a prominent example. To be as light as possible, AR glasses cannot be equipped with heavy batteries. As a result, only highly complex numerical operations are possible on the processors included in AR glasses to allow reasonable computation times. On the other hand, immersive audio is an essential media component of XR services. These services may support adjusting the presented immersive audio / visual scene according to user (head) movements, typically 3 or 6 degrees of freedom (DoF). Implementing the corresponding immersive audio representation with high quality typically requires high numerical complexity.
[0004] One potential solution to address this issue is to perform rendering not on the device itself, but on some entity in the mobile / wireless network to which the end device is connected, or on a powerful mobile user equipment (UE) to which the end device is tethered. In that case, the end device, for example, only receives audio that has already been binaurally rendered. The 3DoF / 6DoF head pose information (head tracking metadata) needs to be transmitted to the rendering entity (network entity / UE). The problem with this is the transmission latency between the end device and the network entity / UE, which can be on the order of 100 ms or more. Therefore, performing rendering on the network entity / UE means that the device must rely on outdated head tracking metadata and that the binauralized audio played by the end rendering device does not match the actual head pose of the head / end device. This latency is called the motion-sound latency. If it is too large, the end user perceives it as quality degradation.
[0005] Regarding the video component of immersive media rendering, this issue is addressed by split rendering techniques, where approximate parts of the video scene are rendered by the network entity / UE and the final video scene adjustment is done on the end device. Regarding audio, this area is currently less researched. Summary of the Invention [Means for solving the problem]
[0006] In various audio services, such as immersive voice and audio services (IVAS), it is desirable to be able to track a user's head movements during audio rendering and adjust the audio accordingly to provide the user with an immersive audio experience. This requires immersive audio decoding and binaural rendering using a set of head-related transfer functions (HRTFs), including the selection of specific HRTFs that may depend on the characteristics of the immersive audio signal and the user's head movement (or head pose). Depending on the immersive audio format, decoding and head-tracked binaural rendering can be computationally complex operations. For example, scene-based audio (e.g., high-order Ambisonics), channel-based audio (e.g., with a 7.1.4 channel layout), or object-based audio with many objects may each rely on a large number of constituent audio components, which are computationally complex to decode and render. This means that decoding and binaural rendering of a bitstream in response to a user's head movements requires a large amount of computational processing. The computational complexity requires power and generates heat, which can be problematic for small, portable devices such as AR glasses.
[0007] It is an object of the present invention to overcome the problems described herein and to provide split rendering in which head pose specific processing can be performed on a second device.
[0008] According to some embodiments, this and other objects are achieved by a method according to claim 1 or claim 14. According to another embodiment, this and other objects are achieved by a user-carried device according to claim 22.
[0009] Techniques for direction-of-arrival (DOA)-based head-tracking split rendering and head-related transfer function (HRTF) personalization are described. Head-tracking audio decoding and binaural rendering may be split between two or more devices. In some examples, a first device may coordinate split decoding and rendering operations with a second device. The first device, e.g., a smartphone, receives a main bitstream representation of encoded audio. The first device decodes and renders the main bitstream into a pre-rendered binaural signal using a main decoder and binaural renderer, and encodes the pre-rendered binaural signal and post-rendering metadata containing information about the HRTFs associated with the binaural rendering. The first device provides the pre-rendered binaural signal and post-renderer metadata as a multiplexed intermediate bitstream to a second device. The second device, e.g., headphones, AR glasses, or earphones, tracks current head pose information. The second device decodes the pre-rendered binaural signal and the post-renderer metadata from the intermediate bitstream and provides the decoded pre-rendered binaural signal and the post-renderer metadata to a lightweight renderer, which renders the pre-rendered binaural signal into binaural audio based on the post-renderer metadata, the current head pose information, the generic HRTF, and the optional personalized HRTF.
[0010] The post-rendering metadata includes at least an indication of the pre-rendered HRTFs used in the binaural pre-rendering. The pre-rendered HRTFs are associated with the directions of arrival (DOAs) of the dominant directional components (typically two degrees) of the audio content relative to the assumed head pose. The indication of the pre-rendered HRTFs may be the DOAs or some kind of index that allows the user-held device to identify the correct HRTF.
[0011] In some implementations, the pre-rendered HRTF metrics also include one or more parameters that may be personalized.
[0012] The rendering may include applying an HRTF compensation operation to the binaural audio signal to compensate for the effect of pre-rendered HRTFs to calculate a compensated stereo audio signal, and applying post-rendered HRTFs to the compensated stereo signal to calculate a binaural output signal. These may be performed in one single operation. The HRTF compensation operation may include the inverse of the pre-rendered HRTFs, obtained, for example, by accessing a lookup table. Other ways of achieving compensation for the pre-rendered HRTFs are possible.
[0013] Binauralization is described herein as being performed using head-related transfer functions (HRTFs), but may equally well be performed using binaural room impulse responses (BRIRs). Note further that all HRTF processing needs to be performed for each time frame and each frequency band, often represented as a time / frequency tile.
[0014] In some applications, the assumed head pose is also included in the metadata. In other implementations, the user-held device is configured to transmit the current head pose to the main device. Optionally, the second device encodes at least a portion of the head pose information into a head pose bitstream and provides the bitstream to the first device. The first device decodes the head pose bitstream to obtain the head pose information and then applies the head pose information to the main decoder and pre-renderer. The main decoder / pre-renderer decodes and pre-renders the main bitstream based on the received head pose information (also called the assumed head pose) and a general HRTF. In this case, the user-held device can estimate the assumed head pose based on an expected transmission delay.
[0015] The assumed head pose information is transmitted to the second device along with other information, unless that device derives the assumed head pose from a priori knowledge, which may be based on (head pose) information previously transmitted to the first device or an assumed head pose previously agreed upon between both devices.
[0016] Additionally, this disclosure relates to further inventive concepts, including techniques for DOA-based head-tracking split rendering with an appropriate prototype signal and optional diffuse signals. A first device decodes a main bitstream using a main decoder and renders the decoded bitstream as a dominant directional component referred to as a prototype signal, zero or more diffuse signals, and post-rendering metadata. The first device then encodes the prototype signal, zero or more diffuse signals (or parameters representing them), and post-rendering metadata and provides them as a multiplexed intermediate bitstream to a second device. The second device decodes the prototype signal, zero or more diffuse signals, and post-rendering metadata from the intermediate bitstream and provides the decoded prototype signal, zero or more diffuse signals, and post-rendering metadata to a lightweight renderer. The lightweight renderer renders the prototype signal and zero or more diffuse signals into binaural audio based on the post-rendering metadata, head pose information, a generic HRTF, and optionally a personalized HRTF.
[0017] The techniques described herein can achieve various technical advantages over conventional rendering techniques. Splitting processing between two devices reduces processing on the wearable device, thereby extending battery life. The wearable device performs lightweight rendering based on the user's current head pose, rather than having to rely solely on binaural representation by a more demanding rendering device that may only have access to delayed / stale head pose information, thereby reducing motion-to-sound latency due to potentially using stale head pose information during rendering. The allocation of processing load between the first device can be flexible, for example, by adjusting the amount of head pose information transmitted from the second device to the first device, from zero to full, thereby enabling compatibility with various wearable devices with different processing capabilities. Other advantages, features, and benefits beyond those explicitly described above will become apparent in light of the detailed description and associated drawings as set forth below.
[0018] This Summary is provided to introduce a selection of concepts in a simplified form, is not intended to identify key features or essential features of the claimed subject matter, nor is it intended for use as an aid in determining the scope of the claimed subject matter. For example, the term "technology" may refer to systems, methods, computer-readable instructions, modules, algorithms, hardware logic, and / or operations, as recognized by the context above and throughout this document. [Brief explanation of the drawings]
[0019] [Figure 1] FIG. 1 is a block diagram of an exemplary system implementing head-tracked split rendering. [Figure 2] 10 is a flowchart showing the processing in the first device, that is, the main device. [Figure 3] 10 is a flowchart showing processing in the second, i.e., user-held, device. [Figure 4]1 illustrates an exemplary technique for DOA-based split rendering using pre-rendered binaural signals. [Figure 5] 1 illustrates an exemplary technique for HRTF personalization. [Figure 6] 1 illustrates an exemplary technique for DOA-based split rendering using a prototype signal. DETAILED DESCRIPTION OF THE INVENTION
[0020] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and which are shown by way of illustration, specific exemplary configurations in which the concepts may be practiced. These configurations are described in sufficient detail to enable those skilled in the art to practice the techniques disclosed herein, with the understanding that other configurations may be utilized and other changes may be made without departing from the spirit or scope of the concepts presented. Accordingly, the following detailed description is not to be taken in a limiting sense, and the scope of the concepts presented is defined only by the appended claims.
[0021] The systems and methods disclosed in this application may be implemented as software, firmware, hardware, or a combination thereof. In a hardware implementation, the division of tasks does not necessarily correspond to the division into physical units; conversely, one physical component may have multiple functions, and one task may be performed by several physical components working together.
[0022] The computer hardware may be, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a switch, or a bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be performed by that computer hardware. Furthermore, the present disclosure relates to any collection of computer hardware that individually or collectively executes instructions to perform any one or more of the concepts described herein.
[0023] Certain or all components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code, which, when executed by the one or more processors, includes a set of instructions that perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken is included. Thus, one example is a typical processing system (e.g., computer hardware) that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem including a hard drive, an SSD, RAM, and / or ROM. A bus subsystem may be included for communication between components. Software may reside in the memory subsystem and / or in the processor during its execution by the computer system.
[0024] One or more processors may operate as stand-alone devices or may be connected to other processors, e.g., networked. Such networks may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0025] Software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes various forms of physical (non-transitory) storage media, such as, but not limited to, ROM, PROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer. Furthermore, communication media (transitory) typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and is well known to those skilled in the art to include any information delivery media.
[0026] This disclosure assumes the presence of an immersive audio codec, such as the IVAS codec, used in some extended reality (XR) applications. The main decoding and pre-rendering may be performed by a first device (user equipment, UE) or the edge or other network node of a 5G system. The second device includes a post-decoder and a (lightweight) post-renderer. Therefore, the overall operation may be divided into the operations of multiple devices. The first device (main device) may be a mobile device such as a laptop, tablet, or smartphone, or a stationary device such as a workstation or server. The first device may also be a combination of multiple processing devices. The second device may be a user-held (e.g., worn) device such as augmented reality (AR) glasses.
[0027] One fundamental assumption and inventive insight applied in the field of partitioned rendering is that for each time-frequency tile, audio consists of one dominant directional component and a diffuse (omnidirectional) component. The directional component is assumed to be a prototype signal S arriving from a certain direction (DOA), while the diffuse component is a decorrelated version of that prototype signal. This concept has proven very powerful in spatial audio coding approaches such as DirAC or Metadata-Assisted Spatial Audio (MASA) coding.
[0028] Based on at least these assumptions, various example implementations may include the following steps.
[0029] 1. The pre-renderer binauralizes the decoded immersive audio using a set of generic HRTFs (or BRIRs) given a head pose P', which may either be transmitted from a lightweight device with a head tracker or may simply be a preset value that does not necessarily correspond to any actual head pose of the user, but rather may be a reasonable default such as a head pose looking straight ahead. The application of HRTFs during the binaural pre-rendering operation may be done with HRTFs specifically selected for each time / frequency tile. The HRTFs are selected based on the direction of arrival (DOA) of the dominant component of the immersive audio content relative to the assumed head pose.
[0030] 2. The first or main device encodes and transmits the binauralized audio channels as well as an indication of the HRTFs and / or DOA angles used and the hypothesized head pose P'.
[0031] 3. The post-renderer of the second device aims to adjust the received left and right binaural signals for the current head pose P (if they deviate from P').
[0032] 4. Given the HRTF or DOA angles and head pose applied by the pre-renderer, left and right HRTF correction signals are calculated by a second device, essentially by inverse HRTF filtering of the left and right audio channels and optionally by linearly combining them.
[0033] 5. The HRTF corrected signal is filtered by a second device with the correct HRTF corresponding to the correct head pose.
[0034] 6. Potential errors in the diffusion component are mitigated by appropriate selection of the weights in the mentioned linear combination.
[0035] The main idea can also be applied to HRTF personalization, where generic HRTFs are used by the pre-renderer and the post-renderer corrects these generic HRTFs and then applies personalized HRTFs.
[0036] In the following, exemplary implementations of the novel concepts described herein are illustrated with reference to FIGS.
[0037] 1 is a block diagram of an exemplary system implementing head-tracked split rendering configured in accordance with various aspects of the present invention. The exemplary system includes a first device 10 and a second device 20. The first device 10 is also referred to as a main device, and the second device 20 is also referred to as a mobile device or a user-held device.
[0038] In Figure 1, a first or main device 10 includes a decoder / renderer 11, an optional head pose decoder 12, encoders 13 and 14, and a multiplexer 15. The decoder / renderer 11, for example an IVAS decoder, receives a main bitstream b1 containing encoded immersive audio content (step S1), decodes the immersive audio content (step S2), and performs binaural rendering of the decoded audio content using HRTFs associated with a direction of arrival (DOA) for a hypothesized head pose P' of the user (step S3). The HRTFs used are typically a set of common HRTFs (for different directions of arrival (DOAs)), H g This processing is typically too computationally complex to perform on mobile or user-held (lightweight) devices. The assumed head pose P' may be a suitable default head pose or the actual user's head pose received from the user-held device, which may optionally be decoded by the head pose decoder 12 (step S21). Such a decoded user head pose may represent the user's recent, but not entirely current, head pose. The renderer 11 outputs the binaural signals L1, R1 as well as post-rendering metadata M. The post-rendering metadata M includes, for example, an indication of the used HRTFs, or an index of the used HRTFs, expressed relative to the assumed head pose P', expressed as the direction of arrival (DOA) of the dominant directional component of the immersive audio content. The post-rendering metadata M may also include an indication of the head pose P' associated with the binaural rendering. The encoders 13, 14 are configured to encode the binaural signals L1, R1 and the post-rendering metadata M (step S4) into encoded signals b11 and b12, respectively. The multiplexer 15 is configured to multiplex or combine the encoded binaural signal b11 and the encoded metadata b12 into an intermediate bitstream b2 (step S5), which is transmitted to the second device 20 (step S6).
[0039] The second device 20, which may be the user-held device 20, includes a demultiplexer 21, decoders 22 and 23, an encoder 25, a renderer 26, and a head tracker 24. The user-held device 20 receives the intermediate bitstream b2 (step S11), and the demultiplexer 21 separates the intermediate bitstream b2 into coded signals b21 and b22, which are received by corresponding decoders 22 and 23. The decoders 22 and 23 accordingly decode the coded signals b21 and b22 (step S12) to obtain decoded binaural signals L2, R2 and decoded metadata M'. As mentioned above, the metadata M' includes an indication of the HRTFs used, e.g., indicated by an index for a forward-looking head pose or a direction of arrival.
[0040] The head tracker 124, which may be included in or connected to the user-held device 120, detects the current head pose P of the user's head (step S13). The encoder 25 optionally encodes the detected head pose P as bP (step S131), and the encoded detected head pose b P may be used to transmit the same to the main device 10.
[0041] The metadata M' may also include an assumed head pose P' for use in the renderer 11. Alternatively, in implementations in which the detected head pose P is transmitted to the main device 10, the user-held device can estimate the assumed head pose based on an expected transmission delay. In principle, the assumed head pose can be assumed to be the head pose detected at the time corresponding to the expected transmission delay.
[0042] Finally, the renderer 26 receives the decoded binaural audio signals L2, R2, the DOA or used HRTFs, the assumed head pose P', and the current head pose P, and calculates the output binaural signals Lout, Rout. This process includes identifying post-rendering HRTFs corresponding to the detected current head pose P (step S14), calculating a corrected stereo audio signal by applying an HRTF correction operation to the binaural audio signal, the HRTF correction operation being configured to correct for the effects of the pre-rendering HRTFs (step S15), and finally applying the identified post-rendering HRTFs (step S16). For this process, as described below, the renderer 26 is provided with HRTF data, typically a set of generic HRTFs Hg (for various directions of arrival (DOAs)). The renderer 26 may also be provided with a set of personalized HRTFs Hp.
[0043] FIG. 2 is a flowchart showing the processing in the first device (or main device), and includes the above-mentioned steps S1 to S6 and an optional step S21.
[0044] The process includes step S1 ("bitstream reception"), receiving a bitstream, and step S2 ("decoding"), decoding the bitstream by a decoder to obtain decoded immersive audio content. In step S21 ("decoding pose"), the process may include the optional step of receiving and decoding an indication of the current user head pose from a second user-held device and determining a hypothesized user head pose based on the current user head pose. In step S3 ("pre-rendering"), the process includes binauralizing the immersive audio content by a pre-renderer to generate a pre-rendered binaural signal, the binauralization using pre-rendered HRTFs from the set of HRTFs and the user's hypothesized head pose. In step S4 ("encoding"), the process includes encoding the pre-rendered binaural signal and encoding post-rendering metadata, the metadata indicating the pre-rendered HRTFs. In step S5 "combining", the process includes combining the encoded binaural audio signal and the encoded post-rendering metadata in a multiplexer to form a bitstream containing the binaural audio representation, and in step S6 "transmitting", the process includes transmitting the bitstream to a second device (or a user-held device).
[0045] FIG. 3 is a flowchart showing the processing in the second device (or the user-held device), which includes the above-mentioned steps S11 to S16, and an optional step S131.
[0046] The process includes, in step S11, "receiving a bitstream" receiving from a first device (or main device) a bitstream containing a binaural pre-rendered representation of the immersive audio content. The binaural pre-rendering has been obtained for a hypothesized head pose P'. In step S12, "decoding," the process includes decoding the bitstream to obtain a binaural audio signal and associated post-rendering metadata. The metadata indicates pre-rendering HRTFs used in the binaural pre-rendering, and the pre-rendering HRTFs are associated with the hypothesized head pose P'. In step S13, "detecting a current pose," the process involves obtaining user head pose information indicating the current head pose P. In step S131, "encoding a pose," the process may include the optional step of encoding the detected head pose P using an encoder and transmitting an indication of the current head pose P to the main device. As described above, the main device can use the current pose received from the second user-held device as the hypothesized pose. Thus, in this case, the second user-held device can estimate a hypothesized head pose P' based on the expected transmission delay (and the previously transmitted current pose). In step S14 "Identify Post-Rendering HRTFs", the process includes identifying post-rendering HRTFs based on the metadata, the hypothesized head pose P', and the current head pose P. In step S15 "Calculate Corrected Audio", the process includes calculating a corrected stereo audio signal by applying an HRTF correction operation to the binaural audio signal, the HRTF correction operation being configured to correct for the effect of the pre-rendering HRTFs.
[0047] Various exemplary HRTF compensation operations suitable for calculating a compensated stereo audio signal are described herein. These operations may include any number of methods for appropriately canceling and adjusting for various effects resulting from the pre-rendered HRTF operations. In some examples, the HRTF compensation may include an inverse operation of the pre-rendered HRTFs. In some other examples, the HRTF compensation may be implemented using a lookup table-type operation, where, to reduce memory needs, the lightweight device may rely on a less dense set of HRTFs than used in the pre-rendered device. Thus, the set of inverse HRTFs available to the lightweight device may include suitable approximations of the inverses of the HRTFs available in the HRTF set in the pre-rendered device. In still other examples, the HRTF compensation may be implemented using a numerical approximation method, which may include linear or nonlinear interpolation, or a combination thereof. In a further example, the HRTF compensation may be implemented using a best-fit type approximation. Combinations of various methods are equally applicable and are considered within the scope of this disclosure.
[0048] 3, in step S16 "Apply Post-Rendered HRTFs", the process includes calculating the binaural output signal by applying the post-rendered HRTFs to the corrected stereo signal. Note that steps S15 and S16 may be performed as a single operation.
[0049] The processes illustrated by Figures 2 and 3 include a collection of blocks that represent a sequence of operations or steps that can be implemented in hardware, software, or a combination thereof. In a software context, the blocks can represent computer-executable instructions that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions can include routines, programs, objects, components, data structures, etc. that perform or implement a function. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described blocks can be combined in any order, separated into additional blocks, and / or operated in parallel to implement a process.
[0050] Pre-renderer 11 approach: The basic assumption is that for each time-frequency tile, the audio consists of one dominant directional component and a diffuse (omnidirectional) component. The directional component is assumed to be a prototype signal S arriving from a specific DOA with azimuth and elevation angles α0, γ0 expressed in some room coordinate system. The diffuse component is a decorrelated version of the prototype signal S.
[0051] Pre-render synthesis is performed here by convolving the directional component (for each time-frequency tile) with the HRTF corresponding to the DOA and adding the diffuse component. Both components are assigned respective weights r dir and r diff Add.
[0052]
number
[0053]
number
[0054] and Ψ L , Ψ R is the decorrelator.
[0055] Here, α′ and γ′ are the azimuth and elevation angles of the directional components for the head pose P′ assumed by the pre-renderer 11. When the head pose P′ is expressed in the same room coordinate system, α′=α0−α P´ So, γ´=γ0-γ P´ is.
[0056] Post-render approach: The post-renderer aims to adjust the received left and right binaural signals L2 and R2 with respect to the current head pose P if P deviates from P'.
[0057] The key insight is that the correct output signal is:
[0058]
number
[0059]
number
[0060] Signal S and decorrelator signal Ψ L , Ψ R is not available. Instead, L * out , R * out are approximated in a parametric approach using the available signals L2, R2. Assuming that the HRTFs applied by the pre-renderer are known, the left and right HRTF-corrected signals can be calculated. One possibility to obtain these signals is to derive them as a weighted combination of the HRTF-corrected left and right channel signals, where κ L , κ R is a suitable weighting factor or operator.
[0061] The HRTF corrected left and right channel signals are expressed as follows for the left and right channel signals:
[0062]
number
[0063]
number
[0064] is.
[0065] So the left and right HRTF corrected signals are:
[0066]
number
[0067]
number
[0068] In one simple example, κ L =κ R = 1 (no weighting),
[0069]
number
[0070]
number
[0071] is.
[0072] Using these signals, the left and right output signals of the post-renderer are obtained as follows:
[0073]
number
[0074]
number
[0075] This approach results in the correct directional component in the output signal with respect to the current head pose, i.e.
[0076]
number
[0077]
number
[0078] is.
[0079] However, there is an error in the diffusion component, which can be quantified as follows:
[0080]
number
[0081]
number
[0082] Given the possible decomposition of HRTFs into delay and gain / shape manipulations, delay changes involving decorrelated diffuse components are perceptually inconsequential, while gain / shape changes can lead to timbral deviation or coloration effects. This error can be reduced by κ depending on the associated set of HRTFs. L , κ R This can be mitigated by choosing κ appropriately. L , κ Rcan be linear and non-linear operators such as a (frequency selective) filter operator or a gain limiter to prevent the output samples from exceeding a predetermined number range.
[0083] Weight κ L , κ R The advantage of choosing γ = γ' is illustrated by an example considering the case where the head pose change is yaw-limited, i.e., rotation occurs only around the z-axis, i.e., the elevation angles of the directional components for the assumed and current head poses are equal (γ = γ').
[0084] First, consider the case where there is almost no yaw rotation. Therefore, α is approximately equal to α'. In this case, the weights are preferably L =κ R =1, which means that the received left and right binaural signals L2, R2 are output almost without modification,
[0085]
number
[0086] Therefore, for (sufficiently) small yaw deviations (e.g., less than 20 degrees), the weight is set to κ L =κ R Setting it to =1 is a good choice.
[0087] Next, if we consider a yaw change of 180, then α is α'±180 degrees. Here, the weights are preferably L =κ R = 0, which means that the received left and right binaural signals L2, R2 are (virtually) swapped as part of the adjustment process. This results in the following diffuse components for the left and right output channels:
[0088]
number
[0089]
number
[0090] Considering that the right ear HRTF can be approximated by the left ear HRTF taken at an azimuth offset of 180 degrees, and similarly the left ear HRTF can be approximated by the right ear HRTF taken at an azimuth offset of 180 degrees, in the above equation:
[0091]
number
[0092]
number
[0093] The term is approximated to 1.
[0094] This leads to the following approximations of the diffuse components for the left and right output channels:
[0095]
number
[0096]
number
[0097] Therefore, if the yaw difference between the current head pose and the assumed head pose is close to 180 degrees, we set the weight κ L =κ R We conclude that setting =0 is a good choice.
[0098] The third considered case is when there is a yaw rotation of 90 degrees, resulting in α being α´±90 degrees. Here, when constructing the HRTF correction signal, there is no reason to favor either of the available signals L2, R2, as this could potentially lead to a solution with asymmetric behavior between the left and right channels. When the yaw rotation is close to 90 degrees, κ L =κ R Choosing =0.5 provides symmetric behavior.
[0099] This discussion is based on the determined yaw rotation Δ yaw , i.e., a weight κ L , κ R This leads to a preferred solution for adaptively selecting
[0100]
number
[0101] Note that a similar embodiment can be formulated based on the roll angles of the assumed and current head poses.
[0102] 4 illustrates an exemplary technique for DOA-based split rendering using pre-rendered binaural signals. The method is shown in which the acoustic wavefronts are assumed to arrive at the listener's head 30 from DOAs with azimuth angles α′, α. The pre-renderer then calculates the angle Only the head pose P' at angle α' is accessible. Therefore, binaural synthesis is performed using the HRTFs corresponding to the head pose P' at angle α'. One main effect is that the interaural time difference (ITD) of the wavefront between the left and right ear is Δ'LR. There is also a corresponding interaural level difference (ILD) and spectral difference. The applied HRTFs mimic this effect by imposing / imprinting the appropriate ITD, ILD and spectrum. The figure further illustrates the assumed situation in a post-renderer with knowledge of the actual head pose P at angle α. If the assumed wavefront arrives from a different DOA, α' instead of α, with respect to the actual head pose P, resulting in a different ITD Δ LR It is shown that the ITD yields Δ'. It should be noted that there are other ILDs and spectra that correspond to actual head poses. One main concept of this disclosure, visualized in Figure 4, is therefore to convert the ITD into Δ' LR From Δ LR First, change it to Δ´ LR Correction is then made for Δ LR Similarly, the ILDs and spectra are modified by correcting those corresponding to head pose P' and applying the ILDs and spectra of HRTFs corresponding to the actual head pose P.
[0103] The above description leads to the following simplified approach (see also Figure 4).
[0104] Assumptions: The rear-front axis A of the listener 30 defines the x-axis of a right-handed coordinate system. Furthermore, in many relevant cases, the user can perform most of their head movement around the yaw axis (z-axis), and most immersive audio content has sound sources near the horizontal plane. Therefore, the elevation angle of the DOA is relatively close to zero degrees (e.g., constrained to the interval [-20, 20] degrees).
[0105] Under these assumptions, according to a simplified formula, only the azimuthal component of the DOA gives rise to significant ITD, which in turn affects the gain and spectral shape.
[0106] The limited elevation component of the DOA (pitch, roll) affects the gain and spectral shape, but not the ITD.
[0107] An HRTF filter can be decomposed into a delay and a gain / shape operation.
[0108]
number
[0109] A head pose P' is assumed during pre-rendering. The pre-renderer renders under the azimuthal component of the DOA α', which deviates from the true azimuthal angle α during playback.
[0110] During playback, the post-renderer renders under the azimuth component of DOA, α, corresponding to the listener's head pose P.
[0111] The interaural time difference (ITD) between the pre-rendered signal and the post-render adjusted signal is calculated as follows:
[0112]
number
[0113]
number
[0114] where d e is the distance between the ears and c is the speed of sound.
[0115] Therefore, the post-renderer calculates the ITD of the directional component in a given time-frequency tile as Δ LR ´ to Δ LR Adjust to.
[0116] Apart from the ITD adjustment, the post-renderer also adjusts the interaural level difference and the spectral shape given the true head pose compared to the head pose assumed by the pre-renderer.
[0117] It is worth noting that a similar formulation is possible without assuming a limited elevation component of the DOA. Even in that case, the post-renderer operations can be decomposed into ITD adjustment, interaural level difference, and spectral shape adjustment. In that case, the amount of ITD adjustment required, however, depends on the azimuth and elevation angles of the DOA assumed during pre-rendering and in effect during post-rendering.
[0118] Figure 5 illustrates an exemplary technique for HRTF personalization. It shows how a hypothesized wavefront arrives at the listener's head 30 from a DOA angle α, resulting in different ITDs depending on the size of the listener's head. Pre-rendering with a generic HRTF can be performed for a generic interaural distance d e We can assume a listener head dimension with Δ LR g , and the corresponding ILDs and spectral shapes of the left and right audio signals. The personalized HRTFs are based on the (more) accurate listener head dimensions. This therefore leads to a more accurately personalized ITD, Δ LR p This results in more accurately corresponding ILD and spectral shapes of the left and right audio signals. The general idea of HRTF personalization is that a post-renderer corrects the general HRTF and applies the effect of the personalized HRTF. The overall concept is very similar to the above-mentioned head pose correction in the post-renderer. Therefore, both concepts are compatible with each other and can be easily combined.
[0119] The same explanations as above for the main device 10 with a pre-renderer 11 apply. However, the pre-renderer 11 may also use a set of generic HRTFs H gThe post-renderer 26 in the user-held device 20 renders the image based on a personalized set of HRTFs H p This makes the post-renderer aware of the general HRTF used by the pre-renderer.
[0120] The post-renderer 26 performs a set of personalized HRTFs H for the current head pose P if P deviates from P'. p The aim is to adjust the received left and right binaural signals L2, R2 with respect to
[0121] The correct output signal is:
[0122]
number
[0123] Signal S and decorrelator signal Ψ L , Ψ R is not available. Instead, L * out , R * out are approximated with a parametric approach using the available signals L2, R2. Assuming the HRTFs applied by the pre-renderer are known, the left and right HRTF-corrected signals are computed as a linear combination of the HRTF-corrected left and right channel signals.
[0124]
number
[0125]
number
[0126] Using these signals, the left and right output signals of the post-renderer are obtained as follows:
[0127]
number
[0128] This approach results in the correct directional component in the output signal with respect to the actual head pose and personalized HRTFs.
[0129]
number
[0130]
number
[0131] However, even in this case, an error is introduced in the diffuse component, which can be quantified as follows:
[0132]
number
[0133]
number
[0134] Given the possible decomposition of HRTFs into delay and gain / shape manipulations, delay changes involving decorrelated diffuse components are perceptually inconsequential, while gain / shape changes can lead to timbral deviations or coloration effects. This error can be reduced by κ depending on the associated set of HRTFs. L , κ R can be mitigated by an appropriate choice of . Embodiments with adaptive selection of weights depending on the yaw and / or roll deviation between the current and assumed head pose remain fully applicable.
[0135] Specific Aspects of the Embodiments The post-renderer receives direction of arrival (DOA) information, which can be expressed as the azimuth and elevation angles (DOA angles) α', γ' of the dominant directional component of the immersive audio content for a hypothesized head pose P'. Note that the DOA is determined for each time-frequency tile. The index of the used HRTF is another way to provide the DOA information to the post-renderer.
[0136] Furthermore, the post-renderer must be aware of the head pose P' assumed in the pre-renderer. The corresponding information can be transmitted to the post-renderer (i.e., in metadata). It is also possible to rely on the fact that P' corresponds to the true head pose at an earlier point in time, which is transmitted from the post-renderer to the pre-renderer. Assuming that the transmission delay from the post-renderer to the pre-renderer is known a priori or can be estimated, this makes transmitting P' to the post-renderer unnecessary. One way to estimate the transmission delay from the post-renderer to the pre-renderer is to base it on a round-trip delay measurement from the post-renderer to the pre-renderer and back, for example using timestamps.
[0137] Parameter r dir , r diff , κ L , κ R are mathematically interconnected. Therefore, there is a possibility to exploit this interdependence, which can be achieved by, for example, L , κ R This helps to find a good choice of . This makes it possible to avoid using a decorrelator in the post-renderer. The advantage of such an approach is that it avoids the complexity of the post-renderer.
[0138] If a decorrelator is to be used in a post-renderer, the appropriate decorrelator input signal is
[0139]
number
[0140] which corrects / removes the directional component as follows:
[0141]
number
[0142] Preferably, the pre-rendered binaural channel signals L1, R1 are transmitted in the complex-valued quadrature mirror filter bank (CQMF) / frequency domain, which avoids performing forward time-CQMF / frequency domain operations in the post-renderer, which is advantageous in terms of complexity and delay.
[0143] A notable difference between this approach and conventional techniques is that the present approach relies on correcting the HRTF filter operations of the pre-renderer and applying the HRTFs that would ideally be used. In contrast, alternative techniques rely on transforming the binaural output channels using a linear transform whose coefficients are obtained following an LMS approach and interpolation.
[0144] An exemplary implementation for DOA-based split rendering using prototype signals includes the following steps.
[0145] 1. A pre-renderer or decoder generates a prototype signal (S). Some example approaches for generating S are as follows:
[0146] 1-a. Obtain the Ambisonics W or omni-directional channel representation from the decoder output using any known technique and use that as S.
[0147] 1-b. Obtain a representation of the dominant eigensignal from the decoder output and use it as S.
[0148] 1-c. Pre-render the decoded immersive audio using a set of generic HRTFs (or BRIRs) to produce S=aL+bR, where L and R are the left and right channels of the pre-rendered bin signal, and a and b are complex or real-only gain factors per time-frequency tile, which can either be dynamically calculated or statically predetermined values, e.g., a=0.5 and b=0.5.
[0149] 2. The main device transmits the coded prototype signal S, the hypothesized head pose P' and / or the hypothesized DOA angle (or equivalent information), and the diffuseness parameters.
[0150] 3. The post-renderer decodes the prototype signal bits and produces S' (which should be the same as S if the codec used to code S has zero delay and is lossless).
[0151] 4. The post-renderer aims to generate left and right binaural signals for the current head pose P (if it deviates from P').
[0152] 4. The post-renderer adjusts the DOA angle sent by the main device based on the difference between P and P'. Together with S', the HRTF in the post-renderer, and the adjusted DOA angle, the post-renderer generates the directional component of the post-rendered binaural signal. The diffuseness parameter is used together with the decorrelated S' to fill in the diffused energy in the post-rendered binaural signal.
[0153] Post-render approach: The post-renderer aims to align the received DOAs to the current head pose P if it deviates from P', and then together with the prototype signal S' and the set of HRTFs Hp. Hp may be personalized or generic. The post-renderer generates the head-tracked binaural signals as follows:
[0154]
number
[0155]
number
[0156]
number
[0157]
number
[0158] r diff is the diffusion parameter transmitted by the main device, r dir is the directional gain that can be calculated using the DOA and spherical harmonics.
[0159] The advantages of this approach are:
[0160] A low bitrate mode can be achieved by only encoding the S channel and sending it to the post-renderer.
[0161] No HRTF correction needs to occur and the error in the diffusion correction can be reduced to zero.
[0162] The drawbacks of this approach are:
[0163] In the potential case where the user-held device only outputs the decoded binaural audio signal without further processing or post-rendering operations, the pre-rendered binaural audio signal is not readily available at the user-held device.
[0164] FIG. 6 illustrates an exemplary technique for DOA-based split rendering using a prototype signal.
[0165] In FIG. 6, the main device 110 includes a decoder / renderer 111, a head pose decoder 112, encoders 113 and 114, and a multiplexer 115. The decoder / renderer 111, e.g., an IVAS decoder, receives a main bitstream b1 and performs rendering synthesis of a prototype signal S having a direction of arrival (DOA) with respect to a hypothesized head pose P′ of the user. The hypothesized head pose P′ may be an appropriate default head pose or an actual user head pose received from a user-held device, which may optionally be decoded by the head pose decoder 112. Such a decoded user head pose represents the user's recent, but not exactly current, head pose. The renderer 111 outputs a prototype signal S. The encoders 113 and 114 encode the prototype signal S and metadata M, which includes at least the direction of arrival (DOA) of the prototype signal, and the multiplexer 115 multiplexes the encoded prototype signal b11 and encoded metadata b12 into one intermediate bitstream b2.
[0166] The user-held device 120 includes a demultiplexer 121, decoders 122 and 123, a head tracker 124, an encoder 125, and a post-renderer 126. The demultiplexer 121 receives the intermediate bitstream and separates it into two encoded signals b21 and b22, and the two decoders 122 and 123 responsively decode these signals to obtain a decoded prototype signal S′ and decoded metadata M′, e.g., a DOA and (optionally) a hypothesized head pose P′, which are used in the renderer 111. The head tracker 124, which may be included in or connected to the user-held device 120, detects the current head pose P of the user's head. The encoder 125 encodes the detected head pose P and transmits it to the main device 110. Finally, the post-renderer 126 receives the decoded prototype signal S′, the DOA, the hypothesized head pose P′, and the current head pose P, and outputs an output binaural signal L out , R out For this process, as explained below, the post-renderer 126 receives HRTF data, typically a set of generic HRTFs (for various directions of arrival DOAs) H g The Renderer 125 provides a set of personalized HRTFs. p may be provided.
[0167] Example use cases that can benefit from split rendering of immersive audio The techniques described herein can be implemented in various use cases: it is assumed that the main audio processing / audio signal enhancement with subsequent pre-rendering is performed in some powerful device or network node, while the post-rendering is performed in a lightweight end device such as AR glasses.
[0168] Some examples are given below:
[0169] 1. AR / MR with Audio Audio Zoom / Magnifier, like a magnifying glass but for sound. Users may zoom in on sounds of interest.
[0170] Overlay of real-world objects with sounds: Real-world objects / items are associated with sounds. Useful for, but not limited to, assistive systems for the visually impaired.
[0171] Dialogue Enhancement / Smart Ambient Noise Reduction: Helps people with cocktail party problems by boosting active voices against ambient noise.
[0172] Mood Sound Ambience: Similar to mood lighting, sounds are associated with real-world environments, items, and personal preferences.
[0173] 2. Use Case Characteristics These use cases will typically rely on audio / visual capture, some scene analysis, and generation of an augmented sound signal, which in some scenarios may also be overlaid with immersive sound from some network node or far end in the communication.
[0174] Use cases typically rely on head-tracked audio / visual rendering.
[0175] 3. Additional non-AR / MR use cases Immersive voice communication (two-party,conferencing) and immersive content streaming with AR glasses as,end devices are likely use cases for IVAS.
[0176] Some of them may rely on head-tracked audio rendering, and some may not.
[0177] Some use cases may involve one-to-many immersive delivery of head-tracked audio.
[0178] Aspects of the systems described herein can be implemented in a computer-based sound processing network environment suitable for processing digital or digitized audio files. Portions of the adaptive audio system can include one or more networks comprising any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route data transmitted between computers. Such networks can be built on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0179] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described as data and / or instructions embodied in various machine-readable or computer-readable media using any number of combinations of hardware, firmware, and / or in terms of their behavior, register transfers, logical components, and / or other characteristics. Computer-readable media on which such formatted data and / or instructions may be embodied include various forms of physical (non-transitory), non-volatile storage media, such as, but not limited to, optical, magnetic, or semiconductor storage media.
[0180] Although one or more embodiments have been described by way of example and in connection with specific embodiments, it is to be understood that the one or more embodiments are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0181] Further details and embodiments of the invention can be seen from the following list of enumerated exemplary embodiments.
[0182] EEE1. A method for processing audio, comprising: receiving, by a first device, a main bitstream representation of encoded audio; acquiring, by a second device, user head pose information; determining, by the first device, from the main bitstream a downmixed signal comprising at least one channel and metadata; providing, by the first device, the downmixed signal and metadata to a second device; and rendering, by a lightweight renderer of the second device, the downmixed signal into output binaural audio based on the metadata and the user head pose information.
[0183] EEE2. The method of EEE1, wherein the downmixed signal includes a pre-rendered binaural signal.
[0184] EEE3. Determining the pre-rendered binaural signal and rendering metadata decoding the main bitstream representation by a main renderer of the first device to generate decoded audio; binauralizing the decoded audio by a pre-renderer of the first device to generate the pre-rendered binaural signal and rendering metadata, wherein the pre-renderer: Generic Head-Related Transfer Function (HRTF) or Binaural Room Impulse Response (BRIR), or The binauralization is performed using at least one of the user head pose information, and the user information is a head tracker of the second device; a storage device for storing preset values, or The method of EEE2, wherein the angle is obtained from at least one of the hypothesized direction of arrival (DOA) angles.
[0185] EEE4. The metadata is: HRTF or BRIR indices used by the pre-renderer; the assumed user head pose used by the pre-renderer, or The method of EEE3, including at least one of the assumed DOA angles used by the pre-renderer.
[0186] EEE5. The method of EEE4, wherein rendering the pre-rendered binaural signals into output binaural audio includes adjusting, by the lightweight renderer, left and right channels of the pre-rendered binaural signals with respect to a current user head pose obtained through the head tracker over the assumed user head pose used by the pre-renderer.
[0187] EEE6. Rendering the pre-rendered binaural signal comprises: inverse HRTF filtering of the left and right channels of the pre-rendered binaural signal according to the HRTF or assumed DOA angles used by the pre-renderer; and linearly combining the inverse HRTF filtered signals.
[0188] EEE7. The method of EEE6, wherein the inverse HRTF filtering includes correcting the HRTF used by the pre-renderer using a current user head pose obtained through the head tracker of the second device.
[0189] EEE8. The method of EEE6 or EEE7, wherein linearly combining the inverse HRTF filtered signals comprises selecting weights for the linear combination to mitigate errors in the diffuse component.
[0190] EEE9. The method of any of EEE2 to EEE8, comprising applying HRTF personalization, wherein the pre-render applies a generic HRTF and the lightweight renderer corrects the generic HRTF and then applies a personalized HRTF.
[0191] EEE10. The method of EEE1, wherein the downmixed signal includes a prototype signal.
[0192] EEE11. The method of EEE10, wherein the prototype signal comprises a single channel.
[0193] EEE12. Computing the prototype signal comprises: decoding the main bitstream representation by a main decoder of the first device to generate decoded audio; applying a gain to the decoded audio; and adding the decoded audio with the applied gain to the decoded audio.
[0194] EEE13. Calculating the prototype signal and DOA angle based on the assumed head pose P' and diffusivity parameters; sending the hypothesized head pose P′, the DOA angle of the hypothesized head pose P′, the diffuseness parameters, and the prototype signal to a post-renderer device; adjusting the DOA angle based on an actual head pose P in the post-renderer device; calculating the directional components using the prototype signal and a set of HRTFs and adjusted DOA angles; calculating a diffuse component using the diffuseness parameters and a decorrelated version of a prototype signal; and adding the directional and diffuse components to produce a post-rendered binaural output.
[0195] EEE14. The method of any of EEE1 to 13, wherein the first device comprises a smartphone, the second device comprises a wearable audio, visual, or AR device, and the main bitstream comprises an immersive audio and video services (IVAS) bitstream.
[0196] EEE15. A system including one or more processors configured to perform the operations recited in any one of EEE1 to 14.
[0197] EEE16. A computer program product configured to cause one or more processors to perform the operations of any one of EEE1 to 14.
Claims
1. 1. A method for processing audio on a user-held processing device, comprising: receiving, from a main device, a bitstream comprising a binaural pre-rendered representation of immersive audio content, said binaural pre-rendering having been obtained for a hypothesized head pose P'; decoding the bitstream to obtain binaural audio signals and associated post-rendering metadata, the metadata indicating pre-rendering HRTFs used in the binaural pre-rendering, the pre-rendering HRTFs being associated with the hypothesized head pose P'; obtaining user head pose information indicating a current head pose P; identifying a post-rendering HRTF based on the metadata, the assumed head pose P′ and the current head pose P; applying an HRTF compensation operation to the binaural audio signal to compensate for the effects of the pre-rendered HRTFs to calculate a compensated stereo audio signal; and calculating a binaural output signal by applying the post-rendering HRTFs to the corrected stereo signal.
2. The method of claim 1 , wherein the steps of calculating the corrected stereo audio signal and calculating the binaural output signal are performed in one single operation.
3. The method of claim 1 or 2, wherein the HRTF correction operation involves an inverse operation of the pre-rendered HRTFs.
4. The method of claim 3 , wherein the inverse of the pre-rendered HRTF is obtained by accessing a lookup into a table containing an approximation of the inverse of the pre-rendered HRTF.
5. The step of calculating the corrected stereo audio signal comprises:
5. A method according to claim 3 or 4, comprising the steps of applying an inverse left HRTF to a left channel of the binaural audio signal to form an inverse filtered left channel, applying an inverse right HRTF to a right channel of the binaural audio signal to form an inverse filtered right channel, and combining the inverse filtered left and right channels to form the left and right channels of the corrected stereo audio signal, respectively.
6. The method of claim 5 , wherein the inverse filtered left and right channels are linearly combined, the weights of the linear combination being selected to mitigate errors in the diffuse component of the binaural output signal.
7. The method of claim 6 , wherein the weights are preferably adaptive depending on the difference between the assumed head pose P′ and the current head pose P.
8. 10. A method according to any one of the preceding claims, wherein the post-rendering HRTFs are personalized to a user of the user-held processing device.
9. 10. The method of any one of the preceding claims, wherein the post-rendering metadata further comprises an indication of the hypothesized head pose P'.
10. 10. The method of claim 1, further comprising transmitting an indication of the current head pose P to the main device and estimating the hypothesized head pose P' based on an expected transmission delay.
11. 10. The method of any one of the preceding claims, wherein the user-carried processing device comprises a wearable audio, visual, or AR device.
12. 10. The method of any one of the preceding claims, wherein the bitstream comprises an Immersive Audio and Video Services (IVAS) bitstream.
13. 10. The method of claim 1, wherein the metadata includes a direction of arrival (DOA) associated with a dominant directional component of the immersive audio content, the DOA being indicative of the pre-rendered HRTF.
14. 1. A method for processing audio, comprising: receiving a bitstream; decoding the bitstream by a decoder to obtain decoded immersive audio content; binauralizing the immersive audio content by a pre-renderer to generate a pre-rendered binaural signal, the binauralization using pre-rendered HRTFs from a set of HRTFs and an assumed head pose of the user; encoding the pre-rendered binaural signal; encoding post-rendering metadata, said metadata indicating said pre-rendering HRTF; combining, in a multiplexer, the encoded binaural audio signal and the encoded post-rendering metadata to form a bitstream comprising a binaural audio representation; transmitting the bitstream to a user-carried device.
15. 135. The method of claim 134, wherein the metadata includes a direction of arrival associated with a dominant directional component of the immersive audio content.
16. The method of claim 14 or 15, wherein the metadata further comprises the hypothesized head pose.
17. receiving an indication of a current user head pose from the user-held device; determining the assumed user head pose based on the current user head pose; 17. The method of any one of claims 14 to 16, further comprising:
18. The method of any one of claims 14 to 17, executed on a smartphone.
19. 19. A system comprising one or more processors configured to perform the method of any one of claims 1 to 18.
20. 19. A computer program product configured to cause one or more processors to perform the operations of any one of claims 1 to 18.
21. 1. A user-held processing device, comprising: a decoder configured to decode a bitstream comprising a binaural pre-rendered representation of immersive audio content to obtain a binaural audio signal and associated post-rendering metadata, the metadata indicating pre-rendering HRTFs used in the binaural pre-rendering, the pre-rendering HRTFs being associated with the assumed head pose; and a head tracker for obtaining user head pose information indicative of a current head pose; A renderer, identifying a post-rendering HRTF based on the metadata, the assumed head pose P′, and the current head pose P; applying an HRTF compensation operation to the binaural audio signal to compensate for the effect of the pre-rendered HRTFs to calculate a compensated stereo audio signal; a renderer that calculates a binaural output signal by applying the post-rendering HRTFs to the corrected stereo signal.
22. The device of claim 21 , wherein the HRTF correction operation involves an inverse operation of the pre-rendered HRTF.
23. 23. The device of claim 22, wherein the renderer is configured to obtain the inverse of the pre-rendering HRTF by accessing a lookup in a table containing an approximation of the inverse of the pre-rendering HRTF.
24. The renderer renders the corrected stereo audio signal as applying an inverse left HRTF to the left channel of the binaural audio signal to form an inverse filtered left channel, applying an inverse right HRTF to the right channel of the binaural audio signal to form an inverse filtered right channel, and combining the inverse filtered left and right channels to form the left and right channels of the compensated stereo audio signal, respectively; 24. A device according to claim 22 or 23, configured to perform a calculation.
25. 25. The device of claim 24, wherein the inverse filtered left and right channels are linearly combined, the weights of the linear combination being selected to mitigate errors in a diffuse component of the binaural output signal.
26. 26. The method of claim 25, wherein the weights are preferably adaptive depending on the difference between the assumed head pose P' and the current head pose P.
27. 27. A device according to any one of claims 21 to 26, wherein the post-rendering HRTFs are personalized to a user of the user-carried processing device.
28. 28. A device according to any one of claims 21 to 27, incorporated into a wearable audio, visual or AR device.