Binaural Rendering
The method of low-complexity, low-bitrate prediction-based split rendering addresses the challenges of high computational requirements and latency in immersive audio on AR glasses by generating a downmix signal and metadata on a high-resource device, ensuring efficient and high-quality audio experiences on lightweight devices.
Patent Information
- Application Number
- JP2025532571
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2024-02-07
- Publication Date
- 2025-12-18
AI Technical Summary
Immersive audio technologies face challenges in providing high-quality experiences on low-power, small form-factor devices like AR glasses due to high computational requirements and transmission latency in split rendering, leading to degraded user experiences.
A method for low-complexity, low-bitrate prediction-based split rendering that involves generating a downmix signal and metadata on a high-resource device, transmitting it to a lightweight device, and adjusting binaural audio based on user pose information to reduce latency and power consumption.
Enables efficient and high-quality immersive audio experiences on lightweight devices by reducing computational load and latency, conserving power, and allowing for more comfortable, cost-effective wearable devices.
Smart Images

Figure 2025541122000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 386,465, filed December 7, 2022, the contents of which are incorporated herein by reference in their entirety.
[0002] The present invention relates generally to audio processing. [Background technology]
[0003] Immersive audio is a key media component of eXtended Reality (XR) applications, which include Augmented Reality (AR), Mixed Reality (MR), and Virtual Reality (VR). To enhance the user experience, immersive audio may support adjusting the presented immersive audio / visual scene according to the user's movements. For example, it may be desirable to track the user's head position and head movements during audio rendering and adjust the audio accordingly. To this end, immersive audio experiences may handle head movements using a three-degrees-of-freedom (3DoF) or six-degrees-of-freedom (6DoF) model.
[0004] Various immersive audio services, such as Immersive Voice and Audio Services (IVAS), can be used to render high-quality audio reproductions on XR devices, including awareness of pose information, which may include metadata about the user's head position relative to their relative or absolute movements. However, making such adjustments in response to pose information can require high computational power to achieve a high-quality immersive audio experience.
[0005] The computational complexity requirements for immersive audio can be problematic for small form-factor devices like AR glasses. To make such AR glasses as practical and user-friendly as possible, designers may avoid using powerful processors and heavy batteries. Otherwise, user-worn devices could become bulky, expensive, heavy, consume more power, and generate significant heat. Therefore, to enable low-power operation with low latency and a reasonable form factor, such AR devices tend to have reduced-complexity, computationally constrained processors.
[0006] The present disclosure recognizes the above problem and explores potential solutions. One potential solution is to reduce audio rendering requirements at an end device (e.g., an AR device operated by a user) through a split-rendering topology that leverages processing from other entities (e.g., network-based devices) in the mobile / wireless network to which the end device is connected or tethered (e.g., via a network or cloud-based connection). For example, a powerful network entity such as a mobile user equipment (e.g., UE, a device used by an end user, a handheld multifunction device, a game console, a cloud-based resource, etc.) may be connected to the end device to assist in split rendering of immersive audio. Pose information based on user movement may be collected at the end device and transmitted to the network entity. The end device may only receive audio that has already been rendered from the network entity. For example, high-complexity computations, such as processing 3DoF / 6DoF pose information (e.g., head tracking metadata), may be performed by the rendering entity (e.g., network entity). One issue with the split rendering topology described above is that the transmission latency between the end device and the network entity can be as high as 100 ms, meaning that the network entity may be relying on outdated pose / head tracking information. Due to this delay, the audio rendered from the network entity may not match the user's current head pose / head position at the end device. If the motion-to-sound latency is too high, the end user will perceive a degradation in the quality of their immersive experience.
[0007] Document US. 63 / 340,181 discloses a novel approach to interactive head tracking. The described approach generates multiple binaural representations corresponding to various head poses in a main device or pre-renderer, and calculates metadata that can be used with a reference binaural signal to reconstruct binaural output corresponding to any given pose in a post-renderer. The reference binaural signal and metadata are transmitted to a post-rendering device. Based on the received binaural signal and metadata and the difference between the reference pose and the user's detected current head pose, the post-renderer determines the binaural audio corresponding to the current head pose. This disclosure recognizes that the metadata requirements for head pose information required in this type of solution can be significant. For example, if the current head pose deviates significantly from the reference head pose, a large amount of metadata will be transmitted to the post-rendering device to cover all possible head poses.
[0008] It is with respect to these and other considerations that the present disclosure has been made. Summary of the Invention
[0009] This disclosure describes techniques for split rendering of immersive audio.
[0010] The object of the present invention is to solve this problem and enable efficient split rendering even in situations where the user's head pose is expected to change significantly.
[0011] In some embodiments, a method of audio processing in a main device is described, the method including: receiving a first bitstream; decoding the first bitstream to obtain decoded immersive audio content; receiving a second bitstream; decoding the second bitstream to obtain pose information about a user of the lightweight processing device; determining a first head pose based on the pose information; rendering a downmix representation of the immersive audio content corresponding to the first head pose; selecting a set of second head poses for the first head pose; rendering a set of binaural representations of the immersive audio content corresponding to the second set of head poses; calculating reconstruction metadata including the first head pose to enable reconstruction of the set of binaural representations from the downmix representations; encoding the downmix representations and the reconstruction metadata into a third bitstream; and outputting the third bitstream.
[0012] In some additional embodiments, a method for processing audio in a lightweight processing device is described, the method including: receiving a bitstream from a main device; decoding the bitstream to obtain a downmix representation of immersive audio content associated with a first head pose; first reconstruction metadata including the first head pose that enables reconstruction of a set of binaural representations associated with a second set of head poses from the downmix representation; and obtaining a set of second head poses associated with the first reconstruction metadata. The method further includes detecting a current head pose of a user of the lightweight processing device, transmitting the current head pose to the main device, and calculating output binaural audio based on the downmix representation, the first reconstruction metadata, the set of second head poses, and a relationship between the first head pose and the current head pose.
[0013] In some further embodiments, the downmix representation is the first binaural representation, hi other embodiments, the downmix representation comprises a mono signal formed by combining channels in a multi-channel representation of the immersive audio content.
[0014] "Lightweight processing device" is intended to include any user device that has limited capabilities and therefore may be unsuitable for real-time binaural rendering. In some examples, "lightweight processing device" refers to the physical weight of the device. In other examples, "lightweight processing device" refers to the processing power of the device. In typical lightweight device examples, battery capacity and processing power may be limited so that the physical device remains in a small form factor.
[0015] Existing techniques for head-tracked split rendering require expensive physical components (e.g., powerful processors that require large heat sinks or active cooling components) that require more processing resources than necessary, waste device energy, and often result in heavy and cumbersome devices. These considerations are particularly important in battery-powered and wearable devices.
[0016] The techniques of this disclosure thus provide electronic devices with a faster, more efficient method for head-tracked split rendering, optionally complementing or replacing other methods of head-tracked split rendering. In battery-powered wearable computing devices, such a method conserves power, extends the time between battery charges, and allows for more comfortable devices to be built at lower cost.
[0017] According to some embodiments, a method is described that is executed on one or more electronic devices. The method includes: a first main processing device receives immersive audio; acquires (current) user posture information; the first device determines a downmix signal from the immersive audio, the downmix signal including at least one channel; determines a set of N (e.g., N≧1) predicted postures based on the acquired user posture information; determines a set of binaural representations from the immersive audio corresponding to the set of N predicted postures; generates metadata from the downmix signal and at least one of the set of binaural representations and a metadata model; and provides the metadata to a second lightweight processing device different from the first device. According to some embodiments, the acquisition of the user posture information is performed at least in part by the second device, and includes providing (e.g., transmitting) data corresponding to the acquired user posture information from the second device to the first device.
[0018] According to some embodiments, the method includes a renderer of the second device rendering a downmix signal into output binaural audio based at least in part on the metadata, the acquired user pose information, and the updated user pose information. According to some embodiments, the downmix signal is a binaural signal generated using a set of HRTFs or a set of BRIRs and the acquired user pose information. According to some embodiments, determining a set of predicted poses includes calculating N poses (hereinafter referred to as yaw angles) corresponding to N predicted angles along a yaw axis by changing a head pose yaw angle derived based on the acquired user pose information by a first predetermined value (e.g., an angle specified in degrees or radians) in a first direction to obtain a first predicted yaw angle among the N predicted yaw angles. According to some embodiments, the method includes obtaining a second predicted yaw angle among the N predicted yaw angles by changing a pose yaw angle derived based on the acquired user pose information by a second predetermined value in a second direction (e.g., counterclockwise, clockwise).
[0019] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more computer programs configured to be executed by one or more processors of a computing device, the one or more computer programs including instructions for: a first device receiving immersive audio; acquiring user posture information; determining a downmix signal from the immersive audio, the downmix signal including at least one channel; determining a set of N (e.g., N≧1) predicted postures based on the acquired user posture information; determining a set of binaural representations from the immersive audio corresponding to the set of N predicted postures; generating metadata from the downmix signal and at least one of the set of binaural representations and a metadata model; and providing the metadata to a second device different from the first device. According to some embodiments, acquiring the user posture information is performed at least in part by the second device and includes providing (e.g., transmitting) data corresponding to the acquired user posture information from the second device to the first device.
[0020] According to some embodiments, the one or more computer programs include instructions for rendering, by a renderer of the second device, a downmix signal into output binaural audio based at least in part on the metadata, the acquired user pose information, and the updated user pose information. According to some embodiments, the downmix signal is a binaural signal generated using a set of HRTFs or a set of BRIRs and the acquired user pose information. According to some embodiments, the one or more computer programs include instructions for determining a set of predicted poses, including calculating N poses corresponding to the N predicted yaw angles by changing a head pose yaw angle derived based on the acquired user pose information in a first direction by a first predetermined value (e.g., an angle specified in degrees or radians) to obtain a first predicted yaw angle among the N predicted yaw angles. According to some embodiments, the one or more computer programs include instructions for obtaining a second predicted yaw angle among the N predicted yaw angles by changing a pose yaw angle derived based on the acquired user pose information in a second direction (e.g., counterclockwise, clockwise) by a second predetermined value.
[0021] According to some embodiments, an apparatus is described. The apparatus comprises one or more processors and a memory storing one or more computer programs configured to be executed by the one or more processors, the one or more programs including instructions for a first device receiving immersive audio, acquiring user pose information, determining a downmix signal from the immersive audio including at least one channel, determining a set of N (e.g., N≧1) predicted poses based on the acquired user pose information, determining a set of binaural representations from the immersive audio corresponding to the set of N predicted poses, generating metadata from the downmix signal and at least one of the set of binaural representations and a metadata model, and providing the metadata to a second device different from the first device. According to some embodiments, acquiring the user pose information is performed at least in part by the second device and includes providing (e.g., transmitting) data corresponding to the acquired user pose information from the second device to the first device.
[0022] According to some embodiments, the one or more computer programs include instructions for rendering, by a renderer of the second device, a downmix signal into output binaural audio based at least in part on the metadata, the acquired user pose information, and the updated user pose information. According to some embodiments, the downmix signal is a binaural signal generated using a set of HRTFs or a set of BRIRs and the acquired user pose information. According to some embodiments, the one or more computer programs include instructions for determining a set of predicted poses, including calculating N poses corresponding to the N predicted yaw angles by changing a head pose yaw angle derived based on the acquired user pose information in a first direction by a first predetermined value (e.g., an angle specified in degrees or radians) to obtain a first predicted yaw angle among the N predicted yaw angles. According to some embodiments, the one or more computer programs include instructions for obtaining a second predicted yaw angle among the N predicted yaw angles by changing a pose yaw angle derived based on the acquired user pose information in a second direction (e.g., counterclockwise, clockwise) by a second predetermined value.
[0023] The embodiments described herein may be generally described as technology, where the term "technology" may refer to systems, devices, methods, computer-readable instructions, modules, components, hardware logic, and / or operations as the context suggests as applied herein.
[0024] Features and technical advantages in addition to those expressly described above will become apparent upon reading the following detailed description and upon reference to the associated drawings. The foregoing Summary of the Invention is provided to introduce some technology in a simplified form and is not intended to identify key or essential features of the claimed invention as defined by the appended claims. [Brief explanation of the drawings]
[0025] The invention will now be explained in more detail with reference to the accompanying drawings. [Figure 1] FIG. 1 is a block diagram illustrating an example of low-complexity, low-bitrate prediction-based split rendering using a downmix signal according to an embodiment of the present invention; [Figure 2] 4 is a flowchart illustrating processing in a main processing device according to an embodiment of the present invention. [Figure 3] 4 is a flowchart illustrating processing in a lightweight processing device according to an embodiment of the present invention. [Figure 4] FIG. 1 is a block diagram illustrating an example of low-complexity, low-bitrate prediction-based split rendering using model-based prediction, according to an embodiment of the present invention. [Figure 5] FIG. 1 is a schematic block diagram of an example device or architecture that may be used to implement embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. The accompanying drawings show specific exemplary configurations in which the inventive concepts may be implemented. These configurations are described in sufficient detail to enable the techniques described in this disclosure to be practiced. It is to be understood that other configurations may be utilized and other changes may be made without departing from the spirit or scope of the concepts presented. Accordingly, the following detailed description is not intended to limit the scope of the inventive concepts, which are defined only by the appended claims.
[0027] The embodiments of the invention disclosed herein contemplate compatibility and consistency with the use of immersive audio codecs such as IVAS in XR applications. In particular, the inventive concepts described in detail below are applicable to systems, devices, architectures, methods, and techniques in which primary decoding and pre-rendering is performed by a high-resource main device (UE) (e.g., an edge node or other network node / server in a 5G system, a high-performance mobile device, etc.) having powerful computational (or processor) resources and having high power or high battery capabilities, and final decoding and post-rendering is performed by a low-resource device (e.g., a lightweight device, a wearable device, AR glasses, a head-mounted display, a head-up display, etc.) relative to the main device.
[0028] Embodiments of the proposed techniques, systems, devices, methods, and computer-readable instructions for low-complexity and low-bitrate prediction-based split rendering may include operations such as: 1. Receiving pose P' (first head pose) information from the post-renderer (lightweight processing device) to the pre-renderer (main device). 2. In the pre-renderer, generating a one or two channel downmix signal from the received immersive audio. The downmix signal can be a binaural signal rendered using a set of HRTFs (or BRIRs) and a pose P', or the downmix signal can be a combination of a prototype signal and zero or more diffuse signals. 3. In the pre-renderer, determine a set of N second poses Pn that are close to the first pose P' (N≧1 and 1≦n≦N). 4. In the pre-renderer, P1 to P N Using a set of poses and HRTFs (or BRIRs), generate N binaural representations. 5. Posture P nOptimized the multi-binauralization process by reusing HRTF-filtered channels that do not change between channels. 6. Calculating prediction gains for predicting correlated components in the N binaural representations with respect to one or more downmix signals. 7. Calculate the diffusion gain parameter to satisfy the uncorrelated energies. 8. Compute the model parameters of the model that approximates the evolution of the metadata (prediction and diffusion gains) as a function of the current (actual) head pose. 9. Encode the first head pose P', the downmix signal and the metadata (prediction and diffusion gain or model parameters) and send the decoded bitstream to a post-renderer device. 10. In the post-renderer, decoding of the first head pose P', the downmix signal and the metadata (prediction and diffusion gain or model parameters). 11. In the post-renderer, adjust the prediction and diffusion gains based on the difference between the current head pose P and the first head pose P', the received metadata, and (optionally) the model. 12. In the post-renderer, reconstruct the binaural output corresponding to the current head pose P by applying the adjusted metadata coefficients to the decoded downmix signal.
[0029] It should be noted that one or more aspects of the proposed techniques, systems, devices, methods, and computer-readable instructions described herein, including those described above, do not require the particular order or sequential order shown to achieve desired results. Also, other steps may be provided or steps described herein may be omitted, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the claims following the main description.
[0030] 1 is a block diagram of an example system for low-complexity, low-bitrate prediction-based split rendering using a downmix signal, according to some embodiments. As shown, the example system may include a first device 10 (or a main processing device) and a second device 20 (or a lightweight processing device).
[0031] The first device 10 (or main processing device, or pre-renderer) includes a decoder 11, a downmixer 12, a head pose decoder 13, a binaural renderer 14, a metadata generator 15, a first encoder 16, a second encoder 17, and a multiplexer 18. The decoder 11 is, for example, an IVAS decoder, and is configured to receive and decode a bitstream b1 to decode immersive audio content A. The downmixer 12 is configured to receive the immersive audio content and provide a downmix representation Dmx of the audio content. The head pose decoder 13 receives a bitstream b1 including head pose information and decodes it to provide a downmix representation Dmx of the audio content. p and generating a first head pose P'. The binaural renderer 14 is configured to receive the first head pose P' and the immersive audio content A, and to render one or more binaural representations corresponding to the first head pose P'. The metadata generator 15 is configured to receive the downmix Dmx and the binaural representations, and to generate reconstruction metadata M that enables reconstructing the binaural representations from the downmix. The metadata M includes the first pose P'. The first encoder 16 receives the downmix Dmx and, in response, encodes the downmix Dmx into an encoded bitstream b 11 The second encoder 17 is configured to receive the reconstruction metadata M (including the first attitude P′) and responsively encode the reconstruction metadata into an encoded bitstream b 12 The multiplexer 18 is configured to encode the output of the two encoders as the coded bit stream b 11 and b 12and responsively combines the coded bits into bitstream b2. The main device may include an interface for outputting bitstream b2 so that the bitstream may be subsequently transmitted or otherwise made available to other devices external to main device 10.
[0032] The second device 20 (lightweight processing device or post-renderer device) includes a demultiplexer 21, a first decoder 22, a second decoder 23, a head tracker 24, a pose information encoder 25, and a binaural reconstruction block 26. The second device, i.e., the lightweight processing device 20, may be a device held by a user. The demultiplexer 21 receives the bitstream b2 and splits it into two encoded bitstreams b 21 and b 22 The decoder 22 is configured to separate the encoded bitstream b 21 and accordingly generates a bitstream b 21 The decoder 23 is configured to decode the coded bitstream b 22 and accordingly generates a bitstream b 22 into metadata M' including a first pose P'. The head tracker 24 is configured to detect the user's head position and, in response, generate pose information including, for example, the current (actual) user's head pose P. The pose information encoder 25 receives the pose information from the head tracker 24 and, in response, generates a bitstream b P The binaural reconstruction block 26 is configured to receive the current user's head pose P, the downmix signal Dmx′, and the metadata M′ including the first pose P′, and to determine a binaural output based on the relationship between the downmix Dmx′, the metadata M′, the current head pose P, and the first head pose P′.
[0033] In FIG. 1 , the lightweight device / post-renderer device 20 (shown as a block on the right side of FIG. 1 ) receives and encodes a pose P (e.g., data representing the pose of a user / wearer of a lightweight device is encoded into a representation suitable for transmission), and transmits the pose P (e.g., the encoded data, bitstream b) to the main / pre-renderer device 10 (shown as a block on the left side of FIG. 1 ) via a data channel (e.g., a back channel). p The main device / pre-renderer device 10 transmits the received posture data b p is received (for example, at decoder 13) and decoded to obtain P′, which may be a delayed and quantized version of pose P.
[0034] In some exemplary embodiments, the pose information received by the lightweight / post-renderer device 20 (shown as a block on the right side of FIG. 1 ) includes not only the pose P but also one or more parameters V associated with the head movement (e.g., rotation including angular velocity, acceleration or deceleration of the user's head rotation, etc.). The pose information encoder 25 then performs n-th order pose prediction (e.g., using motion data and / or pose data associated with a pose at a first time to predict a pose associated with a second time, which is either later or earlier than the first time) to generate a predicted pose P″, and encodes the pose P″ (e.g., into a bitstream b p ) via a data channel (e.g., a back channel) to the main / pre-renderer device 10 (shown as the block on the left side of FIG. 1). The main / pre-renderer device p is decoded (e.g., by decoder 13) to obtain the first head pose P′, which may be a delayed and quantized version of the predicted pose P″. In some exemplary embodiments, the pose information received by the lightweight / post-renderer device 20 (shown as a block on the right side of FIG. 1 ) includes not only the pose P but also one or more parameters V associated with the head movement (e.g., angular velocity, acceleration, or rotation including deceleration of the user's head rotation, etc.). The pose information encoder 25 then encodes the pose P and the parameters V (e.g., data representing the pose and movement of the user / wearer of the lightweight device is encoded into a representation suitable for transmission) and outputs the encoded data (e.g., as a bitstream b p ) via a data channel (e.g., a back channel) to the main / pre-renderer device 10 (shown as the left block in FIG. 1). The main / pre-renderer device then converts the received pose and motion data, e.g., in a decoder 13, into delayed and quantized versions of the pose P and parameters V, respectively, b p In this case, the main device 10 applies n-th order pose prediction based on the received pose and movement data to generate a first head pose P′ (e.g., using movement data and / or pose data associated with the pose at the first time to predict a pose associated with a different time, e.g., a second time, which may be later or earlier than the first time).
[0035] In some exemplary embodiments, the lightweight / post-renderer device 20 (shown as the block on the right in FIG. 1 ) receives a pose P (e.g., data representing the pose of the user / wearer of the lightweight device 20) and does not transmit that pose to the main / pre-renderer device 10 (shown as the block on the left in FIG. 1 ). In such embodiments, the main / pre-renderer device blindly assumes a first head pose P′ based on a default applicable to the operation of that device. This case may apply in a 1-to-n distribution scenario, such as broadcasting to multiple devices (e.g., multiple lightweight / post-renderer devices), when no back channel exists.
[0036] As shown in Figure 1, a main device / pre-renderer 10 receives an immersive audio signal including audio content A (e.g., the output of an immersive decoder 11 such as IVAS, a QMF signal, etc.). The audio content A is converted into a downmix signal Dmx (e.g., by a downmixer 12) using a first head pose P'. In some embodiments, Dmx may include one channel, and in some other embodiments, Dmx may include two or more channels (e.g., at least two channels).
[0037] The renderer 14 generates one or more poses P that are inferred from the pose P' from the audio content A. n One or more binaural representations BIN corresponding to n where 1≦n≦N and N≧1. One or more poses (a set of second head poses) may be determined based on a set of predefined offsets to the first head pose P′. A metadata generator (e.g., generator 15) generates metadata M based on the Dmx signal and the binaural signal BINn such that any of the BINn binaural signals can be reconstructed using the Dmx signal and the metadata M. The downmix representation Dmx is stored in the bitstream b 11 The metadata M is then quantized and encoded (e.g., by encoder 17) to produce a bitstream b 12 Generate the bitstream b 11 and b 12 is combined by multiplexer 18 into bitstream b2.
[0038] In some embodiments, the downmix representation includes two signals. In this case, the metadata should allow reconstruction from two signals (downmix) to two signals (binaural output). A 2x2 matrix is an efficient way to enable such reconstruction. In some embodiments, the metadata M includes a 2x2 matrix for each time unit and each frequency band, i.e., for each time-frequency tile.
[0039] As shown in FIG. 1, in the lightweight device / post-renderer 20, b2 is received and demultiplexed into bitstream b 21 and bitstream b 22 Bitstream b 21 is fed to a decoder (e.g., decoder 22) that reconstructs the downmix signal Dmx and produces a reconstructed downmix representation Dmx′. 22 is provided to an MD decoding and dequantization (un-quant) block (e.g., decoder 23), which reconstructs the metadata M and generates reconstructed metadata M'. As mentioned above, this metadata M' also includes the first pose P'. The downmix representation Dmx' and the metadata M' are then provided to a binaural reconstruction block 26, which generates a head-tracked binaural output using Dmx' and the metadata M', the second set of head poses, and the relationship between the current head pose P and the first pose P'.
[0040] The lightweight device acquires information about a set of second head poses to which the reconstruction metadata is associated in order to enable binaural reconstruction. n In embodiments where the set of m is determined by applying a set of offsets to the first head pose, these offsets may be known in advance (e.g., applied by reconstruction block 26), or may be included in the metadata M received in bitstream b2.
[0041] The reconstruction may involve first computing (e.g., by interpolation) modified reconstruction metadata from the current head pose P and the metadata M′, and then applying this modified metadata to the downmix signal Dmx′.
[0042] In the exemplary embodiment with N=2, the downmixer 12 converts the Dmx signal into a first (reference) binaural signal BIN using a set of HRTFs (or BRIRs) and a first head pose P′.ref This is a binaural renderer that generates the pose P n are P'+X, P'-X', where X and X' are the expected deviations in yaw angle between P' and P. The renderer 14 generates two binaural outputs BIN corresponding to the poses P'+X and P'-X'. n Generate the reference binaural signal BIN ref and posture P n Binaural signal BIN corresponding to n are then fed to a metadata generation block 15 which generates metadata M corresponding to the poses P'+X and P'-X'. The metadata M is quantized and coded by the MD quantization and coding block 17. ref The signal is encoded by the encoder 16. The multiplexed bit stream b2 is ref The signal and metadata M are sent to a post-renderer device 20 which decodes and feeds it to a binaural reconstruction block 26. The reconstruction block 26 interpolates or extrapolates the metadata based on the differences between P', P'+X, and P'-X' and the current head pose P. The interpolation may be based on linear, triangular, or sine or cosine based models, etc. The reconstruction block 26 uses the BIN as proposed in U.S. Provisional Application No. 63 / 340,181 (incorporated herein by reference). ref Apply the interpolated metadata to the head-tracked binaural signal BIN out In one example implementation, BIN ref The decorrelating coefficients are used directly, using the sum of the left and right channels of , avoiding the use of a decorrelator.
number
[0043] where z l,p [n] and z r,p [n] is the n-th sample of the left and right channels of the reconstructed BIN signal according to the current head pose P. M pis the (2x2) prediction coefficient mixing matrix, and y l,po [n] and y r,po [n] is BIN ref are the nth samples of the left and right channels of the signal, and g p,p is the de-correlation coefficient. M p and g p,p The calculation of is similar to that described in US Provisional Application No. 63 / 340,181.
[0044] In some embodiments, the downmixer 12 generates as the Dmx signal a combination of a mono channel (prototype signal) and zero or more diffuse channels (diffuse signal). The mono signal S may for example be formed as a combination of channels of a multi-channel representation of the immersive audio content A, for example a combination of signals of a first binaural representation. The diffuse signal D may be formed as a combination of diffuse components of the same multi-channel representation of the immersive audio content A.
[0045] In some embodiments, such operations may be applied in the time domain, the CQMF domain, the subband domain, or the frequency domain, and all coefficients that are subject to or result from such operations may be complex. ref From the signal, a signal is generated as S=aL+bR, and a diffuse signal is generated as D=cL+dR, where L and R are the BIN ref are the left and right channels of the signal, a and b are dynamically calculated or statically determined gain parameters, e.g., a=0.5, b=0.5; c and d are the BIN ref are dynamically calculated using the covariance of the L and R channels of the signal. S is the prototype signal and D is the spread signal. In one embodiment, a, b, c, d are calculated as follows:
[0046] BIN ref The covariance of
number
number
number
number
number
number
number
number
number
number
[0047] The renderer 12 generates two binaural outputs BIN corresponding to the poses P'+X and P'-X'. nThen, the prototype signal S and the spread signal D and BIN n The signals are fed to a metadata generation block 15 which generates metadata M corresponding to the P'+X and P'-X' signals. x and R x If P′+X are the left and right signals corresponding to P′+X, then the metadata corresponding to the P′+X signal can be calculated as follows:
number
number
number
number
number
number
[0048] From this metadata and the downmix signals S and D, the P′+X channel can be reconstructed by the reconstruction block 26 as follows:
number
number
[0049] Similarly, the metadata of P'-X can be calculated and the P'-X binaural signal can be reconstructed from the metadata, the prototype signal S and the diffuse signal D.
[0050] In some embodiments, due to bitrate limitations, it may be desirable to encode only one channel, in which case only the prototype signal is encoded and metadata is generated as follows:
number
number
number
number
number
number
[0051] From this metadata and the prototype signal, the P'+X channel can be reconstructed by the reconstruction block 26 as follows:
number
number
[0052] In some embodiments, the first head pose P' may be transmitted to the lightweight processing device 20 for better pose synchronization (e.g., as metadata). If the current head pose P differs from P', P'+X, and P'-X', the reconstruction block 26 interpolates or extrapolates the metadata based on the differences between P', P'+X, and P'-X' and the current head pose P. The interpolation can be based, for example, on linear, triangular, or sine- or cosine-based models, etc. The reconstruction block 26 uses the BIN ref Apply the interpolated metadata to the head-tracked binaural signal BIN out Generate.
[0053] In some embodiments, X is equal to X′ and the orientation P n is P'+X, P'-X, where X is the expected deviation in yaw angle between P' and P. In other example implementations, X is not equal to X', and X' may be less or greater than X based, for example, on the angular velocity and acceleration or deceleration of the user's head rotation.
[0054] 2 is a flowchart illustrating processing in the main device 10 (or the first device) according to an embodiment of the present invention. The flowchart may be divided into various blocks or sections, such as S11 to S18. The processing of the various blocks in FIG. 2, which may be described as operations, processes, methods, steps, acts, or functions, may begin with block S11.
[0055] In step S11 (receiving and decoding a bitstream or receiving and decoding a first bitstream), a first bitstream is received and decoded (e.g., by the decoder 11) to obtain decoded immersive audio content A. In step S12 (receiving and decoding posture information or receiving and decoding posture information), a second bitstream is received and decoded (e.g., by the decoder 13) to obtain posture information associated with a user of the lightweight processing device (e.g., 20). In step S13 (determining P' or determining P'), a first head pose P' may be determined (e.g., by the head pose decoder 13) based on the posture information. In step S14 (downmixing audio or downmixing audio), a first downmix of the immersive audio content A may be determined (e.g., by the downmixer 12), where the first downmix is a representation of the immersive audio content corresponding to the first head pose. Step S15 (BIN n Rendering or BIN n In step S16 (rendering M or generating M), a pair of binaural representations of the immersive audio content is rendered (for example, by the renderer 14), where the set of binaural representations corresponds to the second set of poses. In step S17 (encoding or encoding), the downmix representations are encoded (for example, by the encoder 16), and the reconstruction metadata including the first head pose P′ is encoded (for example, by the encoder 17). In step S18 (outputting or outputting), a bitstream b2 including the first downmix representation Dmx and the reconstruction metadata M is output. The outputting step may include transmitting the bitstream b2 to a lightweight processing device (for example, the lightweight processing device 20) that received the pose information.
[0056] 3 is a flowchart illustrating processing in lightweight device 20 (or a second device or a device held by a user) according to an embodiment of the present invention. The flowchart may be divided into various blocks or sections, such as S21 through S25. The processing of the various blocks in FIG. 3, which may be described as operations, processes, methods, steps, acts, or functions, may begin with block S21.
[0057] The process includes, in step S21 (receiving and decoding bitstreams), receiving a bitstream b2 from the main device 10 (e.g., by decoders 22, 23) and obtaining a downmix representation Dmx of the immersive audio content A, a first head pose P', and first reconstruction metadata M' that enables the reconstruction of a set of pairs of binaural representations BINn from the downmix representation Dmx. Prior to step S21, the bitstream is demultiplexed (e.g., by demultiplexer 21) into two or more bitstreams b 21 , b 22 Step S22 (detecting current head pose) involves detecting the current head pose P of the user of the lightweight processing device 20 (e.g., by the head tracker 24). Step S23 (transmitting head pose) involves transmitting the current head pose P to the main device 10 (e.g., by the head pose encoder 25). Finally, step S25 (calculating binaural audio) involves calculating the output binaural audio BIN based on the downmixed presentation Dmx, the first reconstruction metadata M′, and the relationship between the first head pose P′ and the current head pose P (e.g., by the reconstruction block 26). out Step S25 is optionally preceded by step S24 (calculating second reconstruction metadata) of calculating second reconstruction metadata based on the first reconstruction metadata, the first head pose and the current head pose. In this case, step S25 may use this second reconstruction metadata to obtain a binaural output.
[0058] FIG. 4 illustrates an exemplary implementation of low-complexity, low-bitrate prediction-based split rendering according to some embodiments.
[0059] Elements in Figure 4 that correspond to elements in Figure 1 are given the same reference numerals. In addition to these elements, the main device 110 shown in Figure 4 also includes a model-based estimate M of the metadata M. mod The modeling block 19 is configured to receive the immersive audio content A from the decoder 11 and the first pose P′ from the pose decoder 13, and in response to this, to provide a model-based estimate M of the reconstruction metadata M. mod It is configured to generate a model-based estimate M mod is provided to the encoder block 17. Similarly, the lightweight processing device 120 22 from the demultiplexer 21 and in response thereto provides to the decoder 23 a model-based estimate M' mod and a corresponding modelling block 27 configured to generate
[0060] As shown in FIG. 4, the main device / pre-renderer 110 and the lightweight device / post-renderer 120 generate a first model-based estimate M of the predicted metadata parameters. mod or M' mod These are the mathematical models used to generate all the poses P n The prediction coefficient mixing matrix M p In some embodiments, these are estimates of all poses P n Prediction coefficient Pred L and Pred R The metadata quantization and coding block then calculates the metadata parameters M and the corresponding model-based estimates M mod Similarly, in the post-renderer 120, the prediction metadata parameters M' applied in the reconstruction block 26 are simply encoded as the residual between the reconstructed model-based estimate M' modand the reconstructed residual.
[0061] According to one example of a mathematical model for generating predicted metadata estimates19,27, the predicted metadata parameters for a pose P'+X are obtained by facilitating delay and gain / shape operations corresponding to multiplication with the complex predicted parameters in the complex QMF domain. The input parameters of the model are the direction of arrival (DOA) parameters of the dominant sound source in a given QMF band, the azimuth and elevation angles of the pose, and possibly each HRTF (or HRIR or BRIR) coefficient or at least the associated coefficients. Note that the parameter M mod can be efficiently coded by indexing in a codebook of HRTFs (or associated codebook entries).
[0062] A further example implementation of low-complexity, low-bitrate, prediction-based split rendering, according to some embodiments, may rely on a mathematical model of how metadata parameters change when the current head pose P differs from P′ by a certain amount Δ. Compared to the linear or triangular interpolation described above, more sophisticated techniques can rely on specific mathematical properties of the parameter evolution. One such property is symmetry. In the following related discussion, we assume that there is a dominant sound source in a certain frequency or QMF band, and that the DOA of that sound source is known. In that case, the azimuth angle of that DOA can be specified as 0 degrees or 180 degrees, meaning that the x-axis of an assumed Cartesian coordinate system coincides with the DOA.
[0063] For example, assuming that the HRIR / BRIR are symmetrical and the attitude P' is aligned with the x-axis (i.e., the azimuth angle is 0 degrees or 180 degrees), the metadata parameters applied to the left and right output channels for an azimuth attitude shift X are identical to the metadata parameters applied to the swapped output channels (right, left) for the corresponding azimuth attitude shift −X.
[0064] Furthermore, under given assumptions, the parameters or intermediate parameters for deriving the metadata parameters may exhibit odd-symmetry with respect to the parameters of the pose P′, i.e., M(P′+Δ)=−M(P′−Δ) (thereby ignoring possible constant offsets). This symmetry can be exploited, for example, when pre-rendering assuming the adjusted pose P′. The symmetry property allows limiting pre-rendering to the poses P′ and P′+X and skipping pre-rendering for P′-X′. This reduces the complexity of a single rendering process in the pre-renderer 110 and avoids sending metadata parameters for the pose P′-X′.
[0065] Another case is when the (adjusted) pose P' coincides with the y-axis, i.e., pose P' is such that the DOA of the main sound source is left or right. Changing the current head pose by a small amount Δ means that the sound is effectively coming from the left or right, but slightly from the front or back. A good approximation in this case is that the metadata parameters (or intermediate parameters) now exhibit even symmetry, i.e., M(P'+X)=M(P'-X).
[0066] The property of symmetry can be further exploited in modeling the evolution of the metadata (or intermediate) parameters as a function of the pose deviation Δ. For example, this function can be expressed as a Taylor series of the following type:
number
[0067] Further considering the symmetry property in the Taylor series expansion approach, it may be useful to align the pose P' with the x-axis or y-axis, i.e., align P' with the adjusted x-axis or y-axis. In the first case, the even terms (except for i=0) vanish (coefficient a 2j= 0 for any positive integer j). This makes modeling using linear (first-order) terms very accurate, and in many cases, it is not necessary to consider higher-order terms. Similarly, when P' coincides with the y-axis, odd terms vanish due to even symmetry (coefficient a 2j-1 = 0 for any positive integer j), which makes modeling with a single quadratic term very accurate and efficient.
[0068] In summary, by exploiting the symmetry property, the need to pre-render at the three poses P', P'+X, and P'-X' can be reduced, or at least the amount of metadata to be transmitted can be reduced. In practice, rather than transmitting metadata for P'+X and P'-X, it may be more efficient to transmit Taylor series coefficients and DOA angles to indicate the dominant sound direction.
[0069] Another mathematical property of the metadata parameters (or intermediate parameters) is their 360° periodicity with respect to the azimuth angle.
number
[0070] The interaural time difference for a rendered plane wave signal incident from a given azimuth angle α can be modeled by a sinusoidal expression as follows:
number
[0071] Interaural level differences can also be approximately modeled with a similar formula.
[0072] Therefore, a possible approximation of the meta (or intermediate) parameters involves applying the corresponding sinusoidal formula. More generally, these parameters can be efficiently represented by a few low-order harmonics of a discrete Fourier series.
number
[0073] In this equation, the zeroth-order term represents a constant (offset), and the first and second-order (and higher) sinusoids model the variation of a particular periodic metadata parameter. k are generally complex-valued and may depend, for example, on other parameters such as the interaural distance of the assumed listener's head in addition to the first head pose P′ and the DOA of the dominant sound direction. According to the model-based approach described above, the coefficients are determined in the pre-renderer 10 and are used to calculate the approximate metadata parameters M mod are applied to generate , quantized, encoded, and then sent to a post-renderer which decodes and applies them in the model.
[0074] A further embodiment relies solely on the model. In that case, the main / pre-renderer device 110 can send only the model parameters and the first head pose to the post-renderer device 20, thereby significantly reducing the amount of metadata transmitted. In this case, no encoding or residual metadata parameters are transmitted. The renderer 14 and generator 15 still rely on the pose P n , N may be used to generate metadata for the received immersive audio content A. However, in that case, the generated metadata may simply be used to optimize the accuracy of the model parameters. It is also possible to set N to zero, which means not to use the renderer 14 at all. In that case, the model parameters are calculated only from the received immersive audio content A and associated metadata parameters such as DOA angles that may be part of the received immersive audio signal representation.
[0075] It should be noted that in the above example, the letter M generally represents, for example, the prediction gain Pred L or Pred R , Diffusion Gain Diff L or DiffR M may represent an intermediate parameter arising in the calculation of the metadata parameters, such as covariance, as used in the above embodiment.
[0076] Quantization and encoding of metadata parameters A method for quantizing metadata for a prediction-based split-rendering technique is described below.
[0077] In some exemplary embodiments, main device / pre-renderer 10 receives immersive audio signal / content A (e.g., the output of an immersive decoder such as IVAS, a QMF signal, etc.). Audio content A is converted (e.g., by downmixer 12) into a downmix signal Dmx using P'. Main device 10 may receive attitude P' from lightweight device 20, or may assume P' to be a certain attitude value without instruction from lightweight device 20. In some embodiments, Dmx has one channel. In some embodiments, Dmx has two or more channels (e.g., two channels).
[0078] The renderer 14 generates from A one or more poses (a second set of poses) P′ estimated from P′. n corresponds to one or more binaural representations BIN n (where 1≦n≦N and N≧1). A metadata generator (for example, generator 15) generates metadata M based on the Dmx signal and the binaural signal BINn so that any of the BINn binaural signals can be reconstructed using the Dmx signal and the metadata M. The Dmx signal is a bitstream b 11 The metadata M, including the pose P′, is then quantized and encoded (e.g., by encoder 17) to produce a bitstream b 12 Generate the bitstream b 11 and b 12is combined by multiplexer 18 into bitstream b2.
[0079] In the lightweight device / post-renderer 20, b2 is received and demultiplexed by the demultiplexer 21 into bitstream b 21 and bitstream b 22 Bitstream b 21 is fed to a first decoder 22 which reconstructs the Dmx signal and produces a reconstructed downmix representation Dmx' signal. 22 is fed to an MD decoding and unquantization block (e.g., second decoder 23) which reconstructs metadata M including the pose P' and produces reconstructed metadata M'. Dmx' and M' are then fed to a binaural reconstruction block 26 which produces head-tracked binaural output using Dmx' and metadata M' and the current head pose P.
[0080] Posture P n is known to the metadata quantizer (encoder 17), fewer assumptions can be made to more efficiently quantize the metadata corresponding to these poses. n The Dmx signal is constructed from a rotation matrix such that a binaural signal corresponding to P can be reconstructed from the Dmx signal. The Dmx signal is the binaural signal (BIN ref An exemplary representation of the metadata for the signal is shown below:
number
[0081] where z l,p [n] and z r,p [n] is the posture P n are the n-th samples of the left and right channels of the reconstructed BIN signal according to M p is the (2x2) prediction coefficient mixing matrix, and y l,po[n] and y r,po [n] is BIN ref are the nth samples of the left and right channels of the signal, and g p,p is the de-correlation coefficient. M p and g p,p The calculation of is similar to that described in US Provisional Application No. 63 / 340,181.
[0082] Metadata M p and g p,p A technique for efficiently quantizing and encoding is presented below.
[0083] Selecting the origin for quantization of metadata matrices Posture P n and the arrival angle of the reference binaural signal, the rotation matrix M r can be generated, and then M p This can be used as the origin for quantizing the matrix, ensuring that the distribution of quantization points is the same on both sides of the origin. This allows the rotation matrix M r This allows for fine quantization around P, limiting the minimum and maximum values that need to be coded and also limiting the number of quantization points. n If one or more poses from are close to the first head pose P′, then an identity matrix can be assumed as the origin of quantization.
[0084] Furthermore, if the azimuth angle (θ) and elevation angle (φ) of the sound source in the reference BINref signal are known, the attitude P n is the angle along the yaw axis compared to the reference attitude P'.
number
number
number
[0085] where exemplary values for x, x', y, and y' are as follows:
number
number
number
number
number
number
[0086] Exploiting Symmetries in the +X and -X Metadata Matrices for Quantization and Encoding Typical posture P n are symmetrically arranged around the first head pose P'. In one example implementation, N=2, and P'+X and P'-X are expressed as the metadata M as shown in equation (1). pis the pose corresponding to the pose for which P'-X is generated. Here, X may be a tuning parameter that is set based on the expected motion to sound delay of the system. Alternatively, X may be a constant (e.g., 15 degrees along the yaw axis, 0 degrees along the pitch axis, and 0 degrees along the roll axis). When the metadata corresponding to P'+X is calculated, the first head pose and the pose of P'+X can be used to extrapolate intermediate metadata corresponding to P'-X, which can be used to efficiently quantize and encode the actual metadata of P'-X.
[0087] Quantizing and encoding metadata by exploiting symmetries in the left and right channel metadata for a given pose Metadata matrix M p Typically, M has certain symmetries in its left and right channel entries that can be used to efficiently quantize and encode metadata. One symmetry in an example implementation is that the sum of the squares of the real parts of the elements in any row or column is assumed to be close to 1. Another symmetry used in an example implementation is M p For the real part of the matrix, element m ij m ji is expected to be close to M p For the imaginary part of the matrix, element m ij Ga-m ji These symmetries are used in some implementations to save quantization points. Alternatively, these symmetries can be used to reduce M p The set of matrix elements is M p The second set of elements in the matrix is differentially encoded, i.e., the difference between the two sets is encoded. The difference value is likely to be close to 0 in most cases, and can be efficiently encoded using an entropy encoder.
[0088] Differential Encoding Across Subframes and Subbands Posture P nThe metadata for the binaural channels corresponding to the subbands may be calculated in the wideband or band domain. In some implementations, it can also be encoded in the subband domain using a Complex Low Delay Filterbank (CLDFB). The temporal resolution of the metadata calculated by the CLDFB filterbank can be much smaller than the temporal resolution of the codec (e.g., IVAS) or renderer. In one example implementation, the temporal resolution of the renderer or codec is 20 ms, referred to as a frame, and the temporal resolution of the CLDFB domain metadata is 5 ms, referred to as a subframe. Because it can be assumed that the metadata does not change very frequently over time, metadata corresponding to one or more subframes within a frame is differentially encoded with respect to one or more subframes of the same frame. Differential encoding using subframes of the same frame minimizes the impact of packet loss during data transmission to lightweight devices. The encoded differential value is likely to be zero in most cases and can be efficiently encoded using an entropy encoder. In some embodiments, it is determined that the metadata is very infrequent, and therefore metadata corresponding to one or more frequency bands of a frame is differentially encoded with respect to one or more frequency bands of the same frame. The difference value to be coded is likely to be 0 in most cases and can be coded efficiently using an entropy encoder.
[0089] Variations The systems and methods disclosed in this disclosure may be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks does not necessarily correspond to the division into physical units; conversely, one physical component may have multiple functions, and one task may be performed cooperatively by multiple physical components.
[0090] The computer hardware may be, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a switch, a bridge, or any machine capable of executing instructions (sequential or otherwise) that specify operations to be performed by the computer hardware. Furthermore, this disclosure relates to any set of computer hardware that individually or jointly executes instructions to carry out any one or more of the concepts discussed herein.
[0091] FIG. 5 illustrates a schematic block diagram of an exemplary electronic device or architecture 200 (e.g., apparatus 200) suitable for implementing exemplary embodiments of the present disclosure. The architecture 200 may include, but is not limited to, a main processing device and a lightweight processing device, as described in connection with FIGS. 1 and 4 . As illustrated, the architecture 200 includes a central processing unit (CPU) 201 that can execute various processes, for example, according to a program stored in a read-only memory (ROM) 202 or, for example, according to a program read from a storage unit 208 and loaded into a random access memory (RAM) 203. The CPU 201 may be, for example, an electronic processor 201 and may include one or more processor cores; in some examples, the processor 201 may be multiple processors. The RAM 203 also stores data needed by the CPU 201 to perform various processes, as needed. The CPU 201, the ROM 202, and the RAM 203 are connected to one another via a bus 204. An I / O interface 205 is also connected to the bus 204.
[0092] The I / O interface 205 is connected to an input unit 206 which may include a keyboard, a mouse, etc., an output unit 207 which may include a display unit such as a liquid crystal display (LCD) and one or more speakers, a memory unit 208 which may include a storage device such as a hard disk, and a communication unit 209 which may include a network interface card such as a network card (e.g., wired or wireless).
[0093] In some embodiments, input 206 includes one or more microphones in different positions (depending on the host device) that allow for capturing audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0094] In some embodiments, output 207 includes a system with various numbers of speakers. Output 207 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0095] In some embodiments, the communication unit 209 is configured to communicate with other devices (e.g., via a network). The drive 210 is connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or other suitable removable medium, is attached to the drive 210, and a computer program read from the removable medium 211 is installed in the storage unit 208 as needed. In the present disclosure, the device 200 is described as including the above-mentioned components, but it will be understood by those skilled in the art that in actual use, some of these components may be added, deleted, and / or substituted, and all of these modifications or variations are within the scope of the present disclosure.
[0096] According to exemplary embodiments of the present disclosure, the above-described processes may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly stored on a machine-readable medium, the computer program including program code for performing the method. In such an embodiment, the computer program may be downloaded and implemented from a network via communication unit 209 and / or installed from removable media 211, as shown in FIG. 2 .
[0097] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or special purpose circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the various elements of Figures 1 and 4 described above may be executed by control circuitry (e.g., CPU 201 in combination with other components of Figure 5), such that the control circuitry performs the operations described in this disclosure.
[0098] Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, processor, and / or other computing device, which may include control circuitry. Although various aspects of the exemplary embodiments of the present disclosure are shown and described using block diagrams, flowcharts, or other illustrations, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of limited example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controllers, or other computing devices, or any combination thereof.
[0099] Furthermore, the various blocks illustrated in the flowcharts may be viewed as method steps and / or as operations of computer program code and / or as multiple coupled logic circuit elements configured to perform the relevant functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied in a machine-readable medium, the computer program including program code configured to perform the methods described above.
[0100] Computer program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to one or more processors of a general-purpose computer, a special-purpose computer, or other programmable data processing device having control circuitry, so that when executed by the one or more processors of the computer or other programmable data processing device, the functions / acts specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on the computer, on a portion of the computer, as a stand-alone software package, on a portion of the computer, entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0101] One or more processors may operate as stand-alone devices or may be connected to other processors, e.g., networked with other processors. Such networks may be built based on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0102] Software may be distributed across computer-readable media, namely, computer storage media (or non-transitory media) and communication media (or transitory media). The term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data, as is well known to those skilled in the art. Computer storage media includes, but is not limited to, ROM, PROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or other media that can be used to store the desired information and that can be accessed by a computer. Furthermore, communication media (transportable media) typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal (such as a carrier wave or other transmission mechanism), and is well known to those skilled in the art to include any information delivery media.
[0103] The implementations of the technology shown in the drawings are merely exemplary, and the present invention is not limited thereto. For example, the divisions such as blocks shown in FIGS. 1 and 4 are logical divisions for ease of explanation, and may be divided into additional divisions, combined into fewer divisions, supplemented with additional divisions, or eliminated and reduced without departing from the spirit of the present invention. With respect to the flowcharts shown in FIGS. 2 and 3, the divisions of operational steps, which may be referred to as functions, steps, operations, processes, or acts, may be combined into fewer steps, divided into additional steps, reordered, or omitted without departing from the spirit of the disclosure.
[0104] Unless otherwise indicated, and as will be apparent from the discussion that follows, the use of terms such as "processing," "calculating," "computing," "determining," "analyzing," and the like will be understood throughout the discussion of this disclosure to refer to the functions, acts, steps, and / or processes of computer hardware or computing systems or similar electronic computing devices that manipulate and / or transform data represented as physical, e.g., electronic, quantities into other data also represented as physical quantities.
[0105] In the above description of exemplary embodiments of the present invention, it should be understood that various features of the present invention may be grouped together in a single embodiment, drawing, or description to streamline the disclosure and aid in understanding one or more of the various aspects of the present invention. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, aspects of the present invention may reside in fewer than all features of one of the above-disclosed embodiments. Accordingly, the claims following the detailed description are expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention. Furthermore, although some embodiments described herein include some features but not other features included in other embodiments, combinations of features from different embodiments are intended to form different embodiments within the scope of the present invention, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments may be used in any combination.
[0106] Furthermore, some embodiments are described as methods or combinations of elements of methods that may be implemented by a processor of a computer system or other means for performing a function. Thus, a processor with instructions for executing such a method or element of a method forms a means for performing the method or element of a method. It should be noted that, when a method includes several elements, e.g., several steps, no ordering of such elements is implied unless specifically stated. Furthermore, the elements of device embodiments described herein are examples of means for performing the functions performed by the elements to implement an embodiment of the invention. In the description herein, numerous specific details are disclosed. However, it is understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure an understanding of the description.
[0107] It will be apparent to those skilled in the art that the present invention is not limited to the above-described preferred embodiment. However, many modifications and variations are possible within the scope of the appended claims. For example, various downmix representations other than those described above may be employed. Furthermore, the number of second head poses does not need to be two as in the above example, and may be any number. Various aspects and embodiments of the present disclosure can be understood from the following enumerated example embodiments (hereinafter also referred to as "EEE").
[0108] EEE1. 1. A method for processing audio, comprising: a first device (in some embodiments (hereinafter also referred to as "ISE"), a heavy device, a device with high computational or battery resources (e.g., an edge node or network node in a 5G system, a high-performance UE, etc.)) receiving immersive audio (ISE, the immersive audio including audio channels, audio objects, metadata, or a combination thereof (e.g., a QMF signal, the output of an immersive decoder such as IVAS, etc.); acquiring user attitude information (ISE, acquiring user attitude information includes receiving, generating, or accessing data representing an actual or predicted head orientation or head position (e.g., pitch angle, yaw angle, or roll angle, position or translation data, etc.) of a user of the second device at a first time), where the ISE, user attitude information is acquired via one or more sensors (e.g., gyroscope, accelerometer, IMU, camera, LiDAR, etc.), where the ISE, one or more sensors are included on the second device, and the ISE, one or more sensors are included on a device different from both the second device and the first device; The first device determines a downmix signal from the immersive audio, the downmix signal including at least one channel (ISE, the downmix signal is determined based at least in part on the obtained user posture information); The first device determines a set of N (e.g., N≧0) predicted poses based on acquired user pose information (ISE, the acquired user pose information representing a head pose of a user of the second device at a first time); determining, by the first device, from the immersive audio, a set of binaural representations corresponding to the set of N predicted poses; a first device generating metadata (ISE, prediction and diffuseness gain) from the downmix signal and at least one of a set of binaural representations and a metadata model; providing a downmix signal and metadata (ISE, the metadata including data representing the acquired user pose information) to a second device (ISE, the second device being a lightweight device, a wearable device (e.g., an AR / XR headset, earphones, a head-mounted display, etc.), a device with fewer computational or battery resources than the first device, etc.) different from the first device.
[0109] EEE2. obtaining the user posture information is performed at least in part by the second device; The method of EEE1, further comprising providing (e.g., transmitting) data corresponding to the acquired user posture information from the second device to the first device.
[0110] EEE3. The method of EEE1 or EEE2, further comprising: rendering, by a renderer of the second device, a downmix signal into output binaural audio based at least in part on the metadata, the acquired user posture information, and the updated user posture information (ISE, the updated user posture information representing a head pose of a user of the second device at a second time after the first time; ISE, the updated user posture information is acquired in a similar manner as the user posture information is acquired (e.g., via a common set of sensors); ISE, the updated user posture information is acquired in a different manner than the user posture information is acquired (e.g., via a separate set of sensors)).
[0111] EEE4. The downmix signal is A set of HRTFs or a set of BRIRs, The acquired user posture information; The method according to any one of EEE1 to EEE3, wherein the binaural signal is generated using
[0112] EEE5. Determining the set of predicted poses includes: The method of any one of EEE1-EEE4, comprising calculating N attitudes corresponding to the N predicted yaw angles by changing the attitude yaw angle derived based on the acquired user attitude information by a first predetermined value (e.g., an angle specified in degrees or radians) in a first direction (e.g., clockwise, counterclockwise) to obtain a first predicted yaw angle among the N predicted yaw angles (ISE, the attitude yaw angle is directly encoded in the acquired user attitude information; ISE, the attitude yaw angle is derived based in part on data encoded in the acquired user attitude information).
[0113] EEE6. The method according to EEE5, further comprising: obtaining a second predicted yaw angle among the N predicted yaw angles by changing the attitude yaw angle derived based on the obtained user attitude information in a second direction (e.g., counterclockwise, clockwise) by a second predetermined value (ISE, the first predetermined value is different from the second predetermined value; ISE, the first predetermined value and the second predetermined value are the same value; ISE, the first direction is different from the second direction; ISE, the first direction and the second direction are the same value).
[0114] EEE7. The method of any one of EEE5-EEE6, wherein calculating the N postures corresponding to the N predicted yaw angles further comprises generating posture yaw angles derived from the acquired user posture information by modifying posture yaw angles included in the acquired user posture information based on one or more motion data (e.g., angular velocity, acceleration, or deceleration of the user's head rotation).
[0115] EEE8. The method according to any one of EEE5 to EEE6, wherein the attitude yaw angle derived based on the user attitude information corresponds to an angular yaw value represented in the acquired user attitude information.
[0116] EEE9. The method according to any one of EEE1 to EEE3 and EEE5 to EEE8, wherein the downmix signal is a combination of a prototype signal and zero or more spread signals.
[0117] EEE10. The method according to EEE9, wherein the prototype signal and the zero or more diffuse signals are created by applying real or complex gain values to binaural signals generated using a set of HRTFs or BRIRs and the acquired user posture information, and summing the gain-adjusted channels of the binaural signals.
[0118] EEE11. The method according to EEE10, wherein the real or complex gain values are generated based on the normalized covariance of the channels obtained by taking sums and differences of binaural signals generated using a set of HRTFs or BRIRs and using obtained user posture information.
[0119] EEE12. The method of any one of EEE1 to EEE11, wherein the metadata generated by the first device includes real or complex gain values such that a binaural representation corresponding to the predicted pose can be reconstructed by application of the real or complex gain values to the channels of the downmix signal and then adding the gain-adjusted channels of the downmix.
[0120] EEE13. The method of any one of EEE1 to EEE12, wherein the binaural representations corresponding to the N predicted poses are determined by reusing HRTF or BRIR filtered channels that are not expected to change with changes in pose.
[0121] EEE14. The metadata generation is - calculating prediction gains for predicting correlated components in the binaural representation for one or more downmix signals; calculating a diffusion gain parameter to satisfy the uncorrelated energy; The method according to any one of EEE1 to EEE13, including at least one of:
[0122] EEE15. The method according to any one of EEE1 to EEE14, wherein the generation of the metadata includes a quantization and encoding process of the metadata.
[0123] EEE16. The method according to any one of EEE3 to EEE15, wherein the rendering includes a dequantization and decoding process of the metadata.
[0124] EEE17. The first device providing the downmix signal and the metadata to a second device different from the first device is encoding a downmix signal; multiplexing the quantized and coded metadata with the coded downmix signal into a combined bitstream; and transmitting the combined bitstream to the second device.
[0125] EEE18. The method of any one of EEE1 to EEE17, wherein the metadata includes data corresponding to the reference attitude.
[0126] EEE19. On the second device, receiving the combined bitstream; separating the combined bitstream into data corresponding to the downmix signal and data corresponding to the metadata; decoding data corresponding to the downmix signal; The method of any one of EEE1 to EEE18, further comprising decoding and dequantizing data corresponding to the metadata.
[0127] EEE20. The method of any one of EEE1 to EEE19, wherein a model is used to generate first estimates of predictive metadata parameters used in the metadata quantization / dequantization and encoding / decoding processes, the model generating estimates for each pose that differs from the obtained user pose information (e.g., poses received from the second device).
[0128] EEE21. The method according to any one of EEE3 to EEE20, wherein each piece of metadata provided from the second device to the first device is quantized and encoded at the first device using n symmetry orientations corresponding to each piece of metadata and a reference orientation at the first device.
[0129] EEE22. The method according to EEE21, wherein symmetry in the pose corresponding to each calculated metadata and a reference pose at the first device is used to quantize and encode difference values between sets of parameters, such that the overall entropy of the encoded parameters is reduced.
[0130] EEE23. one or more processors; a memory storing instructions that, when executed by one or more processors, cause the computing device to perform any of the methods EEE1-EEE22.
[0131] EEE24. A computer program product configured to cause one or more processors to perform any of the methods EEE1 to EEE22.
[0132] EEE25. A non-transitory computer-readable recording medium storing one or more computer programs configured to be executed by one or more processors of a computing device, the one or more computer programs comprising instructions that cause the computing device to perform a method according to any one of EEE1 to EEE22.
Claims
1. A method of audio processing in a main device (10), comprising: The first bitstream (b 1 ) and The first bitstream (b 1 ) to obtain a decoded immersive audio content (A); The second bitstream (b p ) and The second bitstream (b p ) to obtain pose information (P; P″; P, V) associated with the user of the lightweight processing device; determining a first head pose (P′) based on the pose information (P; P″; P, V); generating a downmix representation (Dmx) of the immersive audio content (A) corresponding to the first head pose (P'); a second head pose (P n ) corresponding to the set of binaural representations (BIN n ) and - calculating reconstruction metadata (M) including said first head pose (P') allowing the reconstruction of said set of binaural representations from said downmix representation (Dmx); The downmix representation (Dmx) and the reconstruction metadata (M) are fed into a third bitstream (b 2 ) and The third bitstream (b 2 ) and A method comprising:
2. the reconstruction metadata comprises a 2x2 matrix for each time-frequency tile; The method of claim 1.
3. encoding the reconstruction metadata using differential encoding between elements of the 2x2 matrix. The method of claim 2.
4. the head poses in the second set of head poses are symmetrically distributed around the first head pose, the method further comprising quantizing and encoding the reconstructed metadata based on symmetry of the reconstructed metadata with respect to the symmetrically distributed head poses. The method according to any one of claims 1 to 3.
5. encoding the reconstruction metadata using differential encoding between metadata for different symmetric poses. The method of claim 4.
6. encoding the reconstruction metadata using differential encoding between successive time frames and / or adjacent frequency bands. The method according to any one of claims 1 to 5.
7. the pose information includes a head pose (P) detected by the lightweight processing device; The method according to any one of claims 1 to 6.
8. the pose information further includes head velocity (V) detected by the lightweight processing device; The method of claim 7.
9. the second set of head poses is determined by adding a set of predefined offsets to the first head poses. The method according to any one of claims 1 to 8.
10. the predefined offset is static; 10. The method of claim 9.
11. the predefined offset is dynamically calculated based on a latency between the main device and the lightweight processing device; 10. The method of claim 9.
12. encoding the set of predefined offsets into the third bitstream. The method according to any one of claims 9 to 11.
13. the downmix representation is a first binaural representation corresponding to the first head pose, and the reconstruction metadata is pose-compensated metadata that enables reconstruction of the set of binaural representations from the first binaural representation. The method according to any one of claims 1 to 12.
14. said downmix representation comprising a mono signal (S) formed by a combination of channels in a multi-channel representation of said immersive audio content, the reconstruction metadata allows the reconstruction of the set of binaural representations from the prototype signal S. The method according to any one of claims 1 to 13.
15. the multi-channel representation is a first binaural representation.
15. The method of claim 14.
16. the reconstruction metadata comprises a 2x2 matrix per time-frequency tile, thereby enabling reconstruction of the set of binaural representations from the monophonic signal (S) and decorrelated versions of the prototype signals.
15. The method of claim 13 or 14.
17. The entries of the 2x2 matrix are [0.001] is calculated as In the formula, Cov SL is the covariance between the prototype signal (S) and the left channel of a particular binaural representation, and Cov SR is the covariance between the mono signal (S) and the right channel of the particular binaural representation, and Cov SS is the variance of the mono signal S, and Cov RR is the variance of the right channel, and Cov LL is the variance of the left channel, and Res RR = Cov RR -Pred R 2 ×Cov SS and Res LL = Cov LL -Pred L 2 ×Cov SS That is, 16. The method of claim 15.
18. the downmix representation further comprises a diffuse signal (D) formed as a combination of diffuse components of the multi-channel representation of the immersive audio content, and the reconstruction metadata comprises a 2x2 matrix per time frame and frequency band that allows the reconstruction of the set of binaural representations from the mono signal (S) and the diffuse signal (D), 15. The method of claim 14.
19. The entries of the 2x2 matrix are [Equation 43] is calculated as In the formula, Cov SL is the covariance between the mono signal S and the left channel of a particular binaural representation, and Cov SR is the covariance between the mono signal (S) and the right channel of a particular binaural representation, and Cov SS is the variance of the mono signal S, and Cov DD is the variance of the diffuse signal (D), and Cov RR is the variance of the right channel, and Cov LL is the variance of the left channel, and Res RR = Cov RR -Pred R 2 ×Cov SS and Res LL = Cov LL -Pred L 2 ×Cov SS That is, 18. The method of claim 17.
20. a processor; a memory coupled to the processor; An apparatus comprising: The processor is configured to cause the device to perform the method of any one of claims 1 to 19. Device.
21. A computer readable recording medium storing a program comprising instructions which, when executed by a processor, cause the processor to carry out the method of any one of claims 1 to 19.
22. A method of audio processing in a lightweight processing device (20), comprising: Bitstream (b) from the main device 2 ) and Decoding the bitstream a downmix representation (Dmx') of an immersive audio content (A) associated with a first head pose (P'); The second head pose (P n ) associated with a set of binaural representations (BIN n first reconstruction metadata (M′) comprising said first head pose (P′), enabling a set of and The second head pose (P n ) and Detecting a current head pose (P) of a user of said lightweight processing device; sending the current head pose to the main device; the downmix representation (Dmx'), the first reconstruction metadata (M'), the second head pose (P n ), and output binaural audio (BIN) based on the relationship between the first head pose (P′) and the current head pose (P). out ) and A method comprising:
23. the lightweight processing device obtains the second set of head poses by adding a set of offsets to the first head poses.
23. The method of claim 22.
24. the lightweight processing device has prior knowledge of the set of offsets; 24. The method of claim 23.
25. the lightweight processing device obtains the set of offsets from the bitstream.
24. The method of claim 23.
26. the first reconstruction metadata includes a 2x2 matrix for each time-frequency tile; The method according to any one of claims 22 to 25.
27. calculating second reconstruction metadata by linearly interpolating or extrapolating the first reconstruction metadata based on a relationship between the current head pose and the first head pose and the set of second head poses; The method according to any one of claims 22 to 25.
28. the downmix representation is a first binaural representation corresponding to the first head pose (P′), and the reconstruction metadata is pose correction metadata that enables reconstruction of a set of binaural representations from the first binaural representation. The method according to any one of claims 22 to 27.
29. said downmix representation comprising a mono signal (S) formed by a combination of channels in a multi-channel representation of said immersive audio content, said reconstruction metadata enabling the reconstruction of said set of binaural representations from said monophonic signal (S), The method according to any one of claims 22 to 27.
30. the multi-channel representation is a first binaural representation.
30. The method of claim 29.
31. obtaining a decorrelated version of the mono signal using a decorrelation function; the first reconstruction metadata further comprises a 2x2 matrix per time frame and frequency band that allows reconstruction of the set of binaural representations from the monophonic signal (S) and the decorrelated version of the monophonic signal.
31. The method of claim 29 or 30.
32. the downmix representation further comprises a diffuse signal (D) associated with the mono signal (S), and the first reconstruction metadata comprises a 2x2 matrix for each time frame and frequency band that enables the set of multiple binaural representations to be reconstructed from the mono signal S and the diffuse signal (D).
31. The method of claim 29 or 30.
33. A computer readable recording medium storing a program comprising instructions which, when executed by a processor, cause the processor to carry out the method of any one of claims 22 to 32.
34. A main device (10), a first decoder (11) for decoding the first bitstream (b1) to obtain a decoded immersive audio content (A); The second bitstream (b p a second decoder (13) for decoding the first head pose (P') to obtain pose information about a user of the lightweight processing device and for determining a first head pose (P') based on the pose information; a downmixer (12) for generating a downmix representation (Dmx) of said immersive audio content (A) corresponding to said first head pose (P'); a renderer (14) for rendering a set of binaural representations of the immersive audio content corresponding to a second set of poses; The downmix representation is converted into the binaural representation (BIN n a metadata generator (15) for calculating reconstruction metadata (M) comprising said first head pose (P') enabling the reconstruction of a set of The downmix representation (Dmx) and the reconstruction metadata (M) are fed into a third bitstream (b 2 ) into an encoder (17); The third bitstream (b 2 an interface (18) for outputting A main device comprising:
35. A lightweight processing device (20), comprising: Bitstream from the main device (b 2 ) to obtain a downmix representation (Dmx') of immersive audio content, said downmix representation being a representation of a first head pose (P') and a second head pose (P n ) associated with a set of binaural representations (BIN n a decoder associated with a first reconstruction metadata (M′) comprising said first head pose (P′), said first reconstruction metadata (M′) enabling a set of head poses (P′) to be reconstructed from said downmix representation; a head tracker (24) for detecting a current head pose (P) of a user of said lightweight processing device; a second encoder (25) for encoding the current head pose (P) and transmitting it to the main device; the downmix representation, the first reconstruction metadata, the second head pose (P n a binaural reconstruction block (26) for reconstructing output binaural audio based on a set of inputs P′ and the relationship between the first head pose (P′) and the current head pose (P); A lightweight processing device comprising:
36. 1. A split-device binaural rendering system, comprising: A main device (10) according to claim 34; A lightweight processing device (20) according to claim 35, The interface (18) receives the third bitstream (b 2 ) to the lightweight processing device (20), A split-device binaural rendering system.
Citation Information
Cited By
Stereo audio signal processing method, communication apparatus, and storage medium
US12725618B2