Split Binaural Rendering
The low-complexity split-rendering technique addresses computational and latency issues in AR glasses by using a limited number of pre-rendered representations and reduced metadata, ensuring high-quality immersive audio experiences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2026-03-10
AI Technical Summary
Immersive audio rendering in small form-factor devices like AR glasses faces challenges due to high computational load and latency issues in split rendering, leading to degraded user experience from stale pose information.
A low-complexity split-rendering technique using a limited number of pre-rendered representations and reduced metadata transmission, enabling efficient attitude compensation for multiple rotation axes.
Significantly reduces computational requirements and metadata transmission, maintaining high-quality immersive audio experiences with reduced latency.
Smart Images

Figure 2026508313000001_ABST
Abstract
Description
[Technical Field]
[0001]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 448,830, filed February 28, 2023, and U.S. Provisional Application No. 63 / 558,596, filed February 27, 2024, the contents of which are incorporated herein by reference in their entireties.
[0002]
[0002] Technical field of the invention The present invention relates generally to audio processing, and more specifically to audio rendering (e.g., binaural rendering) performed on two separate devices ("split rendering"). [Background technology]
[0003]
[0003] Unless otherwise specified in this case, the content described in this section is not prior art to the claims of this application, nor is it admitted as prior art by inclusion in this section.
[0004] Immersive audio is an essential media component of extended reality (XR) applications, including augmented reality (AR), mixed reality (MR), and virtual reality (VR). To enhance the user experience, immersive audio can support adjusting the presented immersive audio / visual scene in response to the user's movements. For example, it may be desirable to track the user's head position and head movements during audio rendering and adjust the audio accordingly. Thus, an immersive audio experience may handle head movements using a model with three degrees of freedom (3DoF) or six degrees of freedom (6DoF).
[0005]
[0005] Various immersive audio services, such as immersive voice and audio services (IVAS), can be used to render high-quality audio renditions in XR devices that include awareness of pose information, which can include metadata about a user's head position along with relative or absolute movement. However, making such adjustments according to pose information can require significant computing power to achieve a high-quality immersive audio experience.
[0006]
[0006] The computational load requirements for immersive audio can be problematic for small form-factor devices such as AR glasses. To make them as practical and user-friendly as possible, such AR glasses may avoid using powerful processors and heavy batteries that would otherwise result in bulkier, more expensive, and heavier user wear, consuming more power and generating significant amounts of heat. Therefore, to enable low-power operation in a reasonable form factor with low latency, such AR devices tend to have processors with reduced computational complexity and limited numerical processing.
[0007] This disclosure recognizes the above problem and explores possible solutions. One possible solution is to reduce the audio rendering requirements at the end device (e.g., an AR device operated by a user) using a split-rendering topology that leverages processing from some other entity (e.g., a network-based device) in the mobile / wireless network to which the end device is connected or tethered (e.g., via a network or cloud-based connection). For example, a powerful network entity such as a mobile user equipment (e.g., a UE, a device used by an end user, a handheld multifunction device, a game console, a cloud-based resource, etc.) may be connected to the end device to assist in split rendering of immersive audio. Pose information based on the user's movements can be collected at the end device and sent to the network entity. The end device can then receive already-rendered audio from the network entity; where high-complexity operations such as processing 3DoF / 6DoF pose information (e.g., head-tracking metadata) can be performed by the rendering entity (e.g., the network entity). One issue with the described split rendering topology is that the latency for transmission between the end device and the network entity can be on the order of 100 ms; this means that the network entity may rely on stale pose / head tracking information. Due to this delay, the rendered audio from the network entity may not match the user's current head pose / head position at the end device. If the motion-to-sound latency is too large, the end user will experience a perceptible degradation in quality in their immersive experience.
[0008]
[0008] US Prov. Appl. No. US 63 / 340,181 discloses a new approach to interactive head tracking. The described approach generates multiple binaural representations (pre-renditions) corresponding to various head poses in a main device or pre-renderer, and calculates metadata that can be used together with a reference binaural representation to reconstruct binaural output corresponding to any given pose in a post-renderer. The reference binaural representation and metadata are sent to a post-rendering device. Based on the reference binaural representation and metadata, and based on the difference between the reference pose and the user's detected current head pose, the post-renderer determines the binaural audio corresponding to the current head pose.
[0009]
[0009] US Prov. Appl. No. US63 / 386,465 discloses a binaural rendering for a reference position P' obtained upstream from a post-renderer device, and N pre-rendered representations P' for "probing" poses close to the reference position P'. 1-N describes a system that relies on Summary of the Invention
[0010]
[0010] In some applications, complexity constraints in the pre-renderer device and metadata bitrate limitations in the transmission interface between the pre-renderer device and the post-renderer device may limit the number of pre-rendered representations (or binaural representations) that need to be computed in the pre-renderer device and may also limit the amount of metadata that needs to be transmitted to the post-renderer for pose correction. Therefore, it may be desirable to carefully choose the pre-rendered representations and head poses in the pre-renderer so that pose correction for any head pose in the post-renderer can be achieved using a limited number of pre-rendered representations in the pre-renderer and a limited amount of metadata transmission.
[0011]
[0011] One source of computational complexity in a pre-renderer device is the number of pre-rendered representations required. A typical case envisaged for split renderer metadata enabling yaw correction is one reference binaural representation and two binaural pre-rendered representations (N=2), where the two probing poses are P'+X and P'-X, where X and X' are the angular displacements from P' around the axis of rotation.
[0012]
[0012] The object of the present invention is to address the problems described and enable efficient split rendering using a reduced number of pre-rendered representations and, consequently, a reduced amount of associated metadata.
[0013] This disclosure describes a low-complexity split-rendering technique that is based on a limited number of pre-rendered representations, thereby significantly reducing the amount of computation required on the pre-rendering side and reducing the amount of transmitted metadata. This disclosure also describes low-complexity solutions for post-render corrections around one, two, or three rotation axes, such as yaw, pitch, and roll displacements.
[0014]
[0014] This and other objects are achieved by various aspects of the present invention, including those defined by the independent claims.
[0015] According to a first aspect, this and other objects are achieved by a method for audio rendering (in a main device) that enables split rendering techniques together with attitude compensation for multiple axes of rotation (in a lightweight device), the method comprising: Obtaining immersive audio content; obtaining a reference posture; rendering the immersive audio content into a first number of binaural pre-rendered representations, the binaural pre-rendered representations corresponding to a set of probing poses, the set of probing poses including poses equal to a reference pose (P') and / or poses displaced from the reference pose by a rotation about at least one of the rotation axes; calculating a second number of approximate binaural representations based on the binaural pre-rendered representations, the approximate binaural representations corresponding to a set of virtual probing poses, the virtual probing poses being displaced from the probing pose by a rotation about at least one of the rotation axes; determining a reference binaural representation based on one or more of the approximation binaural representation and the binaural pre-rendered representation; calculating reconstruction metadata that enables reconstruction of the binaural pre-rendered representation and the approximation binaural representation from the reference binaural representation; encoding the reconstruction metadata and the reference binaural representation into an output bitstream; and Outputting the output bitstream (b2) is included.
[0016] According to a second aspect, this and other objects are achieved by a method for audio rendering (in a main device) that allows split rendering with attitude compensation for yaw and pitch axes (in a lightweight device), the method comprising: Obtaining immersive audio content; obtaining a reference posture; Rendering the immersive audio content into a reference binaural representation corresponding to a reference pose; rendering the immersive audio content into one or more binaural pre-rendered representations, the binaural pre-rendered representations corresponding to one or more probing poses that are displaced from a reference pose by a rotational displacement from the reference pose by a displacement about both a yaw axis and a pitch axis; calculating, for each of the probing attitudes, yaw metadata representing displacement about a yaw axis and pitch metadata representing displacement about a pitch axis; encoding the reference binaural representation, the yaw metadata, and the pitch metadata into an output bitstream; and Outputting the output bitstream is included.
[0017] According to a third aspect, this and other objects are achieved by a method for rendering audio on a main device, enabling split rendering with attitude compensation about at least one axis of rotation (on a lightweight device), the method comprising: Obtaining immersive audio content; receiving head pose information associated with a user of the lightweight processing device; determining a reference pose based on the head pose information; Rendering the immersive audio content into a reference binaural representation corresponding to a reference pose; rendering the immersive audio content into one or more binaural pre-rendered representations, the binaural pre-rendered representations corresponding to one or more probing poses displaced from a reference pose about an axis of rotation; calculating reconstruction metadata that enables reconstruction of the binaural pre-rendered representation from the reference binaural representation, the reconstruction metadata including, for each time-frequency tile, a transformation matrix; Calculating the augmented metadata by multiplying each transformation matrix for a particular probing pose by an additional gain matrix (G), the additional gain matrix being:
number
[0018] According to a fourth aspect, this and other objects are achieved by a method for audio rendering (in a main device) that enables split rendering with attitude compensation for multiple axes of rotation (in a lightweight device), the method comprising: Obtaining immersive audio content; receiving head pose information associated with a user of the lightweight processing device; determining a reference posture and at least one of a head posture rotation axis and a head posture rotation velocity based on the head posture information; Rendering the immersive audio content into a reference binaural representation corresponding to a reference pose; rendering the immersive audio content into a set of binaural pre-rendered representations, the binaural pre-rendered representations corresponding to a set of probing poses rotated relative to a reference pose, the probing poses being selected based on head pose information; calculating reconstruction metadata that enables reconstruction of the binaural pre-rendered representation from the reference binaural representation; encoding the reference binaural representation and the reconstruction metadata into an output bitstream; and Outputting the output bitstream is included.
[0019] According to a fifth aspect, this and other objects are achieved by an audio processing method for attitude compensation about multiple axes of rotation, the method comprising: receiving a bitstream from the main device; decoding the bitstream to obtain a reference binaural representation and first reconstruction metadata associated with a set of probing poses representing displacements from the reference pose due to rotation about a plurality of rotation axes; Detecting the current head pose; For each axis of rotation: selecting a probing pose that is closest to the detected pose with respect to the axis of rotation; determining axis-specific reconstruction metadata based on first reconstruction metadata associated with the selected probing pose and a difference between the current head pose and the reference pose about the rotation axis; and Determining a binaural output corresponding to the current head pose based on the reference binaural representation and axis-specific reconstruction metadata for each of the rotation axes.
[0020] According to a sixth aspect, the various objects are achieved by a main processing device, the main processing device comprising: a decoder configured to decode the first bitstream to obtain decoded immersive audio content; obtaining a reference attitude; Rendering the immersive audio content into a reference binaural representation based on a reference pose; and rendering the immersive audio content into a plurality of binaural pre-rendered representations, the binaural pre-rendered representations corresponding to a set of probing poses associated with a reference pose, the set of probing poses including poses displaced from the reference pose by a rotation about at least one of the rotation axes; a metadata generator configured to calculate reconstruction metadata enabling reconstruction of the binaural pre-rendered representation from the reference binaural representation; an encoder configured to encode the reference binaural representation and the reconstruction metadata into an output bitstream; and An interface (16) configured to output an output bitstream.
[0021] According to a seventh aspect, the various objects are achieved by a lightweight processing device, the lightweight processing device comprising: a decoder configured to decode the bitstream to obtain a reference binaural representation and first reconstruction metadata associated with a set of probing poses representing displacements from the reference pose due to rotation about a plurality of rotation axes; a head tracker configured to detect a current head pose; binaural reconstruction blocks (26); and the binaural reconstruction block (26) includes: for each of the rotation axes, selecting a probing pose that is closest to the detected pose about the rotation axis, and determining axis-specific reconstruction metadata based on first reconstruction metadata associated with the selected probing pose and a difference between the current head pose about the rotation axis and the reference pose; and determining a binaural output corresponding to the current head pose based on the reference binaural representation and the axis-specific reconstruction metadata; The device is configured to: [Brief explanation of the drawings]
[0022]
[0022] The present invention will now be described in more detail with reference to the accompanying drawings. [Figure 1]
[0023] Figure 1 shows a user wearing a smartphone and a pair of headphones. [Figure 2]
[0024] FIG. 2 is a diagram illustrating head pose and rotation around the three axes yaw, pitch, and roll. [Figure 3]
[0025] FIG. 3 is a schematic block diagram illustrating split rendering in a main processing device and a lightweight processing device. [Figure 4]
[0026] FIG. 4 is a flowchart illustrating processing in a main processing device according to an embodiment of the first aspect of the present invention. [Figure 5]
[0027] FIG. 5 is a flowchart illustrating processing in a main processing device according to an embodiment of the second aspect of the present invention. [Figure 6]
[0028] FIG. 6 is a flow chart illustrating processing in a lightweight processing device according to an embodiment of a further aspect of the present invention. [Figure 7]
[0029] FIG. 7 is a flowchart showing processing in a main processing device according to an embodiment of yet another aspect of the present invention. [Figure 8]
[0030] FIG. 8 is a flowchart showing processing in a main processing device according to an embodiment of yet another aspect of the present invention. [Figure 9]
[0031] FIG. 9 shows a schematic block diagram of an exemplary device or architecture that may be used to implement embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023]
[0032] In the following detailed description, reference is made to the accompanying drawings which form a part hereof, and which are shown by way of illustration, in which certain exemplary configurations of the concepts may be practiced. These configurations are described in sufficient detail to enable those skilled in the art to practice the techniques disclosed herein, and it should be understood that other forms may be utilized and other changes may be made without departing from the spirit or scope of the concepts presented. Accordingly, the following detailed description should not be taken in a limiting sense, and the scope of the concepts presented is defined only by the appended claims.
[0024]
[0033] The embodiments of the invention disclosed herein assume compatibility and consistency with the use of immersive audio codecs such as IVAS in XR applications. In particular, the inventive concepts described in detail below are applicable to systems, devices, architectures, methods, and techniques in which primary decoding and pre-rendering are performed by a main device (UE) (e.g., an edge or other network node / server in a 5G system, a high-performance mobile device, etc.) with advanced resources, such as powerful computing (or processor) resources with significant power or battery capacity, and final decoding and post-rendering are performed by a different device (e.g., a lightweight device, a wearable device, AR glasses, a head-mounted display, a head-up display, etc.) with fewer resources than the main device.
[0025]
[0034] Figure 1 shows a schematic representation of a user 1 holding a smartphone 2 and wearing a headset 3. The smartphone, in the context of the present invention, can function as the main, pre-rendering device, while the headset can function as a lightweight, user-held, post-rendering device. It is the post-rendering device that has up-to-date information about the user's head pose P.
[0026]
[0035] 2, the user's head pose P is defined in this context by three degrees of freedom: rotation α about yaw axis 101, rotation β about pitch axis 102, and rotation γ about roll axis 103. Of course, head pose is also related to its position in the room, which is defined by three additional degrees of freedom: spatial coordinates x, y, and z. However, spatial translation is not relevant for the purposes of this disclosure.
[0027]
[0036] FIG. 3 illustrates some example functional blocks that may be implemented in the main device 2 and the lightweight device 3 in an exemplary split rendering system.
[0028]
[0037] Here, the main device or pre-rendering device 2 includes a decoder 11, a binaural renderer 12, a metadata generator 13, a first encoder 14, a second encoder 15, and a multiplexer 16. The main device 2 may also include a pose decoder 17.
[0029]
[0038] A decoder 11, for example an IVAS decoder, is configured to receive and decode a bitstream b1 to decode immersive audio content A. A binaural renderer 12 receives (or obtains) the immersive audio content A and a reference pose P' and, in response, generates a reference binaural representation (rendition) Bin associated with the reference pose P'. ref The reference pose may be an assumed head pose or may be determined based on head pose information received from the lightweight device 3 via the pose decoder 17. In response, the binaural renderer 12 generates a set of probing poses P associated with the reference pose P'. n Bin n The device is further configured to render
[0030]
[0039] The metadata generator 13 generates a reference binaural representation Bin ref and pre-rendered representation Bin n and in response to which generates reconstruction metadata M to enable reconstruction of the pre-rendered representation from the reference binaural representation. Optionally, the reference pose P' can be received by the metadata generator 13 and in response encoded into the reconstruction metadata M.
[0031]
[0040] The attitude decoder 17 is an optional block that may not be required in all implementations. If present, the attitude decoder 17 p and receiving head pose information from the lightweight device 3 via the head pose information receiving unit 100 and generating a reference pose P' in response thereto.
[0032]
[0041] The first encoder 14 outputs a reference binaural representation Bin ref and in response, receiving a reference binaural representation Binref into the encoded bitstream b 11 The second encoder 15 is configured to receive the reconstruction metadata M and, in response, encode the reconstruction metadata (and optionally the pose information) as 12 The multiplexer 16 is configured to receive the coded bitstreams b11 and b12 from the outputs of the two encoders 14 and 15, and in response to this, to encode the coded bitstream b 11 and b 12 into bitstream b2. The main device 2 may also include an interface for outputting bitstream b2, so that the bitstream can be subsequently transmitted or otherwise made available to another device external to the main device 2, here lightweight device 3.
[0033]
[0042] In some embodiments, the encoder 15 encodes the attitude information into a bitstream b 12 and the encoded attitude information is further configured to encode the reference attitude P′ and / or the probing attitude P n Shows.
[0034]
[0043] The lightweight device, or post-renderer device 3, here comprises a demultiplexer 21, a first decoder 22, a second decoder 23, a binaural reconstruction block 24 and a head tracker 25. Optionally, the lightweight device 3 also comprises a pose information encoder 26.
[0035]
[0044] The demultiplexer 21 receives the bit stream b2 from the main device 2 and, in response, splits the received bit stream b2 into two encoded bit streams b 21 and b 22 It is configured to separate into The decoder 22 receives the encoded bitstream b21 and in response, receives the bitstream b 21 Reference binaural signal Bin ref It is configured to decrypt The decoder 23 receives the encoded bitstream b 22 and in response, receives the bitstream b 22 M' (and, if present, the reference pose P' and / or the probing pose P n The binaural reconstruction block 24 is configured to receive the current user head pose P detected by the head tracker 25 and to decode it into a reference binaural signal Bin ref and determining the binaural output based on the metadata M' and the current head pose P relative to the reference pose P'.
[0036]
[0045] As mentioned above, the reference position P′ and / or the probing position P n is the bitstream b received from main device 2 22 This is particularly useful when the reference pose is based on attitude information received by the main device 2 from the lightweight device 3. However, in some implementations, the reference pose' is an assumed pose, and thus the lightweight device already knows the attitude information P'. For example, the reference pose may be a "straight ahead" pose, e.g., gazing straight at the display device.
[0037]
[0046] In a similar manner, information regarding the probing orientation may be received in the bitstream, but may alternatively be predefined and known by the lightweight device, e.g., the probing orientation may be a predefined displacement from a reference orientation.
[0038]
[0047] The encoder 26 is an optional block that may not be required in all implementations. If present, the encoder 26 receives pose information P from the head tracker 25 and, in response, generates a bitstream b P The system is configured to encode the pose information into
[0039]
[0048] In an exemplary implementation, a heavy weight device 2 transmits a reference binaural signal Bin using a posture P′. p’ and metadata (MD), and then a low-cost post-renderer performs pose correction from P' to the actual pose P and uses the metadata M to generate the Bin p’ From Bin p In this case, Bin p contains all spatial cues related to the pose P. In some implementations, the pose P' in the pre-renderer is an assumed pose without any information from the lightweight device. In some other implementations, the pose P' in the pre-renderer is received from the lightweight device via a back channel. Calculating metadata corresponding to various probing poses so that the low-load processing post-renderer can perform pose correction from P' to the actual pose P may require multiple binaural rendering representations (sometimes referred to herein as pre-rendering representations) in the pre-renderer, and these pre-rendering representations may be computationally intensive. Furthermore, computing metadata corresponding to multiple probing poses may significantly increase the metadata bitrate. Therefore, it is desirable to carefully select the probing pose points for the pre-rendering representations so that the total number of pre-rendering representations is limited. Furthermore, the estimated binaural signal Bin corresponding to the actual pose P in the post-renderer may be calculated. pIt is also desirable to reduce the amount of metadata corresponding to these probing pose points while preserving the overall perceptual quality of the image. The following embodiments provide exemplary implementations of such low-load processing, low-metadata-rate implementations.
[0040]
[0049] 1-axis split rendering metadata calculations We will describe an exemplary case where split-rendered metadata is calculated using a single axis (e.g., yaw) correction. This example can use two pre-rendered representations (N=2), and the two probing poses are P'+X and P'-X', where +X and X' are displacements from the reference pose P' about axis 101. More specifically, X and X' may be (yaw) angles α and -α. Note that the displacements from the reference pose are not necessarily equal (±α), but are assumed so here for simplicity.
[0041]
[0050] One possible way to reduce the complexity of the pre-renderer is to skip one rendering representation. This is a viable solution as long as the time trend of the yaw angle is known. In that case, the probing rendering representation can be performed only for either +α or -α, the direction in which the pose is predicted to evolve. However, in general, such a trend may be unknown. For example, if the current pose is stationary, such a trend is unknown because it is not known whether the user will next turn their head to the right or left. In this example, if the probing rendering representation is performed in the wrong direction, the adjustments made by the post-renderer must rely on extrapolated metadata rather than interpolated metadata, which may degrade the quality of the post-rendered output signal.
[0042]
[0051] A first solution to reduce the number of renderings to two is to abandon pre-rendering for the reference pose P'. Instead, pre-rendering representations are generated for poses P'-α and P'+α, and one of these binaural representations, e.g., pose P'-α, is transmitted to the post-render device. Since pose P is static in the post-renderer, the post-renderer only needs to correct for +α if it matches P'. The advantage of this example over a solution with three rendering representations is that the complexity of one rendering representation is saved and only a single set of correction metadata needs to be transmitted. One potential drawback of this solution is that post-rendering corrections are practically always required, even if the pose is static. A further potential drawback is the bias of the solution, which involves a potentially lower-precision post-rendering output signal for pose P'+α compared to the (perfect) post-rendering output signal for pose P'-α. This bias may be overcome by transmitting one binaural channel (e.g., the left channel) of the rendered representation for pose P'-α and one binaural channel (e.g., the right channel) of the rendered representation for pose P'+α. Post-rendering for pose P would therefore involve adjusting the left channel with the renderer metadata for that channel of the rendered representation for pose P'-α, and adjusting the right channel with the renderer metadata for the rendered representation for pose P'+α.
[0043]
[0052] Another solution to this problem is to first generate an approximation of the binaural representation for the reference pose P' in the pre-renderer by interpolating between available rendering representations for poses P'-α and P'+α. This may involve simply averaging two available binaural rendering representations to generate a reference binaural rendering representation for the virtual reference pose P'. Second, it is worth noting that split renderer metadata can be calculated based on the binaural rendering representations for P', P'-α, and P'+α, as described in US 63 / 340,181, whereby the fact that the reference rendering representation is obtained through interpolation results in metadata symmetry (alleviating the need to calculate and transmit metadata associated with one of the probing positions). Third, the reference rendering representation together with the metadata can be transmitted to the post-renderer device, in which case operations can be performed as described in US 63 / 340,181 (incorporated herein by reference).
[0044]
[0053] In the following, the above solution is extended to generating separate renderer metadata for post-render corrections for attitude displacements around two and three axes, e.g., for correction of yaw and pitch displacements, or for correction of yaw, pitch and roll displacements.
[0045]
[0054] 2-axis split rendering metadata calculations Another exemplary case will be described in which split-rendered metadata is calculated using two-axis (e.g., yaw and pitch) correction. In this example, four probing poses may be considered along with the reference pose, and the technique proposed by US 63 / 340,181 may be implemented. Two probing poses may be used to examine the displacement about a first axis (e.g., yaw axis) from the reference pose, for example, by changing the pose relative to the reference pose by ±α deflection angles about the first axis while keeping the second axis (e.g., pitch) constant. Two other probing poses may be used to examine the pitch displacement from the reference pose, for example, by changing the pose relative to the reference pose by ±β pitch deflection angles while keeping the yaw axis constant. In this example, the total number of rendering representations is five.
[0046]
[0055] In the following description, techniques are described for computing split renderer metadata for post-render corrections along several axes. While the description may refer to yaw, pitch, and / or roll corrections, the same techniques may be applied for computationally efficient computation of post-render metadata along any axis or set of axes. Thus, the terms yaw and pitch and / or roll in the following description may be substituted with any other suitable axes.
[0047]
[0056] FIG. 4 is a flowchart illustrating processing in a main device 2 according to an embodiment of the first aspect of the present invention, relating to a method of rendering audio in a main device 2 to enable split rendering with attitude compensation around multiple axes of rotation.
[0048]
[0057] The flowchart may be decomposed into various blocks or partitions, such as blocks S11-S17. Processing for the various blocks in Figure 4, which may be described as operations, processes, methods, steps, acts, or functions, may begin at block S11.
[0049]
[0058] In step S11 (obtaining audio content), a first bitstream is received and decoded (for example by decoder 11) to obtain immersive audio content A. Step S11 may be followed by step S12.
[0050]
[0059] In step S12 (obtaining a reference attitude), a reference attitude P' is obtained. The reference attitude may be an assumed attitude (e.g., facing straight ahead) or may be based on attitude information received from the lightweight device 3. Step S12 may be followed by step S13.
[0051]
[0060] In a (pre-rendering) step S13, a first number of binaural pre-rendered representations are rendered (e.g., by a renderer 12), the binaural pre-rendered representations being represented by a set of probing poses P n , which includes poses that are displaced from the reference pose P' by rotation about at least one of the rotation axes. Optionally, the set of probing poses also includes the reference pose. Step S13 may be followed by step S14.
[0052]
[0061] In step S14 (calculating the approximate representation), a second number of approximate binaural representations Bin' m Binaural pre-rendered representation n An approximate binaural representation is calculated based on Bin' m is a set of virtual probing poses P m and each virtual probing posture (P m ) is a probing position P by rotation around at least one of the rotation axes. n Step S14 may be followed by step S15.
[0053]
[0062] (Bin refIn step S15, the reference binaural representation Bin ref As will be explained in more detail below, the reference binaural representation Bin ref Binaural Pre-rendered Representation Bin n One of these, or an approximate binaural representation, Bin' m The reference binaural representation Bin ref may correspond to the reference pose P', or may be a pre-rendered representation corresponding to the reference pose. ref may be obtained by linearly combining several pre-rendered representations. Step S15 may be followed by step S16.
[0054]
[0063] In step S16 (generating M), the binaural pre-rendered representation Bin n and approximate binaural representation Bin' m The reference binaural representation Bin ref Reconstruction metadata M is calculated that allows reconstruction from
[0055]
[0064] Steps S14-S16 may all be performed by the metadata generator 13 of Figure 3. If a reference binaural representation is to be rendered, such rendering may be performed by the renderer 12, and the reference binaural representation will be one of the binaural pre-rendered representations. Step S16 may be followed by step S17.
[0056]
[0065] In step S17 (encoding and outputting the bitstream), the reference binaural representation Bin refand the reconstructed metadata M are encoded (e.g., by encoders 14, 15 or a single encoder) into an output bitstream (b2) and then output over an appropriate communication channel. The reconstructed metadata may be quantized and coded based on symmetries in the reconstructed metadata. For example, the reconstructed metadata may be coded using differential coding between metadata associated with different (symmetric) probing poses.
[0057]
[0066] The step of calculating reconstruction metadata (step S16) may include calculating axis-specific metadata for each rotation axis. In that case, the method may calculate axis-specific metadata for each rotation axis. n and approximate binaural representation Bin' m selecting a first set of representations from the set, the first set of representations corresponding to probing poses that are offset from one another by rotation about an axis (e.g., a yaw axis or a pitch axis), and calculating axis-specific reconstruction metadata M, H that enables at least one representation in the set to be reconstructed from another representation in the set, the axis-specific reconstruction metadata M, H representing displacement about that particular axis.
[0058]
[0067] According to an exemplary solution, the pre-renderer is capable of rendering binaural representations for three probing poses P1, P2 and P3.
[0059] 1. In the first step, pre-rendering obtains pre-rendered representations Bin1, Bin2, and Bin3 for three probing poses P1, P2, and P3 as follows:
number
[0060] 2. In the second step, the virtual probing posture
number
[0061] 3. In a third step, pitch correction metadata H can be calculated based on the pre-rendered representations Bin1 and Bin2 and the approximate rendered representation Bin'1, using the techniques described in US 63 / 340,181, using Bin'1 as the reference representation.
[0062] 4. In a final step, the yaw correction metadata M can be calculated based on the pre-rendered representation for the probing pose P3 and the approximated rendered representation for the pose P4 using the techniques described in US 63 / 340,181, again using Bin'1 as the reference representation.
[0063]
[0068] It is worth noting that a series of operational steps can be performed to calculate the pitch correction metadata after obtaining the yaw correction metadata, in which case the probing attitudes P1, P2, and P3 and the virtual probing attitude P4 are given as:
number
[0064]
[0069] The yaw and pitch metadata M, H as well as the binaural reference representation are encoded and transmitted to the lightweight device 3, which allows the lightweight device to perform attitude correction. In the above example, the approximate rendering representation Bin'1 is used as the reference for the calculation of both the yaw and pitch metadata, and it would be appropriate to encode and transmit this representation.
[0065]
[0070] However, in principle it is not necessary to use the same reference representation for both sets of metadata, and in fact the binaural reference rendering representation transmitted to the post-renderer (lightweight device 3) can be any of those available for the probing poses P1, P2, P3 and the virtual probing pose P4.
[0066]
[0071] For example, it is possible to choose a different binaural reference rendered representation associated with the virtual reference pose P'. This reference rendered representation may be obtained through a low-intensity processing post-renderer operation based on any (or a combination) of the pre-rendered representations available for the probing positions and using the techniques described in US 63 / 340,181. It is also possible to directly determine the reference rendered representation based on linear or triangular interpolation. Let b1, b2, and b3 denote the binaural pre-rendered representations for the probing positions P1, P2, and P3, and then the interpolated reference rendered representation b for the virtual pose P' is P’ can be calculated by the weighted average of:
number
[0067]
[0072] As described above in the example using yaw-only correction, the fact that the approximate and reference rendering representations for the virtual probing pose P4 are obtained by interpolation can create metadata symmetry, which can alleviate the need to calculate and transmit metadata related to some of the probing positions.
[0068]
[0073] 3-axis split renderer metadata operations Now, considering the three-axis case with yaw, pitch, and roll corrections, a trivial solution would be to consider six probing and reference attitudes and implement the technique proposed by US 63 / 340,181. Compared to the trivial solution for the two-axis case, the two additional probing attitudes probe roll displacements relative to the reference attitude, for example, by varying the attitude relative to the reference attitude by ±γ roll displacements while keeping pitch and yaw constant. This would require a total of seven pre-rendered representations.
[0069]
[0074] Below, a solution for computing split renderer metadata for post-render corrections around three axes is described. The description assumes yaw, pitch, and roll corrections, but the same principles can be applied for the computationally efficient computation of post-render metadata around any three (orthogonal) axes. Therefore, terms used like yaw, pitch, and roll in the following description can be replaced with any three rotational axes of yaw, pitch, and roll.
[0070]
[0075] According to an exemplary solution, the pre-renderer is able to render binaural representations for only four probing poses P1, P2, P3, P4.
[0071] 1. In the first step, pre-rendering obtains pre-rendered representations Bin1, Bin2, and Bin3 for three probing poses P1, P2, and P3 as follows:
number
[0072] 2. In the second step, the virtual probing posture
number
[0073] 3. In a third step, roll correction metadata can be calculated based on the pre-rendered representations Bin1 and Bin2 and the approximate rendered representation Bin'1 using the techniques described in US 63 / 340,181, using Bin'1 for pose P5 as the reference representation.
[0074] 4. In a fourth step, pitch correction metadata can be calculated based on the pre-rendered representation Bin3 for pose P3 and the approximated rendered representation Bin'1 for the virtual probing pose P5 using the techniques described in US 63 / 340,181, again using Bin'1 for pose P5 as the reference representation.
[0075] 5. In the fifth step, the virtual probing posture
number
[0076] 6. In the sixth step, pre-rendering is performed to obtain a fourth pre-rendered representation Bin4 for probing pose P4 as follows:
number
[0077] 7. In a final step, yaw correction metadata H can be calculated based on the pre-rendered representation Bin1 for the probing pose P4 and the approximated rendered representation Bin'2 for the pose P6, using either one of the rendered representations as the reference representation using the techniques described in US 63 / 340,181.
[0078] As described above, the operational steps of obtaining yaw, pitch, and roll correction metadata can be performed in different orders. The order may also be adapted based on the characteristics of the immersive audio signal. Some of the steps and correction metadata calculations may even be omitted based on the characteristics of the immersive audio signal. For example, in the case of an audio signal with a dominant sound coming from the left or right relative to the reference pose P′, pitch attitude correction can be approximated using a table containing gain parameters corresponding to various pitch angles. Such tables can be calculated once during initialization time, and both the pre-renderer and the post-renderer may have prior knowledge of these tables. In these cases, pitch correction metadata is not necessarily required in the bitstream, and the above steps may be adapted to calculate only yaw and roll correction metadata. Another example is the case of an audio signal with a dominant sound coming from the front or rear relative to the reference pose P′. In these cases, roll correction metadata is not necessarily required in the bitstream, and roll attitude correction can be approximated using a table containing gain parameters corresponding to various roll angles.
[0079] As mentioned above, the binaural reference rendering representation (representation) sent to the post-renderer can be any of the pre-rendering representations for the target probing pose or any of the approximated pre-rendering representations at the virtual probing pose. Furthermore, any other approximated binaural pre-rendering representation can be used as the reference rendering representation based on the available pre-rendering representations. The reference rendering representation can be obtained through a low-intensity processing post-renderer operation using the techniques described in US 63 / 340,181 based on any (or a combination) of the available pre-rendering representations for the probing position. An approximation of the pre-rendering representation at the reference pose can be obtained by interpolation between the available pre-rendering representations.
[0080] According to one example, the approximated rendered representation of the reference pose P′ can be obtained from the available pre-rendered representations for the probing poses P1 to P4. Let b1 to b4 denote the binaural pre-rendered representations for the probing poses P1 to P4, and let b1 to b4 denote the stored reference rendered representations b for the virtual pose P′. P’ can be obtained by the weighted average of:
number
[0081]
[0079] Notwithstanding the above, if complexity constraints permit, it may be desirable to additionally generate a pre-rendered representation for the reference pose P' and send this signal to the post-renderer as the reference rendered representation. The availability of a pre-rendered representation for the reference pose P' can also be used to enhance the yaw metadata computation in step 8 above.
[0082]
[0080] Iterative Split Renderer Metadata Extension The above example of reduced processing load and split renderer metadata calculation for post-renderer corrections has a bias. For example, in the three-axis case, roll correction metadata is calculated for yaw and pitch angle displacements from the reference attitude P' about -α and -β. This makes the obtained roll correction metadata less accurate when the yaw and pitch angles of the attitude are more likely to correspond to the yaw and pitch angles of the reference attitude P'. Therefore, calculating roll correction metadata for yaw and pitch angle displacements equal to 0 would be more accurate. Similarly, pitch correction metadata is biased because it is calculated for a yaw displacement angle of -α rather than 0. Bias in metadata calculation can cause inaccuracies in the rendered representation obtained by the post-renderer using the biased metadata.
[0083]
[0081] In the following, we describe an iterative enhancement technique that can mitigate the described bias and the resulting post-render inaccuracies. It is assumed that a reference pose P' is available. We refer to the description of the procedure above for the 3-axis case with yaw, pitch, and roll corrections.
[0084] In a first enhancement step, the roll correction metadata is enhanced using the reference pose P′ and the virtual probing poses P7, P8:
number
[0085] Part of this procedure is the calculation of approximate rendering representations for these virtual probing poses, for example by performing the low-intensity processing post-renderer operations of US 63 / 386,465 or US 63 / 340,181 using the previously calculated pitch and yaw correction metadata from the steps above and pre-rendering representations for the probing poses P1, P2. Along with these approximate rendering representations and the rendering representation of the reference pose P', roll correction metadata is recalculated using, for example, the techniques described in US 63 / 340,181.
[0086] In the corresponding second enhancement step, the pitch correction metadata is 10 This is improved using:
number
[0087] Part of this procedure is the calculation of approximate rendered representations for these virtual probing poses, for example by performing the low-intensity processing post-renderer operations of US 63 / 386,465 or US 63 / 340,181 using previously calculated roll and yaw correction metadata and / or previously performed pre-rendered representations. For example, an approximate rendered representation for pose P9 can be calculated by interpolating (e.g., averaging) between the pre-rendered representations for poses P1 and P2, and then adjusting that representation to vary the yaw displacement angle from -α to 0, while applying post-renderer techniques using the previously calculated yaw correction metadata. Also, the interpolation operation between the pre-rendered representations for poses P1 and P2 may include applying post-renderer techniques using the previously improved roll correction metadata. 10An approximate rendering representation for can be calculated using a pre-rendering representation of the probing pose P3 while applying a post-rendering technique using previously calculated yaw correction metadata.
[0088] In a corresponding third enhancement step, the yaw correction metadata is enhanced using pre-rendered representations for the reference pose P′ and rendered representations for the virtual probing poses. This enhancement step may involve applying post-render operations using the previously enhanced roll and pitch metadata and the pre-rendered representations available for the probing poses P1, P2, and / or P3.
[0089]
[0087] Each of the above metadata enhancement steps relies on previously calculated metadata, so by performing multiple iterations, further improvements can be achieved.
[0090]
[0088] Lightweight processing for yaw and pitch correction, pre-rendering of light side information (metadata) FIG. 5 is a flowchart illustrating processing in the main device 2 according to an embodiment of the second aspect of the present invention, relating to an audio rendering method for facilitating split rendering with attitude compensation about the yaw and pitch axes.
[0091] The flowchart may be decomposed into various blocks or partitions, such as blocks S21-S26. Processing for the various blocks in FIG. 5, which may be described as operations, processes, methods, steps, acts, or functions, may begin in block S21.
[0092]
[0090] In step S21 (obtaining audio content), a first bitstream is received and decoded (e.g., by the decoder 11) to obtain immersive audio content A, and in step S22 (obtaining reference pose), a reference pose P' is obtained. The reference pose may be an assumed pose (e.g., looking straight ahead) or may be based on pose information received from the lightweight device 3. Step S21 may be followed by step S22.
[0093]
[0091] (Bin ref In step S23, the immersive audio content A is rendered (for example by the renderer 12) into a reference binaural representation Bin corresponding to the reference pose P′. ref Step S23 may be followed by step S24.
[0094] In step S24 (pre-rendering), the immersive audio content A is converted (for example by the renderer 12) into two binaural pre-rendered representations Bin n The binaural pre-rendered representation is then rendered at two probing poses P, which are displaced from the reference pose by rotations about both the yaw and pitch axes. n In other words, each probing attitude is displaced from the reference attitude by rotation about a probing axis 104 (see FIG. 2) that has the same origin as the yaw and pitch axes and extends between the yaw axis 101 and the pitch axis 102. The displacements in yaw and pitch may be equal, in which case the probing 104 extends symmetrically between the yaw and pitch axes (i.e., at 45 degrees from each axis). Step S24 may be followed by step S25.
[0095]
[0093] In step S25 (calculating M and H), the yaw metadata M representing the displacement around the yaw axis and the pitch metadata H representing the displacement around the pitch axis are calculated for each probing attitude Pn is calculated (for example, by the metadata generating unit 13). Step S25 may be followed by step S26.
[0096]
[0094] In step S26 (encoding and outputting the bitstream), the reference binaural representation Bin ref The yaw metadata M and pitch metadata H are encoded (e.g., by encoders 14, 15) into an output bitstream b2, which is then output over an appropriate communication channel. The yaw and pitch metadata can be quantized and coded based on symmetries in the reconstructed metadata. For example, the metadata may be coded using differential coding between the metadata for different (symmetric) probing poses.
[0097] As will be explained in more detail below, the step of calculating yaw and pitch metadata involves the binaural pre-rendered representation Bin n The reference binaural representation Bin ref To enable reconstruction from the complete reconstruction metadata M ^ and then calculating yaw and pitch metadata based on the complete reconstruction metadata, which includes, for each time-frequency tile, a complex or real 2 × 2 transformation matrix M ^ may contain.
[0098] The yaw metadata may include, for each time-frequency tile, a complex or real 2x2 yaw correction matrix M. The pitch metadata may include, for each time-frequency tile, a real 2x2 diagonal pitch correction matrix H.
[0099] In an exemplary implementation, the low-load processing post-renderer device 3 transmits the reference head pose P′ via a back channel to the heavy weight pre-renderer device 2. The heavy weight device uses the pose P′ to generate the reference binaural signal Bin p’ and metadata, so that the post-renderer performs pose correction from P' to the actual pose P, and uses the metadata to generate Bin p’ From Bin p where Bin p contains all spatial cues to posture P. As explained here, the discrepancy between P' and P depends on the latency from movement to sound.
[0100]
[0098] The 3DOF attitude angles for the yaw, pitch, and roll axes at the attitude P' are respectively defined as α ref ,β ref ,γ ref and the angular displacements about the yaw, pitch, and roll axes between P and P' are α d ,β d ,γ d (Pitch and roll displacement) β d and γ d than α d It has been observed that compensating for pitch (or yaw) displacement is perceptually more important. Furthermore, pitch displacement is more relevant to user applications than roll displacement.
[0101] In an exemplary implementation, a low processing load and low metadata rate solution is proposed to correct αd and βd, and the heavy weight device uses a reference posture P′ to generate a reference binaural signal Bin p’ and two additional rendering representations are computed for each probing head pose as shown below:
number
[0102]
[0100] Prediction parameters are calculated as in US 63 / 340,181, with modifications as shown below: First, the prediction matrix is calculated as in US 63 / 340,181:
number
number
[0103]
[0102] If α is small, it can be assumed that the total energy of the binaural signal rendered at pose P'+(α,0,0) is the same as the total energy of the binaural signal rendered at pose P', and any energy difference between P and P' arises from the pitch angle β. Based on this assumption, the prediction matrix M p1 ^ can be scaled, so that the total energy of the post-prediction matrix is p’ and one or more additional real-only scale factors can be calculated to compensate for the pitch gain. Using these parameters, Binp1 can be estimated as:
number
number
[0104]
[0103] It can be shown that:
number
[0105]
[0104] H p1 and M p1 =G p1 M ^ p1 are quantized and coded. Similarly, parameters corresponding to the P2 posture are calculated and coded. These coded bits are then used to calculate the coded Bin p’ The data is multiplexed with the H bit and sent to the post-renderer. p1 ,M p1 ,H p2 ,M p2 ,Bin p’ [2x1]. Furthermore, if the actual pose P in the post-renderer is not equal to any of P1 or P2 or P', the parameters corresponding to P1 or P2 or both are interpolated or extrapolated using linear interpolation, which involves selecting two pose points among P', P1, and P2 that are close to the pose P, where the two pose points may differ in terms of yaw and pitch interpolation or extrapolation. The parameters M and H are then interpolated or extrapolated between the two selected pose points using linear interpolation. In an exemplary implementation, the interpolated parameters are H p and M p If so, Bin p The binaural signal corresponding to can be calculated as follows:
number
[0106] In some embodiments, this may be as follows, which means that only one pitch gain parameter needs to be coded for both the left and right channels in a given time-frequency tile:
number
[0107]
[0106] Selection of pre-rendered representations based on perceptual importance In an exemplary implementation, the number of pre-rendered representations may be controlled by selecting pose points based on perceptual importance, as follows: Let α be the 3DOF pose angles about the yaw, pitch, and roll axes at pose P', respectively. ref ,β ref ,γ ref and the angular displacements about the yaw, pitch, and roll axes between P and P' are α d ,β d ,γ d (Pitch and roll displacement) β d and γ d than α d It has been observed that correcting for pitch (or displacements in the yaw axis) is perceptually more important, and it is desirable to make attitude corrections about the yaw axis as accurate as possible. For this reason, it is desirable to have a larger number of probing attitude points for generating yaw-only related side information compared to the number of probing attitude points for pitch- and roll-related side information. Furthermore, displacements related to the pitch axis are likely to be more relevant to user applications than displacements related to the roll axis, and therefore it may be desirable to have a larger number of probing attitude points for generating pitch-related side information compared to the number of probing attitude points for generating roll-related side information. In an exemplary implementation, the following probing attitude points have been selected for generating side information for rotations about the yaw, pitch, and roll axes:
number
[0108]
[0107] The side information corresponding to P1 and P2 can be calculated as in US 63 / 340,181. To calculate the side information corresponding to P3, it can be assumed that interaural time difference (ITD) cues do not vary with pitch angle, and that posture displacements about the pitch axis can be modeled using one or more real-only gain parameters. These gain parameters can be calculated as follows:
number
[0109] H3 is quantized and coded, and contains yaw and roll related side information and coded Bin p’ Together with the coded bits for the signal, they are multiplexed into a bitstream.
[0110] In some implementations,
number
[0111]
[0110] Roll angle displacements can change ITD cues, and therefore it may be desirable to model roll displacements using complex gain parameters at low frequencies (e.g., 0-2 kHz) and real-only gain parameters at high frequencies (e.g., above 2 kHz). The side information corresponding to the roll-probing attitude P4 = P' + (0,0,γ) can be calculated as follows:
[0112]
[0111] The prediction parameters can be calculated as in US63 / 340,181, with the corrections as shown below. In an exemplary implementation, the same corrections are applied to the side information corresponding to the yaw probing attitude (M p1 and M p2 ). First, the prediction matrix is calculated as follows:
number
[0113] Using the above prediction matrix, the post-prediction matrix is calculated:
number
[0114] To further match the energy of the post-prediction matrix, an additional gain matrix G p4 Calculate:
number
[0115]
[0114] M p4 =G p4 M ^ p4 is quantized and coded, and the yaw and pitch related side information and coded Bin p’ Together with the coded bits for the signal, they are multiplexed into a bitstream.
[0116]
[0115] Figure 6 is a flow chart illustrating processing in a lightweight device 3 according to an embodiment of another aspect of the present invention, relating to a split rendering method with pose correction about multiple rotation axes.
[0117]
[0116] The flowchart may be decomposed into various blocks or partitions, such as blocks S41-S45. Processing for the various blocks in Figure 6, which may be described as operations, processes, methods, steps, acts, or functions, may begin in block S41.
[0118] In step S41 (receiving and decoding a bitstream), a bitstream is received from a main device (e.g., main device 2) and decoded (e.g., by decoders 22, 23) to generate a reference binaural representation Bin ref and a set of probing poses P that represent the displacements from the reference pose P' due to rotations around multiple rotation axes. n Step S41 may be followed by step S42.
[0119]
[0118] In step S42 (detecting current head pose), the current head pose P is detected (eg by the head tracker 25). Step S42 may be followed by step S43.
[0120] Step S43 (for each axis) is the beginning of a loop that includes one or more of steps S44-S45. A loop is performed for each of the rotational axes, e.g., yaw, pitch, and roll, respectively. Note that step S43 may be followed by step S44 if additional processing is required for the additional rotational axes. Alternatively, step S43 may be followed by step S46 if processing is not required for any additional rotational axes.
[0121]
[0120] In step S44 (selecting a probing pose), the probing pose that is closest to the detected pose for a particular axis of rotation is selected. Step S44 may be followed by step S45.
[0122] In step S45 (generating second metadata), axis-specific reconstruction metadata (M) is generated based on the first reconstruction metadata associated with the selected probing pose and the difference between the current head pose and the reference pose for a particular rotation axis. α ,M β ,M γ ) is determined. If the processing loop is completed, step S46 may be followed by step S43 or step S46.
[0123]
[0122] (Bin out In step S46, the binaural output Bin corresponding to the current head posture is determined. out is the reference binaural representation Bin ref and second, axis-specific reconstruction metadata for each axis of rotation.
[0124] An indication of the reference pose (P') may be obtained from the bitstream. Alternatively, in embodiments where an indication of the current head pose (P) is transmitted to the main device, it is possible to determine the reference pose (P') based on the expected delay of transmission to the main device.
[0125] The set of probing poses may be obtained from the bitstream, or by adding a set of offsets to the reference poses, where such offsets may be predefined (e.g., known a priori) or obtained from the bitstream.
[0126] Returning to our concrete example, in the post-renderer, the estimates of the left and right channels of the binaural signal corresponding to the actual pose P are given as:
number
[0127]
[0126] Note that the additional gain matrix G calculated here for pose P4 may be calculated for any pre-rendered representation including P1 and P2. The matrix M calculated as disclosed in US 63 / 340,181 for a particular pose ^ The combination of ,with an additional gain matrix G for the same pose is an advantageous way of generating metadata for segmented rendering that is new to the art.
[0128]
[0127] Figure 7 is a flowchart illustrating processing in the main device 2 according to another embodiment of the present invention. The flowchart may be decomposed into various blocks or partitions, such as blocks S51-S58. Processing for the various blocks in Figure 7, which may be described as operations, processes, methods, steps, acts, or functions, may begin in block S51.
[0129]
[0128] In step S51 (obtaining audio content), a first bitstream is received and decoded (for example by the decoder 11) to obtain immersive audio content A. Step S51 may be followed by step S52.
[0130] In step S52 (receiving head pose information), the second bitstream is received and decoded (e.g., by decoder 17) to receive head pose information P associated with the user of the lightweight processing device. Step S52 may be followed by step S53.
[0131]
[0130] In step S53 (determining a reference pose), a reference pose P' is determined (eg in the decoder 17) based on the received head pose information. Step S53 may be followed by step S54.
[0132]
[0131] (Bin refIn step S54, the immersive audio content A is rendered (for example by the renderer 12) into a reference binaural representation Bin corresponding to the reference pose. ref Step S54 may be followed by step S55.
[0133] In step S55 (pre-rendering), the immersive audio content A is rendered (for example by the renderer 12) into one or more binaural pre-rendered representations Bin n The pre-rendered representation is then rendered into one or more probing poses P, which are displaced from the reference pose around the axis of rotation. n Step S55 may be followed by step S56.
[0134]
[0133] In step S56 (generating metadata), the binaural pre-rendered representation Bin n The reference binaural representation Bin ref The reconstruction metadata is calculated (by the metadata generator 13) to enable reconstruction from the transformation matrix M ^ Step S56 may be followed by step S57.
[0135] In step S57 (of improving the metadata), improved metadata M is calculated (for example by the metadata generator 13) by multiplying each reconstruction matrix M̂ for a particular pose Ps by an additional gain matrix G, which has the following form:
number
[0136] It may be advantageous to have at least two pre-rendered representations for each axis of rotation. The axis of rotation may be the yaw axis and / or the roll axis. For the yaw and roll axes, it may be possible to show that the variance of the left and right channels after application of the improved metadata M is substantially equal to the variance of the left and right channels of the reference binaural representation.
[0137]
[0136] In step S58 (encoding and outputting the bitstream), the reference binaural representation Bin ref and the reconstructed metadata M is encoded (e.g., by encoders 14, 15) into an output bitstream b2, which is then output over an appropriate communication channel. The reconstructed metadata may be quantized and coded based on symmetries in the reconstructed metadata. For example, the metadata may be coded using differential coding between metadata associated with different (symmetric) probing poses.
[0138] The approach described in steps S51-58 may be combined with pitch metadata as described above. In that case, the set of probing poses includes pitch probing poses that are displaced from the reference pose only by rotation about the pitch axis. Furthermore, pitch reconstruction metadata is calculated to generate binaural pre-rendered representations Bin corresponding to the pitch probing poses. n The reference binaural representation Bin ref The pitch reconstruction metadata includes a real-diagonal 2x2 pitch correction matrix H for each time-frequency tile.
[0139] As described above, the pitch correction matrix for attitude P1 is
number
number
[0140] Alternatively, the pitch correction matrix for attitude P1 is
number
number
[0141]
[0140] Selection of pre-rendered representations based on head movement speed and direction 8 is a flowchart illustrating processing in the main device 2 according to another embodiment of the present invention. The flowchart may be broken down into various blocks or partitions, such as blocks S31-S37. Processing for the various blocks in FIG. 8, which may be described as operations, processes, methods, steps, acts, or functions, may begin in block S31.
[0142]
[0141] In step S31 (obtaining audio content), a first bitstream is received and decoded (for example by the decoder 11) to obtain immersive audio content A. Step S31 may be followed by step S32.
[0143] In step S32 (receiving head pose information), the second bitstream is received and decoded (for example by decoder 17) to provide head pose information (P,Ω) associated with the user of the lightweight processing device. P ,ω P Step S32 may be followed by step S33.
[0144]
[0143] In step S33 (determining the reference posture), the reference posture P' and the head rotation axis Ω P and head posture rotation speed ω P is determined (eg, in the decoder 17) based on the received head pose information. Step S33 may be followed by step S34.
[0145]
[0144] (Bin ref In step S34, the immersive audio content A is rendered (for example by the renderer 12) into a reference binaural representation Bin corresponding to the reference pose. ref Step S34 may be followed by step S35.
[0146] In step S35 (pre-rendering), the immersive audio content (A) is converted (for example by the renderer 12) into a set of binaural pre-rendered representations Bin n The binaural pre-rendered representation is then rendered at a set of probing poses P, which are rotated relative to the reference pose. n The probing posture corresponds to the head posture information (P, Ω P ,ω P ) Step S35 may be followed by step S36.
[0147] In step S36 (generating M), the binaural pre-rendered representation Bin n The reference binaural representation Bin refReconstruction metadata M is calculated (by the metadata generator 13) that allows reconstruction from Step S36. Step S36 may be followed by Step S37.
[0148]
[0147] In step S37 (encoding and outputting the bitstream), the reference binaural representation Bin ref and the reconstruction metadata M is encoded (e.g., by encoders 14, 15) into an output bitstream b2, which is then output over an appropriate communication channel. The reconstruction metadata may be quantized and coded based on symmetries in the reconstruction metadata. For example, the reconstruction metadata may be coded using differential coding between metadata associated with different (symmetric) probing poses.
[0149]
[0148] As further explained below, the head pose rotation axis Ω P Once is determined, the probing posture P n is the head posture rotation axis Ω P The probing position P can be displaced from the reference position P' by rotating around the n may be symmetrically distributed around the reference pose P'. Also, in this example, the probing pose P n may include only one probing pose about each rotational degree of freedom, for example, probing poses may be selected about the yaw and pitch axes.
[0150]
[0149] Furthermore, head posture rotation speed ω P When the head pose rotation velocity ω P When is below a predetermined threshold, the probing posture P n is the first angle ω lower may be selected to have a smaller displacement from the reference posture, and the head posture rotation velocity ω P When exceeds the threshold, the probing posture P n is the second angle ω upper may be displaced from the reference position by a larger angle, where the first angle ω lower is the second angle ω upperSmaller than.
[0151] In an exemplary implementation, the lightweight post-renderer device 3 transmits the reference head pose P' to the heavyweight pre-rendering device via a back channel. The heavyweight device uses P' to generate the reference binaural signal Bin p’ and generate metadata M, so that the post-renderer performs pose correction from P' to the actual pose P, and uses the metadata to generate Bin p’ From Bin p where Bin p contains all spatial cues related to posture P. The discrepancy between P' and P depends on the motion-sound latency.
[0152]
[0151] Here, the number of pre-rendering representations is determined by the head movement speed (rotation speed, ω P ) or head movement direction (head posture rotation axis, Ω P ) or both. Using pose information from the post-render device, head movement speed and direction can be calculated in the pre-renderer. In some implementations, the post-renderer may provide the speed and direction of movement along with the pose information. Using this information, the pre-renderer can significantly reduce the number of probing pose points for the pre-rendered representation by selecting probing pose points along the head movement axis.
[0153] For example, if the pose from the post-render is P' and the head is rotating around the yaw axis, the following probing pose points can be selected to generate side information:
number
[0154]
[0153] Here, P' is the reference pose from the post-renderer, and P'+(α,0,0) is the pose rotated around the head movement axis (the yaw axis in this case) in the head movement direction.
[0155] In general, the following probing pose points may be used to generate side information:
number
[0156] In an exemplary implementation, the head movement velocity and acceleration may be used to further limit the number of pose points to one to generate side information. If the delay between the post-renderer and the pre-renderer is known and below a threshold, and the head is accelerating, the pre-rendered representation corresponding to the following probing pose point is sufficient to generate side information:
number
[0157]
[0156] The side information can be calculated according to US 63 / 340,181.
[0158]
[0157] When the head is stationary, i.e. the velocity is zero, the following probing pose points can be chosen to generate side information:
number
[0159]
[0158] The head is stationary (or, more generally, the head pose rotation velocity ω P is below a given threshold), we can safely assume that the angular deviations about the yaw, pitch, and roll axes between the reference pose P' in the pre-renderer and the actual pose P in the post-renderer will be smaller than the deviations about the yaw, pitch, and roll axes if the head was not stationary (or rotated faster). This assumption allows α, β, and γ at the probing pose point to be small (e.g., the lower bound ω lower The values of α, β, and γ can be set to a value of n smaller than α, β, and γ, which will be sufficient to estimate the side information corresponding to -α, -β, and -γ. In an exemplary implementation, the values of α, β, and γ are controlled based on the head velocity and acceleration. The side information of P1, P2, and P3 can be calculated as in the above section.
[0160]
[0159] The systems and methods disclosed in this disclosure may be implemented as software, firmware, hardware, or a combination thereof. In a hardware implementation, the division of tasks does not necessarily correspond to the division into physical units; conversely, one physical component may have multiple functions, and one task may be performed by several physical components working together.
[0161]
[0160] Computer hardware may be, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a smartphone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) that specify operations to be performed by that computer hardware. Furthermore, this disclosure relates to any collection of computer hardware that individually or collectively executes instructions to perform any one or more of the concepts described herein.
[0162] FIG. 9 illustrates a schematic block diagram of an exemplary electronic device or architecture 200 (e.g., apparatus 200) suitable for implementing exemplary embodiments of the present disclosure. The architecture 200 includes, but is not limited to, a main processing device and a lightweight processing device, such as those described in connection with FIG. 3. As illustrated, the architecture 200 includes a central processing unit (CPU) 201 capable of executing various processes according to programs stored, for example, in read-only memory (ROM) 202 or loaded, for example, from a storage unit 208 into random access memory (RAM) 203. The CPU 201 may be, for example, an electronic processor 201 that may include one or more processor cores, and in some examples, the processor 201 may be multiple processors. The RAM 203 also stores data needed by the CPU 201 to perform various processes, as needed. The CPU 201, the ROM 202, and the RAM 203 are interconnected via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204 .
[0163]
[0162] The following components are connected to the I / O interface 205: an input unit 206, which may include a keyboard, a mouse, or the like; an output unit 207, which may include a display, such as a liquid crystal display (LCD), and one or more speakers; a storage unit 208, which may include a hard disk or another suitable storage device; and a communication unit 209, which may include a network interface card, such as a network card (e.g., wired or wireless).
[0164]
[0163] In some implementations, the input unit 206 includes one or more microphones in different positions (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0165] In some implementations, the output unit 207 includes a system with a variable number of speakers and is capable of rendering audio signals in a variety of formats (e.g., mono, stereo, immersive, binaural, and other suitable formats) (depending on the capabilities of the host device).
[0166]
[0165] In some implementations, the communication unit 209 is configured to communicate with other devices (e.g., via a network). The drive 210 is also connected to the I / O interface 205, as needed. A removable medium 211, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium, is attached to the drive 210, so that computer programs read therefrom are installed in the storage unit 208, as needed. Although the device 200 has been described as including the above components, those skilled in the art will understand that in actual applications, some of these components may be added, removed, and / or substituted, and all such modifications or variations fall within the scope of the present disclosure.
[0167] According to exemplary embodiments of the present disclosure, the above-described processes may be implemented as a computer software program or in a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied in a machine-readable medium, the computer program including program code for performing the method. In such embodiments, the computer program may be downloaded and installed from a network via a communication unit 209 and / or from a removable medium 211, as shown in FIG. 9.
[0168] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the various elements of FIG. 3 described above may be executed by control circuitry (e.g., CPU 201 in combination with other components of FIG. 9), and thus the control circuitry may perform the operations described in the present disclosure.
[0169]
[0168] Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be executed by a controller, processor, and / or other computing device, which may include control circuitry. Although various aspects of the exemplary embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other depicted representation, it will be appreciated that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuitry or logic, general purpose hardware or controller or other computing device, or any combination thereof.
[0170]
[0169] Furthermore, the various blocks illustrated in the flowcharts may be considered as method steps and / or as operations resulting from the operation of computer program code and / or as a plurality of coupled logic circuit elements configured to perform the related functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the method described above.
[0171]
[0170] Computer program code for carrying out the methods of the present disclosure may be written in any combination of one or more programming languages. The computer program code may be provided to one or more processors of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that the program code, when executed by one or more processors of the computer or other programmable data processing apparatus, performs the functions / acts specified in the flowcharts and / or block diagrams. The program code may run entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer, partially on a remote computer, entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0172]
[0171] One or more processors may operate as standalone devices or may be connected, e.g., networked, to other processors. Such a network may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0173]
[0172] Software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes various forms of physical (non-transitory) storage media, such as, but not limited to, ROM, PROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store the desired information and that can be accessed by a computer. Additionally, those skilled in the art will recognize that communication media (transitory) typically embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and include any information delivery media.
[0174]
[0173] The implementations of the technology disclosed in the drawings are merely illustrative examples, and the present invention is not so limited. For example, the illustrated divisions, such as the blocks in FIG. 3, are merely exemplary logical divisions for ease of explanation, and such divisions may be divided into additional divisions, combined into fewer divisions, supplemented with additional divisions, or reduced by eliminating divisions without departing from the spirit of the present invention. With respect to the illustrative flowcharts in FIGS. 4-8, divisions of operational steps, which may also be referred to as functions, steps, operations, processes, or acts, may be combined into fewer steps or divided into additional steps, where steps may be rearranged or eliminated, in whole or in part, without departing from the spirit of the present disclosure.
[0175]
[0174] Unless otherwise specified, as will be apparent from the following description, it is understood that throughout this disclosure, descriptions using terms such as "processing," "operating," "calculating," "calculating," "determining," "analyzing," or the like may refer to the functions, operations, steps, and / or processes of computer hardware or computing systems or similar electronic computing devices that manipulate and / or convert data expressed as physical quantities, such as electronic quantities, into other data similarly expressed as physical quantities.
[0176] In the foregoing description of exemplary embodiments of the present invention, it should be appreciated that various features of the invention are often grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in understanding one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment. Accordingly, the claims following the detailed description are expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention. Furthermore, some embodiments described herein include some features but not other features included in other embodiments, and combinations of features from different embodiments are intended to form different embodiments within the scope of the present invention, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0177] Furthermore, some embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or by other means for performing a function. Thus, a processor with instructions for executing such a method or method elements forms a means for performing the method or method elements. It should be noted that when a method includes several elements, e.g., several steps, no ordering of such elements is implied unless specifically stated. Furthermore, elements described herein of apparatus embodiments are examples of means for performing the functions performed by the elements for the purposes of implementing embodiments of the present invention. Numerous specific details are set forth in the description provided herein. However, it is understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0178]
[0177] Those skilled in the art will appreciate that the present invention is by no means limited to the preferred embodiment described above. On the contrary, many modifications and variations are possible within the scope of the appended claims. For example, as described above, the probing poses may be asymmetrically distributed around the reference pose. Also, the selection of the actual pre-rendering representation and the approximate representation may differ from those proposed above. Also, various additional techniques for encoding metadata not disclosed herein may be used.
Claims
1. 1. A method of audio rendering that enables split rendering with attitude compensation for multiple axes of rotation, comprising: obtaining immersive audio content (A); A step of obtaining a reference attitude (P'); The immersive audio content (A) is converted into a first number of binaural pre-rendered representations (Bin n ), wherein the binaural pre-rendered representation (Bin n ) is a set of probing postures (P) associated with the reference posture (P'). n ) and the set of probing postures (P n ) includes a position equal to the reference position (P′) and / or a position displaced from the reference position (P′) by a rotation about at least one of the rotation axes; The binaural pre-rendered representation (Bin n ) based on which we can obtain an approximate binaural representation of the second number (Bin' m ) by calculating the approximate binaural representation (Bin' m ) is a set of virtual probing postures (P m ) and corresponds to the virtual probing attitude (P m ) is a rotation about at least one of the rotation axes that changes the probing position (P n ) displaced from the step; The approximate binaural representation (Bin' m ) and the binaural pre-rendered representation (Bin n ) based on one or more of the reference binaural representations (Bin ref ) determining the The binaural pre-rendered representation (Bin n ) and the approximate binaural representation (Bin' m ) of the reference binaural representation (Bin ref calculating reconstruction metadata (M) that allows reconstruction from the The reconstruction metadata (M) and the reference binaural representation (Bin ref ) and the output bitstream (b 2 ) encoding the The output bitstream (b 2 ) outputting the A method comprising:
2. The method according to claim 1, wherein the approximate binaural representation (Bin') m ) is the binaural pre-rendered representation (Bin n ) is calculated by interpolating
3. 3. The method according to claim 1, wherein the approximate binaural representation (Bin') m ) is calculated by modifying the reconstruction metadata and applying the modified reconstruction metadata to one of the binaural pre-rendered representations.
4. 4. The method according to claim 1, wherein the reference binaural representation (Bin ref ) is the binaural pre-rendered representation (Bin n ) and the approximate binaural representation (Bin' m ) A method.
5. 4. The method according to claim 1, wherein the reference binaural representation (Bin ref ) corresponds to said reference pose (P') and is obtained by linearly combining said binaural pre-rendered representations.
6. 4. The method according to claim 1, wherein the reference binaural representation (Bin ref ) corresponds to said reference pose (P') and is obtained by binaural rendering of said immersive audio content (A).
7. 7. The method according to claim 1, wherein the step of calculating the reconstruction metadata (M) comprises the steps of: The binaural pre-rendered representation (Bin n ) and the approximate binaural representation (Bin' m selecting a first set of representations from the set of representations, the first set of representations corresponding to probing poses that are displaced from one another by rotation about the rotation axis; and calculating axis-specific reconstruction metadata (M,H), the axis-specific reconstruction metadata (M,H) enabling reconstruction of at least one representation in the set from another representation in the set, the axis-specific reconstruction metadata (M,H) representing displacement about the axis of rotation; A method comprising:
8. 2. The method of claim 1, wherein at least one of the axis-specific reconstruction metadata (M,H) is a representation of all approximate binaural representations (Bin' m ) is calculated before the approximate binaural representation (Bin' m ) is calculated by applying the axis-specific reconstruction metadata to one of the binaural pre-rendered representations.
9. 9. The method of claim 1, wherein the first number is D+1 and the second number is D-1, where D is the number of axes of rotation.
10. 10. The method of claim 9, wherein D=2 and: The first probing position (P 1 ) is displaced from the reference orientation (P′) by a first angle (−α) about a first axis and a second angle (−β) about a second axis; The second probing position (P 2 ) is displaced from the reference orientation (P′) by a first angle (−α) about the first axis and by a third angle (+β) about the second axis; The third probing position (P 3 ) is displaced from the reference position (P′) by a fourth angle (+α) about the first axis and is equal to the reference position (P′) with respect to the second axis; First virtual probing position (P 4 ) is displaced from the reference orientation (P') by a first angle (-α) about the first axis and is equal to the reference orientation (P') with respect to the second axis.
11. 11. The method of claim 10, wherein the fourth value (+α) is the negative of the first value (-α), and the third value (+β) is the negative of the second value (-β).
12. 10. The method of claim 9, wherein D=3: The first probing position (P 1 ) is displaced from the reference orientation (P′) by a first angle (−α) about a first axis, a second angle (−β) about a second axis, and a third angle (−γ) about a third axis; The second probing position (P 2 ) is displaced from the reference orientation (P′) by a first angle (−α) about the first axis, a second angle (−β) about the second axis, and a fourth angle (+γ) about the third axis; The third probing position (P 3 ) is displaced from the reference orientation (P') by a first angle (-α) about the first axis and a fifth angle (+β) about the second axis, and is equal to the reference orientation (P') with respect to the third axis; The fourth probing position (P 4 ) is displaced from the reference position (P′) by a sixth angle (+α) about the first axis and is equal to the reference position (P′) with respect to the second and third degrees of freedom; First virtual probing position (P 5 ) is displaced from the reference orientation (P′) by the first angle (−α) about the first axis and the second angle (−β) about the second axis, and is equal to the reference orientation (P′) with respect to the third axis; and The second virtual probing position (P 6 ) is displaced from the reference orientation (P') by the first angle (-α) about the first axis and is equal to the reference orientation (P') with respect to the second and third degrees of freedom.
13. 13. The method of claim 12, wherein the sixth angle (+α) is the negative of the first angle (-α), the fifth angle (+β) is the negative of the second angle (-β), and the fourth angle (+γ) is the negative of the third angle (-γ).
14. 14. The method of any one of claims 10 to 13, wherein the first axis is a yaw axis, the second axis is a pitch axis, and the third axis is a roll axis.
15. 14. The method according to any one of claims 10 to 13, wherein the reference pose (P') is based on a user head pose obtained from a user's mobile device.
16. 16. The method of claim 15, wherein the attitude information representing the reference attitude (P') is encoded and the output bitstream (b 2 ) a method included in
17. 17. The method according to claim 1, wherein the probing position (P n ) and virtual probing attitude (P m ) is encoded into the output bitstream (b 2 ) a method included in
18. 18. The method according to any one of claims 1 to 17, wherein the reconstruction metadata (M) comprises a 2x2 matrix for each time-frequency tile.
19. 20. The method of claim 18, further comprising the step of quantizing and encoding the reconstructed metadata (M) based on symmetries of the reconstructed metadata.
20. 20. The method of claim 18, further comprising encoding the reconstructed metadata (M) using differential coding between metadata relating to different probing poses.
21. 21. An apparatus comprising a processor and a memory coupled to the processor, the processor configured to cause the apparatus to perform a method according to any one of claims 1 to 20.
22. 21. A computer readable storage medium storing a program including instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 20.
23. 1. A method of audio rendering that allows split rendering with attitude compensation for yaw and pitch axes, comprising: obtaining immersive audio content (A); A step of obtaining a reference attitude (P'); The immersive audio content (A) is converted into a reference binaural representation (Bin ref ); The immersive audio content (A) is converted into one or more binaural pre-rendered representations (Bin n ), wherein the binaural pre-rendered representation (Bin n ) corresponds to one or more probing attitudes displaced from the reference attitude by rotation about both the yaw axis and the pitch axis; calculating, for each probing attitude, yaw metadata (M) representing displacement about the yaw axis and pitch metadata (H) representing displacement about the pitch axis; The reference binaural representation (Bin ref ), the yaw metadata (M), and the pitch metadata (H) are output to the output bitstream (b 2 ) encoding the The output bitstream (b 2 ) outputting the A method comprising:
24. 24. The method of claim 23, wherein the pitch metadata (H) is calculated based on an energy difference between the reference binaural representation and the binaural pre-rendered representation.
25. 25. The method according to claim 23 or 24, wherein the step of calculating the yaw metadata (M) and the pitch metadata (H) is performed by using the binaural pre-rendered representation (Bin' m ) of the reference binaural representation (Bin ref Full reconstruction metadata (M) allows reconstruction from ^ ) and then calculating yaw and pitch metadata (H) based on the full reconstruction metadata.
26. 26. The method of claim 25, wherein the perfect reconstruction metadata comprises, for each time-frequency tile, a 2×2 transformation matrix (M ^ ) a method comprising:
27. 26. The method of claim 25, wherein the yaw metadata includes a 2x2 yaw correction matrix (M) for each time-frequency tile.
28. 28. The method of claim 27, wherein the yaw correction matrix (M) is a transformation matrix (M ^ ) by a 2x2 diagonal gain matrix (G).
29. 29. The method of claim 28, wherein the specific probing position P s The gain matrix (G) for [Equation 1] where: R y,p’,p’ is the reference binaural representation Bin p’ represents the covariance matrix of R y^,ps,ps is the specific probing posture P s represents the 2×2 covariance matrix of the reconstructed binaural pre-rendered representation for
30. 30. The method of any one of claims 23-29, wherein the pitch metadata comprises, for each time-frequency tile, a real diagonal 2x2 pitch correction matrix (H).
31. The method according to claim 30, wherein the orientation P 1 The pitch correction matrix H for P1 The elements of are obtained by the following formula: [Equation 2] R y,p’,p’ is the reference binaural representation Bin p’ is the covariance matrix of R y,p1,p1 is posture P 1 The method is a covariance matrix of the binaural pre-rendered representation for
32. 32. The method according to any one of claims 23 to 31, wherein the probing position (P n ) is displaced from the reference attitude (P') by equal rotations about both the yaw axis and the pitch axis.
33. 33. The method according to any one of claims 23-32, wherein the reference pose (P') is based on a user head pose obtained from a user's mobile device.
34. 34. The method of claim 33, wherein the attitude information representing the reference attitude (P') is encoded and the output bitstream (b 2 ) a method included in
35. 35. The method according to any one of claims 23 to 34, wherein the probing position (P n ) is encoded into the output bitstream (b 2 ) a method included in
36. 36. The method of any one of claims 23 to 35, wherein the yaw metadata (M) comprises a 2x2 matrix for each time-frequency tile.
37. 37. The method of any one of claims 23 to 36, wherein the pitch metadata (H) comprises a 2x2 matrix for each time-frequency tile.
38. 38. The method of claim 36 or 37, further comprising the step of quantizing and encoding the yaw and / or pitch metadata (M,H) based on symmetry of the reconstructed metadata.
39. 39. The method according to any one of claims 36 to 38, further comprising encoding the yaw and / or pitch metadata (M,H) using differential coding between metadata relating to different probing poses.
40. 40. An apparatus comprising a processor and a memory coupled to said processor, said processor configured to cause said apparatus to perform a method according to any one of claims 23-39.
41. 40. A computer readable storage medium storing a program including instructions that, when executed by a processor, cause the processor to perform a method according to any one of claims 23-39.
42. A main processing device (2) comprising: The first bitstream (b 1 a decoder (11) configured to decode the audio content (A) to obtain a decoded immersive audio content (A); obtaining a reference attitude (P'); The immersive audio content (A) is converted into a reference binaural representation (Bin ref ) and The immersive audio content (A) is divided into a number of binaural pre-rendered representations (Bin n and a renderer (12) configured to render said binaural pre-rendered representation (Bin n ) is a set of probing postures (P) associated with the reference posture (P'). n ) and the set of probing postures (P n ) includes a pose displaced from said reference pose (P′) by a rotation about at least one of the rotation axes; The binaural pre-rendered representation (Bin n ) of the reference binaural representation (Bin ref a metadata generator (13) configured to calculate reconstruction metadata (M) enabling reconstruction from the The reference binaural representation (Bin ref ) and the reconstructed metadata (M) are output as an output bitstream (b 2 an encoder (14, 15) configured to encode the The output bitstream (b 2 an interface (16) configured to output the Devices containing:
43. 1. A method for rendering audio on a main device, enabling split rendering with attitude compensation about at least one axis of rotation, comprising: obtaining immersive audio content (A); receiving head pose information (P) associated with a user of a lightweight processing device; determining a reference pose (P') based on the head pose information; The immersive audio content (A) is converted into a reference binaural representation (Bin ref ); The immersive audio content (A) is converted into one or more binaural pre-rendered representations (Bin n ) wherein the binaural pre-rendered representation is rendered at one or more probing poses (P) displaced from the reference pose with respect to the rotation axis. n ) corresponding to the step; The binaural pre-rendered representation (Bin n ) of the reference binaural representation (Bin ref ), the reconstruction metadata comprising, for each time-frequency tile, a transformation matrix (M ^ ) the step of: Each transformation matrix (M ^ calculating augmented metadata (M) by multiplying the augmented metadata (M) by an additional gain matrix (G), said additional gain matrix being: [Equation 3] where: R y,ps,ps is a specific probing posture (P s represents a 2×2 covariance matrix of said binaural pre-rendered representation with respect to R y^,ps,ps is the specific probing posture (P s ) expressing a 2×2 covariance matrix of the reconstructed binaural pre-rendered representation for the The reference binaural representation (Bin ref ) and the enriched metadata (M) are output as an output bitstream (b 2 ) encoding the The output bitstream (b 2 ) outputting the A method comprising:
44. 44. The method of claim 43, wherein the immersive audio content is rendered into at least two pre-rendered representations corresponding to at least two probing poses.
45. 45. The method of claim 43 or 44, wherein the axis of rotation is a yaw axis and / or a roll axis.
46. 44. The method of claim 43, wherein the at least one axis of rotation includes a yaw axis and a roll axis; one set of the probing attitudes includes a yaw probing attitude displaced from the reference attitude by rotation only about the yaw axis, and a low probing attitude displaced from the reference attitude by rotation only about the low axis; The reconstruction metadata: A binaural pre-rendered representation (Bin) corresponding to the yaw probing pose is generated. n ) of the reference binaural representation (Bin ref ) yaw reconstruction metadata that allows reconstruction from Binaural pre-rendered representations (Bin) corresponding to the low probing postures n ) of the reference binaural representation (Bin ref and raw reconstruction metadata that allows reconstruction from the The method, wherein the augmented metadata is calculated based on the yaw reconstruction metadata and / or the ro reconstruction metadata.
47. 10. The method of claim 1, wherein the at least one axis of rotation further comprises a pitch axis; One set of the probing attitudes includes a pitch probing attitude that is displaced from the reference attitude by rotation only about the pitch axis, and the method includes: Binaural pre-rendered representations (Bin) corresponding to the pitch probing postures n ) of the reference binaural representation (Bin ref ), wherein the pitch reconstruction metadata includes a real diagonal 2x2 pitch correction matrix (H) for each time-frequency tile.
48. 48. The method of claim 47, wherein the orientation P 1 The pitch correction matrix H for P1 The elements of are obtained by the following formula: [Equation 4] R y,p’,p’ is the reference binaural representation Bin p’ is the covariance matrix of R y,p1,p1 is posture P 1 The method is a covariance matrix of the binaural pre-rendered representation for
49. 48. The method of claim 47, wherein the orientation P 1 The pitch correction matrix H for P1 The elements of are obtained by the following formula: [Equation 5] R y,p’,p’ is the reference binaural representation Bin p’ is the covariance matrix of R y,p1,p1 is posture P 1 The method is a covariance matrix of the binaural pre-rendered representation for
50. 50. The method according to any one of claims 43 to 49, wherein the attitude information representing the reference attitude (P') is coded and the output bitstream (b 2 ) a method included in
51. 51. The method according to any one of claims 43 to 50, wherein the probing position (P n ) is encoded into the output bitstream (b 2 ) a method included in
52. 52. The method of any one of claims 43 to 51, further comprising the step of quantizing and encoding the reconstructed metadata based on symmetries of the reconstructed metadata.
53. 53. The method of any one of claims 43 to 52, further comprising encoding the reconstructed metadata using differential coding between metadata relating to different probing poses.
54. 54. An apparatus comprising a processor and a memory coupled to the processor, the processor configured to cause the apparatus to perform a method according to any one of claims 43 to 53.
55. 54. A computer readable storage medium storing a program including instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 43 to 53.
56. 1. A method of audio rendering that enables split rendering with attitude compensation for multiple axes of rotation, comprising: obtaining immersive audio content (A); Head pose information (P,Ω) associated with the user of a lightweight processing device P ,ω P ) receiving the signal; Based on the head posture information, the reference posture (P') and the head posture rotation axis (Ω P ) and head posture rotational speed (ω P ) determining at least one of the following: The immersive audio content (A) is converted into a reference binaural representation (Bin ref ); The immersive audio content (A) is converted into a set of binaural pre-rendered representations (Bin n ), wherein the binaural pre-rendered representation is rendered at a set of probing poses (P n ), and the probing posture corresponds to the head posture information (P,Ω P ,ω P ) selected based on the The binaural pre-rendered representation (Bin n ) of the reference binaural representation (Bin ref calculating reconstruction metadata (M) that allows reconstruction from the The reference binaural representation (Bin ref ) and the reconstructed metadata (M) are output as an output bitstream (b 2 ) encoding the The output bitstream (b 2 ) outputting the A method comprising:
57. 57. The method of claim 56, wherein the head pose rotation axis (Ω P ) is determined, and the probing posture (P n ) is the head posture rotation axis (Ω P ) and displaced from the reference orientation (P').
58. 58. The method of claim 57, wherein the probing poses are symmetrically distributed around the reference pose (P').
59. 58. The method of claim 57, wherein the head pose rotation velocity (ω P ) is determined, The head posture rotation speed (ω P ) is below the threshold, the probing posture (P n ) is the first angle (ω lower ) is displaced from the reference attitude by less than The head posture rotation speed (ω P ) exceeds the threshold, the probing posture (P n ) is the second angle (ω upper ) from the reference attitude, and the first angle (ω lower ) is the second angle (ω upper ) smaller, way.
60. 60. The method of any one of claims 56 to 59, wherein the head pose rotation velocity (ω P ) is determined, The head posture rotation speed (ω P ) is below the threshold, the probing posture (P n ) includes only one probing pose for each rotational degree of freedom.
61. 61. The method according to any one of claims 56 to 60, wherein the attitude information representing the reference attitude (P') is coded and the output bitstream (b 2 ) a method included in
62. 62. The method according to any one of claims 56 to 61, wherein the probing position (P n ) is encoded into the output bitstream (b 2 ) a method included in
63. 63. The method of any one of claims 56 to 62, wherein the reconstruction metadata (M) comprises a 2x2 matrix for each time-frequency tile.
64. 64. The method of claim 63, further comprising the step of quantizing and encoding the reconstructed metadata (M) based on symmetries of the reconstructed metadata.
65. 63. The method of claim 61 or 62, further comprising encoding the reconstructed metadata (M) using differential coding between metadata relating to different probing poses.
66. 66. An apparatus comprising a processor and a memory coupled to the processor, the processor configured to cause the apparatus to perform a method according to any one of claims 56 to 65.
67. 66. A computer readable storage medium storing a program including instructions that, when executed by a processor, cause the processor to perform a method according to any one of claims 56 to 65.
68. 1. A method for audio processing with attitude compensation about multiple axes of rotation, comprising: receiving a bitstream (b2) from the main device; The bitstream is decoded to obtain a reference binaural representation (Bin ref ) and first reconstruction metadata, which is a set of probing poses (P') representing displacements from a reference pose (P') due to rotation about the plurality of rotation axes. n ) associated with the Detecting the current head pose (P); For each of said axes of rotation: selecting a probing pose that is closest to the detected pose with respect to the rotation axis; axis-specific reconstruction metadata (M) based on the first reconstruction metadata associated with a selected probing pose and the difference between the current head pose and the reference pose about the rotation axis; α ,M β ,M γ ) determining the The reference binaural representation (Bin ref determining a binaural output corresponding to the current head pose based on the axis-specific reconstruction metadata; A method comprising:
69. 69. The method of claim 68, wherein the set of probing poses is obtained from the bitstream.
70. 69. The method of claim 68, wherein the set of probing poses is obtained by adding a set of offsets to the reference pose.
71. 71. The method of claim 70, wherein the set of offsets is predetermined.
72. 71. The method of claim 70, wherein the set of offsets is obtained from the bitstream.
73. 73. The method of any one of claims 68 to 72, wherein the indication of the reference attitude (P') is obtained from the bitstream.
74. 74. A method according to any one of claims 68 to 73, further comprising transmitting an indication of the current head pose (P) to the main device.
75. 75. The method of claim 74, wherein the reference attitude (P') is estimated based on expected transmission delays.
76. 76. The method of any one of claims 68 to 75, wherein the first reconstruction metadata comprises a 2x2 matrix for each time-frequency tile.
77. 77. The method of claim 76, wherein the axis-specific reconstruction metadata is a 2x2 matrix for each time-frequency tile, and the binaural output for the time-frequency tile is obtained by multiplying a reference binaural representation for the time-frequency tile by the corresponding 2x2 matrix for each axis-specific reconstruction metadata.
78. 78. An apparatus comprising a processor and a memory coupled to the processor, the processor configured to cause the apparatus to perform a method according to any one of claims 68 to 77.
79. 78. A computer readable storage medium storing a program including instructions that, when executed by a processor, cause the processor to perform a method according to any one of claims 68 to 77.
80. A lightweight processing device (3) comprising: Bitstream (b 2 ) and decode it into a reference binaural representation (Bin ref ) and first reconstruction metadata, which is a set of probing poses (P') representing displacements from the reference pose (P') due to rotations about multiple rotation axes. n a decoder configured to obtain a signal associated with the a head tracker (25) configured to detect a current head pose (P); binaural reconstruction block (26); wherein the binaural reconstruction block (26) comprises: For each of the rotation axes, a probing pose that is closest to the detected pose for the rotation axis is selected, and axis-specific reconstruction metadata (M) is calculated based on the first reconstruction metadata associated with the selected probing pose and the difference between the current head pose for the rotation axis and the reference pose. α ,M β ,M γ ) determining the The reference binaural representation (Bin ref determining a binaural output corresponding to the current head pose based on the axis-specific reconstruction metadata; A device that is configured to: