Method, apparatus and system for three degrees of freedom (3DOF+) extension of MPEG-H 3D audio

The method and apparatus address the limitation of MPEG-H 3D Audio by processing audio object positions to account for both translational and rotational head movements, enhancing the immersive experience by allowing listeners to perceive audio objects from various angles, thus extending the standard to 6DoF.

JP2026035666APending Publication Date: 2026-03-04DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing methods and apparatus for processing audio object positions and information fail to handle the small translational movements of a user's head in a three-degrees-of-freedom (3DoF) environment, limiting the immersive experience.

Method used

A method and apparatus for processing audio object position information that accounts for small translational movements of a user's head, in conjunction with rotational movements, by modifying object positions based on listener displacement and orientation information, ensuring audio objects are perceived from a fixed position relative to the listener's head, thus enhancing the immersive experience.

Benefits of technology

Enables a more realistic and immersive listening experience by allowing listeners to approach audio objects from different angles, even from the side, by correcting audio object positions for head displacements and orientations, thereby providing a limited six degrees of freedom (6DoF) experience within the MPEG-H 3D Audio standard.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035666000001_ABST
    Figure 2026035666000001_ABST
Patent Text Reader

Abstract

A method and apparatus for processing audio object position information that takes into account small translational movements of a user's head is provided. A method for making object positions usable for rendering audio objects includes obtaining listener orientation information indicative of a listener's head orientation, obtaining listener displacement information indicative of a listener's head displacement, determining an object position from the position information, modifying the object position based on the listener displacement information by applying a translation to the object position, and further modifying the modified object position based on the listener orientation information. In a corresponding device for processing position information indicative of the object positions of audio objects, the object positions are usable for rendering the audio objects.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: U.S. Provisional Application No. 62 / 654,915 (Reference No. D18045USP1), filed April 9, 2018; U.S. Provisional Application No. 62 / 695,446 (Reference No. D18045USP2), filed July 9, 2018; and U.S. Provisional Application No. 62 / 823,159 (Reference No. D18045USP3), filed March 25, 2019. These applications are incorporated herein by reference.

[0002] Technical Field The present disclosure relates to methods and apparatus for processing position information indicative of audio object positions and information indicative of positional displacement of a listener's head. [Background technology]

[0003] The ISO / IEC23008-3 MPEG-H 3D Audio standard, version 1 (October 15, 2015) and amendments 1 to 4, do not provide for allowing for small translational movements of the user's head in a three-degrees-of-freedom (3DoF) environment. Summary of the Invention [Means for solving the problem]

[0004] The ISO / IEC 23008-3 MPEG-H 3D Audio standard, version 1 (October 15, 2015) and amendments 1 to 4, provide functionality for the possibility of a 3DoF environment in which the user (listener) performs head rotational movements. However, such functionality only supports, at best, the signaling and corresponding rendering of rotational scene displacements. This means that the audio scene can remain spatially stationary under changes in the listener's head orientation that correspond to 3DoF attributes. However, within the current MPEG-H 3D Audio ecosystem, there is no possibility to take into account small translational movements of the user's head.

[0005] Thus, there is a need for a method and apparatus for processing audio object position information that can take into account small translational movements of a user's head, potentially in conjunction with rotational movements of the user's head.

[0006] The present disclosure provides an apparatus and a system for processing location information having the features of the respective independent and dependent claims.

[0007] According to an aspect of the present disclosure, a method for processing position information indicating a position of an audio object is described. The processing may be compliant with the MPEG-H 3D Audio standard. The object position may be usable for rendering the audio object. The audio object may be included in object-based audio content along with its position information. The position information may be (part of) metadata for the audio object. The audio content (e.g., the audio object with its position information) may be conveyed in an encoded audio bitstream. The method may include receiving the audio content (e.g., the encoded audio bitstream). The method may include obtaining listener orientation information indicating a listener's head orientation. The listener may be referred to, for example, as a user of an audio decoder that executes the method. The listener's head orientation (listener orientation) may be the listener's head orientation relative to a nominal orientation. The method may further include obtaining listener displacement information indicating a listener's head displacement. The listener's head displacement may be a displacement relative to a nominal listening position. The nominal listening position (or nominal listener position) may be a default position (e.g., a predetermined position, an expected position for the listener's head, or a sweet spot of a speaker arrangement). The listener orientation information and listener displacement information may be obtained via an MPEG-H 3D Audio decoder input interface. The listener orientation information and listener displacement information may be derived based on sensor information. The combination of the orientation information and the position information may be referred to as posture information. The method may further include determining an object position from the position information. For example, the object position may be extracted from the position information. The determination (e.g., extraction) of the object position may further be based on information about the geometry of a speaker arrangement of one or more speakers in the listening environment. The object position may also be referred to as a channel position of an audio object.The method may further include modifying the object positions based on the listener displacement information by applying a translation to the object positions. Modifying the object positions may involve correcting the object positions for a displacement of the listener's head from a nominal listening position. In other words, modifying the object positions may involve applying a position displacement correction to the object positions. The method may further include further modifying the modified object positions based on the listener orientation information, for example, by applying a rotational transformation to the modified object positions (e.g., a rotation about the listener's head or the nominal listening position). Further modifying the modified object positions to render audio objects may involve a rotational audio scene displacement.

[0008] Configured as described above, the proposed method provides a more realistic listening experience, especially for audio objects located near the listener's head. In addition to the three (rotational) degrees of freedom typically provided to a listener in a 3DoF environment, the proposed method can also take into account the translational movement of the listener's head. This allows the listener to approach a nearby audio object from different angles, even from the side. For example, by slightly moving their head in addition to rotating it, the listener can hear a "mosquito" audio object close to the listener's head from various angles. As a result, the proposed method can enable an improved, more realistic, and immersive listening experience for the listener.

[0009] In some embodiments, modifying the object position and further modifying the modified object position may be performed such that the audio object, after being rendered onto one or more real or virtual speakers according to the modified object position, is psychoacoustically perceived by a listener as emanating from a fixed position relative to a nominal listening position, regardless of the listener's head displacement from the nominal listening position or the listener's head orientation relative to the nominal orientation. Thus, the audio object may be perceived to move relative to the listener's head when the listener's head undergoes a displacement from the nominal listening position. Similarly, the audio object may be perceived to rotate relative to the listener's head when the listener's head undergoes an orientation change from the nominal orientation. The one or more speakers may, for example, be part of a headset or part of a speaker arrangement (e.g., a 2.1, 5.1, 7.1, etc. speaker arrangement).

[0010] In some embodiments, modifying the object position based on the listener displacement information may be performed by translating the object position by a vector that is positively correlated to the absolute value and negatively correlated to the direction of the listener's head displacement vector from the nominal listening position.

[0011] This ensures that nearby audio objects are perceived by the listener to move in unison with the listener's head movements, contributing to a more realistic listening experience for those audio objects.

[0012] In some embodiments, the listener displacement information may indicate the displacement of the listener's head from the nominal listening position due to small positional displacements. For example, the absolute value of the displacement may be 0.5 m or less. The displacement may be expressed in Cartesian coordinates (e.g., x, y, z) or spherical coordinates (e.g., azimuth, elevation, radius).

[0013] In some embodiments, the listener displacement information may indicate a displacement of the listener's head from a nominal listening position that can be achieved by the listener moving their upper body and / or head. Thus, the displacement may be achievable for the listener without moving their lower body. For example, the displacement of the listener's head may be achievable when the listener is sitting in a chair.

[0014] In some embodiments, the position information may include an indication of the distance of the audio object from the nominal listening position. The distance (radial) may be less than 0.5 m. For example, the distance may be less than 1 cm. Alternatively, the distance of the audio object from the nominal listening position may be set to a default value by the decoder.

[0015] In some embodiments, the listener orientation information may include information regarding the yaw, pitch, and roll of the listener's head, which may be given relative to a nominal orientation (e.g., a reference orientation) of the listener's head.

[0016] In some embodiments, the listener displacement information may include information about the listener's head displacement from a nominal listening position expressed in Cartesian or spherical coordinates, where the displacement may be expressed in x, y, z coordinates for Cartesian coordinates, or in azimuth, elevation, and radial coordinates for spherical coordinates.

[0017] In some embodiments, the method may further include detecting the listener's head orientation by wearable and / or static equipment. Similarly, the method may further include detecting the listener's head displacement from a nominal listening position by wearable and / or static equipment. The wearable equipment may be, correspond to, and / or include, for example, a headset or an augmented reality (AR) / virtual reality (VR) headset. The static equipment may be, correspond to, and / or include, for example, a camera sensor. This allows for accurate information regarding the listener's head displacement and / or orientation, thereby enabling realistic processing of nearby audio objects depending on the orientation and / or displacement.

[0018] In some embodiments, the method may further include rendering the audio object to one or more real or virtual speakers according to the modified object position, for example, the audio object may be rendered to the left or right speaker of a headset.

[0019] In some embodiments, rendering may be performed based on head-related transfer functions (HRTFs) for the listener's head to take into account sonic occlusion for audio objects at small distances from the listener's head, so that rendering of nearby audio objects is perceived as much more realistic by the listener.

[0020] In some embodiments, the modified object positions may further be adjusted to the input format used by an MPEG-H 3D Audio renderer. In some embodiments, the rendering may be performed using an MPEG-H 3D Audio renderer. In some embodiments, the processing may be performed using an MPEG-H 3D Audio decoder. In some embodiments, the processing may be performed by a scene displacement unit of an MPEG-H 3D Audio decoder. Thus, the proposed method allows implementing a limited six degrees of freedom (6DoF) experience (i.e., 3DoF+) within the framework of the MPEG-H 3D Audio standard.

[0021] According to another aspect of the present disclosure, a further method is described for processing position information indicative of an object position of an audio object. The object position may be usable for rendering the audio object. The method may include obtaining listener displacement information indicative of a displacement of a listener's head. The method may further include determining the object position from the position information. The method may further include modifying the object position based on the listener displacement information by applying a translation to the object position.

[0022] Configured as described above, the proposed method provides a more realistic listening experience, especially for audio objects located near the listener's head. By taking into account small translational movements of the listener's head, the proposed method allows the listener to approach nearby audio objects from different angles, even from the side. As a result, the proposed method can enable an improved, more realistic, and immersive listening experience for the listener.

[0023] In some embodiments, modifying the object position based on the listener displacement information may be performed such that the audio object, after being rendered to one or more real or virtual loudspeakers according to the modified object position, is psychoacoustically perceived by the listener as emanating from a fixed position relative to the nominal listening position, regardless of the displacement of the listener's head from the nominal listening position.

[0024] In some embodiments, modifying the object position based on the listener displacement information may be performed by translating the object position by a vector that is positively correlated to the absolute value and negatively correlated to the direction of the listener's head displacement vector from the nominal listening position.

[0025] According to another aspect of the present disclosure, a further method of processing position information indicative of an object position of an audio object is described. The object position may be usable for rendering the audio object. The method may include obtaining listener orientation information indicative of an orientation of a listener's head. The method may further include determining an object position from the position information. The method may further include modifying the object position based on the listener orientation information, for example, by applying a rotational transformation to the object position (e.g., a rotation about the listener's head or a nominal listening position).

[0026] Configured as described above, the proposed method is able to take into account the listener's head orientation and provide a more realistic listening experience to the listener.

[0027] In some embodiments, modifying the object position based on the listener orientation information may be performed such that the audio object, after being rendered to one or more real or virtual loudspeakers according to the modified object position, is psychoacoustically perceived by the listener as emanating from a fixed position relative to the nominal listening position, regardless of the orientation of the listener's head with respect to the nominal orientation.

[0028] According to another aspect of the present disclosure, an apparatus is described that processes position information indicating an object position of an audio object. The object position may be usable for rendering the audio object. The apparatus may include a processor and a memory coupled to the processor. The processor may be adapted to obtain listener orientation information indicating an orientation of a listener's head. The processor may further be adapted to obtain listener displacement information indicating a displacement of the listener's head. The processor may further be adapted to determine the object position from the position information. The processor may further be adapted to modify the object position based on the listener displacement information by applying a translation to the object position. The processor may further be adapted to further modify the modified object position based on the listener orientation information, for example, by applying a rotational transformation to the modified object position (e.g., a rotation about the listener's head or a nominal listening position).

[0029] In some embodiments, the processor may be adapted to modify the object position and further modify the modified object position such that the audio object, after being rendered to one or more real or virtual loudspeakers according to the further modified object position, is psychoacoustically perceived by a listener as emanating from a fixed position relative to the nominal listening position, regardless of the displacement of the listener's head from the nominal listening position or the orientation of the listener's head with respect to the nominal orientation.

[0030] In some embodiments, the processor may be adapted to modify the object position based on the listener displacement information by translating the object position by a vector that is positively correlated to the absolute value and negatively correlated to the direction of the listener's head displacement vector from the nominal listening position.

[0031] In some embodiments, the listener displacement information may indicate the displacement of the listener's head from the nominal listening position due to small positional variations.

[0032] In some embodiments, the listener displacement information may indicate the displacement of the listener's head from the nominal listening position that can be achieved by the listener moving their upper body and / or head.

[0033] In some embodiments, the position information may include an indication of the distance of the audio object from the nominal listening position.

[0034] In some embodiments, the listener orientation information may include information regarding the yaw, pitch, and roll of the listener's head.

[0035] In some embodiments, the listener displacement information may include information regarding the listener's head displacement from a nominal listening position expressed in Cartesian or spherical coordinates.

[0036] In some embodiments, the device may further include wearable and / or static equipment for detecting the listener's head orientation. In some embodiments, the device may further include wearable and / or static equipment for detecting the listener's head displacement from a nominal listening position.

[0037] In some embodiments, the processor may be further adapted to render the audio objects to one or more real or virtual loudspeakers according to the further modified object positions.

[0038] In some embodiments, the processor may be adapted to perform said rendering based on HRTFs for the listener's head and taking into account sonic occlusion for small distances of audio objects from the listener's head.

[0039] In some embodiments, the processor may be further adapted to adjust the modified object positions to an input format used by an MPEG-H 3D Audio renderer. In some embodiments, the rendering may be performed using an MPEG-H 3D Audio renderer. In some embodiments, the processor may be adapted to implement an MPEG-H 3D Audio decoder. In some embodiments, the processor may be adapted to implement a scene displacement unit of an MPEG-H 3D Audio decoder.

[0040] According to another aspect of the present disclosure, a further apparatus is described for processing position information indicative of an object position of an audio object. The object position may be usable for rendering the audio object. The apparatus may include a processor and a memory coupled to the processor. The processor may be adapted to obtain listener displacement information indicative of a displacement of a listener's head. The processor may be further adapted to determine the object position from the position information. The processor may be further adapted to modify the object position based on the listener displacement information by applying a translation to the object position.

[0041] In some embodiments, the processor may be adapted to modify the object positions based on the listener displacement information such that the audio objects, after being rendered to one or more real or virtual loudspeakers according to the modified object positions, are psychoacoustically perceived by the listener as emanating from a fixed position relative to the nominal listening position, regardless of the displacement of the listener's head from the nominal listening position.

[0042] In some embodiments, the processor may be adapted to modify the object position based on the listener displacement information by translating the object position by a vector that is positively correlated to the absolute value and negatively correlated to the direction of the listener's head displacement vector from the nominal listening position.

[0043] According to another aspect of the present disclosure, a further apparatus is described for processing position information indicative of an object position of an audio object. The object position may be usable for rendering the audio object. The apparatus may include a processor and a memory coupled to the processor. The processor may be adapted to obtain listener orientation information indicative of an orientation of a listener's head. The processor may be further adapted to determine the object position from the position information. The processor may be further adapted to modify the object position based on the listener orientation information, for example, by applying a rotational transformation to the modified object position (e.g., a rotation about the listener's head or a nominal listening position).

[0044] In some embodiments, the processor may be adapted to modify the object positions based on the listener orientation information such that the audio objects, after being rendered to one or more real or virtual loudspeakers according to the modified object positions, are psychoacoustically perceived by the listener as emanating from a fixed position relative to a nominal listening position, regardless of the orientation of the listener's head with respect to the nominal orientation.

[0045] According to yet another aspect, a system is described that may include an apparatus according to any of the above aspects and wearable and / or static equipment capable of detecting a listener's head orientation and detecting a listener's head displacement.

[0046] It will be understood that method steps and apparatus features may be interchanged in many ways. In particular, details of the disclosed methods can be implemented as an apparatus adapted to perform part or all of the method or steps, and vice versa, as will be understood by those skilled in the art. In particular, an apparatus according to the present disclosure may relate to an apparatus for realizing or performing the method according to the above embodiments and variations thereof, and it will be understood that each statement made with respect to the method also applies to the corresponding apparatus. Similarly, a method according to the present disclosure may relate to a method of operation of the apparatus according to the above embodiments and variations thereof, and it will be understood that each statement made with respect to the apparatus also applies to the corresponding method. [Brief explanation of the drawings]

[0047] The invention is described below, by way of example, with reference to the accompanying drawings, in which: [Figure 1] 1 illustrates a schematic diagram of an example MPEG-H 3D Audio system. [Figure 2] 1 illustrates schematically an example of an MPEG-H 3D Audio system according to the present invention; [Figure 3]1 illustrates schematically an example of an audio rendering system according to the present invention; [Figure 4] 1 illustrates schematically an exemplary set of Cartesian coordinate axes and their relationship to spherical coordinates. [Figure 5] 3 is a flow chart that schematically illustrates an example of a method for processing position information for audio objects according to the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0048] As used herein, 3DoF refers to a system that can correctly handle a user's head movements, especially head rotation, typically specified by three parameters (e.g., yaw, pitch, and roll). Such systems are often available in various gaming systems, such as virtual reality (VR) / augmented reality (AR) / mixed reality (MR) systems, or in other acoustic environments of that type.

[0049] As used herein, a user (e.g., a user of an audio decoder or a playback system that includes an audio decoder) is also referred to as a "listener."

[0050] As used herein, 3DoF+ means that in addition to the user's head movements that can be properly handled by a 3DoF system, small translational movements can also be handled.

[0051] As used herein, "small" indicates that movement is limited to less than a threshold, typically 0.5 meters. This means that movement is 0.5 meters or less from the user's original head position. For example, the user's movement may be constrained by the user being seated in a chair.

[0052] As used herein, "MPEG-H 3D Audio" refers to the specifications standardized in ISO / IEC 23008-3 and / or any future amendments, editions or other versions of the ISO / IEC 23008-3 standard.

[0053] In the context of the audio standards provided by the MPEG organization, the distinction between 3DoF and 3DoF+ can be defined as follows: · 3DoF: allows the user to experience yaw, pitch and roll movements (for example, of the user's head); 3DoF+: The user can experience yaw, pitch, and roll movements (for example, of the user's head) and limited translational movement, such as when sitting in a chair.

[0054] Limited (small) head translational movement may be movement that is constrained to a certain radius of movement. For example, the movement may be constrained because the user is in a seated position, e.g., not using the lower body. Small head translational movement may be related to or correspond to the displacement of the user's head relative to a nominal listening position. The nominal listening position (or nominal listener position) may be a default position (e.g., a predetermined position, an expected position of the listener's head, or a sweet spot of a speaker placement, etc.).

[0055] A 3DoF+ experience can be comparable to a constrained 6DoF experience, where translational motion can be described as limited or small head movements. In one example, audio is also rendered based on the position and orientation of the user's head, including possible sound masking. Rendering may be performed to take into account sound masking for small distances of audio objects from the listener's head, for example, based on head-related transfer functions (HRTFs) for the listener's head.

[0056] For methods, systems, apparatus and other devices that conform to the functionality specified by the MPEG-H 3D Audio standard, this means that 3DoF+ will be enabled in any future versions of the MPEG standard, for example, for future versions of Omnidirectional Media Formats (e.g., as standardized in future versions of MPEG-I), and / or in any updates to MPEG-H Audio (e.g., amendments or newer standards based on the MPEG-H 3D Audio standard), or any other related or supporting standards that may require updates (e.g., standards that specify certain types of metadata and SEI messages).

[0057] For example, the audio renderer that is mandatory for the audio standard specified in the MPEG-H 3D Audio standard may be extended to include rendering of the audio scene, in order to accurately take into account the user's interaction with the audio scene, for example when the user moves their head slightly sideways.

[0058] The present invention provides various technical advantages, including the advantage of providing MPEG-H 3D Audio that can handle 3DoF+ use cases. The present invention extends the MPEG-H 3D Audio standard to support 3DoF+ functionality.

[0059] To support 3DoF+ functionality, the audio rendering system should take into account limited / small positional displacement of the user / listener's head. The positional displacement should be determined based on a relative offset from an initial position (i.e., default position / nominal listening position). In one example, the magnitude of this offset (e.g., P0 is the nominal listening position and P1 is the displaced position of the listener's head) can be calculated as follows: offset= ||P0-P1||) is at most about 0.5 m. In another example, the magnitude of the offset is limited to the offset that can be achieved while the user is sitting in a chair and performing no lower body movements (but the head is moving relative to the body). This (small) offset distance makes little (perceptual) difference in level and pan for distant audio objects. However, for close objects, even such small offset distances can be perceptually significant. In fact, the listener's head movements can have a perceptual effect on the perception of where the correct audio object localization is. This perceptual effect can be caused by (i) the displacement of the user's head (e.g., r offset = ||P0-P1||) and the distance to the audio object (e.g., r) remains significant as long as the ratio between the angle (e.g., r) trigonometrically results in an angle that is within the range of the user's psychoacoustic ability to detect the sound direction (i.e., perceptually detectable by the user / listener). Such range may vary for different audio renderer settings, audio material, and playback configurations. For example, if the localization accuracy range is, say, ±3° and the listener's head has a freedom of lateral movement of ±0.25 m, this corresponds to an object distance of approximately 5 m.

[0060] For objects close to the listener (e.g., objects less than 1 m from the user), it is important for 3DoF+ scenarios to properly handle positional displacements of the listener's head, as this has a large perceptual effect during both panning and level changes.

[0061] An example of processing an object close to the listener is when an audio object (e.g., a mosquito) is positioned very close to the listener's face. An audio system, such as one that provides VR / AR / MR functionality, should allow the user to perceive this audio object from all sides and angles, even while the user makes small translational head movements. For example, the user should be able to accurately perceive the object (e.g., a mosquito) even while moving their head without moving their lower body.

[0062] However, systems compatible with the current MPEG-H 3D Audio specification cannot handle this correctly. Instead, when using a system compatible with MPEG-H 3D Audio, the "mosquito" is perceived from the wrong position relative to the user. In scenarios involving 3DoF+ performance, small translational movements should result in a significant difference in the perception of the audio object (for example, moving the head to the left should cause the "mosquito" audio object to be perceived from the right side of the user's head).

[0063] The MPEG-H 3D Audio standard includes a bitstream syntax that allows signaling of object distance information (starting from 0.5m) via the bitstream syntax, for example via the object_metadata() syntax element.

[0064] The syntax element prodMetadataConfig() may be introduced into the bitstream provided by the MPEG-H 3D Audio standard. It can be used to signal that the object distance is very close to the listener. For example, the syntax prodMetadataConfig() may signal that the distance between the user and the object is less than a certain threshold distance (e.g., <1 cm).

[0065] 1 and 2 illustrate the present invention based on headphone rendering (i.e., the loudspeakers move with the listener's head).

[0066] FIG. 1 shows an example of system behavior 100 compliant with the MPEG-H 3D Audio system. This example assumes that a listener's head is at position P0 103 at time t0 and moves to position P1 104 at time t1>t0. The dashed circle around positions P0 and P1 indicates the allowable 3DoF+ movement region (e.g., radius 0.5 m). Position A 101 indicates the signal object position (position at time t0 and time t1; i.e., the signaled object position is assumed to be constant over time). Position A also indicates the object position rendered by the MPEG-H 3D Audio renderer at time t0. Position B 102 indicates the object position rendered by MPEG-H 3D Audio at time t1. The vertical lines extending upward from positions P0 and P1 indicate the respective orientations (e.g., looking direction) of the listener's head at time t0 and time t1. The displacement of the user's head between positions P0 and P1 is r offset =||P0-P1|| 106. If the listener is located at the default position (nominal listening position) P0 103 at time t0, the listener will perceive the audio object (e.g., a mosquito) at the correct position A 101. If the user moves to position P1 104 at time t1, and MPEG-H 3D Audio processing is applied as currently standardized, the error δ AB 105, the user will perceive the audio object at position B 102. That is, despite the listener's head movement, the audio object (e.g., a mosquito) will still be perceived as being located directly in front of the listener's head (i.e., substantially co-moving with the listener's head). Notably, the introduced error δ AB 105 occurs regardless of the listener's head orientation.

[0067] FIG. 2 illustrates an example of system behavior for an MPEG-H 3D Audio system 200 in accordance with the present invention. In FIG. 2, the listener's head is located at position P0 203 at time t0 and moves to position P1 204 at time t1>t0. The dashed circle around positions P0 and P1 again indicates the allowable 3DoF+ movement region (e.g., radius 0.5 m). In 201, position A=B is shown, i.e., the signaled object position (position at times t0 and t1; i.e., the signaled object position is assumed to be constant over time). Position A=B 201 also indicates the position of the object rendered by MPEG-H 3D Audio at times t0 and t1. The vertical arrows extending upward from positions P0 203 and P1 204 indicate the respective orientations (e.g., looking direction) of the listener's head at times t0 and t1. At time t0, the listener is located at an initial / default position (nominal listening position) P0 203, so the listener perceives the audio object (e.g., a mosquito) at the correct position A 201. If the user moves to position P1 203 at time t1, under the present invention, the user will still perceive the audio object at position B 201, which is similar (e.g., substantially equal) to position A 201. In this way, the present invention allows the user's position to change over time (e.g., from position P0 203 to position P1 204) while still perceiving sound from the same (spatially fixed) position (e.g., position A = B 201, etc.). In other words, the audio object (e.g., a mosquito) moves relative to the listener's head in response to (e.g., negatively correlated with) the listener's head movement. This allows the user to move around the audio object (e.g., a mosquito) and perceive the audio object from different angles or sides. The displacement of the user's head between positions P0 and P1 is r offset =||P0-P1|| 206.

[0068] FIG. 3 illustrates an example of an audio rendering system 300 according to the present invention. The audio rendering system 300 may correspond to or include a decoder, such as an MPEG-H 3D Audio decoder. The audio rendering system 300 may include an audio scene displacement unit 310 having a corresponding audio scene displacement processing interface (e.g., an interface for scene displacement data according to the MPEG-H 3D Audio standard). The audio scene displacement unit 310 may output object positions 321 for rendering respective audio objects. For example, the scene displacement unit may output object position metadata for rendering respective audio objects.

[0069] The audio rendering system 300 may further include an audio object renderer 320. For example, the renderer may consist of hardware, software, and / or any partial or complete process running via cloud computing, including various services on the Internet, often referred to as the "cloud," such as software development platforms, servers, storage, and software, that are compatible with the specifications defined by the MPEG-H 3D Audio standard. The audio object renderer 320 may render audio objects to one or more speakers (real or virtual) according to their respective object positions (which may be modified or further modified object positions described below). The audio object renderer 320 may also render audio objects to headphones and / or loudspeakers. That is, the audio object renderer 320 may generate object waveforms according to a given playback format. To this end, the audio object renderer 320 may utilize compressed object metadata. Each object may be rendered to a certain output channel according to its object position (e.g., modified object position or further modified object position). The object positions may therefore also be referred to as the channel positions of those audio objects. The audio object positions 321 may be included in the object position metadata or scene displacement metadata output by the scene displacement unit 310.

[0070] The processing of the present invention may conform to the MPEG-H 3D Audio standard. Accordingly, the processing may be performed by an MPEG-H 3D Audio decoder, or more specifically, an MPEG-H scene displacement unit and / or an MPEG-H 3D Audio renderer. Accordingly, the audio rendering system 300 of FIG. 3 may correspond to or include an MPEG-H 3D Audio decoder (i.e., a decoder conforming to the specifications defined by the MPEG-H 3D Audio standard). In one example, the audio rendering system 300 may be a device having a processor and a memory coupled to the processor, where the processor is adapted to implement an MPEG-H 3D Audio decoder. In particular, the processor may be adapted to implement an MPEG-H scene displacement unit and / or an MPEG-H 3D Audio renderer. Accordingly, the processor may be adapted to perform the processing steps described in this disclosure (e.g., steps S510-S560 of method 500 described below with reference to FIG. 5). In another example, the processing or audio rendering system 300 may be executed in the cloud.

[0071] The audio rendering system 300 may acquire (e.g., receive) listening position data 301. The audio rendering system 300 may obtain the listening position data 301 via an MPEG-H 3D Audio decoder input interface. The listening position data 301 may indicate the orientation and / or position (e.g., displacement) of a listener's head. Thus, the listening position data 301 (also referred to as posture information) may include listener orientation information and / or listener displacement information.

[0072] The listener displacement information may indicate the displacement of the listener's head (e.g., displacement from a nominal listening position). The listener displacement information may be expressed as the absolute value r of the listener's head displacement from a nominal listening position, as shown in FIG. offset=||P0-P1|| 206 or may include an indication thereof. In the context of the present invention, listener displacement information indicates a small positional displacement of the listener's head from the nominal listening position. For example, the absolute value of the displacement may be 0.5 m or less. Typically, this is a displacement of the listener's head from the nominal listening position that is achievable by the listener moving their upper body and / or head. That is, the displacement may be achievable for the listener without moving their lower body. For example, the displacement of the listener's head may be achievable when the listener is sitting in a chair, as described above. The displacement may be expressed in various coordinate systems, such as, for example, Cartesian coordinates (e.g., in terms of x, y, and z) or spherical coordinates (e.g., in terms of azimuth, elevation, and radius). It should be understood that alternative coordinate systems for expressing the listener's head displacement are also feasible and are encompassed by the present disclosure.

[0073] The listener orientation information may indicate the orientation of the listener's head (e.g., the orientation of the listener's head with respect to a nominal / reference orientation of the listener's head). For example, the listener orientation information may include information regarding the yaw, pitch, and roll of the listener's head, where the yaw, pitch, and roll may be given with respect to the nominal orientation.

[0074] The listening position data 301 may be continuously collected from a receiver, which may provide information about the translational movement of the user. For example, the listening position data 301 used at a given time may be the most recently collected data from the receiver. The listening position data may be derived / collected / generated based on sensor information. For example, the listening position data 301 may be derived / collected / generated by wearable and / or static equipment having appropriate sensors. That is, the listener's head orientation may be detected by the wearable and / or static equipment. Similarly, the listener's head displacement (e.g., displacement from a nominal listening position) may be detected by the wearable and / or static equipment. The wearable equipment may be, correspond to, and / or include, for example, a headset (e.g., an AR / VR headset). The static equipment may be, correspond to, and / or include, for example, a camera sensor. The static equipment may be included in, for example, a TV set or a set-top box. In some embodiments, the listening position data 301 may be received from an audio encoder (e.g., an MPEG-H 3D Audio compliant encoder), which may have acquired (e.g., received) the sensor information.

[0075] In one example, wearable and / or static equipment for detecting listening position data 301 may be referred to as a tracking device that supports head position estimation / detection and / or head orientation estimation / detection. There are various solutions that enable accurate tracking of a user's head movements using a computer or smartphone camera (e.g., "FaceTrackNoIR," "opentrack," which are based on face recognition and tracking). Also, some head-mounted display (HMD) virtual reality systems (e.g., HTC VIVE, Oculus Rift) have integrated head tracking technology. Any of these solutions may be used in the context of the present disclosure.

[0076] It is also important to note that head displacement distances in the physical world do not necessarily correspond one-to-one to the displacements indicated by listening position data 301. To achieve hyper-realistic effects (e.g., an over-amplified user motion parallax effect), certain applications may use different sensor calibration settings or specify different mappings between movements in real and virtual spaces. Thus, in some use cases, small physical movements can be expected to result in larger displacements in virtual reality. In any case, the absolute values ​​of the displacements in the physical world and virtual reality (i.e., the displacements indicated by listening position data 301) can be said to be positively correlated. Similarly, the directions of the displacements in the physical world and virtual reality are also positively correlated.

[0077] The audio rendering system 300 may further receive (object) position information (e.g., object position data) 302 and audio data 322. The audio data 322 may include one or more audio objects. The position information 302 may be part of the metadata of the audio data 322. The position information 302 may indicate the object position of each of the one or more audio objects. For example, the position information 302 may include an indication of the distance of each audio object relative to a nominal listening position of a user / listener. The distance (radial) may be less than 0.5 m. For example, the distance may be less than 1 cm. If the position information 302 does not include an indication of the distance of a given audio object from the nominal listening position, the audio rendering system may set the distance of this audio object from the nominal listening position to a default value (e.g., 1 m). The position information 302 may further include an indication of the elevation and / or azimuth angle of each audio object.

[0078] Each object position may be usable to render a corresponding audio object. Thus, the position information 302 and audio data 322 may be included in or form object-based audio content. The audio content (e.g., audio objects / audio data 322 and their position information 302) may be conveyed in an encoded audio bitstream. For example, the audio content may be in the form of a bitstream received from transmission over a network. In this case, the audio rendering system may be said to receive the audio content (e.g., from an encoded audio bitstream).

[0079] In one example of the present invention, metadata parameters may be used to correct the processing of certain use cases of backward-compatible enhancements for 3DoF and 3DoF+. The metadata may include listener displacement information in addition to listener orientation information. Such metadata parameters may be utilized by the systems shown in Figures 2 and 3, as well as any other embodiments of the present invention.

[0080] Backward-compatible enhancements may allow for correcting use case processing (e.g., implementations of the present invention) based on the canonical MPEG-H 3D Audio scene displacement interface. This means that legacy MPEG-H 3D Audio decoders / renderers will still generate output, even if incorrect. However, an enhanced MPEG-H 3D Audio decoder / renderer according to the present invention will correctly apply the extension data (e.g., extended metadata) and processing, and thus be able to handle scenarios with objects located close to the listener in a correct manner.

[0081] In one example, the present invention relates to providing data for small translational movements of a user's head in a format different from that outlined below, and these formulas can be adapted accordingly. For example, the data may be provided in a format such as x, y, z coordinates (in a Cartesian coordinate system) instead of azimuth, elevation, and radius (in a spherical coordinate system). An example of the relationship of these coordinate systems to each other is shown in Figure 4.

[0082] In one example, the present invention is directed to providing metadata for inputting translational movement of a listener's head (e.g., listener displacement information included in listening position data 301 shown in FIG. 3). The metadata may be used, for example, for interfacing with scene displacement data. The metadata (e.g., listener displacement information) may be obtained by deployment of a tracking device that supports 3DoF+ or 6DoF tracking.

[0083] In one example, metadata (e.g., listener displacement information, in particular listener head displacement, or equivalently, scene displacement) may be represented by three parameters sd_azimuth, sd_elevation, and sd_radius, which relate to the azimuth, elevation, and radius (spherical coordinates) of the listener's head displacement (or scene displacement).

[0084] The syntax for these parameters is given by the following table: [Table 1]

[0085] In another example, the metadata (e.g., listener displacement information) may be represented by the following three parameters sd_x, sd_y, and sd_z in Cartesian coordinates, which reduces the processing of the data from spherical coordinates to Cartesian coordinates. The metadata may be based on the following syntax: [Table 2] As mentioned above, the above syntax or its equivalents may signal information related to rotation around the x, y, and z axes.

[0086] In one example of the present invention, the processing of scene displacement angles for channels and objects may be improved by extending the formula to take into account changes in the position of the user's head, i.e., the processing of object positions may take into account listener displacement information (e.g., may be based at least in part on listener shift information).

[0087] An example of a method 500 for processing position information indicating object positions of audio objects is shown in the flowchart of Figure 5. This method may be performed by a decoder, such as an MPEG-H 3D audio decoder. The audio rendering system 300 of Figure 3 may be an example of such a decoder.

[0088] As a first step (not shown in FIG. 5), audio content including audio objects and corresponding position information is received, for example, from an encoded audio bitstream. The method may then further include decoding the encoded audio content to obtain the audio objects and position information.

[0089] Step S510 In, listener orientation information is obtained (e.g., received). The listener orientation information may indicate an orientation of the listener's head.

[0090] Step S520 In, listener displacement information is obtained (e.g., received). The listener displacement information may indicate a displacement of a listener's head.

[0091] Step S530In the audio decoder, the object position is determined from the position information. For example, the object position (e.g., azimuth, elevation, radius, or x, y, z, or their equivalents) may be extracted from the position information. The determination of the object position may also be based, at least in part, on information about the speaker geometry of one or more (real or virtual) speakers in the listening environment. If the radius is not included in the position information for that audio object, the decoder may set the radius to a default value (e.g., 1 m).

[0092] In some embodiments, the default value may depend on the geometry of the speaker arrangement.

[0093] In particular, steps S510, S520, and S520 can be performed in any order.

[0094] Step S540 In step S530, the object position determined in step S530 is modified based on the listener displacement information. This may be done by applying a translation to the object position according to the displacement information (e.g., in response to the listener's head displacement). Therefore, modifying the object position can be said to involve correcting the object position for the listener's head displacement (e.g., displacement from a nominal listening position). In particular, modifying the object position based on the listener displacement information may be performed by translating the object position by a vector that is negatively correlated with the direction and positively correlated with the absolute value of the listener's head displacement vector from the nominal listener position. An example of such a translation is shown schematically in FIG. 2.

[0095] Step S550Then, the modified object positions obtained in step S540 are further modified based on the listener orientation information. For example, this may be done by applying a rotation transformation to the modified object positions according to the listener orientation information. This rotation may be, for example, a rotation relative to the listener's head or a nominal listening position. The rotation transformation may be performed by a scene displacement algorithm.

[0096] As mentioned above, user offset compensation (i.e., modifying object positions based on listener displacement information) is taken into account when applying rotation transforms. For example, applying a rotation transform may include: Calculation of rotation transformation matrices (based on user orientation, e.g., listener orientation information), Transformation of object positions from spherical to Cartesian coordinates, Applying a rotation transformation to the audio object (i.e., to correct object position) with user position offset compensation, Converts the object position after rotation transformation from Cartesian coordinates back to spherical coordinates.

[0097] Further Step S560 Method 500 may further include rendering the audio objects to one or more real or virtual speakers according to the modified object positions (as shown in FIG. 5).

[0098] To this end, the further modified object positions may be adjusted to the input format used by an MPEG-H 3D Audio renderer (e.g., the audio object renderer 320 described above). The one or more (real or virtual) speakers may, for example, be part of a headset or may be part of a speaker arrangement (e.g., a 2.1 speaker arrangement, a 5.1 speaker arrangement, a 7.1 speaker arrangement, etc.). In some embodiments, the audio objects may, for example, be rendered to the left and right speakers of a headset.

[0099] The aim of steps S540 and S550 described above is as follows: modifying the object position and further modifying the modified object position are performed so that the audio object, after being rendered to one or more (real or virtual) speakers according to the modified object position, is psychoacoustically perceived by the listener as originating from a fixed position relative to the nominal listening position. This fixed position of the audio object is psychoacoustically perceived regardless of the displacement of the listener's head from the nominal listening position and regardless of the orientation of the listener's head with respect to the nominal orientation. In other words, the audio object may be perceived as moving (translating) relative to the listener's head when the listener's head undergoes a displacement from the nominal listening position. Similarly, the audio object may be perceived as moving (rotating) relative to the listener's head when the listener's head undergoes a change in orientation from the nominal orientation. This allows the listener to perceive nearby audio objects from various angles and distances by moving their head.

[0100] The modification of the object positions and further modification of the modified object positions in steps S540 and S550 may be performed in the context of a (rotational / translational) audio scene displacement, for example by the audio scene displacement unit 310 described above.

[0101] It should be noted that certain steps may be omitted depending on the particular use case being addressed. For example, if the listening position data 301 includes only listener displacement information (but not listener orientation information, or only listener orientation information indicating that the listener's head orientation has not deviated from the nominal orientation), step S550 may be omitted. The rendering in step S560 is then performed according to the corrected object position determined in step S540. Similarly, if the listener position information 301 includes only listener orientation information (but not listener displacement information, or only listener displacement information indicating that the listener's head position has not deviated from the nominal listener position), step S540 may be omitted. Then, step S550 relates to correcting the object position determined in step S530 based on the listener orientation information. The rendering in step S560 is performed according to the corrected object position determined in step S550.

[0102] Broadly speaking, the present invention proposes position updating of object positions (e.g., position information 302 accompanying audio data 322) received as part of object-based audio content based on listening position data 301 for a listener.

[0103] First, the object position (or channel position) p(az,el,r) is determined, which may be performed in the context of (e.g., as part of) step 530 of method 500.

[0104] For channel-based signals, the radial r may be determined as follows: If the intended loudspeaker (for the channel of the channel-based input signal) is present in a playback loudspeaker setup and the distance of the playback setup is known, then the radius r is set to the loudspeaker distance (e.g., in cm). If the intended loudspeaker is not present in the playback speaker setup, but the distance of the playback speakers (e.g. from the nominal listening position) is known, then radius r is set to the maximum playback speaker distance. If the intended loudspeaker is not present in the playback speaker setup and there is no known playback speaker distance, the radius r is set to a default value (e.g. 1023 cm).

[0105] For object-based signals, the radial coordinate r is determined as follows: If the object distance is known (e.g. communicated in prodMetadataConfig() from the production tools and format), the radius r is set to the known object distance (e.g. signaled by goa_bsObjectDistance[] (in cm) according to table AMD5.7 of the MPEG-H 3D Audio standard). [Table 3] If the object distance is known from the position information (e.g., from object metadata, conveyed in object_metadata()), then radius r is set to the object distance signaled in the position information (e.g., radius[] in cm conveyed in the object metadata). Radius r may be signaled according to the "Object Metadata Scaling" and "Object Metadata Restrictions" sections below.

[0106] Scaling object metadata As an optional step in the context of determining object positions, the object positions p=(az,el,r) determined from the position information may be scaled. This may involve applying a scaling factor to invert the encoder scaling of the input data for each component. This may be performed for all objects. The actual scaling of the object positions may be implemented along the lines of the following pseudocode: [Table 4]

[0107] Object Metadata Restrictions As a further optional step in the context of determining the object position p=(az,el,r), the (possibly scaled) object position determined from the position information may be constrained. This may involve applying constraints to the decoded values ​​for each component to keep the values ​​within a valid range. This may be performed for all objects. The actual constraints on the object position may be implemented according to the following pseudocode function: [Table 5]

[0108] The determined (and optionally scaled and / or constrained) object position may then be transformed into a predetermined coordinate system, for example a "common convention" coordinate system where 0° azimuth is at the right ear (positive values ​​counterclockwise) and 0° elevation is at the top of the head (positive values ​​downwards). Thus, the object position p may be transformed into a "common" convention position. This gives the object position p' as follows: p'(az',el',r) az'=az+90° el'=90°-el The radius r remains unchanged.

[0109] At the same time, listener displacement information (az offset ,el offset ,r offset ) may be transformed into a given coordinate system. Using "common convention", this becomes:

[0110] az' offset =az offset +90° el' offset =90°-el offset The radius r remains unchanged.

[0111] In particular, the transformation of both the object position and the listener's head displacement into said predetermined coordinate system may be performed in the context of step S530 or step S540.

[0112] The actual location update may be performed in the context of (e.g., as part of) step S540 of method 500. The location update may include the following steps:

[0113] As a first step, the position p, or the position p' if a transformation to a predetermined coordinate system has been performed, is transformed into Cartesian coordinates (x, y, z). In the following, without any intention of limitation, the process is described for the position p' in the predetermined coordinate system. Also, without any intention of limitation, the following orientation / direction of the coordinate axes may be assumed: x-axis points to the right (as viewed from the listener's head when in the nominal orientation), y-axis points straight ahead, and z-axis points straight up. At the same time, the listener's displacement information (az' offset ,el' offset ,r offset ) is transformed into Cartesian coordinates.

[0114] As a second step, the object position in Cartesian coordinates is shifted (translated) according to the listener's head displacement (scene displacement) in the manner described above. This may proceed as follows:

number

[0115] The shifted object position in Cartesian coordinates may be transformed into spherical coordinates and referred to as p". Following common convention, the shifted object position can be expressed in a given coordinate system as p" = (az",el",r).

[0116] Small resulting in a change in the radial parameter

number

[0117] In another example, this can result in significant radial parameter changes (i.e., r'≫r). large In the presence of a listener's head displacement, the modified position p" of the object can also be defined as p"=(az",el",r') instead of p"=(az",el",r) with a modified radial parameter r'.

[0118] The corresponding value of the modified radial parameter r' is determined by the listener's head displacement distance (i.e., r offset =||P0-P1||) and the initial radial parameter (i.e., r = ||P0-A||) (see, e.g., Figures 1 and 2). For example, the corrected radial parameter r' can be determined based on the following trigonometric relationship:

number

[0119] This mapping of the modified radial parameter r' to object / channel gains and its subsequent application for audio rendering can significantly improve the perceptual effect of level changes due to user movement. Allowing such modification of the radial parameter r' enables an "adaptive sweet spot," meaning that the MPEG rendering system dynamically adjusts the sweet spot position depending on the listener's current position. In general, rendering of audio objects according to modified (or further modified) object positions may be based on the modified radial parameter r'. In particular, object / channel gains for rendering audio objects may be based on (e.g., modified based on) the modified radial parameter r'.

[0120] Another example is during loudspeaker playback setup and rendering (e.g. Step S560 above In some cases, scene displacement can be disabled. However, optional enabling of scene displacement may be available, allowing the 3DoF+ renderer to generate a dynamically adjustable sweet spot depending on the listener's current position and orientation.

[0121] In particular, the step of converting object positions and listener head displacements into Cartesian coordinates is optional, and the translation / shift (correction) according to listener head displacement (scene displacement) can be performed in any suitable coordinate system, in other words, the choice of Cartesian coordinates above should be understood as a non-limiting example.

[0122] In some embodiments, the scene displacement process (including modification of object positions and / or further modification of modified object positions) is performed using flags (fields, elements, set bits) in the bitstream (e.g., useTrackingModeThe subclauses "17.3 Interface for local loudspeaker setup and rendering" and "17.4 Interface for binaural room impulse response (BRIR)" in ISO / IEC 23008-3 activate scene displacement processing. useTrackingMode In the context of this disclosure, useTrackingMode The element defines whether processing of scene displacement values ​​sent via the mpegh3daSceneDisplacementData() and mpegh3daPositionalSceneDisplacementData() interfaces is performed (subclause 17.3). Alternatively or additionally (subclause 17.4): useTrackingMode The field defines if a tracker device is connected and binaural rendering is processed in a special head tracking mode, which means processing of scene displacement values ​​sent via the mpegh3daSceneDisplacementData() and mpegh3daPositionalSceneDisplacementData() interfaces.

[0123] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented, for example, as software running on a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware and / or as an application-specific integrated circuit. Signals encountered in the methods and systems described above may be stored on media such as random access memory or optical storage media. The signals may be transmitted over a network, such as an airwave network, a satellite network, a wireless network, or a wired network such as the Internet. Typical devices utilizing the methods and systems described herein are portable electronic devices or other consumer equipment used to store and / or render audio signals.

[0124] While this document refers to MPEG, and in particular MPEG-H 3D Audio, this disclosure should not be construed as limited to these standards. Rather, as will be understood by those skilled in the art, this disclosure may find advantageous application in other standards for audio coding. Furthermore, while this document frequently refers to small positional displacements of a listener's head (e.g., from a nominal listening position), this disclosure is not limited to small positional displacements and is generally applicable to any positional displacements of a listener's head.

[0125] It should be noted that the specification and drawings merely illustrate the principles of the proposed methods, systems, and devices. Those skilled in the art will be able to implement various configurations, not explicitly described or shown herein, which embody the principles of the present invention and fall within its spirit and scope. Moreover, all examples and embodiments outlined herein are expressly intended solely for illustrative purposes, primarily to aid the reader in understanding the principles of the proposed methods. Furthermore, all statements herein providing principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass equivalents thereof.

[0126] In addition to the above, various example implementations and exemplary embodiments of the present invention will be apparent from the following enumerated example embodiments (EEE), which are not claims.

[0127] A first EEE relates to a method for decoding an encoded audio signal bitstream, the method comprising the steps of: receiving, by an audio decoding device 300, an encoded audio signal bitstream (302, 322), the encoded audio signal bitstream including encoded audio data (322) and metadata corresponding to at least one object audio signal (302); decoding, by the audio decoding device (300), the encoded audio signal bitstream (302, 322) to obtain a representation of multiple sound sources; receiving, by the audio decoding device (300), listening position data (301); and generating, by the audio decoding device (300), audio object position data (321), the audio object position data (321) describing multiple sound sources relative to a listening position based on the listening position data (301).

[0128] The second EEE relates to the method of the first EEE, wherein the listening position data (301) is based on a first set of first translational position data and a second set of second translational position and orientation data.

[0129] A third EEE relates to the method of the second EEE, wherein either the first translational position data or the second translational position data is based on at least one of a set of spherical coordinates or a set of Cartesian coordinates.

[0130] The fourth EEE relates to the method of the first EEE, in which the listening position data (301) is obtained through an MPEG-H 3D Audio decoder input interface. A fifth EEE relates to the method of the first EEE, wherein the encoded audio signal bitstream includes an MPEG-H 3D Audio bitstream syntax element, and the MPEG-H 3D Audio bitstream syntax element includes encoded audio data (322) and metadata (302) corresponding to at least one object audio signal.

[0131] A sixth EEE relates to the method of the first EEE, further comprising the step of rendering the multiple sound sources to multiple loudspeakers by the audio decoding device (300), wherein the rendering process complies with at least the MPEG-H 3D Audio standard.

[0132] A seventh EEE relates to the method of the first EEE, further comprising a step of converting, by the audio decoding device (300), a position (302) corresponding to the at least one object audio signal into a second position (321) corresponding to an audio object position based on a translation of listening position data (301).

[0133]

[0023] An eighth aspect of the present invention relates to the method of the seventh aspect, wherein the position p' of the audio object positions is determined in a predetermined coordinate system (e.g., according to common convention) as follows: p'=(az',el',r) az'=az+90° el'=90°-el az' offset =az offset +90° el' offset =90°-el offset is determined based on where az corresponds to a first azimuth parameter, el corresponds to a first elevation parameter, r corresponds to a first radial parameter, where az' corresponds to a second azimuth parameter, el' corresponds to a second elevation parameter, r' corresponds to a second radial parameter, and az offset corresponds to the third azimuthal parameter, and el offset corresponds to the third elevation parameter, and az offset corresponds to the fourth azimuthal parameter, and el offset corresponds to the fourth elevation parameter.

[0134] The 9th EEE relates to the method of the 8th EEE, wherein the shifted audio object position (321) of the audio object position (302) is expressed in Cartesian coordinates (x, y, z):

number

[0135] The tenth EEE relates to the method of the ninth EEE, where the parameter x offset ,y offset ,z offset teeth,

number

[0136] The eleventh EEE relates to the seventh EEE method, and the orientation parameters az offset is related to the scene displacement azimuthal position,

number

number

number

[0137] A twelfth EEE relates to the method of the tenth EEE, wherein x offset The parameters relate the scene displacement offset position sd_x to the x-axis direction; offset The parameter relates the scene displacement offset position sd_y to the y-axis direction; z offset The parameter relates the scene displacement offset position sd_z to the direction of the z axis.

[0138] A thirteenth EEE relates to the method of the first EEE, further comprising the step of interpolating, by the audio decoding device, the first position data related to the listening position data (301) and the object audio signal (102) at an update rate.

[0139] A fourteenth aspect relates to the method of the first aspect, further comprising determining, by the audio decoding device 300, efficient entropy coding of the listening position data (301).

[0140] A fifteenth EEE relates to the first EEE method, wherein the location data (301) regarding the listening position is derived based on sensor information.

Claims

1. 1. A method for processing position information indicative of an object position of an audio object, the processing being performed using an MPEG-H 3D Audio decoder, the object position being usable for rendering of the audio object, the method comprising: obtaining listener orientation information indicative of an orientation of the listener's head, the listener orientation information including information regarding yaw, pitch, and roll of the listener's head; obtaining listener displacement information indicating a displacement of the listener's head relative to a nominal listening position via an MPEG-H 3D Audio decoder input interface; determining the object position from the position information; modifying the object position based on the listener displacement information by applying a translation to the object position; Further modifying the modified object position based on the listener orientation information, When the listener displacement information indicates a small positional displacement of the listener's head from the nominal listening position, and the small positional displacement has an absolute value of 0.5 meters or less, the distance between the audio object position and the listening position after the listener's head displacement is equal to the distance between the corrected object position and the nominal listening position; A method comprising:

2. 10. A non-transitory computer-readable medium having instructions that, when executed by a digital signal processor or microprocessor, cause the digital signal processor or microprocessor to perform the method of claim 1.

3. 1. An MPEG-H 3D Audio decoder for processing position information indicative of object positions of audio objects, the object positions usable for rendering the audio objects, the decoder comprising a processor and a memory coupled to the processor, the processor comprising: obtaining listener orientation information indicative of an orientation of the listener's head, the listener orientation information including information regarding yaw, pitch, and roll of the listener's head; obtaining listener displacement information indicating a displacement of the listener's head relative to a nominal listening position via an MPEG-H 3D Audio decoder input interface; determining the object position from the position information; modifying the object position based on the listener displacement information by applying a translation to the object position; further modifying the modified object position based on the listener orientation information, When the listener displacement information indicates a small positional displacement of the listener's head from the nominal listening position, and the small positional displacement has an absolute value of 0.5 meters or less, the distance between the audio object position and the listening position after the listener's head displacement is equal to the distance between the corrected object position and the nominal listening position; adapted to carry out decoder.