Method, apparatus, and system for three degrees of freedom (3DOF+) extension for MPEG-H 3D audio
By handling the rotation and translation movement of the user's head within the framework of the MPEG-H 3D audio standard and modifying the position of the audio object, the problem that the existing technology cannot effectively handle the translation movement of the user's head, and achieving a more realistic audio object position experience and immersive listening effect.
Patent Information
- Application Number
- CN202111295025.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-25
- Filing Date
- 2019-04-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-04-09
AI Technical Summary
The existing MPEG-H 3D audio standards cannot effectively handle a certain small translational movement of the user's head, resulting in the inability to provide a more realistic audio object position experience in the 3DoF environment.
By processing position information indicating the position of the audio object, in combination with rotation and translational movement of the user's head, the position of the audio object is modified to achieve a limited six-degree of freedom (6DoF) experience within the framework of the MPEG-H 3D audio standard.
A more realistic listening experience is achieved, especially for audio objects close to the listener's head, allowing listeners to approach the audio objects from different angles and sides, enhancing the immersive listening experience.
Smart Images

Figure CN113993062B_ABST
Abstract
Description
[0001] Relevant information of divisional application
[0002] This is a divisional application. The parent case of this divisional application is a patent application for invention with the application date of April 9, 2019, application number 201980018139.X, and invention title "Method, apparatus, and system for three-degree-of-freedom (3DOF+) expansion for MPEG-H 3D audio".
[0003] Cross-reference to related applications
[0004] This application claims priority from the following priority applications: U.S. Provisional Application 62 / 654,915 filed on April 9, 2018 (Reference: D18045USP1); U.S. Provisional Application 62 / 695,446 filed on July 9, 2018 (Reference: D18045USP2); and U.S. Provisional Application 62 / 823,159 filed on March 25, 2019 (Reference: D18045USP3), which are hereby incorporated by reference herein. Technical Field
[0005] The present disclosure relates to methods and apparatuses for processing position information indicating the position of an audio object and information indicating a displacement of a listener's head position. Background Art
[0006] The first version (October 15, 2015) and Amendments 1-4 of the ISO / IEC 23008-3 MPEG-H 3D audio standard do not provide for a certain small translational movement of the user's head in a three-degree-of-freedom (3DoF) environment. Summary of the Invention
[0007] The first version (October 15, 2015) and Amendments 1-4 of the ISO / IEC 23008-3 MPEG-H 3D audio standard provide functionality for the possibilities in a 3DoF environment, where the user (listener) performs head rotation actions. However, such functionality only supports rotational scene displacement signaling and corresponding rendering at most. This means that in the case of a change in the orientation of the listener's head, the audio scene can remain spatially fixed, which corresponds to the 3DoF nature. However, in the current MPEG-H 3D audio ecosystem, it is not possible to consider a certain small translational movement of the user's head.
[0008] Therefore, there is a need for methods and apparatuses for processing the position information of audio objects, which can potentially consider a certain small translational movement of the user's head in combination with the rotational movement of the user's head.
[0009] The present disclosure provides devices and systems for processing location information, the devices and systems having the features of the respective independent claims and dependent claims.
[0010] According to one aspect of the present disclosure, a method for processing location information indicating the position of an audio object is described, wherein the processing may conform to the MPEG-H 3D audio standard. The object position may be used to render the audio object. The audio object may be included in object-based audio content together with its location information. The location information may be (part of) the metadata of the audio object. The audio content (e.g., the audio object and its location information) may be transmitted in an encoded audio bitstream. The method may include receiving the audio content (e.g., the encoded audio bitstream). The method may include obtaining listener orientation information indicating the orientation of the listener's head. The listener may be referred to as the user (e.g., of the audio decoder performing the method). The orientation of the listener's head (listener orientation) may be the orientation of the listener's head relative to a nominal orientation. The method may further include obtaining listener displacement information indicating the displacement of the listener's head. The displacement of the listener's head may be a displacement relative to a nominal listening position. The nominal listening position (or nominal listener position) may be a default position (e.g., a predetermined position, the expected position of the listener's head, or the optimal point of the speaker arrangement). The listener orientation information and the listener displacement information may be obtained through the MPEG-H 3D audio decoder input interface. The listener orientation information and the listener displacement information may be derived based on sensor information. The combination of the orientation information and the location information may be referred to as pose information. The method may further include determining the object position according to the location information. For example, the object position may be extracted from the location information. The determination (e.g., extraction) of the object position may be further based on information about the geometry of the speaker arrangement of one or more speakers in the listening environment. The object position may also be referred to as the channel position of the audio object. The method may further include modifying the object position based on the listener displacement information by applying a translation to the object position. Modifying the object position may involve correcting the object position for the displacement of the listener's head from the nominal listening position. In other words, modifying the object position may involve applying a position displacement compensation to the object position. The method may further include further modifying the modified object position based on the listener orientation information, for example, by applying a rotation transformation (e.g., a rotation relative to the listener's head or the nominal listening position) to the modified object position. Further modifying the modified object position for rendering the audio object may involve rotating the audio scene displacement.
[0011] Configured as described above, the proposed method provides a more realistic listening experience, especially for audio objects positioned close to the listener's head. In addition to the three (rotational) degrees of freedom conventionally provided to the listener in a 3DoF environment, the proposed method can also account for translational movements of the listener's head. This enables the listener to approach the nearby audio object from different angles and even from the side. For example, the listener may be able to listen to a "mosquito" audio object close to the listener's head from different angles by not only rotating their head but also by slightly moving their head. Thus, the proposed method can achieve an improved and more realistic immersive listening experience for the listener.
[0012] In some embodiments, modifying the object position and further modifying the modified object position can be performed such that after rendering to one or more physical speakers or virtual speakers according to the further modified object position, the audio object is perceived by the listener psychoacoustically as originating from a position fixed relative to the nominal listening position, regardless of the displacement of the listener's head from the nominal listening position and the orientation of the listener's head relative to the nominal orientation. Thus, when the listener's head undergoes a displacement from the nominal listening position, the audio object can be perceived as moving relative to the listener's head. Similarly, when the listener's head undergoes an orientation change from the nominal orientation, the audio object can be perceived as rotating relative to the listener's head. For example, the one or more speakers can be part of a headset or can be part of a speaker arrangement (e.g., 2.1 speaker arrangement, 5.1 speaker arrangement, 7.1 speaker arrangement, etc.).
[0013] In some embodiments, modifying the object position based on listener displacement information can be performed by translating the object position by a certain vector that is positively correlated with the magnitude and negatively correlated with the direction of the vector of the displacement of the listener's head from the nominal listening position.
[0014] Thereby, ensuring that the nearby audio object is perceived by the listener as moving according to their head movement. This helps to provide a more realistic listening experience for these audio objects.
[0015] In some embodiments, the listener displacement information can indicate a small positional displacement of the listener's head from the nominal listening position. For example, the absolute value of the displacement may not exceed 0.5 m. The displacement can be expressed in Cartesian coordinates (e.g., x, y, z) or spherical coordinates (e.g., azimuth, elevation, radius).
[0016] In some embodiments, the listener displacement information may indicate the displacement of the listener's head from a nominal listening position, which displacement can be achieved by the listener moving their upper body and / or head. Thus, the listener can achieve displacement without moving their lower body. For example, when the listener is sitting on a chair, displacement of the listener's head can be achieved.
[0017] In some embodiments, the position information may include an indication of the distance of the audio object from the nominal listening position. The distance (radius) can be less than 0.5 m. For example, the distance can be less than 1 cm. Alternatively, the decoder may set the distance of the audio object from the nominal listening position to a default value.
[0018] In some embodiments, the listener orientation information may include information about the yaw, pitch, and roll of the listener's head. The yaw, pitch, and roll may be given relative to a nominal orientation (e.g., a reference orientation) of the listener's head.
[0019] In some embodiments, the listener displacement information may include information about the displacement of the listener's head expressed in Cartesian coordinates or in spherical coordinates from the nominal listening position. Thus, for Cartesian coordinates, the displacement can be expressed in x, y, and z coordinates, and for spherical coordinates, the displacement can be expressed in azimuth, elevation, and radius coordinates.
[0020] In some embodiments, the method may further include detecting the orientation of the listener's head by a wearable device and / or a stationary device. Similarly, the method may further include detecting the displacement of the listener's head from the nominal listening position by a wearable device and / or a stationary device. The wearable device can be, correspond to, and / or include, for example, a headset or an augmented reality (AR) / virtual reality (VR) headset. For example, the stationary device can be, correspond to, and / or include a camera sensor. This allows obtaining accurate information about the displacement and / or orientation of the listener's head, and thereby enables realistic processing of approaching audio objects based on the orientation and / or displacement.
[0021] In some embodiments, the method may further include rendering the audio object to one or more physical speakers or virtual speakers according to the further modified object position. For example, the audio object can be rendered to the left and right speakers of a headset.
[0022] In some embodiments, the rendering may be performed to account for sound occlusion at a small distance of the audio object from the listener's head based on a head-related transfer function (HRTF) of the listener's head. Thus, rendering of approaching audio objects will be perceived by the listener in an even more realistic form.
[0023] In some embodiments, the further modified object position may be adjusted to an input format used by an MPEG-H 3D audio renderer. In some embodiments, the rendering may be performed using an MPEG-H 3D audio renderer. In some embodiments, the processing may be performed using an MPEG-H 3D audio decoder. In some embodiments, the processing may be performed by a scene displacement unit of an MPEG-H 3D audio decoder. Accordingly, the proposed method allows for the implementation of a limited six degrees of freedom (6DoF) experience (i.e., 3DoF+) within the framework of the MPEG-H 3D audio standard.
[0024] According to another aspect of the present disclosure, an additional method of processing position information indicative of an object position of an audio object is described. The object position may be used to render the audio object. The method may include obtaining listener displacement information indicative of a displacement of the listener's head. The method may further include determining the object position based on the position information. The method may further include modifying the object position based on the listener displacement information by applying a translation to the object position.
[0025] Configured as described above, the proposed method provides a more realistic listening experience, especially for audio objects positioned close to the listener's head. By being able to account for a certain small translational movement of the listener's head, the proposed method enables the listener to approach a nearby audio object from different angles and even from the side. Accordingly, the proposed method can achieve an improved and more realistic immersive listening experience for the listener.
[0026] In some embodiments, modifying the object position based on the listener displacement information is performed such that after rendering to one or more physical speakers or virtual speakers according to the modified object position, the audio object is perceptually, by the listener, to originate from a position fixed relative to a nominal listening position, regardless of the displacement of the listener's head from the nominal listening position.
[0027] In some embodiments, modifying the object position based on listener displacement information may be performed by translating the object position by a certain vector that is positively correlated with the magnitude and negatively correlated with the direction of the vector of the displacement of the listener's head from the nominal listening position.
[0028] According to another aspect of the present disclosure, an additional method of processing position information indicating an object position of an audio object is described. The object position can be used to render the audio object. The method can include obtaining listener orientation information indicating an orientation of a listener's head. The method can further include determining the object position based on the position information. The method can further include modifying the object position based on the listener orientation information, for example, by applying a rotation transformation to the object position (e.g., a rotation relative to the listener's head or the nominal listening position).
[0029] Configured as described above, the proposed method can take into account the orientation of the listener's head to provide a more realistic listening experience for the listener.
[0030] In some embodiments, modifying the object position based on the listener orientation information can be performed such that after rendering to one or more physical speakers or virtual speakers according to the modified object position, the audio object is perceived by the listener to originate from a position fixed relative to the nominal listening position, regardless of the orientation of the listener's head relative to the nominal orientation.
[0031] According to another aspect of the present disclosure, a device for processing position information indicating an object position of an audio object is described. The object position can be used to render the audio object. The device can include a processor and a memory coupled to the processor. The processor can be adapted to obtain listener orientation information indicating an orientation of a listener's head. The processor can further be adapted to obtain listener displacement information indicating a displacement of the listener's head. The processor can further be adapted to determine the object position based on the position information. The processor can further be adapted to modify the object position based on the listener displacement information by applying a translation to the object position. The processor can further be adapted to further modify the modified object position based on the listener orientation information, for example, by applying a rotation transformation (e.g., a rotation relative to the listener's head or the nominal listening position) to the modified object position.
[0032] In some embodiments, the processor can be adapted to modify the object position and further modify the modified object position such that after rendering to one or more physical speakers or virtual speakers according to the further modified object position, the audio object is perceived by the listener to originate from a position fixed relative to the nominal listening position, regardless of the displacement of the listener's head from the nominal listening position and the orientation of the listener's head relative to the nominal orientation.
[0033] In some embodiments, the processor may be adapted to modify the object position based on the listener displacement information by translating the object position by a certain vector, the vector being positively correlated with the magnitude and negatively correlated with the direction of the vector by which the listener's head is displaced from the nominal listening position.
[0034] In some embodiments, the listener displacement information may indicate a certain small position displacement of the listener's head from the nominal listening position.
[0035] In some embodiments, the listener displacement information may indicate the displacement of the listener's head from the nominal listening position, which can be achieved by the listener moving their upper body and / or head.
[0036] In some embodiments, the position information may include an indication of the distance of the audio object from the nominal listening position.
[0037] In some embodiments, the listener orientation information may include information about the yaw, pitch, and roll of the listener's head.
[0038] In some embodiments, the listener displacement information may include information about the displacement of the listener's head from the nominal listening position expressed in Cartesian coordinates or spherical coordinates.
[0039] In some embodiments, the device may further include a wearable device and / or a fixed device for detecting the orientation of the listener's head. In some embodiments, the device may further include a wearable device and / or a fixed device for detecting the displacement of the listener's head from the nominal listening position.
[0040] In some embodiments, the processor may further be adapted to render the audio object to one or more physical speakers or virtual speakers according to the further modified object position.
[0041] In some embodiments, the processor may be adapted to perform rendering that takes into account sound occlusion at a small distance of the audio object from the listener's head based on the HRTF of the listener's head.
[0042] In some embodiments, the processor may be adapted to adjust the further modified object position to an input format used by an MPEG-H 3D audio renderer. In some embodiments, the rendering may be performed using an MPEG-H 3D audio renderer. That is, the processor may implement an MPEG-H 3D audio renderer. In some embodiments, the processor may be adapted to implement an MPEG-H 3D audio decoder. In some embodiments, the processor may be adapted to implement a scene displacement unit of an MPEG-H 3D audio decoder.
[0043] According to another aspect of the present disclosure, an additional device for processing position information indicating an object position of an audio object is described. The object position may be used to render the audio object. The device may include a processor and a memory coupled to the processor. The processor may be adapted to obtain listener displacement information indicating a displacement of the listener's head. The processor may further be adapted to determine the object position based on the position information. The processor may further be adapted to modify the object position based on the listener displacement information by applying a translation to the object position.
[0044] In some embodiments, the processor may be adapted to modify the object position based on the listener displacement information such that after rendering to one or more physical speakers or virtual speakers according to the modified object position, the audio object is perceived by the listener psychoacoustically as originating from a position fixed relative to a nominal listening position, regardless of the displacement of the listener's head from the nominal listening position.
[0045] In some embodiments, the processor may be adapted to modify the object position based on the listener displacement information by translating the object position by a certain vector, the vector being positively correlated with a magnitude and negatively correlated with the direction of the vector of the displacement of the listener's head from the nominal listening position.
[0046] According to another aspect of the present disclosure, an additional device for processing position information indicating an object position of an audio object is described. The object position may be used to render the audio object. The device may include a processor and a memory coupled to the processor. The processor may be adapted to obtain listener orientation information indicating the orientation of the listener's head. The processor may further be adapted to determine the object position based on the position information. The processor may further be adapted to modify the object position based on the listener orientation information, for example, by applying a rotation transformation (e.g., a rotation relative to the listener's head or the nominal listening position) to the modified object position.
[0047] In some embodiments, the processor may be adapted to modify the object position based on the listener orientation information such that, after rendering the audio object to one or more physical speakers or virtual speakers according to the modified object position, the audio object is perceived by the listener psychoacoustically as originating from a position fixed relative to a nominal listening position, regardless of the orientation of the listener's head relative to the nominal orientation.
[0048] According to yet another aspect, a system is described. The system may include a device and a wearable device and / or a stationary device according to any of the above aspects, the wearable device and / or the stationary device being capable of detecting the orientation of the listener's head and detecting the displacement of the listener's head.
[0049] It should be understood that method steps and device features may be interchanged in various ways. Specifically, as understood by those skilled in the art, the details of the disclosed method may be implemented as a device adapted to perform some or all of the steps of the method, and vice versa. Specifically, it should be understood that a device according to the present disclosure may relate to a device for implementing or performing a method according to the above embodiments and their variations, and the corresponding statements made with respect to the method similarly apply to the corresponding device. Similarly, it should be understood that a method according to the present disclosure may relate to a method of operating a device according to the above embodiments and their variations, and the corresponding statements made with respect to the device similarly apply to the corresponding method. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The present invention will be explained in a demonstrative manner with reference to the following drawings, in which
[0051] Figure 1 an example of an MPEG-H 3D audio system is schematically shown;
[0052] Figure 2 an example of an MPEG-H 3D audio system according to the present invention is schematically shown;
[0053] Figure 3 an example of an audio rendering system according to the present invention is schematically shown;
[0054] Figure 4 an example set showing the Cartesian coordinate axes and their relationship to spherical coordinates is schematically shown; and
[0055] Figure 5 is a flowchart schematically showing an example of a method of processing position information of an audio object according to the present invention. DETAILED DESCRIPTION
[0056] As used herein, 3DoF generally refers to a system that can correctly handle the head movement (especially head rotation) of a user specified by three parameters (e.g., yaw, pitch, roll). Such systems are typically used in various gaming systems, such as virtual reality (VR) / augmented reality (AR) / mixed reality (MR) systems, or in other acoustic environments of this type.
[0057] As used herein, a user (e.g., of an audio decoder or a reproduction system including an audio decoder) may also be referred to as a "listener".
[0058] As used herein, 3DoF+ shall mean that in addition to the head movement of a user that can be correctly handled in a 3DoF system, a certain small translational movement can also be handled.
[0059] As used herein, "a certain small" shall indicate that the movement is restricted to below a threshold value typically of 0.5 meters. This means that the movement from the user's original head position is not greater than 0.5 meters. For example, the user's movement is restricted by his / her sitting on a chair.
[0060] As used herein, "MPEG-H 3D Audio" shall refer to the specification standardized in ISO / IEC 23008-3 and / or any future amendment, version or other version of the ISO / IEC 23008-3 standard.
[0061] In the context of the audio standards provided by the MPEG organization, the difference between 3DoF and 3DoF+ can be defined as follows:
[0062] ● 3DoF: Allows the user to experience yaw movement, pitch movement, roll movement (e.g., of the user's head);
[0063] ● 3DoF+: Allows the user to experience yaw movement, pitch movement, roll movement and limited translational movement (e.g., of the user's head) when sitting on a chair, for example.
[0064] The limited (a certain small) head translational movement can be a movement restricted by a certain movement radius. For example, since the user is in a seated position, the movement may be restricted, for example, without using the lower body. A certain small head translational movement can involve or correspond to the displacement of the user's head relative to the nominal listening position. The nominal listening position (or nominal listener position) can be a default position (e.g., a predetermined position, the expected position of the listener's head or the optimal point of the speaker arrangement).
[0065] The 3DoF+ experience can be comparable to a restricted 6DoF experience where translational movement can be described as limited or a certain small head movement. In one instance, audio is also rendered based on the user's head position and orientation, including possible sound occlusion. Rendering can be performed, for example, considering sound occlusion of an audio object at a small distance from the listener's head based on the head-related transfer function (HRTF) of the listener's head.
[0066] Regarding methods, systems, devices, and other apparatuses that are compatible with the functionality set forth by the MPEG-H 3D audio standard, it can mean that 3DoF+ can be used for one or more any future versions of the MPEG standard, such as a future version of the omnidirectional media format (e.g., standardized in a future version of MPEG-I); and / or any update to MPEG-H audio (e.g., based on an amendment or updated standard of the MPEG-H 3D audio standard); or any other related standard or companion standard that may need to be updated (e.g., a standard specifying certain types of metadata messages and SEI messages).
[0067] For example, an audio renderer that is normative for the audio standard set forth in the MPEG-H 3D audio specification can be extended to include rendering of an audio scene to accurately account for, for example, user interaction with the audio scene when the user slightly laterally moves their head.
[0068] The present invention provides various technical advantages, including the advantage of providing MPEG-H 3D audio capable of handling 3DoF+ use cases. The present invention extends the MPEG-H 3D audio standard to support 3DoF+ functionality.
[0069] To support 3DoF+ functionality, an audio rendering system should consider the limited / certain small position displacement of the user / listener's head. The position displacement should be determined based on the relative offset from the initial position (i.e., the default position / nominal listening position). In one instance, the magnitude of this offset (e.g., a radius offset that can be determined based on r offset = ||P0 - P1|| where P0 is the nominal listening position and P1 is the displaced position of the listener's head) of the listener's head is at most about 0.5 m. In another instance, the magnitude of the offset is limited to an offset that can only be achieved when the user is sitting on a chair and not performing lower body movement (but their head is moving relative to their body). This (certain small) offset distance results in a very small (perceived) horizontal difference and translational difference for far audio objects. However, for close objects, even such a certain small offset distance can become perceptually relevant. In fact, the listener's head movement can have a perceptual effect on the localization of the audio object for correct perception. As long as (i) the displacement of the user's head (e.g., r offsetThe ratio of =||P0 - P1||) to the distance to the audio object (e.g., r) trigonometrically generates an angle within the psychoacoustic capabilities of the user to detect the sound direction, and this perceptual effect can remain significant (i.e., perceptually detectable by the user / listener). For different audio renderer settings, audio materials, and playback configurations, such ranges can vary. For example, assuming the localization accuracy range is, for example, + / -3°, where the left - right movement freedom of the listener's head is + / -0.25m, this would correspond to an object distance of ~5m.
[0070] For an object close to the listener (e.g., an object at a distance of <1m from the user), correctly handling the position displacement of the listener's head is crucial for 3DoF+ scenarios because there are significant perceptual effects during both translational and horizontal changes.
[0071] An example of the handling of an object close to the listener is, for example, when an audio object (e.g., a mosquito) is positioned very close to the listener's face. Audio systems such as those providing VR / AR / MR capabilities should allow the user to perceive this audio object from all sides and angles, even if the user is making a small translational head movement. For example, the user should be able to accurately perceive the object (e.g., the mosquito), even when the user moves their head without moving their lower body.
[0072] However, systems currently compatible with the current MPEG - H 3D audio specification cannot correctly handle this problem. Instead, using a system compatible with the MPEG - H 3D audio system results in perceiving the "mosquito" from an incorrect position relative to the user. In scenarios involving 3DoF+ performance, a small translational movement should result in a significant difference in the perception of the audio object (e.g., when moving a user's head to the left, the "mosquito" audio object should be perceived from the right side relative to the user's head, etc.).
[0073] The MPEG - H 3D audio standard includes a bit - stream syntax that allows, through the bit - stream syntax, for example, through the object_metadata() - syntax element (starting from 0.5m), to signal object - distance information.
[0074] The syntax element prodMetadataConfig() can be introduced into the bit - stream provided by the MPEG - H 3D audio standard, which can be used to signal that the object distance is very close to the listener. For example, the syntax prodMetadataConfig() can signal that the distance between the user and the object is less than a certain threshold distance (e.g., <1cm).
[0075] Figure 1 and Figure 2Shows the present invention based on headphone rendering (i.e., where the speaker moves with the listener's head).
[0076] Figure 1 Illustrates an example of the system behavior 100 compliant with the MPEG-H 3D audio system. This example assumes that the listener's head is located at position P0 103 at time t0 and moves to position P1 104 at time t1 > t0. The dashed circles around positions P0 and P1 indicate the allowable 3DoF+ movement area (e.g., with a radius of 0.5 m). Position A 101 indicates the object position being signaled (at time t0 and time t1, i.e., assuming the object position being signaled is constant over time). Position A also indicates the object position rendered by the MPEG-H 3D audio renderer at time t0. Position B 102 indicates the object position rendered by the MPEG-H 3D audio at time t1. The vertical lines extending upward from positions P0 and P1 indicate the respective orientations (e.g., viewing directions) of the listener's head at times t0 and t1. The displacement of the user's head between position P0 and position P1 can be represented by r offset = ||P0 - P1|| 106. In the case where the listener is located at the default position (nominal listening position) P0 103 at time t0, he / she will perceive the audio object (e.g., a mosquito) at the correct position A 101. If the user moves to position P1 104 at time t1 and the MPEG-H 3D audio processing is applied in the current standardized form, he / she will perceive the audio object at position B 102, which introduces the shown error δ AB 105. That is, although the listener's head moves, the audio object (e.g., a mosquito) will still be perceived as being located directly in front of the listener's head (i.e., moving substantially with the listener's head). It is worth noting that the introduced error δ AB 105 occurs regardless of the orientation of the listener's head.
[0077] Figure 2 Illustrates an example of the system behavior of system 200 according to the present invention with respect to MPEG-H 3D audio. In Figure 2In this case, the listener's head is located at position P0 203 at time t0 and moves to position P1 204 at time t1 > t0. The dashed circles around positions P0 and P1 again indicate the allowable 3DoF+ movement areas (e.g., with a radius of 0.5 m). At 201, indicating position A = B means the object position being signaled (at time t0 and time t1, i.e., assuming the object position being signaled is constant over time). Position A = B 201 also indicates the object positions rendered by MPEG-H 3D audio at time t0 and time t1. The vertical arrows extending upward from positions P0 203 and P1 204 indicate the respective orientations (e.g., viewing directions) of the listener's head at times t0 and t1. In the case where the listener is located at the initial / default position (nominal listening position) P0 203 at time t0, he / she will perceive the audio object (e.g., a mosquito) at the correct position A 201. If the user moves to position P1 203 at time t1, he / she will still perceive the audio object at position B 201, which is similar to (e.g., substantially equal to) position A 201 according to the present invention. Thus, the present invention allows the user's position to change over time (e.g., from position P0 203 to position P1 204), while still perceiving the sound from the same (spatially fixed) localization (e.g., position A = B 201, etc.). In other words, the audio object (e.g., a mosquito) moves relative to the listener's head according to the movement of the listener's head (e.g., negatively correlated with the head movement). This enables the user to move around the audio object (e.g., a mosquito) and perceive the audio object from different angles or even from the side. The displacement of the user's head between position P0 and position P1 can be represented by r offset = ||P0 - P1|| 206.
[0078] Figure 3 An example of an audio rendering system 300 according to the present invention is shown. The audio rendering system 300 may correspond to or include a decoder, e.g., an MPEG-H 3D audio decoder. The audio rendering system 300 may include an audio scene displacement unit 310 having a corresponding audio scene displacement processing interface (e.g., an interface for scene displacement data according to the MPEG-H 3D audio standard). The audio scene displacement unit 310 may output object positions 321 for rendering the respective audio objects. For example, the scene displacement unit may output object position metadata for rendering the respective audio objects.
[0079] The audio rendering system 300 may further include an audio object renderer 320. For example, the renderer may consist of hardware, software, and / or any part or all of the processing performed via cloud computing, which includes various services on the Internet commonly referred to as "the cloud" that are compatible with the specifications set forth by the MPEG-H 3D audio standard, such as software development platforms, servers, storage, and software. The audio object renderer 320 may render audio objects to one or more (real or virtual) speakers according to the corresponding object positions (which may be the modified object positions or further modified object positions described below). The audio object renderer 320 may render audio objects to headphones and / or speakers. That is, the audio object renderer 320 may generate object waveforms according to a given reproduction format. To this end, the audio object renderer 320 may utilize compressed object metadata. Each object may be rendered to certain output channels according to its object position (e.g., the modified object position, or the further modified object position). Thus, the object position may also be referred to as the channel position of its audio object. The audio object position 321 may be included in the object position metadata or scene displacement metadata output by the scene displacement unit 310.
[0080] The processing of the present invention may comply with the MPEG-H 3D audio standard. Thus, the processing may be performed by an MPEG-H 3D audio decoder, or more specifically, by an MPEG-H scene displacement unit and / or an MPEG-H 3D audio renderer. Therefore, Figure 3 the audio rendering system 300 may correspond to or include an MPEG-H 3D audio decoder (i.e., a decoder that complies with the specifications set forth by the MPEG-H 3D audio standard). In one example, the audio rendering system 300 may be a device including a processor and a memory coupled to the processor, where the processor is adapted to implement the MPEG-H 3D audio decoder. Specifically, the processor may be adapted to implement the MPEG-H scene displacement unit and / or the MPEG-H 3D audio renderer. Thus, the processor may be adapted to perform the processing steps described in the present disclosure (e.g., steps S510 to S560 of method 500 described below with reference to Figure 5 ). In another example, the processing or the audio rendering system 300 may be performed in the cloud.
[0081] The audio rendering system 300 may obtain (e.g., receive) listening location data 301. The audio rendering system 300 may obtain the listening location data 301 through the MPEG-H 3D audio decoder input interface.
[0082] The listening position data 301 can indicate the orientation and / or position (e.g., displacement) of the listener's head. Thus, the listening position data 301 (which may also be referred to as pose information) can include listener orientation information and / or listener displacement information.
[0083] The listener displacement information can indicate the displacement of the listener's head (e.g., from a nominal listening position). The listener displacement information can correspond to or include an indication of the magnitude of the displacement of the listener's head from the nominal listening position, r offset = ||P0 - P1||206, as Figure 2 shown. In the context of the present invention, the listener displacement information indicates a certain small positional displacement of the listener's head from the nominal listening position. For example, the absolute value of the displacement may not exceed 0.5 m. Generally, this is the displacement of the listener's head from the nominal listening position, which can be achieved by the listener moving their upper body and / or head. That is, the listener can achieve the displacement without moving their lower body. For example, as indicated above, when the listener is sitting on a chair, the displacement of the listener's head can be achieved. The displacement can be represented in various coordinate systems, for example, in Cartesian coordinates (represented by x, y, z) or spherical coordinates (e.g., represented by azimuth, elevation, radius). Alternative coordinate systems for representing the displacement of the listener's head are also feasible and should be understood to be covered by the present disclosure.
[0084] The listener orientation information can indicate the orientation of the listener's head (e.g., the orientation of the listener's head relative to the nominal orientation / reference orientation of the listener's head). For example, the listener orientation information can include information about the yaw, pitch, and roll of the listener's head. Here, the yaw, pitch, and roll can be given relative to the nominal orientation.
[0085] Listening location data 301 can be continuously collected from a receiver that can provide information about the translational movement of a user. For example, the listening location data 301 used at a certain time instance may have been recently collected from the receiver. The listening location data can be derived / collected / generated based on sensor information. For example, the listening location data 301 can be derived / collected / generated by a wearable device and / or a stationary device having appropriate sensors. That is, the orientation of the listener's head can be detected by the wearable device and / or the stationary device. Similarly, the displacement of the listener's head (e.g., from a nominal listening position) can be detected by the wearable device and / or the stationary device. For example, the wearable device can be, correspond to, and / or include a headset (e.g., an AR / VR headset). For example, the stationary device can be, correspond to, and / or include a camera sensor. For example, the stationary device can be included in a television or a set-top box. In some embodiments, the listening location data 301 can be received from an audio encoder (e.g., an encoder compliant with MPEG-H 3D audio) that may have obtained (e.g., received) the sensor information.
[0086] In one instance, the wearable device and / or the stationary device for detecting the listening location data 301 can be referred to as a tracking device that supports head position estimation / detection and / or head orientation estimation / detection. There are various solutions that allow the accurate tracking of a user's head movement using a computer or a smartphone camera (e.g., based on face recognition and tracking such as "FaceTrackNoIR", "opentrack"). Moreover, several head-mounted display (HMD) virtual reality systems (e.g., HTC VIVE, Oculus Rift) have integrated head tracking technology. Any of these solutions can be used in the context of the present disclosure.
[0087] It is also important to note that the head displacement distance in the physical world does not have to correspond one-to-one with the displacement indicated by the listening location data 301. To achieve a surreal effect (e.g., an overly magnified user motion parallax effect), certain applications can use different sensor calibration settings or specify different mappings between the motion in the real space and the motion in the virtual space. Therefore, it can be expected that in some use cases, a certain small physical movement results in a larger displacement in virtual reality. In any case, it can be said that the magnitudes of the displacements in the physical world and in virtual reality (i.e., the displacements indicated by the listening location data 301) are positively correlated. Similarly, the directions of the displacements in the physical world and in virtual reality are positively correlated.
[0088] The audio rendering system 300 can further receive object location information (e.g., object location data) 302 and audio data 322. The audio data 322 can include one or more audio objects. The location information 302 can be part of the metadata of the audio data 322. The location information 302 can indicate the respective object locations of the one or more audio objects. For example, the location information 302 can include an indication of the distance of the respective audio object relative to the nominal listening position of the user / listener. The distance (radius) can be less than 0.5 m. For example, the distance can be less than 1 cm. If the location information 302 does not include an indication of the distance of a given audio object from the nominal listening position, the audio rendering system can set the distance of this audio object from the nominal listening position to a default value (e.g., 1 m). The location information 302 can further include an indication of the elevation angle and / or azimuth angle of the respective audio object.
[0089] Each object location can be used to render its corresponding audio object. Thus, the location information 302 and the audio data 322 can be included in or form object-based audio content. The audio content (e.g., the audio object / audio data 322 and its location information 302) can be transmitted in an encoded audio bitstream. For example, the audio content can be in a format of a bitstream received through network transmission. In this case, it can be said that the audio rendering system receives the audio content (e.g., from the encoded audio bitstream).
[0090] In an example of the present invention, metadata parameters can be used to correct the handling of use cases with backward-compatible enhancements for 3DoF and 3DoF+. In addition to the listener orientation information, the metadata can further include listener displacement information. Such metadata parameters can be utilized by Figure 2 and 3 the systems shown in any other embodiments of the present invention.
[0091] The backward-compatible enhancements can allow the handling of use cases (e.g., embodiments of the present invention) to be corrected based on the normative MPEG-H 3D audio scene displacement interface. This means that traditional MPEG-H 3D audio decoders / renderers will still produce output, even if it is incorrect. However, the enhanced MPEG-H 3D audio decoder / renderer according to the present invention will correctly apply the extended data (e.g., extended metadata) and processing, and thus may handle the scenes of objects located close to the listener in a correct manner.
[0092] In one example, the present invention relates to providing data for a small translational movement of a user's head in a format different from the format outlined below, and the formula may be adapted accordingly. For example, the data may be provided in a format such as x-coordinate, y-coordinate, z-coordinate (in a Cartesian coordinate system), rather than in a format of azimuth, elevation, and radius (in a spherical coordinate system). Examples of the relative positions of these coordinate systems with respect to each other are as Figure 4 shown.
[0093] In one example, the present invention relates to providing metadata for inputting a translational movement of a listener's head (e.g., listener displacement information included in Figure 3 the listener positioning data 301 shown). The metadata may be used, for example, for an interface of scene displacement data. The metadata (e.g., listener displacement information) may be obtained by deploying a tracking device that supports 3DoF+ or 6DoF tracking.
[0094] In one example, the metadata (e.g., listener displacement information, specifically the displacement of the listener's head, or equivalently, scene displacement) may be represented by three parameters: sd_azimuth, sd_elevation, and sd_radius, which relate to the azimuth, elevation, and radius (spherical coordinates) of the displacement of the listener's head (or scene displacement).
[0095] The syntax of these parameters is given in the following table.
[0096] Table 264b - Syntax of mpegh3daPositionalSceneDisplacementData()
[0097]
[0098] The sd_azimuth field defines the azimuth position of the scene displacement. This field can take values from -180 to 180.
[0099] az offset =(sd_azimuth - 128)·1.5
[0100] az offset =min(max(az offset , -180), 180)
[0101] The sd_elevation field defines the elevation position of the scene displacement. This field can take values from -90 to 90.
[0102] el offset =(sd_elevation - 32)·3.0
[0103] el offset= min(max(el offset , -90), 90)
[0104] The sd_radius field defines the scene displacement radius. This field can take values from 0.015626 to 0.25.
[0105] r offset = (sd_radius + 1) / 16
[0106] In another example, metadata (e.g., listener displacement information) can be represented by the following three parameters in Cartesian coordinates: sd_x, sd_y, and sd_z, which reduces the processing of the data from spherical coordinates to Cartesian coordinates. The metadata can be based on the following syntax:
[0107]
[0108] As described above, the above syntax or its equivalent syntax can convey information related to rotation about the x-axis, y-axis, and z-axis.
[0109] In one example of the present invention, the processing of the scene displacement angles of the channels and objects can be enhanced by expanding the equations that describe the position changes of the user's head. That is, the processing of the object positions can take into account (e.g., can be at least partially based on) the listener displacement information.
[0110] Figure 5 An example of a method 500 for processing position information indicating the position of an audio object is shown in the flowchart of. This method can be performed by a decoder, such as an MPEG-H 3D audio decoder. Figure 3 The audio rendering system 300 of can be an example of such a decoder.
[0111] As a first step ( Figure 5 not shown in), for example, audio content including audio objects and corresponding position information is received from a bitstream of encoded audio. Then, the method can further include decoding the encoded audio content to obtain the audio objects and position information.
[0112] At Step S510 , listener orientation information is obtained (e.g., received). The listener orientation information can indicate the orientation of the listener's head.
[0113] At Step S520 , listener displacement information is obtained (e.g., received). The listener displacement information can indicate the displacement of the listener's head.
[0114] At Step S530At [a certain point], the object position is determined based on the position information. For example, the object position can be extracted from the position information (e.g., represented by azimuth, elevation angle, radius, or x, y, z or their equivalents). The determination of the object position can also be at least partially based on information about the geometry of the speaker arrangement of one or more (real or virtual) speakers in the listening environment. If the radius is not included in the position information of the audio object, the decoder can set the radius to a default value (e.g., 1m). In some embodiments, the default value can depend on the geometry of the speaker arrangement.
[0115] It is noted that steps S510, S520, and S520 can be executed in any order.
[0116] At Step S540 At [a certain point], the object position determined at step S530 is modified based on the listener displacement information. This can be done by applying a translation to the object position according to the displacement information (e.g., according to the displacement of the listener's head). Thus, it can be said that modifying the object position involves correcting the object position for the displacement of the listener's head (e.g., the displacement from the nominal listening position). Specifically, modifying the object position based on the listener displacement information can be performed by translating the object position by a certain vector, which is positively correlated with the magnitude and negatively correlated with the direction of the vector of the listener's head displacement from the nominal listening position. Figure 2 An example of such a translation is schematically shown in
[0117] At Step S550 At [a certain point], the modified object position obtained at step S540 is further modified based on the listener orientation information. For example, this can be done by applying a rotation transformation to the modified object position according to the listener orientation information. This rotation can be, for example, a rotation relative to the listener's head or the nominal listening position. The rotation transformation can be performed by a scene displacement algorithm.
[0118] As pointed out above, when applying the rotation transformation, user offset compensation is considered (i.e., the modification of the object position based on the listener displacement information). For example, applying the rotation transformation can include:
[0119] ● Calculating a rotation transformation matrix (based on the user orientation, e.g., the listener orientation information),
[0120] ● Converting the object position from spherical coordinates to Cartesian coordinates;
[0121] ● Applying the rotation transformation to the audio object compensated for user - position - offset (i.e., applying it to the modified object position), and
[0122] ● After the rotation transformation, converting the object position back from Cartesian coordinates to spherical coordinates.
[0123] As a further Step S560 ( Figure 5 (not shown in), method 500 may include rendering an audio object to one or more physical or virtual speakers according to a further modified object position. To this end, the further modified object position may be adjusted to an input format used by an MPEG-H 3D audio renderer (e.g., the audio object renderer 320 described above). The one or more (physical or virtual) speakers may be part of, for example, a pair of headphones, or may be part of a speaker arrangement (e.g., a 2.1 speaker arrangement, a 5.1 speaker arrangement, a 7.1 speaker arrangement, etc.). In some embodiments, for example, an audio object may be rendered to the left and right speakers of a pair of headphones.
[0124] The purposes of steps S540 and S550 described above are as follows. That is, modifying the object position and further modifying the modified object position are performed such that after rendering the audio object to one or more (physical or virtual) speakers according to the further modified object position, the audio object is perceived by a listener to originate from a position fixed relative to a nominal listening position in psychoacoustics. This fixed position of the audio object should be perceived in psychoacoustics regardless of the displacement of the listener's head from the nominal listening position and regardless of the orientation of the listener's head relative to the nominal orientation. In other words, when the listener's head experiences a displacement from the nominal listening position, the audio object may be perceived to move (translate) relative to the listener's head. Similarly, when the listener's head experiences a change in orientation from the nominal orientation, the audio object may be perceived to move (rotate) relative to the listener's head. Thus, the listener can perceive an approaching audio object from different angles and distances by moving their head.
[0125] Modifying the object position and further modifying the modified object position at steps S540 and S550, respectively, may be performed, for example, by the audio scene displacement unit 310 described above in the context of (rotation / translation) audio scene displacement.
[0126] It should be noted that certain steps can be omitted according to the specific use case at hand. For example, if the listening location data 301 only contains the listener displacement information (but does not contain the listener orientation information, or only contains the listener orientation information indicating that the orientation of the listener's head has no deviation from the nominal orientation), then step S550 can be omitted. Then, the rendering at step S560 will be performed according to the modified object position determined at step S540. Similarly, if the listening location data 301 only contains the listener orientation information (but does not contain the listener displacement information, or only contains the listener displacement information indicating that the position of the listener's head has no deviation from the nominal listening position), then step S540 can be omitted. Then, step S550 will involve modifying the object position determined at step S530 based on the listener orientation information. The rendering at step S560 will be performed according to the modified object position determined at step S550.
[0127] Broadly speaking, the present invention proposes a position update for an object position received as part of object-based audio content (e.g., position information 302 and audio data 322) based on the listener's listening location data 301.
[0128] First, the object position (or channel position) p = (az, el, r) is determined. This can be performed in the context of step 530 of method 500 (e.g., as part of said step).
[0129] For a channel-based signal, the radius r can be determined as follows:
[0130] — If there is an expected speaker in the reproduction speaker setup (for the channel of the channel-based input signal) and the distance of the reproduction setup is known, then the radius r is set to the speaker distance (e.g., in cm).
[0131] — If there is no expected speaker in the reproduction speaker setup, but the distance of the reproduction speaker (e.g., from the nominal listening position) is known, then the radius r is set to the maximum reproduction speaker distance.
[0132] — If there is no expected speaker in the reproduction speaker setup and the reproduction speaker distance is not known, then the radius r is set to a default value (e.g., 1023 cm).
[0133] For an object-based signal, the radius r is determined as follows:
[0134] — If the object distance is known (e.g., known from the production tool and production format and conveyed in prodMetadataConfig()), then set the radius r to the known object distance (e.g., signaled by goa_bsObjectDistance[] (in cm) according to Table AMD5.7 of the MPEG-H 3D Audio standard).
[0135] Table AMD5.7 — Syntax of goa_Production_Metadata()
[0136]
[0137]
[0138] — If the object distance is known from the position information (e.g., known from the object metadata and conveyed in object_metadata()), then set the radius r to the object distance signaled in the position information (e.g., set to radius[] (in cm) conveyed together with the object metadata). The radius r can be signaled according to the sections shown below: "Scaling of Object Metadata" and "Limiting Object Metadata".
[0139] Scaling of Object Metadata
[0140] As an optional step in the context of determining the object position, the object position p = (az, el, r) determined from the position information can be scaled. This can involve applying a scaling factor to reverse the encoder scaling of the input data for each component. This can be done for each object. The actual scaling of the object position can be implemented according to the following pseudocode:
[0141]
[0142]
[0143] Limiting Object Metadata
[0144] As an additional optional step in the context of determining the object position, the (possibly scaled) object position p = (az, el, r) determined from the position information can be limited. This can involve imposing limits on the decoded values for each component to keep the values within a valid range. This can be done for each object. The actual limiting of the object position can be implemented according to the functionality of the following pseudocode:
[0145]
[0146]
[0147] Afterwards, the determined (and optionally, scaled and / or limited) object position p = (az, el, r) can be transformed into a predefined coordinate system, e.g., a coordinate system according to the "common convention", where the 0° azimuth is at the right ear (positive values in the counterclockwise direction), and the 0° elevation is at the top of the head (positive values downwards). Thus, the object position p can be transformed into the position p' according to the "common" convention. This generates the object position p' using the following:
[0148] p' = (az', el', r)
[0149] az′ = az + 90°
[0150] el′ = 90° - el
[0151] where the radius r remains unchanged.
[0152] At the same time, the displacement of the listener's head indicated by the listener displacement information (az offset , el offset , r offset ) can be transformed into a predefined coordinate system. Using the "common convention", this is equivalent to
[0153] az′ offset = az offset + 90°
[0154] el′ offset = 90° - el offset
[0155] where the radius r offset remains unchanged.
[0156] It is worth noting that the transformation to the predefined coordinate system for both the object position and the listener's head displacement can be performed in the context of step S530 or step S540.
[0157] The actual position update can be performed in the context of step S540 of method 500 (e.g., as part of said step). The position update can include the following steps:
[0158] As a first step, the position p or, in the case where the transfer to the predefined coordinate system has already been performed, the position p' is transferred to Cartesian coordinates (x, y, z). Hereinafter, without loss of generality, the process will be described for the position p' in the predefined coordinate system. Also, without loss of generality, the following orientation / direction of the coordinate axes can be assumed: the x-axis points to the right (when viewed from the listener's head in the nominal orientation), the y-axis points straight ahead, and the z-axis points straight up. At the same time, the displacement of the listener's head indicated by the listener displacement information (az′ offset , el′ offset , roffset )Convert the indicated displacement into Cartesian coordinates.
[0159] As a second step, offset (translate) the object position in Cartesian coordinates according to the displacement of the listener's head (scene displacement) in the above-described manner. This can be done by:
[0160] x = r·sin(el′)·cos(az′)+r offset ·sin(el′ offset )·cos(az′ offset )
[0161] y = r·sin(el′)·sin(az′)+r offset ·sin(el′ offset )·sin(az′ offset )
[0162] z = r·cos(el′)+r offset ·cos(el′ offset )
[0163] The above translation is an example of modifying the object position based on the listener displacement information in step S540 of method 500.
[0164] The offset object position in Cartesian coordinates is converted into spherical coordinates and can be referred to as p″. The offset object position can be expressed as p″ = (az″, el″, r′) in a predetermined coordinate system according to a common convention.
[0165] When there is a listener head displacement that produces a certain small change in the radius parameter (i.e., r′≈r), the modified object position p″ can be redefined as p″ = (az″, el″, r).
[0166] In another example, when there is a large listener head displacement that can produce a significant change in the radius parameter (i.e., r′>>r), the modified object position p″ can also be defined as p″ = (az″, el″, r′) instead of p″ = (az″, el″, r) with a modified radius parameter r′.
[0167] The corresponding value of the modified radius parameter r′ can be obtained from the listener's head displacement distance (i.e., r offset = ||P0 - P1||) and the initial radius parameter (i.e., r = ||P0 - A||) (see, for example, Figure 1 and 2 ). For example, the modified radius parameter r′ can be determined based on the following trigonometric relationship:
[0168]
[0169] Mapping this modified radius parameter r' to the object / channel gain and its application in subsequent audio rendering can significantly improve the perceived effects of the horizontal changes due to user movement. Allowing such modification of the radius parameter r' enables an "adaptive sweet spot". This would mean that the MPEG rendering system dynamically adjusts the sweet spot position according to the current location of the listener. Generally, the rendering of audio objects based on the modified (or further modified) object positions can be based on the modified radius parameter r'. Specifically, the object / channel gain for rendering the audio object can be based on the modified radius parameter r' (e.g., modified based on the modified radius parameter).
[0170] In another instance, during speaker reproduction setup and rendering (e.g., at The above Step S560 ), scene displacement can be disabled. However, optional enabling of scene displacement can be available. This enables the 3DoF+ renderer to create a dynamically adjustable sweet spot according to the current location and orientation of the listener.
[0171] It is noted that the step of converting the object position and the displacement of the listener's head to Cartesian coordinates is optional, and the translation / offset (modification) according to the displacement of the listener's head (scene displacement) can be performed in any suitable coordinate system. In other words, the choice of Cartesian coordinates above should be understood as a non-limiting example.
[0172] In some embodiments, scene displacement processing (including modifying object positions and / or further modifying modified object positions) can be enabled or disabled by flags (fields, elements, set bits) in a bitstream (e.g., the useTrackingMode element). Subclauses "17.3 Interfaces for Local Speaker Setup and Rendering" and "17.4 Interfaces for Binaural Room Impulse Response (BRIR)" in ISO / IEC 23008-3 contain descriptions of the useTrackingMode element that activates scene displacement processing. In the context of the present disclosure, the useTrackingMode element should define (subclause 17.3) whether processing of scene displacement values sent via the mpegh3daSceneDisplacementData() interface and the mpegh3daPositionalSceneDisplacementData() interface occurs. Alternatively or additionally, (subclause 17.4) the useTrackingMode field should define whether a tracker device is connected and whether binaural rendering should be processed in a special head tracking mode, which means that processing of scene displacement values sent via the mpegh3daSceneDisplacementData() interface and the mpegh3daPositionalSceneDisplacementData() interface should occur.
[0173] The methods and systems described herein can be implemented as software, firmware, and / or hardware. Certain components can be implemented, for example, as software running on a digital signal processor or a microprocessor. Other components can be implemented, for example, as hardware and / or an application specific integrated circuit. Signals encountered in the described methods and systems can be stored on media such as random access memory or optical storage media. The signals can be transmitted via a network, such as a radio network, a satellite network, a wireless network, or a wired network, e.g., the Internet. A typical device that utilizes the methods and systems described herein is a portable electronic device or other consumer device for storing and / or rendering audio signals.
[0174] Although this document refers to MPEG, and specifically MPEG-H 3D Audio, the present disclosure should not be construed as limited to these standards. Instead, as will be understood by those skilled in the art, the present disclosure can also find advantageous applications in other audio coding standards.
[0175] Furthermore, although this document frequently refers to a small positional displacement of the listener's head (e.g., from a nominal listening position), the present disclosure is not limited to a small positional displacement and can generally be applied to any positional displacement of the listener's head.
[0176] It should be noted that the description and the drawings only illustrate the principles of the proposed methods, systems, and devices. Those skilled in the art will be able to implement various arrangements, although not explicitly described or shown herein, which embody the principles of the present invention and are within the spirit and scope of the present invention. In addition, all examples and embodiments outlined in this document are explicitly intended, in principle, solely for explanatory purposes to assist the reader in understanding the principles of the proposed methods. Moreover, all statements providing the principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to cover their equivalents.
[0177] In addition to the above, various exemplary embodiments and example embodiments of the present invention will become apparent from the following listed enumerated example embodiments (EEE), which are not claims.
[0178] The first EEE relates to a method for decoding an encoded audio signal bitstream, the method comprising: receiving, by an audio decoding device 300, the encoded audio signal bitstream 302, 322, wherein the encoded audio signal bitstream includes encoded audio data 322 and metadata corresponding to at least one object-audio signal 302; decoding, by the audio decoding device 300, the encoded audio signal bitstream 302, 322 to obtain a representation of a plurality of sound sources; receiving, by the audio decoding device 300, listening location data 301; generating, by the audio decoding device 300, audio object position data 321, wherein the audio object position data 321 describes a plurality of sound sources relative to the listening location based on the listening location data 301.
[0179] The second EEE relates to the method of the first EEE, wherein the listening location data 301 is based on a first set of first translational displacement data and a second set of second translational and orientation data.
[0180] The third EEE relates to the method of the second EEE, wherein the first translational displacement data or the second translational displacement data is based on at least one of a set of spherical coordinates or a set of Cartesian coordinates.
[0181] The fourth EEE relates to the method of the first EEE, wherein the listening location data 301 is obtained through an MPEG-H 3D audio decoder input interface.
[0182] The fifth EEE relates to the method of the first EEE, wherein the encoded audio signal bitstream contains MPEG-H 3D audio bitstream syntax elements, and wherein the MPEG-H 3D audio bitstream syntax elements contain the encoded audio data 322 and the metadata corresponding to at least one object-audio signal 302.
[0183] The sixth EEE relates to the method of the first EEE, the method further comprising rendering, by the audio decoding device 300, the plurality of sound sources to a plurality of speakers, wherein the rendering process complies at least with the MPEG-H 3D audio standard.
[0184] The seventh EEE relates to the method of the first EEE, the method further comprising converting, by the audio decoding device 300 based on the translation of the listening position data 301, a position p corresponding to the at least one object-audio signal 302 to a second position p″ corresponding to the audio object position 321.
[0185] The eighth EEE relates to the method of the seventh EEE, wherein a position p' of the audio object position in a predetermined coordinate system is determined (e.g., according to a common convention) based on:
[0186] P' = (az', el', r)
[0187] az′ = az + 90°
[0188] el′ = 90° - el
[0189] az′ offset = az offset + 90°
[0190] el′ offset = 90° - el offset
[0191] wherein az corresponds to a first azimuth parameter, el corresponds to a first elevation parameter, and r corresponds to a first radius parameter, herein az′ corresponds to a second azimuth parameter, el′ corresponds to a second elevation parameter and r′ corresponds to a second radius parameter, wherein az offset corresponds to a third azimuth parameter, el offset corresponds to a third elevation parameter, and wherein az′ offset corresponds to a fourth azimuth parameter, el′ offset corresponds to a fourth elevation parameter.
[0192] The ninth EEE relates to the method of the eighth EEE, wherein an offset audio object position p″ 321 of the audio object position 302 is determined in Cartesian coordinates (x, y, z) based on:
[0193] x = r·sin(el′)·cos(az′) + x offset
[0194] y = r·sin(el′)·sin(az′) + y offset
[0195] z = r·cos(el′) + z offset
[0196] wherein the Cartesian position (x, y, z) consists of an x parameter, a y parameter, and a z parameter, and wherein x offset relates to a first x-axis offset parameter, y offset relates to a first y-axis offset parameter, and z offset relates to a first z-axis offset parameter.
[0197] The tenth EEE relates to the method of the ninth EEE, wherein the parameter x offset , y offset , and z offset are based on the following:
[0198] x offset = r offset ·sin(el′ offset )·cos(az′ offset )
[0199] y offset = r offset ·sin(el′ offset )·sin(az′ offset )
[0200] z offset = r offset ·cos(el′ offset )
[0201] The eleventh EEE relates to the method of the seventh EEE, wherein the azimuth parameter az offset relates to a scene displacement azimuth position and is based on the following:
[0202] az offset = (sd_azimuth - 128)·1.5
[0203] az offset = min(max(az offset , -180), 180)
[0204] where sd_azimuth is an azimuth metadata parameter indicating the MPEG-H 3DA azimuth scene displacement, and wherein the elevation parameter el offset relates to a scene displacement elevation position and is based on the following:
[0205] el offset = (sd_elevation - 32)·3
[0206] el offset= min(max(el offset , -90), 90)
[0207] where sd_elevation is an elevation metadata parameter indicating the elevation scene displacement of MPEG-H 3DA, and where the radius parameter r offset relates to the scene displacement radius and is based on the following:
[0208] r offset = (sd_radius + 1) / 16
[0209] where sd_radius is a radius metadata parameter indicating the radius scene displacement of MPEG-H 3DA, and where the parameters X and Y are scalar variables.
[0210] The twelfth EEE relates to the method of the tenth EEE, where the x offset parameter relates to the scene displacement offset sd_x in the x-axis direction; the y offset parameter relates to the scene displacement offset sd_y in the y-axis direction; and the z offset parameter relates to the scene displacement offset sd_z in the z-axis direction.
[0211] The thirteenth EEE relates to the method of the first EEE, the method further comprising interpolating, by the audio decoding device, the first position data associated with the listening location data 301 and the object audio signal 302 at an update rate.
[0212] The fourteenth EEE relates to the method of the first EEE, the method further comprising determining, by the audio decoding device 300, the efficient entropy coding of the listening location data 301.
[0213] The fifteenth EEE relates to the method of the first EEE, where the position data associated with the listening location data 301 is derived based on sensor information.
Claims
1. A method for processing position information indicating the position of an audio object, wherein the processing is performed using an MPEG-H 3D audio decoder, and wherein the object position can be used to render the audio object, the method comprising: obtaining listener orientation information indicating the orientation of the listener's head, wherein the listener orientation information includes information about the yaw, pitch, and roll of the listener's head; obtaining listener displacement information indicating the displacement of the listener's head relative to a nominal listening position; determining the object position based on the position information; modifying the object position by applying a translation to the object position based on the listener displacement information; and further modifying the modified object position based on the listener orientation information, wherein when the listener displacement information indicates that the listener's head is displaced from the nominal listening position by a certain small position displacement, the absolute value of the certain small position displacement is 0.5 meters or less than 0.5 meters, and after the displacement of the listener's head, the distance between the modified audio object position and the listening position remains equal to the original distance between the audio object position and the nominal listening position.
2. The method according to claim 1, wherein: modifying the object position and further modifying the modified object position are performed such that after rendering to one or more real speakers or virtual speakers according to the further modified object position, the audio object is perceived by the listener psychoacoustically as originating from a position fixed relative to the nominal listening position, regardless of the displacement of the listener's head from the nominal listening position and the orientation of the listener's head relative to the nominal orientation.
3. The method according to claim 1, wherein: modifying the object position based on the listener displacement information is performed by translating the object position by the same displacement as the listener's head from the nominal listening position, but in the opposite direction to the displacement of the listener's head from the nominal listening position.
4. The method according to any one of claims 1 to 3, wherein: the listener displacement information indicates the displacement of the listener's head from the nominal listening position, and the displacement can be achieved by the listener moving their upper body and / or head.
5. The method according to any one of claims 1 to 3, further comprising: detecting the orientation of the listener's head by a wearable device and / or a fixed device.
6. The method according to any one of claims 1 to 3, further comprising: detecting the displacement of the listener's head from the nominal listening position by a wearable device and / or a fixed device.
7. The method according to any one of claims 1 to 3, wherein the distance between the modified audio object position and the listening position after displacement is mapped to a gain to modify the audio level.
8. An MPEG-H 3D audio decoder for processing position information indicating the object positions of audio objects, wherein the object positions can be used to render the audio objects, the decoder comprising a processor and a memory coupled to the processor, wherein the processor is adapted to: obtain listener orientation information indicating the orientation of the listener's head, wherein the listener orientation information includes information about the yaw, pitch, and roll of the listener's head; obtain listener displacement information indicating the displacement of the listener's head relative to a nominal listening position; determine the object positions based on the position information; modify the object positions by applying a translation to the object positions based on the listener displacement information; and further modify the modified object positions based on the listener orientation information, wherein when the listener displacement information indicates that the listener's head is displaced from the nominal listening position by a certain small positional displacement, the absolute value of the certain small positional displacement is 0.5 meters or less than 0.5 meters, the processor is configured to, after the listener's head is displaced, maintain the distance between the modified audio object position and the listening position equal to the original distance between the audio object position and the nominal listening position.
9. A computer storage medium comprising instructions that, when executed by a digital signal processor or a microprocessor, cause the digital signal processor or the microprocessor to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-dimensional parametric audio system and method
CN104737557A
Methods, apparatus and systems for three degrees of freedom (3dof+) extension of mpeg-h 3D audio
CN111886880A
Methods, apparatus and systems for three degrees of freedom (3DOF+) extension of MPEG-h 3D audio
CN113993058A
Methods, apparatus and systems for three degrees of freedom (3DOF+) extension of MPEG-h 3D audio
CN113993059A
Methods, apparatus and systems for three degrees of freedom (3DOF+) extension of MPEG-h 3D audio
CN113993060A