Method, apparatus, and system for 6DOF audio rendering, data representation, and bitstream structure for 6DOF audio rendering

The method of encoding 3DoF and 6DoF audio signals in separate bitstream portions addresses the lack of 6DoF audio rendering capabilities, achieving efficient and compatible audio output across different systems.

JP7704330B2Active Publication Date: 2025-07-08DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024000945
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-04-11
Filing Date
2024-01-09
Publication Date
2025-07-08
Estimated Expiration
2039-04-09

AI Technical Summary

Technical Problem

Current audio rendering technologies lack the capability to efficiently process audio signals for six degrees of freedom (6DoF) movement, which includes yaw, pitch, roll, and translational movement, while maintaining compatibility with three degrees of freedom (3DoF) systems.

Method used

A method and apparatus for encoding and decoding audio signals that separate 3DoF and 6DoF data into distinct bitstream portions, allowing for efficient 6DoF audio rendering with backward compatibility to 3DoF systems, using a bitstream structure that encapsulates 6DoF metadata within a 3DoF format, such as MPEG-H 3DA, to ensure consistent audio output across different systems.

Benefits of technology

Enables efficient 6DoF audio rendering with improved bitrate and audio quality, ensuring seamless compatibility with existing 3DoF systems by using a single bitstream that can be decoded by both 3DoF and 6DoF systems, maintaining audio fidelity and reducing redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704330000010
    Figure 0007704330000010
  • Figure 0007704330000011
    Figure 0007704330000011
  • Figure 0007704330000012
    Figure 0007704330000012
Patent Text Reader

Abstract

To provide methods, apparatus, and systems for 6DOF audio rendering and data representations and bitstream structures for 6DOF audio rendering.SOLUTION: The present disclosure relates to methods, apparatus and systems for encoding an audio signal into a bitstream, in particular at an encoder, including: encoding or including audio signal data associated with 3DoF audio rendering into one or more first bitstream parts of the bitstream; and encoding or including metadata associated with 6DoF audio rendering into one or more second bitstream parts of the bitstream. The present disclosure further relates to methods, apparatus, and systems for decoding an audio signal and audio rendering based on the bitstream.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the benefit of U.S. Provisional Application No. 62 / 655,990, filed Apr. 11, 2018, which is hereby incorporated by reference in its entirety.

[0002] Technical Field The present disclosure relates to apparatuses, systems, and methods for six degrees of freedom (6DoF) audio rendering, particularly in connection with data representations and bitstream structures for 6DoF audio rendering.

Background Art

[0003] Currently, there is no adequate solution for rendering audio in combination with a user's six degrees of freedom (6DoF) movement. There are solutions for rendering channel signals, object signals, and first / higher order ambisonics (HOA) signals in combination with three degrees of freedom (3DoF) movement (yaw, pitch, roll), but there is no support for processing such signals in combination with a user's six degrees of freedom (6DoF) movement (yaw, pitch, roll, and translational movement).

[0004] Generally, 3DoF audio rendering provides an acoustic field in which one or more audio sources are rendered at angular positions surrounding a predetermined listener position (referred to as a 3DoF position). An example of 3DoF audio rendering is included in the MPEG-H 3D Audio standard (abbreviated MPEG-H 3DA).

[0005] MPEG-H 3DA was developed to support channel signals, object signals, and HOA signals for 3DoF, but it still cannot process true 6DoF audio. The envisioned MPEG-I 3D audio implementation preferably provides backward compatibility for 3DoF rendering while extending 3DoF (and 3DoF+) capabilities to 6DoF 3D audio devices in an efficient manner (preferably including efficient signal generation, encoding, decoding, and / or rendering).

Summary of the Invention

Problems to be Solved by the Invention

[0006] In view of the above, an object of the present disclosure is to provide a method, apparatus, data representation, and / or bitstream structure for 3D audio encoding and / or 3D audio rendering that allows for efficient 6DoF audio encoding and / or rendering, preferably together with backward compatibility for 3DoF audio rendering based on, for example, the MPEG-H 3DA standard.

[0007] Another object of the present disclosure may be to provide a data representation and / or bitstream structure for 3DoF audio encoding and / or 3D audio rendering that allows for efficient 6DoF audio encoding and / or rendering, preferably together with backward compatibility for 3DoF audio rendering based on, for example, the MPEG-H 3DA standard, and / or to provide an encoding and / or rendering apparatus for efficient 6DoF audio encoding and / or rendering, preferably together with backward compatibility for 3DoF audio rendering based on, for example, the MPEG-H 3DA standard.

Means for Solving the Problems

[0008] According to exemplary aspects, a method for encoding an audio signal into a bitstream, particularly in an encoder, includes encoding and / or including audio signal data related to 3DoF audio rendering into one or more first bitstream portions of the bitstream; and / or encoding and / or including metadata related to 6DoF audio rendering into one or more second bitstream portions of the bitstream.

[0009] According to exemplary aspects, the audio signal data related to 3DoF audio rendering includes the audio signal data of one or more audio objects.

[0010] According to exemplary aspects, the one or more audio objects are located on one or more spheres surrounding a default 3DoF listener position.

[0011] According to exemplary aspects, the audio signal data related to 3DoF audio rendering includes direction data of one or more audio objects and / or distance data of one or more audio objects.

[0012] According to exemplary aspects, the metadata related to 6DoF audio rendering indicates one or more default 3DoF listener positions.

[0013] According to exemplary aspects, the metadata related to 6DoF audio rendering includes or indicates at least one of: a description of a 6DoF space optionally including object coordinates; the audio object direction of one or more audio objects; a virtual reality (VR) environment; and / or parameters related to distance attenuation, occlusion, and / or reverberation.

[0014] According to exemplary aspects, the method may further include: receiving an audio signal from one or more audio sources; and / or generating audio signal data related to 3DoF audio rendering based on the audio signal and a transfer function from the one or more audio sources.

[0015] According to exemplary aspects, audio signal data related to 3DoF audio rendering is generated by converting the audio signal from the one or more audio sources to a 3DoF audio signal using the transfer function.

[0016] According to exemplary aspects, the transfer function maps or projects the audio signal of the one or more audio sources to respective audio objects located on one or more spheres surrounding a default 3DoF listener position.

[0017] According to exemplary aspects, the method may further include determining a parameterization of the transfer function based on environmental characteristics and / or parameters related to distance attenuation, occlusion, and / or reverberation.

[0018] According to exemplary aspects, the bitstream is an MPEG-H 3D Audio bitstream or a bitstream using the MPEG-H 3D Audio syntax.

[0019] According to exemplary aspects, the one or more first bitstream portions of the bitstream represent the payload of the bitstream and / or the one or more second bitstream portions represent one or more extension containers of the bitstream.

[0020] According to yet another exemplary aspect, a method for decoding and / or audio rendering, particularly in a decoder or a renderer, may be provided. The method includes: receiving a bitstream, wherein the bitstream includes audio signal data related to 3DoF audio rendering in one or more first bitstream portions of the bitstream and further includes metadata related to 6DoF audio rendering in one or more second bitstream portions of the bitstream, and / or performing at least one of 3DoF audio rendering and 6DoF audio rendering based on the received bitstream.

[0021] According to exemplary aspects, when performing 3DoF audio rendering, the 3DoF audio rendering is performed based on the audio signal data related to 3DoF audio rendering in the one or more first bitstream portions of the bitstream, while the metadata related to 6DoF audio rendering in the one or more second bitstream portions of the bitstream is discarded.

[0022] According to exemplary aspects, when performing 6DoF audio rendering, the 6DoF audio rendering is performed based on the audio signal data related to 3DoF audio rendering in the one or more first bitstream portions of the bitstream and the metadata related to 6DoF audio rendering in the one or more second bitstream portions of the bitstream.

[0023] According to exemplary aspects, the audio signal data related to 3DoF audio rendering includes the audio signal data of one or more audio objects.

[0024] According to exemplary aspects, the one or more audio objects are positioned on one or more spheres surrounding a default 3DoF listener position.

[0025] According to exemplary aspects, audio signal data related to 3DoF audio rendering includes direction data of the one or more audio objects and / or distance data of the one or more audio objects.

[0026] According to exemplary aspects, metadata related to 6DoF audio rendering indicates one or more default 3DoF listener positions.

[0027] According to exemplary aspects, metadata related to 6DoF audio rendering includes or indicates at least one of: a description of a 6DoF space optionally including object coordinates; an audio object direction of the one or more audio objects; a virtual reality (VR) environment; and / or parameters related to distance attenuation, occlusion, and / or reverberation.

[0028] According to exemplary aspects, audio signal data related to 3DoF audio rendering is generated based on the audio signals from the one or more audio sources and a conversion function.

[0029] According to exemplary aspects, audio signal data related to 3DoF audio rendering is generated by converting the audio signals from the one or more audio sources into 3DoF audio signals using the conversion function.

[0030] According to exemplary aspects, the conversion function maps or projects the audio signals of the one or more audio sources to respective audio objects positioned on one or more spheres surrounding a default 3DoF listener position.

[0031] According to exemplary aspects, the bitstream is an MPEG-H 3D Audio bitstream or a bitstream using MPEG-H 3D Audio syntax.

[0032] According to exemplary aspects, the one or more first bitstream portions of the bitstream represent the payload of the bitstream and / or the one or more second bitstream portions of the bitstream represent one or more extension containers of the bitstream.

[0033] According to exemplary aspects, performing 6DoF audio rendering based on audio signal data related to 3DoF audio rendering in the one or more first bitstream portions of the bitstream and metadata related to 6DoF audio rendering in the one or more second bitstream portions of the bitstream includes generating audio signal data related to 6DoF audio rendering based on the audio signal data related to 3DoF audio rendering and an inverse conversion function.

[0034] According to exemplary aspects, the audio signal data related to 6DoF audio rendering is generated by converting the audio signal data related to 3DoF audio rendering using the inverse conversion function and the metadata related to 6DoF audio rendering.

[0035] According to exemplary aspects, the inverse conversion function is the inverse function of a conversion function that maps or projects the audio signals of the one or more audio sources to respective audio objects located on one or more spheres surrounding a default 3DoF listener position.

[0036] According to exemplary aspects, performing 3DoF audio rendering based on audio signal data related to 3DoF audio rendering in the one or more first bitstream portions of the bitstream results in the same generated sound field as performing 6DoF audio rendering at a default 3DoF listener position, based on the audio signal data related to 3DoF audio rendering in the one or more first bitstream portions of the bitstream and metadata related to 6DoF audio rendering in the one or more second bitstream portions of the bitstream.

[0037] According to yet another exemplary aspect, a bitstream for audio rendering may be provided. The bitstream includes audio signal data related to 3DoF audio rendering in one or more first bitstream portions of the bitstream, and further includes metadata related to 6DoF audio rendering in one or more second bitstream portions of the bitstream. This aspect may be combined with any one or more of the above exemplary aspects.

[0038] According to yet another exemplary aspect, there may be provided an apparatus, particularly an encoder, comprising a processor configured to encode and / or include audio signal data related to 3DoF audio rendering in one or more first bitstream portions of a bitstream, encode and / or include metadata related to 6DoF audio rendering in one or more second bitstream portions of the bitstream, and / or output the encoded bitstream. This aspect may be combined with any one or more of the above exemplary aspects.

[0039] According to yet another exemplary aspect, an apparatus, particularly a decoder or an audio renderer, comprising: a processor configured to receive a bitstream that includes audio signal data related to 3DoF audio rendering in one or more first bitstream portions of the bitstream and further includes metadata related to 6DoF audio rendering in one or more second bitstream portions of the bitstream, and / or to perform at least one of 3DoF audio rendering and 6DoF audio rendering based on the received bitstream may be provided. This aspect may be combined with any one or more of the above exemplary aspects.

[0040] According to exemplary aspects, when performing 3DoF audio rendering, the processor is configured to perform 3DoF audio rendering based on the audio signal data related to 3DoF audio rendering in the one or more first bitstream portions of the bitstream, while discarding the metadata related to 6DoF audio rendering in the one or more second bitstream portions of the bitstream.

[0041] According to exemplary aspects, when performing 6DoF audio rendering, the processor is configured to perform 6DoF audio rendering based on the audio signal data related to 3DoF audio rendering in the one or more first bitstream portions of the bitstream and the metadata related to 6DoF audio rendering in the one or more second bitstream portions of the bitstream.

[0042] According to yet another exemplary aspect, particularly in an encoder, there may be provided a non - transitory computer program product including instructions which, when executed by a processor, cause the processor to execute a method of encoding an audio signal into a bitstream. The method includes: encoding or including audio signal data associated with 3DoF audio rendering into one or more first bitstream portions of the bitstream; and / or encoding or including metadata associated with 6DoF audio rendering into one or more second bitstream portions of the bitstream. This aspect may be combined with any one or more of the above - mentioned exemplary aspects.

[0043] According to yet another exemplary aspect, particularly in a decoder or an audio renderer, there may be provided a non - transitory computer program product including instructions which, when executed by a processor, cause the processor to execute a method for decoding and / or audio rendering. The method includes receiving a bitstream including audio signal data associated with 3DoF audio rendering in one or more first bitstream portions of the bitstream and further including metadata associated with 6DoF audio rendering in one or more second bitstream portions of the bitstream, and / or performing at least one of 3DoF audio rendering and 6DoF audio rendering based on the received bitstream. This aspect may be combined with any one or more of the above - mentioned exemplary aspects.

[0044] A further aspect of the present disclosure relates to corresponding computer programs and computer - readable storage media.

[0045] It will be understood that the method steps and apparatus features may be interchanged in many ways. In particular, the details of the disclosed method may, as will be understood by those skilled in the art, be implemented as an apparatus adapted to perform some or all or steps of the method, and vice versa. In particular, it is understood that each description made with respect to the method equally applies to the corresponding apparatus, and vice versa.

Brief Description of the Drawings

[0046] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings. Like reference numerals may indicate like or similar elements.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 7C

Figure 8

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0047] Hereinafter, with reference to the accompanying drawings, preferred exemplary aspects will be described in more detail. The same or similar features in different drawings and embodiments may be referred to by the same reference numerals. It should be understood that the following detailed description of various preferred exemplary aspects is not intended to limit the scope of the present invention.

[0048] As used in this document, "MPEG-H 3D Audio" refers to the specifications standardized in ISO / IEC 23008-3 and / or any past and / or future amendments, editions, or other versions of the ISO / IEC 23008-3 standard.

[0049] As used in this document, an MPEG-I 3D audio implementation desirably extends 3DoF (and 3DoF+) functionality towards 6DoF 3D audio while preferably providing 3DoF rendering backward compatibility.

[0050] As used in this document, 3DoF is a system that can typically handle the movement of a user's head, particularly head rotation, specified by three parameters (e.g., yaw, pitch, roll). Such a system is often available in various gaming systems, such as virtual reality (VR) / augmented reality (AR) / mixed reality (MR) systems, or other such types of acoustic environments.

[0051] As used in this document, 6DoF is a system that can typically handle 3DoF and translational movement correctly.

[0052] Exemplary aspects of the present disclosure relate to an audio system (e.g., an audio system compatible with the MPEG-I audio standard), where the audio renderer extends functionality towards 6DoF by converting relevant metadata into a 3DoF format, such as an audio renderer input format compatible with the MPEG standard (e.g., the MPEG-H 3DA standard).

[0053] Figure 1 shows an exemplary system 100 configured to use metadata extension and / or audio renderer extension in addition to an existing 3DoF system to enable a 6DoF experience. The system 100 includes an original environment 101 (which may include one or more audio sources 101a by way of example), a content format 102 (e.g., a bitstream including 3D audio data), an encoder 103, and a proposed metadata encoder extension 106. The system 100 may also include a 3D audio renderer 105 (e.g., a 3DoF renderer) and a proposed renderer extension 107 (e.g., a 6DoF renderer extension for the reproduced environment 108).

[0054] In a method of 3D audio rendering by 3DoF, only the angles of the user's angular orientation (e.g., yaw angle y, pitch angle p, roll angle r) at a given 3DoF position can be input to the 3DoF audio renderer 105. With the extended 6DoF functionality, the user's position coordinates (e.g., x, y, and z) can additionally be input to a 6DoF audio renderer (extended renderer).

[0055] Advantages of the present disclosure include bitrate improvement for a bitstream transmitted between an encoder and a decoder. The bitstream may be encoded and / or decoded in accordance with a standard, e.g., the MPEG-I Audio standard and / or the MPEG-H 3D Audio standard, or at least be backward compatible with a standard such as the MPEG-H 3D Audio standard.

[0056] In some examples, exemplary aspects of the present disclosure are directed to the processing of a single bitstream (e.g., an MPEG-H 3D Audio (3DA) bitstream (BS), or a bitstream using the syntax of an MPEG-H 3DA BS) that is compatible with multiple systems.

[0057] For example, in some exemplary aspects, an audio bitstream may be compatible with two or more different renderers, such as a 3DoF audio renderer that may be compatible with a certain standard (e.g., the MPEG-H 3D Audio standard) and a newly defined 6DoF audio renderer or renderer extension that may be compatible with a second different standard (e.g., the MPEG-I Audio standard).

[0058] Exemplary aspects of the present disclosure are preferably directed to different decoders configured to perform decoding and rendering of the same audio bitstream to generate the same audio output.

[0059] For example, exemplary aspects of the present disclosure relate to 3DoF decoders and / or 3DoF renderers and / or 6DoF decoders and / or 6DoF renderers configured to generate the same output for the same bitstream (e.g., a 3D ABS or a bitstream using 3D ABS). As an example, the bitstream may include information regarding defined positions of a listener in a VR / AR / MR (virtual reality / augmented reality / mixed reality) space, for example as part of 6DoF metadata.

[0060] The present disclosure also relates, by way of example, to an encoder and / or a decoder configured to encode and / or decode 6DoF information respectively (e.g., compatible with an MPEG-I Audio environment). Here, the encoder and / or decoder of the present disclosure provides one or more of the following advantages: · Good representation of VR / AR / MR-related audio data quality and bitrate efficiency, and its encapsulation into an audio bitstream syntax (e.g., MPEG-H 3D Audio BS); · Backward compatibility between different systems (e.g., the MPEG-H 3D Audio standard and the contemplated MPEG-I Audio standard).

[0061] Backward compatibility is highly beneficial in order to preferably avoid competition between 3DoF and 6DoF solutions and provide a smooth transition between current and future technologies.

[0062] For example, backward compatibility between 3DoF audio systems and 6DoF audio systems is highly beneficial, and in a 6DoF audio system such as MPEG-I Audio, backward compatibility to a 3DoF audio system such as MPEG-H 3D Audio is provided.

[0063] According to exemplary aspects of the present disclosure, this can be achieved by providing backward compatibility, for example at the bitstream level, for a 6DoF related system consisting of: · Encoded data of 3DoF audio material and related metadata; and · 6DoF related metadata The exemplary aspects of the present disclosure relate to a standard 3DoF bitstream syntax that encapsulates 6DoF bitstream elements, such as a first type of audio bitstream (e.g., MPEG-H 3D ABS) syntax. Such 6DoF bitstream elements are MPEG-I Audio bitstream elements within one or more extension containers of a first type of audio bitstream (e.g., MPEG-H 3D ABS).

[0064] To provide a system that guarantees backward compatibility at the performance level, the following systems and / or structures may be significant and may exist:

[0065] The following systems and / or structures may be significant and may exist to provide a system that guarantees backward compatibility at the performance level: 1a. A 3DoF system (e.g., a system compatible with the MPEG-H 3D Audio standard) must be able to ignore all 6DoF-related syntax elements (e.g., ignore MPEG-I Audio bitstream syntax elements based on the functionality of "mpegh3daExtElementConfig()" or "mpegh3daExtElement()" in the MPEG-H 3D Audio bitstream syntax). That is, the 3DoF system (decoder / renderer) may preferably be configured to ignore additional 6DoF-related data and / or metadata (e.g., by not reading 6DoF-related data and / or metadata); 2a. The remaining part of the bitstream payload (e.g., an MPEG-I Audio bitstream payload containing data and / or metadata compatible with an MPEG-H 3D Audio bitstream parser) must be decodable by a 3DoF system (e.g., a legacy MPEG-H 3D Audio system) to generate the desired audio output. That is, the 3DoF system (decoder / renderer) may preferably be configured to decode the 3DoF part of the BS; 3a. A 6DoF system (e.g., an MPEG-I Audio system) must be able to process both the 3DoF-related and 6DoF-related parts of the audio bitstream and generate an audio output that matches the audio output of the 3DoF system (e.g., an MPEG-H 3D Audio system) at a predefined backward-compatible 3DoF position(s) in the VR / AR / MR space. That is, the 6DoF system (decoder / renderer) may preferably be configured to render an audio field / audio output that matches the 3DoF-rendered audio field / audio output at the default 3DoF position(s); A 4a.6DoF system (e.g., an MPEG-I Audio system) provides a smooth change (transition) of the audio output around a predefined backward-compatible 3DoF position(s) (i.e., provides a continuous sound field in a 6DoF space). That is, a 6DoF system (decoder / renderer) may be configured to render a sound field / audio output that smoothly transitions to a 3DoF-rendered sound field / audio output at a default 3DoF position(s) around the default 3DoF position(s).

[0066] In some examples, the present disclosure relates to providing a 6DoF audio renderer (e.g., an MPEG-I audio renderer) that generates the same audio output as a 3DoF audio renderer (e.g., an MPEG-H 3D Audio renderer) at one, more than one, or several 3DoF positions.

[0067] Currently, when directly transferring 3DoF-related audio signals and metadata to a 6DoF audio system, there are the following drawbacks: 1. Increase in bitrate (i.e., in addition to 6DoF-related audio signals and metadata, 3DoF-related audio signals and metadata are transmitted); 2. Limited effectiveness (i.e., 3DoF-related audio signal(s) and metadata are only effective for the 3DoF position(s)).

[0068] Exemplary aspects of the present disclosure relate to overcoming the above drawbacks.

[0069] In some examples, the present disclosure is directed to: 1. Use the 1.3DoF interchangeable audio signal(s) and metadata (e.g., signals and metadata compatible with MPEG-H 3D Audio) instead of (or as supplementary addition to) the original audio source signal and metadata; and / or 2. Increase the application range (for use in 6DoF rendering) from 3DoF positions to a 6DoF space (defined by the content producer) while maintaining a high level of sound field approximation.

[0070] Exemplary aspects of the present disclosure are directed to efficiently generating, encoding, decoding, and rendering such signal(s) to achieve these goals and to provide a 6DoF rendering function.

[0071] FIG. 2 shows an exemplary floor plan 202 of an exemplary room 201. As shown in FIG. 2, an exemplary listener stands in the center of a room having several audio sources and non-trivial wall geometry. In a 6DoF device (e.g., a system providing provisions for 6DoF functionality), the exemplary listener can move around, but in some instances, the default 3DoF position 206 may be assumed to correspond to the intended area for the best VR / AR / MR audio experience (e.g., by the content producer's settings or intentions).

[0072] Specifically, FIG. 2 shows a wall 203, a 6DoF space 204, an exemplary (optional) directional vector 205 (e.g., when one or more sound sources emit sound directionally), a 3DoF listener position 206 (the default 3DoF position 206), and an audio source 207 shown exemplarily as a star in FIG. 2.

[0073] FIG. 3 shows an exemplary 6DoF VR / AR / MR scene such as, for example, FIG. 2, as well as an audio object (audio data + metadata) 320 included in a 3DoF audio bitstream 302 (e.g., an MPEG-H 3D Audio bitstream), and an extended container 303. The audio bitstream 302 and the extended container 303 may be encoded via a device or system compatible with the MPEG standard (e.g., MPEG-H or MPEG-I) (e.g., via software, hardware, or the cloud).

[0074] Exemplary aspects of the present disclosure relate to reproducing a sound field at a “3DoF position” in a manner corresponding to a 3DoF audio renderer output signal (which may or may not be consistent with the physical law of sound propagation) when using a 6DoF audio renderer (e.g., an MPEG-I Audio renderer). This sound field is preferably based on the original “audio source” and should reflect the influence of the complex geometry of the corresponding VR / AR / MR environment (e.g., effects such as “walls”, structures, sound reflections, reverberations, and / or occlusions).

[0075] Exemplary aspects of the present disclosure relate to parameterization by an encoder of all relevant information describing this scenario in such a way as to ensure that one, several, or preferably all of the corresponding requirements (1a)-(4a) above are met.

[0076] Such an approach is not optimal because it requires the following when two audio rendering modes (i.e., 3DoF and 6DoF) are executed in parallel and an interpolation algorithm is applied to the corresponding output in 6DoF space: · Parallel execution of two different rendering algorithms (i.e., one for a specific 3DoF position and the other for 6DoF space); · A large amount of audio data (for transferring additional audio data for the 3DoF Audio renderer).

[0077] Exemplary aspects of the present disclosure preferably avoid the above drawbacks in that only a single audio rendering mode is executed (instead of parallel execution of, for example, two audio rendering modes) and / or preferably 3DoF audio data, together with additional metadata for restoring and / or approximating the original sound source(s) signal(s), is used for 6DoF audio rendering (instead of transmitting, for example, 3DoF Audio data and the original sound source data).

[0078] Exemplary aspects of the present disclosure relate to (1) a single 6DoF audio rendering algorithm (e.g., compatible with MPEG-I Audio) that preferably produces exactly the same output as a 3DoF audio rendering algorithm (e.g., compatible with MPEG-H 3DA) at a particular location(s), and / or (2) representing audio (e.g., 3DoF audio data) and 6DoF-related audio metadata so as to minimize redundancy in the 3DoF-related and VR / AR / MR-related portions of 6DoF audio bitstream data (e.g., MPEG-I audio bitstream data).

[0079] Exemplary aspects of the present disclosure use a bitstream (e.g., MPEG-H 3DA BS) syntax of a first standardized format to encapsulate a bitstream (future standard, e.g., MPEG-I) of a second standardized format or a part thereof and 6DoF-related metadata: ·When decoded by a preferably 3DoF audio system, transfer the audio source signal and metadata (e.g., in the core part of the 3DoF audio bitstream syntax) that preferably approximate the desired sound field well enough at the preferably (default) 3DoF position(s); ·Transfer the 6DoF related metadata and / or further data (e.g., parametric and / or signal data) (e.g., in the extended part of the 3DoF audio bitstream syntax) used to approximate (restore) the original audio source signal for 6DoF audio rendering Regarding this.

[0080] One aspect of the present disclosure relates to the determination of the desired "3DoF position(s)" and signals compatible with a 3DoF audio system (e.g., the MPEG-H 3DA system) on the encoder side.

[0081] For example, as shown in relation to FIG. 3, the virtual 3DA object signal for 3DA can generate the same sound field at a specific 3DoF position (based on signal x 3DA ). Since some 3DoF systems (e.g., the MPEG-H 3DA system) cannot incorporate the effects of the VR / AR / MR environment (e.g., occlusion, reverberation, etc.), the sound field should preferably include the effects of the VR environment for a specific 3DoF position(s) ("wet" signal). The methods and processes shown in FIG. 3 can be implemented via various systems and / or products.

[0082] Inverse function A -1 Should preferably "dewet" these signals (i.e., remove the influence of the VR environment) in some exemplary aspects. It should be good because it is necessary to approximate the original "dry" signal (without the effects of the VR environment).

[0083] Audio signal for 3DoF rendering ((x 3DA )) is preferably defined, for example, based on the following, to provide the same / similar output for both 3DoF and 6DoF audio rendering:

Number

[0084] BS may include information regarding object signals, object directions, and object distances.

[0085] Figure 3 further illustratively shows an extended container 303 that may include extended metadata within the BS, for example. The extended container 303 of the BS may include the following metadata: (i) 3DoF (default) position parameters; (ii) 6DoF spatial description parameters (object coordinates); (iii) (optional) object directivity parameters; (iv) (optional) VR / AR / MR environment parameters; and / or (v) at least one of (optional) distance attenuation parameters, occlusion parameters, and / or reverberation parameters, etc.

[0086] There may be an approximation of the desired audio rendering based on the following:

Number

[0087] Additionally or optionally, smoothness for 6DoF audio renderer (for example, MPEG-I audio renderer) output may preferably be provided based on the following:

Number

[0088] Exemplary aspects of the present disclosure are directed to defining an encoder-side 3DoF audio object (e.g., an MPEG-H 3DA object), preferably based on the following:

Number

[0089] One aspect of the present disclosure relates to recovering the original object on a decoder based on the following:

Number

[0090] The approximated sound source / object signal is preferably regenerated using a 6DoF audio renderer at the "3DoF position" in a manner corresponding to the 3DoF audio renderer output signal.

[0091] The sound source / object signal is preferably approximated based on the sound field that reflects the influence of the complex geometry (such as "walls", structures, reverberations, occlusions, etc.) of the corresponding VR / AR / MR environment, based on the original "audio source".

[0092] That is, the virtual 3D object signal for 3D A preferably generates the same sound field including the effects of the VR environment at a specific 3DoF position (based on signal x 3DA ) for the specific 3DoF position(s).

[0093] On the rendering side, the following may be available (for example, for a decoder compliant with a standard such as the MPEG-H or MPEG-I standard): · Audio signal(s) for 3DoF audio rendering: x 3DA · Rendering function for either 3DoF or 6DoF audio: F 3DoF (x 3DA ) or F 6DoF (x) Equation (6)

[0094] For 6DoF audio rendering, additionally, there may be 6DoF metadata available on the rendering side for the 6DoF audio rendering function (for example, to approximate / recover the audio signal x of the one or more audio sources based on the 3DoF audio signal and 6DoF metadata).

[0095] Exemplary aspects of the present disclosure relate to (i) the definition of 3DoF audio objects (such as MPEG-H 3D objects), and / or (ii) the restoration (approximation) of the original audio object.

[0096] The audio object may be included, for example, in a 3DoF audio bitstream (such as an MPEG-H 3D BS).

[0097] The bitstream may include information regarding an object audio signal, object direction, and / or object distance.

[0098] An extended container (such as a bitstream of MPEG-H 3D ABS) may include the following metadata: (i) 3DoF (default) position parameters; (ii) 6DoF spatial description parameters (object coordinates); (iii) (optional) object directivity parameters; (iv) (optional) VR / AR / MR environment parameters; and / or (v) at least one of (optional) distance attenuation parameters, occlusion parameters, reverberation parameters, etc.

[0099] The present disclosure may provide the following advantages: · For 3DoF audio decoding and rendering (such as MPEG-H 3D decoding and rendering) Rearward compatibility : The 6DoF audio renderer (such as an MPEG-I audio renderer) output corresponds to the 3DoF rendering output of a 3DoF rendering engine (such as an MPEG-H 3D rendering engine, etc.) for a given 3DoF position(s). · Symbolization efficiency : For this approach, the legacy 3DoF audio bitstream syntax (such as the MPEG-H 3D bitstream syntax) structure can be efficiently reused. · At a given (3DoF) position(s) Audio quality control : The best perceptual audio quality can be explicitly guaranteed by the encoder for any position(s) and corresponding 6DoF space.

[0100] Exemplary aspects of the present disclosure may relate to the following signaling in a format compatible with an MPEG standard (such as the MPEG-I standard) bitstream: ·Tacit 3DoF audio - system (e.g., MPEG - H 3DA) compatibility signaling via an extended container mechanism (e.g., MPEG - H 3D ABS). This enables a 6DoF audio (e.g., MPEG - I Audio - compatible) processing algorithm to restore the original audio - object signal. ·Parameterization to describe data for approximation of the original audio - object signal.

[0101] A 6DoF audio renderer can specify how to restore the original audio - object signal in, for example, an MPEG - compatible system (e.g., an MPEG - I audio system).

[0102] This proposed concept is: ·General with respect to the definition of the approximation function (i.e., A(x)); ·It may be arbitrarily complex, but on the decoder side, there should exist a corresponding approximation (i.e., ∃A -1 ); ·Approximately, mathematically "well - defined" (e.g., algorithmically stable, etc.); ·General with respect to the type of the approximation function (i.e., A(x)); ·The approximation function may be based on the following approximation types or any combination of these approaches (listed in ascending order of bit - rate consumption): - Parameterized audio effects applied to the signal x 3DA (e.g., parametrically controlled levels, reverberation, reflections, masking, etc.) - Parametrically - encoded modifications (e.g., time / frequency variant modification gains for the transmitted signal x 3DA ) - Signal - encoding modifications (e.g., an encoded signal that approximates the residual waveform (x - x 3DA )) · General sound fields and sound source representations (and combinations thereof): Extensible and applicable to objects, channels, FOA, HOA.

[0103] A in FIG. 6 schematically shows an exemplary data representation and / or bitstream structure according to exemplary aspects of the present disclosure. The data representation and / or bitstream structure may be encoded via a device or system (e.g., software, hardware, or cloud) compatible with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0104] The bitstream BS includes, as an example, a first bitstream portion 302 that includes 3DoF-encoded audio data (e.g., in a main or core portion of the bitstream). Preferably, the bitstream syntax of the bitstream BS is compatible with or conforms to a BS syntax for 3DoF audio rendering, such as the MPEG-H 3DA bitstream syntax. The 3DoF-encoded audio data may be included as a payload in one or more packets of the bitstream BS.

[0105] As previously described, for example, in connection with FIG. 3 above, the 3DoF-encoded audio data may include audio object signals of one or more audio objects (e.g., on a sphere around a default 3DoF position). For directional audio objects, the 3DoF-encoded audio data may optionally further include an object direction and / or may optionally further indicate an object distance (e.g., by use of a gain and / or one or more attenuation parameters).

[0106] As an example, the BS includes, illustratively, a second bitstream portion 303 that includes 6DoF metadata for 6DoF audio encoding (e.g., in the metadata portion or the extension portion of the bitstream). Preferably, the bitstream syntax of the bitstream BS is compatible with or compliant with the BS syntax for 3DoF audio rendering, such as, for example, the MPEG-H 3DA bitstream syntax. The 6DoF metadata may be included as extended metadata in one or more packets of the bitstream BS (e.g., in one or more extension containers already provided by the MPEG-H 3DA bitstream structure).

[0107] As described above in connection with FIG. 3 for example, the 6DoF metadata may include position data (e.g., coordinates) of one or more 3DoF (default) positions, further optionally 6DoF space descriptions (e.g., object coordinates), further optionally object orientations, and further optionally metadata that describes and / or parameterizes the VR environment, and / or, further optionally, parameter information and / or parameters regarding attenuation, occlusion, and / or reverberation, etc.

[0108] FIG. 6B schematically shows an exemplary 3DoF audio rendering based on the data representation and / or bitstream structure of FIG. 6A according to exemplary aspects of the present disclosure. As in FIG. 6A, the data representation and / or bitstream structure may be encoded via a device or system (e.g., software, hardware, or cloud) that is compatible with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0109] Specifically, in B of FIG. 6, it is illustratively shown that 3DoF audio rendering can be achieved by a 3DoF audio renderer that can perform 3DoF audio rendering based only on the 3DoF-encoded audio data obtained from the first bitstream portion 302 by discarding the 6DoF metadata. That is, for example, in the case of MPEG-H 3DA backward compatibility, the MPEG-H 3DA renderer can efficiently and surely ignore / discard the 6DoF metadata in the extended portion of the bitstream (e.g., extended container(s)) so as to perform efficient normal MPEG-H 3DA 3DoF (or 3DoF+) audio rendering based only on the 3DoF-encoded audio data obtained from the first bitstream portion 302.

[0110] FIG. 6C schematically shows an exemplary 6DoF audio rendering according to exemplary aspects of the present disclosure, based on the data representation and / or bitstream structure of FIG. 6A. As in FIG. 6A, the data representation and / or bitstream structure may be encoded via a device or system (e.g., software, hardware, or cloud) compatible with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0111] Specifically, in C of FIG. 6, it is illustratively shown that 6DoF audio rendering can be achieved by a novel 6DoF audio renderer (e.g., compliant with MPEG-I or a subsequent standard) that performs 6DoF audio rendering using the 3DoF-encoded audio data obtained from the first bitstream portion 302 together with the 6DoF metadata obtained from the second bitstream portion 303, based on the 3DoF-encoded audio data obtained from the first bitstream portion 302 and the 6DoF metadata obtained from the second bitstream portion 303.

[0112] Thus, without redundancy in the bitstream, or at least with reduced redundancy, the same bitstream can be used by a legacy 3DoF audio renderer that allows for simple and useful backward compatibility for 3DoF audio rendering, and a new 6DoF audio renderer for 6DoF audio rendering.

[0113] FIG. 7A schematically shows a 6DoF audio encoding conversion A based on 3DoF audio signal data, according to exemplary aspects of the present disclosure. The conversion (and any inverse conversion) can be performed according to a method, process, apparatus, or system (e.g., software, hardware, or cloud) that is compatible with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0114] Exemplarily, similar to FIGS. 2 and 3 above, FIG. 7A shows an exemplary top view 202 of a room that includes a plurality of audio sources 207 (which may be located behind the wall 203 or whose sound signals may be obstructed by other structures, resulting in attenuation, reverberation, and / or occlusion effects).

[0115] For the purpose of 3DoF audio rendering, the audio signal x of the plurality of audio sources 207 is converted to obtain a 3DoF audio signal (audio object) on a sphere S around a default 3DoF position 206 (e.g., the listener position in a 3DoF sound field). As described above, the 3DoF audio signal is referred to as x 3DA and may be obtained using a conversion function A as X 3DA =A(x) Equation (6) as shown.

[0116] In the above equation, x represents the sound source(s) / object signal(s), x 3DA represents the corresponding virtual 3DA object signal for 3DA that generates the same sound field at the default 3DoF position 206, and A is the audio signal x based on the audio signal x3DA represents a conversion function that approximates. The inverse conversion function A -1 may be used to restore / approximate the sound source signal for 6DoF audio rendering. This has been discussed above and will be further discussed below. AA -1 = 1 and A -1 A = 1, or at least [Number] note that it is.

[0117] In a general way, the conversion function A may be regarded as a mapping / projection function that projects, or at least maps, the audio signal x onto a sphere S around the default 3DoF position 206 in some exemplary aspects of the present disclosure.

[0118] Furthermore, note that 3DoF audio rendering does not recognize the VR environment (such as existing walls 203 or other structures that may lead to attenuation, reverberation, masking effects, etc.). Thus, the conversion function A may preferably include effects based on such VR environment characteristics.

[0119] FIG. 7B schematically shows a 6DoF audio decoding conversion A for approximating / restoring 6DoF audio signal data based on 3DoF audio signal data according to exemplary aspects of the present disclosure. -1 to show schematically.

[0120] The inverse conversion function A -1 and the approximated 3DoF audio signal x obtained as in FIG. 7A above 3DA By using, the original audio signal x* of the original audio source 207 can be restored / approximated as follows: x* = A -1 (x 3DA ) Equation (7) Therefore, the audio signal x* of the audio object 320 in FIG. 7B can be restored to be the same as or similar to the audio signal x of the original source 207, particularly at the same position as the original source 207.

[0121] FIG. 7C schematically shows an exemplary 6DoF audio rendering based on the approximated / restored 6DoF audio signal data of FIG. 7B according to exemplary aspects of the present disclosure.

[0122] The audio signal x* of the audio object 320 in FIG. 7B can be used in 6DoF audio rendering, in which the position of the listener is also variable.

[0123] Assuming that the listener position of the listener is position 206 (the same position as the default 3DoF position), the 6DoF audio rendering renders the same sound field as the 3DoF audio rendering based on the audio signal x 3DA Therefore, the 6DoF rendering F at the default 3DoF position, which is the assumed listener position, of (x*) is equal to (or at least approximately equal to) the 3DoF rendering F 6DoF (x 3DoF ( 3DA ). Furthermore, when the listener position is shifted to, for example, position 206' in FIG. 7C, the sound field generated in the 6DoF audio rendering becomes different, but may preferably occur smoothly.

[0124] As another example, a third listener position 206" may be assumed, and the sound field generated in the 6DoF audio rendering becomes different, particularly for the audio signal in the upper left, which is not blocked by the wall 203 for the third listener position 206". Preferably, this is possible because the inverse function A -1 restores the original sound source (without environmental effects such as VR environment characteristics).

[0125] Figure 8 schematically shows an exemplary flowchart of a method for 3DoF / 6DoF bitstream encoding according to exemplary aspects of the present disclosure. It should be noted that the order of the steps is not limiting and may be changed according to the situation. Also, it should be noted that some steps of this method are optional. This method may be performed, for example, by a decoder, an audio decoder, an audio / video decoder, or a decoder system.

[0126] In step S801, the method receives (e.g., on the decoder side) the original audio signal x of one or more audio sources.

[0127] In step S802, the method optionally determines environmental characteristics (such as room shape, walls, wall sound reflection characteristics, objects, obstacles, etc.) and / or determines parameters (parameters such as attenuation, gain, masking, reverberation, etc. that parameterize the effects).

[0128] In step S803, the method optionally determines the parameterization of the conversion function A, for example, based on the result of step S802. Preferably, step S803 provides a parameterized or pre-set conversion function A.

[0129] In step S804, the method converts the original audio signal(s) x of one or more audio sources into the corresponding approximated 3DoF audio signal(s) x 3DA based on the conversion function A.

[0130] In step S805, the method determines 6DoF metadata (the metadata may include one or more 3DoF positions, VR environment information, and / or parameters and parameterizations of environmental effects such as attenuation, gain, masking, reverberation, etc.).

[0131] In step S806, the method 3DAInclude (embed) it in the first bitstream portion (or multiple first bitstream portions).

[0132] In step S807, the method includes (embeds) the 6DoF metadata in the second bitstream portion (or multiple second bitstream portions).

[0133] Next, in step S808, the method encodes a bitstream based on the first bitstream portion and the second bitstream portion, and provides an encoded bitstream including the 3DoF audio signal x 3DA in the first bitstream portion (or multiple first bitstream portions) and the 6DoF metadata in the second bitstream portion (or multiple second bitstream portions).

[0134] The encoded bitstream is then provided to a 3DoF decoder / renderer for 3DoF audio rendering based only on the 3DoF audio signal x 3DA in the first bitstream portion (or multiple first bitstream portions), or can be provided to a 6DoF decoder / renderer for 6DoF audio rendering based on the 3DoF audio signal x 3DA in the first bitstream portion (or multiple first bitstream portions) and the 6DoF metadata in the second bitstream portion (or multiple second bitstream portions).

[0135] Figure 9 schematically shows an exemplary flowchart of a method for 3DoF and / or 6DoF audio rendering according to exemplary aspects of the present disclosure. Note that the order of the steps is not limiting and may be changed according to the situation. Also note that some steps of the method are optional. This method may be performed, for example, by an encoder, a renderer, an audio encoder, an audio renderer, an audio / video encoder, or an encoder system or a renderer system.

[0136] In step S901, an encoded bitstream is received that includes a 3DoF audio signal x in a first bitstream portion (or a plurality of first bitstream portions) and 6DoF metadata in a second bitstream portion (or a plurality of second bitstream portions). 3DA

[0137] In step S902, the 3DoF audio signal x 3DA is obtained from the first bitstream portion (or a plurality of first bitstream portions). This can be done by a 3DoF decoder / renderer or also by a 6DoF decoder / renderer.

[0138] If the decoder / renderer is a legacy device for 3DoF audio rendering purposes (or a new 3DoF / 6DoF decoder / renderer switched to the 3DoF audio rendering mode), the method proceeds to step S903, where the 6DoF metadata is discarded / ignored, and then a 3DoF audio rendering operation is performed to render 3DoF audio based on the 3DoF audio signal x 3DA obtained from the first bitstream portion (or a plurality of first bitstream portions). That is, backward compatibility is advantageously ensured.

[0139] If the other party, the decoder / renderer, is for 6DoF audio rendering purposes (for example, a new 6DoF decoder / renderer or a 3DoF / 6DoF decoder / renderer switched to 6DoF audio rendering mode), this method proceeds to step S905 to obtain 6DoF metadata from the second bitstream portion.

[0140] In step S906, this method is based on the 6DoF metadata obtained from the second bitstream portion (or multiple second bitstream portions) and the inverse transformation function A -1 to approximate / recover the audio signal x* of the audio object / source from the 3DoF audio signal x 3DA obtained from the first bitstream portion (or multiple first bitstream portions).

[0141] Then, in step S907, this method proceeds to perform 6DoF audio rendering based on the approximated / recovered audio signal x* of the audio object / source and based on the listener position (which may be variable within the VR environment).

[0142] In the above exemplary aspects, an efficient and reliable method, apparatus, and data representation and / or bitstream structure for 3D audio encoding and / or 3D audio rendering can be provided, thereby enabling efficient 6DoF audio encoding and / or rendering with beneficial backward compatibility for 3DoF audio rendering, for example, according to the MPEG-H 3DA standard. Specifically, it is possible to provide a data representation and / or bitstream structure for 3DoF audio encoding and / or 3D audio rendering, thereby enabling efficient 6DoF audio encoding and / or rendering with preferably backward compatibility for 3DoF audio rendering, for example, according to the MPEG-H 3DA standard. Also, a corresponding encoding and / or rendering apparatus for efficient 6DoF audio encoding and / or rendering with backward compatibility for 3DoF audio rendering, for example, according to the MPEG-H 3DA standard, is provided.

[0143] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software operating on a digital signal processor or a microprocessor. Other components may be implemented as hardware and / or as application specific integrated circuits. Signals arising from the methods and systems described above may be stored in a medium such as random access memory or an optical storage medium. They may be transferred via a network such as a wireless network, a satellite network, a wireless network, or a wired network, for example, the Internet. A typical apparatus utilizing the methods and systems described herein is a portable electronic device or other consumer device used to store and / or render audio signals.

[0144] Exemplary implementations of the methods and apparatus according to the present disclosure will be apparent from the following enumerated example embodiments (EEE), but these are not the claims.

[0145] EEE1 is, by way of example, a method for encoding audio, 3DoF-related data, and 6DoF-related data including an audio source signal, comprising: encoding an audio source signal that approximates a desired sound field at a 3DoF position(s), for example by an audio source device such as particularly within an encoder, to determine 3DoF data; and / or encoding 6DoF-related data, for example by an audio source device such as particularly within an encoder, to determine 6DoF metadata, the metadata being usable to approximate the original audio source signal for 6DoF rendering.

[0146] EEE2 is, by way of example, related to the method of EEE1, wherein the 3DoF data relates to at least one of an object audio signal, an object direction, and an object distance.

[0147] EEE3 is, by way of example, related to the method of EEE1 or EEE2, wherein the 6DoF data relates to at least one of a 3DoF (default) position parameter, a 6DoF spatial description (object coordinates) parameter, an object directivity parameter, a VR environment parameter, a distance attenuation parameter, a masking parameter, and a reverberation parameter.

[0148] EEE4 relates, by way of example, to a method for transferring data, in particular 3DoF and 6DoF renderable audio data, the method comprising: transferring an audio source signal that preferably approximates a desired sound field at 3DoF position(s), for example when decoded by a 3DoF audio system, for example in an audio bitstream syntax; and / or transferring 6DoF related metadata for approximating and / or restoring the original audio source signal for 6DoF rendering, for example in an extended part of an audio bitstream syntax, where the 6DoF related metadata may be parametric data and / or signal data.

[0149] EEE5 relates, by way of example, to the method of EEE4, where an audio bitstream syntax comprising, for example, 3DoF metadata and / or 6DoF metadata conforms to at least a certain version of the MPEG-H Audio standard.

[0150] EEE6 relates, by way of example, to a method for generating a bitstream, the method comprising: determining 3DoF metadata based on an audio source signal that approximates a desired sound field at 3DoF position(s); determining 6DoF related metadata, which may be used to approximate the original audio source signal for 6DoF rendering; and / or inserting the audio source signal and the 6DoF related metadata into a bitstream.

[0151] EEE7 relates, by way of example, to a method of audio rendering. The method comprises: A stage of preprocessing 6DoF metadata of an approximate audio signal of an original audio signal at 3DoF position(s), and the 6DoF rendering can provide the same output as the 3DoF rendering of the transferred audio source signal for approximating a desired sound field at 3DoF position(s).

[0152] EEE8, illustratively with respect to the method of EEE7, the audio rendering is:

Number

[0153] EEE9, illustratively with respect to the method of EEE8, the approximate audio signal of the original audio signal is: X* := A -1 (x 3DA ) based on this, where A -1 relates to the inverse function of the approximation function A.

[0154] EEE10, illustratively with respect to the method of EEE8 or EEE9, the metadata used to obtain the approximate audio signal of the original audio source signal using the approximation method is

Number

[0155] Exemplary aspects and embodiments of the present disclosure may be implemented in hardware, firmware, software, or a combination of both (e.g., as a programmable logic array). Unless otherwise specified, algorithms or processes included as part of the present disclosure are not inherently related to any particular computer or other device. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized device (e.g., an integrated circuit) to perform the required method steps. Thus, the present disclosure may be implemented in one or more computer programs (e.g., implementation of any of the elements of the figures) executed on one or more programmable computer systems each including at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.

[0156] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or object-oriented programming languages) to communicate with a computer system. In any case, the language can be a compiled language or an interpreted language.

[0157] For example, when implemented by a computer software instruction sequence, the various functions and stages of the embodiments of the present disclosure may be implemented by a multi-threaded software instruction sequence executed by suitable digital signal processing hardware, in which case the various devices, stages, and functions of the embodiments may correspond to portions of the software instructions.

[0158] Each such computer program is preferably stored or downloaded to a storage medium or device readable by a general-purpose or special-purpose programmable computer (such as solid-state memory or media, or magnetic or optical media) to configure and operate a computer when read by the computer system to execute the procedures described herein. The system of the present invention may be implemented as a computer-readable storage medium configured (i.e., storing) a computer program, and such a configured storage medium causes the computer system to operate in a specific predefined manner to execute the functions described herein.

[0159] Some exemplary aspects and exemplary embodiments of the present disclosure have been described above. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present invention of the present disclosure. Many modifications and variations of the present invention are possible in light of the above teachings. It should be understood that within the scope of the appended claims, the invention of the present disclosure may be practiced otherwise than as specifically described herein.

Claims

1. A method for decoding a bitstream, the method comprising: receiving the bitstream comprising encoded audio signal data related to three degrees of freedom (3DoF) audio rendering and metadata related to six degrees of freedom (6DoF) audio rendering; decoding the encoded audio signal data related to 3DoF to obtain a 3DoF audio signal; rendering the 3DoF audio signal based on one of 3DoF audio rendering or 6DoF audio rendering to generate a sound field, wherein the 6DoF audio rendering generates 6DoF audio signal data based on the 3DoF audio signal and the metadata related to 6DoF, wherein the 3DoF audio rendering: ignores the metadata related to 6DoF; includes executing the 3DoF audio rendering using the 3DoF audio signal, wherein the 6DoF audio rendering: From the 3DoF audio signal, the original audio signal x -1 of one or more audio sources is restored using the inverse transformation function A * ; Restored original audio signal x * and performing the 6DoF audio rendering using the metadata related to 6DoF, Method.

2. The encoded audio signal data related to 3DoF is generated by mapping the original audio signal x of the one or more audio sources to corresponding audio objects located on one or more spheres around a default 3DoF listener position using a conversion function A. The method according to claim 1.

3. The encoded audio signal data related to 3DoF audio rendering includes at least one of one or more audio objects, direction data of the one or more audio objects, and distance data of the one or more audio objects. The method according to claim 1.

4. The one or more audio objects are located on one or more spheres around a default 3DoF listener position. The method according to claim 3.

5. The metadata related to 6DoF audio rendering indicates one or more default 3DoF listener positions. The method according to claim 1.

6. The metadata related to 6DoF audio rendering is: a description of the 6DoF space, at least one parameter related to at least one of the audio object directions of one or more audio objects, a virtual reality environment, distance attenuation, occlusion, and reverberation, the method according to claim 1, showing at least one of these.

7. The encoded audio signal data related to 3DoF audio rendering is determined based on the original audio signal from one or more audio sources and the conversion function A, the method according to claim 2.

8. The encoded audio signal data related to 3DoF audio rendering is determined by converting the original audio signal from the one or more audio sources into the 3DoF audio signal using the conversion function A, and the conversion function A maps the original audio signal of the one or more audio sources to each audio object located on the one or more spheres around the default 3DoF listener position, the method according to claim 7.

9. The bitstream is compatible with the MPEG-H 3D Audio standard, the method according to claim 1.

10. The encoded audio signal data related to 3DoF audio rendering is part of the payload of the bitstream, The metadata related to 6DoF audio rendering is part of one or more extension containers of the bitstream, the method according to claim 1.

11. A computer program for causing one or more processors to execute the method according to claim 1.

12. An audio decoder device for decoding a bitstream, the device comprising: A receiver for receiving the bitstream containing encoded audio signal data related to three degrees of freedom (3DoF) audio rendering and metadata related to six degrees of freedom (6DoF) audio rendering; A decoder for decoding the encoded audio signal data related to 3DoF to obtain a 3DoF audio signal; A renderer that renders the 3DoF audio signal based on one of 3DoF audio rendering or 6DoF audio rendering to generate a sound field, wherein the 6DoF audio rendering generates 6DoF audio signal data based on the 3DoF audio signal and the metadata related to 6DoF, and has a renderer, The 3DoF audio rendering is: Ignore the metadata related to 6DoF; Including executing the 3DoF audio rendering using the 3DoF audio signal, The 6DoF audio rendering is: From the 3DoF audio signal, the original audio signal x -1 of one or more audio sources is restored using the inverse transformation function A * ; Restored original audio signal x * and performing the 6DoF audio rendering using the metadata related to 6DoF, Device.

Citation Information

Patent Citations

  • Distance panning using near / far-field rendering

    WO2017218973A1