Method for decoding a bitstream, a machine-readable storage medium, and an audio decoding device

The method and device transform 3DoF audio signals with 6DoF metadata to generate 6DoF audio, addressing the lack of 6DoF support in existing systems and ensuring compatibility with 3DoF standards, enhancing audio quality in immersive environments.

RU2865406C2Active Publication Date: 2026-07-02DOLBY INTERNATIONAL AB
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
RU · RU
Patent Type
Patents
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2019-04-09
Publication Date
2026-07-02

AI Technical Summary

Technical Problem

Current audio generation technologies lack support for six degrees of freedom (6DoF) user movement, specifically in combination with three degrees of freedom (3DoF) audio systems, and there is a need for efficient encoding, decoding, and shaping methods that ensure backward compatibility with existing 3DoF standards like MPEG-H 3DA.

Method used

A method and device for decoding a bitstream that includes 3DoF audio signal data and 6DoF metadata, allowing for the generation of 6DoF audio by transforming 3DoF audio signals using inverse transformation and incorporating 6DoF-related metadata, while ensuring compatibility with MPEG-H 3DA standards.

Benefits of technology

Enables efficient encoding and decoding of 6DoF audio with backward compatibility to 3DoF systems, improving bit rate efficiency and maintaining high-quality audio output in virtual reality, augmented reality, and mixed reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000020_ABST
    Figure 00000020_ABST
Patent Text Reader

Abstract

FIELD: sound engineering.SUBSTANCE: invention relates to a device, system and method for generating sound with six degrees of freedom (6DoF), in particular in connection with the representation of data and structures of bit streams for generating 6DoF sound, in particular to methods, a device and systems for decoding an audio signal and generating sound based on a bit stream. The claimed method for decoding a bitstream includes receiving a bitstream comprising encoded audio signal data associated with the formation of sound with three degrees of freedom (3DoF) and metadata associated with the formation of sound with six degrees of freedom (6DoF), and decoding the encoded audio signal data associated with 3DoF, with the receipt of an audio signal with 3DoF. Next, a 3DoF audio signal is generated with the generation of a sound field based on one of the 3DoF audio generation and the 6DoF audio generation, wherein the 6DoF audio generation generates 6DoF audio signal data based on the 3DoF audio signal and metadata associated with 6DoF. 3DoF audio generation involves ignoring the 6DoF related metadata and performing 3DoF audio generation using the 3DoF audio signal. The formation of 6DoF audio includes the restoration of the original audio signals x* of one or more audio sources from the 3DoF audio signal using the function of inverse transform and performing 6DoF audio generation using the reconstructed original audio signals x* and 6DoF-related metadata.EFFECT: efficient implementation of coding and / or generation of 6DoF audio, preferably with backward compatibility for generating 3DoF audio, for example, according to the MPEG-H 3DA standard.12 cl, 14 dwg
Need to check novelty before this filing date? Find Prior Art

Description

[0001] RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 655,990, filed April 11, 2018, which is incorporated herein by reference in its entirety.

[0003] FIELD OF TECHNOLOGY

[0004] The present invention relates to the provision of a device, system and method for generating six degrees of freedom (6DoF) audio, in particular in connection with data representation and bitstream structures for generating 6DoF audio.

[0005] BACKGROUND OF THE INVENTION

[0006] Currently, there is no proper solution for generating audio in combination with six degrees of freedom (6DoF) user movement. While solutions exist for generating channel, object, and first / higher-order ambiphony (HOA) signals in combination with three degrees of freedom (3DoF) (yaw, pitch, and roll), there is no support for processing such signals in combination with six degrees of freedom (6DoF) user movement (yaw, pitch, roll, and translation).

[0007] In general, 3DoF audio generation provides a sound field in which one or more sound sources are generated at angular positions surrounding a given listener position, called a 3DoF position. One example of 3DoF audio generation is included in the MPEG-H 3D Audio standard (abbreviated as MPEG-H 3DA).

[0008] While MPEG-H 3DA was designed to support channel, object, and HOA signals for 3DoF, it cannot yet handle true 6DoF audio. It is desirable that the intended implementation of 3D audio in MPEG-I extend the functionality of 3DoF (and 3DoF+) to 3D 6DoF audio applications in an efficient manner (preferably including efficient signal generation, encoding, decoding, and / or shaping), while preferably ensuring backward compatibility with 3DoF shaping.

[0009] Considering the above, the object of the present invention is to provide methods, a device and a representation of data and / or bitstream structures for coding 3D audio and / or generating 3D audio, which allows for efficient coding and / or generating 6DoF audio, preferably with backward compatibility for generating 3DoF audio, for example according to the MPEG-H 3DA standard.

[0010] Another object of the present invention may be to provide a representation of data and / or bitstream structures for encoding 3D audio and / or generating 3D audio, which allows for efficient encoding and / or generating 6DoF audio, preferably with backward compatibility with generating 3DoF audio, for example according to the MPEG-H 3DA standard, and an encoding and / or generating device designed for efficient encoding and / or generating 6DoF audio, preferably with backward compatibility for generating 3DoF audio, for example according to the MPEG-H 3DA standard.

[0011] BRIEF DESCRIPTION OF THE INVENTION

[0012] In a first aspect of the present invention, there is provided a method for decoding a bit stream, including:

[0013] receive a bitstream containing encoded audio signal data related to three degrees of freedom (3DoF) audio generation and metadata related to six degrees of freedom (6DoF) audio generation;

[0014] decoding the encoded audio signal data associated with 3DoF to obtain an audio signal with 3DoF; and

[0015] generating a 3DoF audio signal by generating a sound field based on one of the 3DoF audio generation and the 6DoF audio generation, wherein the 6DoF audio generation generates 6DoF audio signal data based on the 3DoF audio signal and metadata related to the 6DoF;

[0016] while the formation of sound with 3DoF includes:

[0017] Ignore 6DoF related metadata; and

[0018] Perform 3DoF audio generation using 3DoF audio signal;

[0019] while the sound generation with 6DoF includes:

[0020] Restore original audio signalsx*of one or more audio sources from a 3DoF audio signal using the A function -1 inverse transformation; and

[0021] Perform 6DoF audio generation using reconstructed original audio signals x* and 6DoF-related metadata.

[0022] In one embodiment, encoded audio signal data associated with 3DoF is generated by mapping original audio signals input from one or more audio sources to corresponding audio objects located on one or more spheres surrounding a default 3DoF listener position using a transform function A.

[0023] In another embodiment, encoded audio signal data associated with generating 3DoF audio comprises at least one of: one or more audio objects, directional data of the one or more audio objects, and distance data of the one or more audio objects.

[0024] In another embodiment, one or more audio objects are positioned on one or more spheres surrounding the default 3DoF listener position.

[0025] In another embodiment, the metadata associated with generating 6DoF audio indicates one or more default 3DoF listener positions.

[0026] In another embodiment, the metadata associated with the generation of 6DoF audio indicates at least one of: a description of a 6DoF space, directions of audio objects of one or more audio objects, a virtual reality environment, at least a parameter related to at least one of attenuation with increasing range, absorption and reverberations.

[0027] In another embodiment, encoded audio signal data associated with generating 3DoF audio is determined based on original audio signals from one or more audio sources and a transform function A.

[0028] In another embodiment, encoded audio signal data associated with generating 3DoF audio is determined by transforming original audio signals from one or more audio sources into 3DoF audio signals using a transform function A, wherein the transform function A maps the original audio signals of the one or more audio sources to respective audio objects located on one or more spheres surrounding the default 3DoF listener position.

[0029] In another embodiment, the bitstream is compatible with the MPEG-H 3D Audio standard.

[0030] In another embodiment, encoded audio signal data associated with generating 3DoF audio is part of the payload data of the bitstream, and

[0031] Metadata related to 6DoF audio generation is part of one or more bitstream expansion containers.

[0032] In a second aspect of the present invention, a non-transitory computer-readable storage medium is provided that stores a computer program containing instructions that, when executed by a processor, cause the processor to perform the method according to the first aspect.

[0033] In a third aspect, the present invention provides an audio decoding device for decoding a bit stream, comprising:

[0034] a receiving module for receiving a bitstream containing encoded audio signal data related to three degrees of freedom (3DoF) audio generation and metadata related to six degrees of freedom (6DoF) audio generation;

[0035] a decoding module for decoding the encoded audio signal data associated with 3DoF to obtain an audio signal with 3DoF; and

[0036] a generating module for generating a 3DoF audio signal with sound field generation based on one of 3DoF audio generation and 6DoF audio generation, wherein the 6DoF audio generation generates 6DoF audio signal data based on the 3DoF audio signal and metadata related to the 6DoF:

[0037] while the formation of sound with 3DoF includes:

[0038] Ignore 6DoF related metadata; and

[0039] Perform 3DoF audio generation using 3DoF audio signal;

[0040] while the sound generation with 6DoF includes:

[0041] Restore original audio signalsx*of one or more audio sources from a 3DoF audio signal using the A function -1 inverse transformation; and

[0042] Perform 6DoF audio generation using reconstructed original audio signals x* and 6DoF-related metadata.

[0043] It should be understood that the method steps and the device features can be interchanged in various ways. In particular, the details of the disclosed method can be implemented as a device adapted to perform some or all of the method steps, and vice versa, as will be understood by those skilled in the art. In particular, it should be understood that the corresponding statements made with respect to the methods are similarly applicable to the corresponding device, and vice versa.

[0044] BRIEF DESCRIPTION OF FIGURES

[0045] Illustrative embodiments of the present invention are described below with reference to the accompanying drawings, in which like reference numerals may designate like or similar elements, and in which:

[0046] Fig. 1 is a schematic diagram of an exemplary system including MPEG-H 3D Audio decoder / encoder interfaces according to exemplary aspects of the present invention.

[0047] Fig. 2 schematically shows an illustrative top view of the 6DoF room environment (6DoF space).

[0048] Fig. 3 is a schematic diagram illustrating a top view of the 6DoF environment of Fig. 2, as well as 3DoF audio data and 6DoF extension metadata according to illustrative aspects of the present invention.

[0049] Fig. 4A is a schematic diagram of an exemplary system for processing 3DoF, 6DoF and audio data according to exemplary aspects of the present invention.

[0050] Fig. 4B schematically illustrates exemplary decoding and generating methods for generating 6DoF audio and generating 3DoF audio according to exemplary aspects of the present invention.

[0051] Fig. 5 schematically illustrates an exemplary condition for matching the generation of a 6DoF sound and the generation of a 3DoF sound at a 3DoF position in the system according to one or more of Figs. 2–4B.

[0052] Fig. 6A is a schematic diagram illustrating an exemplary data representation and / or bitstream structure according to exemplary aspects of the present invention.

[0053] Fig. 6B is a schematic diagram illustrating an exemplary generation of 3DoF audio based on the data representation and / or bitstream structure of Fig. 6A according to exemplary aspects of the present invention.

[0054] Fig. 6C is a schematic diagram illustrating an exemplary 6DoF audio generation based on the data representation and / or bitstream structure of Fig. 6A according to exemplary aspects of the present invention.

[0055] Fig. 7A is a schematic diagram illustrating an encoding transformation A of a 6DoF audio based on 3DoF audio signal data according to exemplary aspects of the present invention.

[0056] Fig. 7B schematically shows the encoding transformation A -16DoF audio for approximating / reconstructing 6DoF audio signal data based on 3DoF audio signal data according to exemplary aspects of the present invention.

[0057] Fig. 7C is a schematic diagram illustrating an exemplary 6DoF audio generation based on the approximated / reconstructed 6DoF audio signal data of Fig. 7B according to exemplary aspects of the present invention.

[0058] Fig. 8 is a schematic diagram of an exemplary flow chart of a method for encoding a 3DoF / 6DoF bitstream according to exemplary aspects of the present invention.

[0059] Fig. 9 is a schematic block diagram of illustrative methods for generating 3DoF and / or 6DoF audio according to illustrative aspects of the present invention.

[0060] DETAILED DESCRIPTION

[0061] Next, preferred illustrative aspects will be described in more detail with reference to the accompanying figures. The same or similar features in different drawings and different embodiments may be designated by the same reference numerals. It should be understood that the detailed description presented below regarding various preferred illustrative aspects should not be construed as limiting the scope of the present invention.

[0062] In the context of this document, the term "MPEG-H 3D Audio" refers to the technical description specified in ISO / IEC 23008-3 and / or any past and / or future editions, revisions or other versions of ISO / IEC 23008-3.

[0063] In the context of this document, it is desirable that an implementation of MPEG-I 3D audio extend the functionality of 3DoF (and 3DoF+) to 3D 6DoF audio, while preferably providing backward compatibility with 3DoF generation.

[0064] In the context of this document, 3DoF typically refers to a system that can accurately process user head motion, specifically head rotation, characterized by three parameters (e.g., yaw, pitch, and roll). Such systems are often available in various gaming systems, such as virtual reality (VR) / augmented reality (AR) / mixed reality (MR) systems, or other similar acoustic environments.

[0065] In the context of this document, 6DoF generally refers to a system that can correctly handle 3DoF and translational motion.

[0066] Illustrative aspects of the present invention relate to an audio system (e.g., an audio system compliant with the MPEG-I Audio standard), wherein an audio generation module extends functionality to 6DoF by converting corresponding metadata into a 3DoF format, such as an input format of the audio generation module compliant with the MPEG standard (e.g., the MPEG-H 3DA standard).

[0067] Figure 1 shows an example system 100 configured to use metadata extensions and / or audio generation module extensions in addition to existing 3DoF systems to enable a 6DoF experience. The system 100 includes a source environment 101 (which, as an example, may include one or more audio sources 101a), a content format 102 (for example, a bitstream containing 3D audio data), an encoding device 103, and a proposed metadata encoding device extension 106. The system 100 may also include a 3D audio generation module 105 (for example, a 3DoF generation module), and proposed generation module extensions 107 (for example, 6DoF generation module extensions for the rendered environment 108).

[0068] In the 3D audio generation method with 3DoF, only the angles (e.g., yaw angle, pitch angle, roll angle) of the user's angular orientation at a given 3DoF position can be input into the 3DoF audio generation module 105. With the functionality extended to 6DoF, the user's location coordinates (e.g., x, y, and z) can be additionally input into the 6DoF audio generation module (extension generation module).

[0069] The advantage of the present invention lies in improving the bit rate of the bitstream transmitted between the encoding device and the decoding device. The bitstream can be encoded and / or decoded in accordance with a standard, such as the MPEG-I Audio standard and / or the MPEG-H 3D Audio standard, or at least backward compatible with a standard, such as the MPEG-H 3D Audio standard.

[0070] In some examples, illustrative aspects of the present invention relate to processing a single bitstream (e.g., an MPEG-H 3D Audio (3DA) bitstream (BS) or a bitstream using MPEG-H 3DA BS syntax) compatible with multiple systems.

[0071] For example, in some exemplary aspects, an audio bitstream may be compatible with two or more different generation modules, such as a 3DoF audio generation module that may be compatible with one standard (e.g., the MPEG-H 3D Audio standard), and a newly defined 6DoF audio generation module or an extension of the generation module that may be compatible with a second, different standard (e.g., the MPEG-I Audio standard).

[0072] Illustrative aspects of the present invention relate to different decoding devices configured to perform decoding and generate the same audio bitstream, preferably to create the same audio output.

[0073] For example, illustrative aspects of the present invention relate to a 3DoF decoding device and / or a 3DoF generation module and / or a 6DoF decoding device and / or a 6DoF generation module, configured to generate the same output for the same bitstream (e.g., a 3DA BS or a bitstream using a 3DA BS). As an example, the bitstream may contain information related to certain listening positions in a VR / AR / MR (virtual reality / augmented reality / mixed reality) space, for example, as part of 6DoF metadata.

[0074] As an example, the present invention further relates to encoding devices and / or decoding devices configured to encode and / or decode, respectively, 6DoF information (e.g., compatible with the MPEG-I Audio environment), wherein such encoding devices and / or decoding devices according to the present invention provide one or more of the following advantages:

[0075] representation with effective quality and bit rate of VR / AR / MR related audio data and its encapsulation in audio bitstream syntax (e.g. MPEG-H 3D Audio BS);

[0076] Backward compatibility between different systems (such as the MPEG-H 3DA standard and the MPEG-I Audio standard).

[0077] Preferably in order to avoid competition between 3DoF and 6DoF solutions and to ensure a smooth transition between current and future technologies, backward compatibility has many advantages.

[0078] For example, backward compatibility between a 3DoF audio system and a 6DoF audio system can have many benefits, such as providing a 6DoF audio system such as MPEG-I Audio with backward compatibility with a 3DoF audio system such as MPEG-H 3D Audio.

[0079] According to illustrative aspects of the present invention, this may be achieved by providing backward compatibility, for example at the bitstream level, for systems related to 6DoF and consisting of:

[0080] encoded data and corresponding metadata of 3DoF audio material; and

[0081] metadata related to 6DoF.

[0082] Illustrative aspects of the present invention relate to a standard 3DoF bitstream syntax, such as a first type of audio bitstream syntax (e.g., MPEG-H 3DA BS), that encloses 6DoF bitstream elements, such as MPEG-I Audio bitstream elements, for example, in one or more extension containers of the first type of audio bitstream (e.g., MPEG-H 3DA BS).

[0083] In order to provide a system that ensures backward compatibility at the performance level, the following systems and / or frameworks may be applicable and may be used:

[0084] 1a. The 3DoF system (e.g., systems compliant with the MPEG-H 3DA standards) shall be able to ignore all syntax elements related to 6DoF (e.g., ignore the MPEG-I Audio bitstream syntax elements based on the "mpegh3daExtElementConfig()" or "mpegh3daExtElement()" functionality of the MPEG-H 3D Audio bitstream syntax), i.e., the 3DoF system (decoder / generator module) may preferably be designed to ignore additional data and / or metadata related to 6DoF (e.g., by not reading the data and / or metadata related to 6DoF); and

[0085] 2a. The remaining portion of the bitstream payload (e.g., the MPEG-I Audio bitstream payload containing data and / or metadata compatible with the MPEG-H 3DA bitstream parser) shall be decodable by the 3DoF system (e.g., the legacy MPEG-H 3DA system) to produce the desired audio output, i.e., the 3DoF system (decoder / generator module) may preferably be configured to decode the 3DoF-related portion of the BS; and

[0086] 3a. The 6DoF system (e.g., MPEG-I Audio system) shall be capable of processing both 3DoF-related parts and 6DoF-related parts of the audio bitstream and producing an audio output corresponding to the audio output of the 3DoF system (e.g., MPEG-H 3DA systems) at a given backward compatible 3DoF position(s) in the VR / AR / MR space, i.e., the 6DoF system (decoder / generator module) may preferably be configured to generate, at the default 3DoF position(s), a sound field / sound output corresponding to the generated 3DoF sound field / sound output; and

[0087] 4a. The 6DoF system (e.g. MPEG-I Audio system) shall provide a smooth change (transition) of the audio output around the specified backward compatible 3DoF position(s) (i.e., providing a continuous audio field in the 6DoF space), i.e., the 6DoF system (decoder / generator module) may preferably be configured to generate, in the surroundings of the default 3DoF position(s), an audio field / audio output that smoothly transitions in the default 3DoF position(s) into the audio field / audio output generated by 3DoF.

[0088] In some examples, the present invention relates to providing a 6DoF audio generation module (e.g., an MPEG-I Audio generation module) that produces the same audio output as a 3DoF audio generation module (e.g., an MPEG-H 3D Audio generation module) in one or more or some 3DoF position(s).

[0089] Currently, there are drawbacks to directly transmitting 3DoF related audio signals and metadata directly to a 6DoF audio system, which include:

[0090] Increase the bit rate (i.e., 3DoF-related audio signals and metadata are sent in addition to 6DoF-related audio signals and metadata); and

[0091] Limited validity (i.e. the audio signal(s) and 3DoF related metadata are only valid for the 3DoF position(s).

[0092] Illustrative aspects of the present invention relate to overcoming the above disadvantages.

[0093] In some examples, the present invention relates to:

[0094] using 3DoF-compatible audio signal(s) and metadata (e.g. MPEG-H 3D Audio-compatible signals and metadata) instead of (or in addition to) the original audio source signals and metadata; and / or

[0095] increase the applicability range (usable for 6DoF formation) from 3DoF position(s) to 6DoF space (defined by the content author), while maintaining a high level of sound field approximation.

[0096] Illustrative aspects of the present invention relate to the efficient creation, encoding, decoding and generation of such signal(s) to achieve these objectives and to provide 6DoF generation functionality.

[0097] Figure 2 depicts an illustrative view 202 from above of an illustrative room 201. As shown in Figure 2, the illustrative listener stands in the middle of the room with multiple sound sources and non-trivial geometric wall shapes. In 6DoF devices (e.g., systems that provide 6DoF capabilities), the illustrative listener can move, but in some examples, it is assumed that the default 3DoF position 206 may correspond to the intended sweet spot of VR / AR / MR audio (e.g., according to the setting or intent of the content author).

[0098] In particular, Fig. 2 depicts illustrative walls 203, a 6DoF space 204, illustrative (optional) directional vectors 205 (e.g., if one or more sound sources are emitting sound in a directionally direction), a 3DoF listener position 206 (the default 3DoF position 206), and sound sources 207, depicted as an example in the shape of a star in Fig. 2.

[0099] Figure 3 shows an exemplary VR / AR / MR 6DoF environment, such as in Figure 2, and audio objects (audio data + metadata) 320 contained in a 3DoF audio bitstream 302 (such as, for example, an MPEG-H 3D Audio bitstream) and an extension container 303. The audio bitstream 302 and the extension container 303 can be encoded using a device or system (for example, software, hardware, or via a cloud solution) that is compliant with the MPEG standard (for example, MPEG-H or MPEG-I).

[0100] Illustrative aspects of the present invention relate to the reconstruction of a sound field using a 6DoF audio generation module (e.g., an MPEG-I Audio generation module) in a "3DoF position" in such a way as to match the output signal (which may or may not correspond to the sound propagation according to the laws of physics) of a 3DoF audio generation module (e.g., an MPEG-H Audio generation module). This sound field should preferably be based on the original "sound sources" and reflect the influence of complex geometric shapes of the corresponding VR / AR / MR environment (e.g., the effect of "walls", structures, sound reflections, reverberations and / or absorptions, etc.).

[0101] Illustrative aspects of the present invention relate to parameterizing by the encoding device all relevant information describing the scenario in such a way as to ensure that one or more, or preferably all, of the relevant requirements (1a) to (4a) described above are met.

[0102] If two audio generation modes are performed (i.e. 3DoF and 6DoF) in parallel and the interpolation algorithm is applied to the corresponding output data in 6DoF space, this approach will be approximately optimal since it will require:

[0103] parallel execution of two different generation algorithms (i.e. one for a specific 3DoF position and one for 6DoF space);

[0104] Large amount of audio data (to transmit additional audio data for the 3DoF audio generation module).

[0105] Illustrative aspects of the present invention avoid the above-mentioned disadvantages in that only one audio generation mode is preferably performed (e.g., instead of performing two audio generation modes in parallel), and / or 3DoF audio data is preferably used to generate 6DoF audio with additional metadata for reconstructing and / or approximating the original signal(s) of the audio source(s) (e.g., instead of transmitting 3DoF audio data and the original data of the audio source(s).

[0106] Illustrative aspects of the present invention relate to (1) one 6DoF audio generation algorithm (e.g., MPEG-I Audio compliant) that preferably produces exactly the same output as a 3DoF audio generation algorithm (e.g., MPEG-H 3DA compliant) at a particular position(s), and / or (2) a representation of audio (e.g., 3DoF audio data) and 6DoF-related audio metadata to minimize redundancy in the 3DoF-related and VR / AR / MR-related portions of a 6DoF audio bitstream data (e.g., MPEG-I Audio bitstream data).

[0107] Illustrative aspects of the present invention relate to the use of syntax of a first bitstream of a standardized format (e.g., MPEG-H 3DA BS) to enclose a second bitstream of a standardized format (e.g., future standards such as MPEG-I) or portions thereof and metadata related to 6DoF for:

[0108] transmitting (e.g., in the central part of a 3DoF audio bitstream syntax) audio source signals and metadata that the 3DoF audio system preferably decodes and that preferably approximate the desired sound field reasonably well at the 3DoF position(s) (default); and

[0109] transmitting (e.g., in terms of the 3DoF audio bitstream syntax extension) 6DoF-related metadata and / or additional data (e.g., parametric data and / or signal data) that are used to approximate (reconstruct) the original audio source signals to form 6DoF audio.

[0110] One aspect of the present invention relates to determining a desired "3DoF position(s)" and signals compatible with a 3DoF audio system (e.g., an MPEG-H 3DA system) on the encoding device side.

[0111] For example, as shown in relation to Fig. 3, the 3DA virtual object signals for 3DA can create the same sound field at a specific 3DoF position (based on the x signals 3DA ), which should preferably contain the VR environment effects for a specific 3DoF position(s) ("processed" signals), since some 3DoF systems (such as the MPEG-H 3DA system) cannot take into account the VR / AR / MR environment effects (e.g. absorption, reverberation, etc.). The methods and processes depicted in Fig. 3 may be performed using various systems and / or products.

[0112] In some illustrative aspects, the inverse functionA -1 , which preferentially "raw-casts" (i.e. removes VR environment effects) these signals, will be useful since this is necessary to approximate the original "raw" signals x (which do not contain VR environment effects).

[0113] Audio signal(s) for 3DoF generation ((x 3DA)) may be preferably defined to provide the same / similar output for both 3DoF audio generation and 6DoF audio generation, for example based on the following:

[0114] Equation No. (1)

[0115] Audio objects can be contained in a standardized bitstream. This bitstream can be encoded according to various standards, such as MPEG-H 3DA and / or MPEG-I.

[0116] BS can contain information about object signals, object directions and object distances.

[0117] Figure 3 further shows an example of an extension container 303 that may contain extension metadata, such as in a BS. The BS extension container 303 may contain at least one of the following metadata: (i) 3DoF position parameters (default); (ii) 6DoF space description parameters (object coordinates); (iii) (optional) object directivity parameters; (iv) (optional) VR / AR / MR environment parameters; and / or (v) (optional) range attenuation parameters, absorption parameters and / or reverberation parameters, etc.

[0118] The desired sound formation can be approximated based on the following:

[0119] Equation No. (2)

[0120] The approximation can be based on the VR environment, and the environment characteristics can be included in the metadata of the extension container.

[0121] Additionally or optionally, the output smoothness of the 6DoF audio generation module (such as the MPEG-I Audio generation module) may be provided, preferably based on the following:

[0122] - class of geometric continuity. Equation No. (3)

[0123] Illustrative aspects of the present invention relate to determining 3DoF audio objects (e.g. MPEG-H 3DA objects) on the encoding device side, preferably based on the following:

[0124] Equation No. (4)

[0125] One aspect of the present invention relates to restoring original objects on a decoding device based on the following:

[0126] , Equation No. (5)

[0127] at the same time refers to the signals of the sound source / object, refers to the approximation of sound source / object signals, F(x) for 3DoF / for 6DoF refers to the sound generation function for 3DoF / 6DoF listener position(s), 3DoF refers to the given position(s) with reference compatibility ∈ 6DoF space; 6DoF refers to the arbitrary allowed position(s) ∈ VR environment;

[0128] F6 DoF (x)refers to decoder-driven 6DoF audio generation (e.g. MPEG-I Audio generation);

[0129] F3 DoF (x3 DA )refers to decoder-induced 3DoF generation (e.g. MPEG-H 3DA generation); and

[0130] A, A -1 refer to the function (A) approximating signals x3 DA based on signals x, and function (A -1 ), its inverse.

[0131] The approximated signals of the sound sources / object are preferably reproduced using the 6DoF sound generation module in the “3DoF position” in a manner that matches the output signal of the 3DoF sound generation module.

[0132] Sound source / object signals are preferably approximated based on a sound field that is based on the original “sound sources” and reflects the influence of complex geometric shapes of the corresponding VR / AR / MR environment (e.g. “walls”, structures, reverberations, absorptions, etc.).

[0133] In other words, the 3DA virtual object signals for 3DA preferentially create the same sound field at a specific 3DoF position (based on the x3 signals DA ), which contains the effects of the VR environment for a specific 3DoF position(s).

[0134] The following may be available on the generation side (e.g., to a decoding device that complies with a standard such as MPEG-H or MPEG-I):

[0135] audio signal(s) for 3DoF sound generation:x3 DA

[0136] 3DoF audio generation or 6DoF audio generation functionality:

[0137] Equation No. (6)

[0138] For 6DoF audio generation, there may additionally be 6DoF metadata available on the generation side for 6DoF audio generation functionality (e.g., to approximate / reconstruct audio signals of one or more audio sources, e.g., based on audio signalsx3 DA 3DoF and 6DoF metadata.

[0139] Illustrative aspects of the present invention relate to (i) determining 3DoF audio objects (e.g., MPEG-H 3DA objects) and / or (ii) reconstructing (approximating) original audio objects.

[0140] As an example, audio objects can be contained in a 3DoF audio bitstream (such as MPEG-H 3DA BS).

[0141] The bitstream may contain information about the sound signals of objects, the directions of objects, and / or the distances to objects.

[0142] An extension container (e.g., a bitstream such as MPEG-H 3DA BS) may contain at least one of the following metadata: (i) 3DoF position parameters (default); (ii) 6DoF space description parameters (object coordinates); (iii) (optional) object directivity parameters; (iv) (optional) VR / AR / MR environment parameters; and / or (v) (optional) range attenuation parameters, absorption parameters, reverberation parameters, etc.

[0143] The present invention can provide the following advantages:

[0144] Backward compatibilitywith 3DoF audio decoding and generation (e.g., with MPEG-H 3DA decoding and generation): the output of the 6DoF audio generation module (e.g., MPEG-I Audio generation module) corresponds to the 3DoF generation output of the 3DoF generation engine (e.g., MPEG-H 3DA generation engine) for a given 3DoF position(s).

[0145] Coding efficiency: This approach can effectively reuse the legacy structure of 3DoF audio bitstream syntax (e.g. MPEG-H 3DA bitstream syntax).

[0146] Sound quality control at a given position(s) (3DoF): the best perceived audio quality can be explicitly provided by the encoder for any arbitrary position(s) and the corresponding 6DoF space.

[0147] Illustrative aspects of the present invention may relate to the following transmission of signals in a format compatible with the MPEG standard bitstream (e.g., the MPEG-I standard):

[0148] It is assumed that a 3DoF audio system (e.g., MPEG-H 3DA) provides signal transmission compatibility through an extension container mechanism (e.g., MPEG-H 3DA BS), which allows a 6DoF audio processing algorithm (e.g., compatible with MPEG-I Audio) to restore the original signals of the audio object.

[0149] Parameterization describes the data for approximating the original signals of the sound object.

[0150] The 6DoF audio generation module can specify how to restore the original signals of an audio object, such as in an MPEG-compliant system (such as MPEG-I Audio system).

[0151] This proposed concept:

[0152] is general in terms of the definition of the approximation function (i.e. A(x));

[0153] can be arbitrarily complex, but there must be an appropriate approximation on the decoder side

[0154] is approximately "uniquely defined" mathematically (e.g., algorithmically stable, etc.);

[0155] is general in application to the types of approximation function (i.e. A(x));

[0156] The approximation function can be based on the following approximation types or any combination of these approaches (listed in order of increasing bit rate consumption):

[0157] Parameterized sound effect(s) applied to signal x3 DA (e.g. parametrically controlled level, reverb, reflection, absorption, etc.)

[0158] parametrically encoded modification(s) (e.g. time / frequency variable gain modifications for transmitted signalx3 DA )

[0159] signal coded modification(s) (e.g. coded signals approximating the residual oscillation shape (x-x3 DA )); And

[0160] is extensible and applicable to common representations of sound field and sound sources (and their combinations): objects, channels, FOA, HOA.

[0161] Fig. 6A schematically depicts an exemplary data representation and / or bitstream structure according to exemplary aspects of the present invention. The data representation and / or bitstream structure may be encoded using a device or system (e.g., software, hardware, or via a cloud solution) compliant with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0162] As an example, the BS bitstream comprises a first bitstream portion 302 containing encoded 3DoF audio data (e.g., in the main portion or central portion of the bitstream). Preferably, the syntax of the BS bitstream is compatible with or corresponds to the BS syntax for generating 3DoF audio, such as, for example, the MPEG-H 3DA bitstream syntax. The encoded 3DoF audio data can be included as payload data in one or more packets of the BS bitstream.

[0163] As described previously, for example in connection with Fig. 3 above, the encoded 3DoF audio data may include signals of one or more audio objects (e.g., on a sphere around the default 3DoF position). For directional audio objects, the encoded 3DoF audio data may additionally optionally include directions of the objects and / or may additionally optionally indicate distances to the objects (e.g., by using gain and / or one or more attenuation parameters).

[0164] As an example, the BS comprises a second part 303 of the bitstream containing 6DoF metadata for encoding 6DoF audio (e.g., in the metadata part or the extension part of the bitstream). Preferably, the syntax of the BS bitstream is compatible with or corresponds to the syntax of the BS for generating 3DoF audio, such as, for example, the syntax of the MPEG-H 3DA bitstream. The 6DoF metadata can be included as extension metadata in one or more packets of the BS bitstream (e.g., in one or more extension containers, which, for example, are already provided by the MPEG-H 3DA bitstream structure).

[0165] As described previously, for example in connection with Fig. 3 above, the 6DoF metadata may include position data (e.g., coordinate(s)) of one or more 3DoF positions (by default), optionally further a description of the 6DoF space (e.g., coordinates of objects), optionally further a directionality of objects, optionally further metadata describing and / or parameterizing the VR environment and / or optionally further including parameterization information and / or parameters related to attenuation, absorptions and / or reverberations, etc.

[0166] Fig. 6B is a schematic diagram illustrating an exemplary generation of 3DoF audio based on the data representation and / or bitstream structure of Fig. 6A according to exemplary aspects of the present invention. As in Fig. 6a, the data representation and / or bitstream structure may be encoded using a device or system (e.g., software, hardware, or via a cloud solution) compliant with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0167] In particular, Fig. 6B illustratively shows that 3DoF audio generation can be achieved by using a 3DoF audio generation module that can exclude 6DoF metadata in order to perform 3DoF audio generation based only on the encoded 3DoF audio data obtained from the first part 302 of the bitstream. That is, for example, in the case of backward compatibility with MPEG-H 3DA, the MPEG-H 3DA generation module can effectively and reliably ignore / exclude 6DoF metadata in the extension part (for example, extension container(s)) of the bitstream in order to perform efficient conventional MPEG-H 3DA 3DoF (or 3DoF+) audio generation based only on the encoded 3DoF audio data obtained from the first part 302 of the bitstream.

[0168] Fig. 6C is a schematic diagram illustrating an exemplary generation of 6DoF audio based on the data representation and / or bitstream structure of Fig. 6A according to exemplary aspects of the present invention. As in Fig. 6a, the data representation and / or bitstream structure may be encoded using a device or system (e.g., software, hardware, or via a cloud solution) compliant with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0169] In particular, Fig. 6C illustratively shows that 6DoF audio generation can be achieved by using a new 6DoF audio generation module (for example, according to MPEG-I or later standards), which uses encoded 3DoF audio data obtained from the first part 302 of the bitstream, together with 6DoF metadata obtained from the second part 303 of the bitstream, to perform 6DoF audio generation based on the encoded 3DoF audio data obtained from the first part 302 of the bitstream and the 6DoF metadata obtained from the second part 303 of the bitstream.

[0170] Accordingly, with no or at least reduced redundancy in the bitstream, the same bitstream can be used by legacy 3DoF audio generation modules, which provides simple and useful backward compatibility, to generate 3DoF audio and by new 6DoF audio generation modules to generate 6DoF audio.

[0171] Fig. 7A is a schematic diagram of a coding transformA of 6DoF audio based on 3DoF audio signal data according to exemplary aspects of the present invention. The transform (and any inverse transforms) can be performed according to methods, processes, devices or systems (e.g., software, hardware or via a cloud solution) compliant with the MPEG standard (e.g., MPEG-H or MPEG-I).

[0172] As an example, similar to Fig. 2 and Fig. 3 above, Fig. 7A shows an illustrative view 202 from above of a room, including, as an example, a plurality of sound sources 207 (which may be located behind walls 203 or their sound signals may be obstructed by other structures, which may result in attenuation, reverberation and / or absorption effects).

[0173] To generate 3DoF audio, audio signalsxof a plurality of audio sources 207 are transformed to produce 3DoF audio signals (audio objects) on a sphere S around a default 3DoF position 206 (e.g., the position of a listener in a 3DoF audio field). As indicated above, 3DoF audio signals are denoted asx 3DA and can be obtained using the transformation function A, so that:

[0174] x 3DA = A(x).Equation No. (6)

[0175] In the above expression, x denotes the sound source(s) / signal(s) of the object, x 3DA denotes the corresponding 3DA virtual object signals for 3DA that produce the same sound field at the default 206 3DoF position, and A denotes a transfer function that approximates the sound signalsx 3DA based on sound signalsx. FunctionA -1The inverse transform can be used to reconstruct / approximate sound source signals to generate 6DoF audio, as discussed above and will be discussed below. It should be noted thatA A -1 = 1 and A -1 A = 1 or at leastA A -1 ≈ 1 iA -1 A ≈ 1.

[0176] In general, the transform function A can be regarded as a mapping / projection function that projects or at least displays the audio signals x on a sphere S surrounding the default 3DoF position 206 in some illustrative aspects of the present invention.

[0177] It should also be noted that 3DoF sound generation is unaware of the VR environment (such as existing walls 203 or the like, or other structures that may cause attenuation, reverberation, absorption effects, or the like). Accordingly, the A-transformation function can preferably include effects based on such characteristics of the VR environment.

[0178] Fig. 7B schematically shows the decoding transformation A -1 6DoF audio for approximating / reconstructing 6DoF audio signal data based on 3DoF audio signal data according to exemplary aspects of the present invention.

[0179] By using function A -1 inverse transform and approximated audio signalsx 3DA 3DoF obtained as shown above in Fig. 7A, the original audio signals x*of the original audio sources 207 can be reconstructed / approximated as:

[0180] x* = A -1 (x 3DA ).Equation No. (7)

[0181] Accordingly, the sound signalsx* of the sound objects 320 in Fig. 7B can be reconstructed in a similar or the same manner as the sound signalsx of the original sources 207, in particular in the same places as the original sources 207.

[0182] Fig. 7C is a schematic diagram illustrating an exemplary 6DoF audio generation based on the approximated / reconstructed 6DoF audio signal data of Fig. 7B according to exemplary aspects of the present invention.

[0183] The audio signals x* of the audio objects 320 in Fig. 7B can in this case be used to generate 6DoF audio, in which the position of the listener also becomes variable.

[0184] When the listener position is assumed to be position 206 (the same position as the default 3DoF position), 6DoF sound generation generates the same sound field as 3DoF sound generation based on audio signalsx 3DA .

[0185] Accordingly, the formation of 6DoF F 6DoF (x*) at the default 3DoF position, which is the intended position of the listener, is equal to (or at least approximately equal to) the 3DoF formation F 3DoF (x 3DA ).

[0186] Furthermore, if the position of the listener shifts, such as to position 206' in Fig. 7C, the sound field generated in the 6DoF sound formation changes, but this can preferably happen smoothly.

[0187] As another example, a third listener position 206" may be assumed, and the sound field generated in the 6DoF sound formation is changed specifically for the upper left sound signal that is not obstructed by the wall 203 at the third listener position 206". This is preferably made possible by the fact that the inverse functionA -1 Restores the original sound source (without environmental effects such as VR environment characteristics).

[0188] Figure 8 schematically depicts an illustrative flowchart of a method for encoding a 3DoF / 6DoF bitstream according to illustrative aspects of the present invention. It should be noted that the order of the steps is not limiting and can be changed according to the circumstances. It should also be noted that some steps of the method are optional. For example, the method can be performed by a decoding device, an audio decoding device, an audio / video decoding device, or a decoding system.

[0189] In step S801, the method (for example, on the decoding device side) includes receiving an original audio signal(s) from one or more audio sources.

[0190] In step S802, the method optionally comprises determining characteristics of the environment (such as the shape of the room, walls, characteristics of sound reflection by walls, objects, obstacles, etc.) and / or determining parameters (parameterization effects such as attenuation, amplification, absorption, reverberation, etc.).

[0191] In step S803, the method optionally comprises determining a parameterization of the transformation function A, for example based on the results of step S802. Preferably, in step S803, a parameterized or predefined transformation function A is provided.

[0192] In step S804, the method comprises converting the original audio signal(s)x of one or more audio sources into a corresponding one or more approximated audio signal(s)x 3DA 3DoF based on the transform function.

[0193] In step S805, the method comprises determining 6DoF metadata (which may include one or more 3DoF positions, VR environment information and / or environment effect parameters and parameterization such as attenuation, amplification, absorption, reverberations, etc.).

[0194] In step S806, the method comprises turning on (implementing) an audio signal(s)x 3DA 3DoF into the first part of the bitstream (or first few parts of the bitstream).

[0195] In step S807, the method comprises including (embedding) 6DoF metadata into the second part of the bitstream (or several second parts of the bitstream).

[0196] Then, in step S808, the method provides for continuing to encode the bit stream based on the first and second portions of the bit stream to provide an encoded bit stream that contains the audio signal(s)x 3DA3DoF in the first part of the bitstream (or several first parts of the bitstream) and 6DoF metadata in the second part of the bitstream (or several second parts of the bitstream).

[0197] The encoded bitstream can then be fed to a decoder / 3DoF generating module to generate 3DoF audio based on the audio signal(s)x 3DA 3DoF only in the first part of the bitstream (or first few parts of the bitstream) or to a 6DoF decoder / generator to generate 6DoF audio based on the audio signal(s)x 3DA 3DoF in the first part of the bitstream (or several first parts of the bitstream) and 6DoF metadata in the second part of the bitstream (or several second parts of the bitstream).

[0198] FIG. 9 schematically depicts an illustrative flowchart of methods for generating 3DoF and / or 6DoF audio according to illustrative aspects of the present invention. It should be noted that the order of the steps is not limiting and can be changed according to the circumstances. It should also be noted that some steps of the methods are optional. For example, the method can be performed by an encoding device, a generating module, an audio encoding device, an audio generating module, an audio / video encoding device, or an encoding system or system of generating modules.

[0199] In step S901, an encoded bit stream is received that contains an audio signal(s)x 3DA 3DoF in the first part of the bitstream (or several first parts of the bitstream) and 6DoF metadata in the second part of the bitstream (or several second parts of the bitstream).

[0200] At step S902, the sound signal(s)x 3DA3DoF is obtained from the first part of the bitstream (or the first few parts of the bitstream). This can be accomplished using a 3DoF decoder / generator, as well as a 6DoF decoder / generator.

[0201] Then, if the decoding device / generating module is a legacy device for the purpose of 3DoF audio generation (or a new 3DoF / 6DoF decoding device / generating module that has been switched to the 3DoF audio generation mode), the method includes proceeding to step S903, in which the 6DoF metadata is excluded / ignored, and then proceeding to the 3DoF audio generation operation to generate 3DoF audio based on the audio signal(s)x 3DA 3DoF obtained from the first part of the bitstream (or several first parts of the bitstream).

[0202] In other words, backward compatibility is mainly guaranteed.

[0203] On the other hand, if the decoding device / generating module is intended for the purpose of generating 6DoF audio (such as a new 6DoF decoding device / generating module or a 3DoF / 6DoF decoding device / generating module switched to the 6DoF audio generating mode), the method provides for proceeding to step S905 to obtain 6DoF metadata from the second part(s) of the bitstream.

[0204] In step S906, the method provides for approximating / restoring audio signalsx*audio objects / sources from the audio signal(s)x 3DA 3DoF obtained from the first part of the bitstream (or several first parts of the bitstream), based on 6DoF metadata obtained from the second part of the bitstream (or several second parts of the bitstream), and function A -1 inverse transformation.

[0205] Then, in step S907, the method provides for proceeding to performing 6DoF sound generation based on the approximated / reconstructed sound signalsx*of the sound objects / sources and based on the position of the listener (which may be variable in the VR environment).

[0206] In the exemplary aspects presented above, it is possible to provide efficient and reliable methods, a device and a representation of data and / or a bitstream structure for encoding 3D audio and / or generating 3D audio, which allows for efficient encoding and / or generating 6DoF audio, preferably with backward compatibility for generating 3DoF audio, for example, according to the MPEG-H 3DA standard. In particular, it is possible to provide a representation of data and / or a bitstream structure for encoding 3D audio and / or generating 3D audio, which allows for efficient encoding and / or generating 6DoF audio, preferably with backward compatibility for generating 3DoF audio, for example, according to the MPEG-H 3DA standard, and a corresponding encoding and / or generating device for efficiently encoding and / or generating 6DoF audio with backward compatibility for generating 3DoF audio, for example, according to the MPEG-H 3DA standard.

[0207] The methods and systems described in this document can be implemented as software, firmware, and / or hardware. Some components can be implemented as software executed by a digital signal processor or a microprocessor. Other components can be implemented as hardware or as specialized integrated circuits. The signals that occur in the described methods and systems can be stored on media such as random access memory or optical storage media. They can be transmitted over networks such as radio networks, satellite networks, wireless networks, or wired networks such as the Internet. Typical devices using the methods and systems described in this document are portable electronic devices or other consumer equipment that is used to store and / or generate audio signals.

[0208] Examples of the implementation of the methods and apparatus according to the present invention will become apparent from the following numbered examples of embodiments (EEE), which are not claims.

[0209] EEE1 illustratively relates to a method for encoding audio containing audio source signals, 3DoF-related data, and 6DoF-related data, comprising: encoding, for example by a device in the form of an audio source, such as, for example, an encoding device, audio source signals that approximate a desired sound field at a 3DoF position(s) to determine 3DoF data; and / or encoding, for example by a device in the form of an audio source, such as, for example, an encoding device, data related to 6DoF, to determine 6DoF metadata, wherein the metadata can be used to approximate the original audio source signals to form 6DoF.

[0210] EEE2 illustratively refers to the method of EEE1, wherein the 3DoF data refers to at least one of sound signals of objects, directions of objects, and distances to objects.

[0211] EEE3 illustratively refers to the method of EEE1 or EEE2, wherein the 6DoF data refers to at least one of the following: 3DoF position parameters (default), 6DoF space description parameters (object coordinates), object directivity parameters, VR environment parameters, attenuation parameters with increasing range, absorption parameters, and reverberation parameters.

[0212] EEE4 illustratively relates to a method for transmitting data, in particular audio data used to generate 3DoF and 6DoF, wherein the method comprises: transmitting, for example in the syntax of an audio bitstream, audio source signals that can preferably approximate a desired sound field at the position(s) of the 3DoF, for example when decoded by a 3DoF audio system; and / or transmitting, for example in part of an extension of the syntax of the audio bitstream, metadata related to 6DoF for approximating and / or reconstructing the original audio source signals to generate 6DoF; wherein the metadata related to 6DoF may be parametric data and / or signal data.

[0213] EEE5 illustratively refers to the method of EEE4, wherein the syntax of the audio bitstream, for example including 3DoF metadata and / or 6DoF metadata, complies with at least a version of the MPEG-H Audio standard.

[0214] EEE6 illustratively relates to a method for generating a bitstream, wherein the method includes: determining 3DoF metadata that is based on audio source signals that approximate a desired sound field at a 3DoF position(s); determining metadata related to 6DoF, wherein the metadata can be used to approximate the original audio source signals to form 6DoF; and / or introducing the audio source signal and the metadata related to 6DoF into the bitstream.

[0215] EEE7 illustratively relates to a method of generating a sound, wherein said method comprises:

[0216] Pre-processing of 6DoF approximated audio signal metadatax * source audio signalsx in 3DoF position(s), while 6DoF generation can provide the same output as 3DoF generation of transmitted audio source signalsx3 DAto generate 3DoFs that approximate the desired sound field at the 3DoF position(s).

[0217] EEE8 illustratively refers to the method of EEE7, wherein the sound generation is determined based on the following:

[0218]

[0219] where refers to the sound generation function for the 6DoF listener position(s), refers to the sound generation functions for the 3DoF listener position(s), are audio signals containing the effects of the VR environment for a specific 3DoF position(s), * refers to approximated sound signals.

[0220] EEE9 illustratively refers to the method of EEE8, wherein the approximated sound signalsx * original audio signals are based on the following:

[0221]

[0222] at the same timeA -1refers to the inverse function of the approximation function A.

[0223] EEE10 illustratively refers to the method of EEE8 or EEE9, wherein the metadata used to obtain approximated audio signalsx * the original signals of the sound source x, using the approximation method A, are determined on the basis of the following:

[0224]

[0225] where the amount of metadata is less than the amount of audio data required to transmit the original audio source signalsx,

[0226] in this case, the sound formation is determined based on the following:

[0227]

[0228] where refers to the sound generation function for the 6DoF listener position(s), refers to the sound generation functions for the 3DoF listener position(s), are audio signals containing the effects of the VR environment for a specific 3DoF position(s), * refers to approximated sound signals.

[0229] Illustrative aspects and embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination thereof (e.g., in the form of a programmable logic array). Unless otherwise indicated, the algorithms or processes included as part of the invention are not inherently related to any particular computer or other device. In particular, various general-purpose machines may be used in conjunction with programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized device (e.g., integrated circuits) to perform the necessary method steps.Thus, the invention can be implemented in one or more computer programs running on one or more programmable computer systems (e.g., an implementation of any of the elements in the figures), each of which comprises at least one processor, at least one data storage system (including volatile and non-volatile memory devices and / or storage elements), at least one input device or port, and at least one output device or port. The program code is applied to the input data to perform the functions described herein and generate output information. The output information is applied in a known manner to one or more output devices.

[0230] Each such program may be implemented in any desired computer language (including machine language, assembly language, high-level procedural language, logical language, or object-oriented programming language) to communicate with the computer system. In any case, the language may be compiled or interpreted.

[0231] For example, when implemented by sequences of computer software instructions, the various functions and steps of the embodiments of the invention may be implemented by multi-threaded sequences of software instructions running on suitable digital signal processing hardware, in which case the various devices, steps and functions of the embodiments may correspond to portions of the software instructions.

[0232] Each such computer program is preferably stored or loaded onto a storage medium or device (e.g., a solid-state memory device or media, or magnetic or optical media) readable by a general-purpose or special-purpose programmable computer, for configuring and operating the computer, when the storage medium or device is read by the computer system to perform the procedures described herein. The system of the invention can also be implemented as a computer-readable storage medium equipped with (i.e., storing) a computer program, wherein the storage medium so equipped causes the computer system to operate in a specified and predetermined manner to perform the functions described herein.

[0233] A number of illustrative aspects and illustrative embodiments of the present invention have been described above. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present invention. In light of the above teachings, numerous modifications and variations of the present invention are possible. It should be understood that, within the scope of the appended claims, the present invention may be practiced in a manner other than as specifically described herein.

Claims

1. A method for decoding a bit stream, comprising: receiving a bitstream containing encoded audio signal data associated with three degrees of freedom (3DoF) audio generation and metadata associated with six degrees of freedom (6DoF) audio generation; decoding the encoded audio signal data associated with 3DoF to obtain an audio signal with 3DoF; and generating a 3DoF audio signal with generating a sound field based on one of generating a 3DoF audio and generating a 6DoF audio, wherein the 6DoF audio generation generates 6DoF audio signal data based on the 3DoF audio signal and metadata associated with the 6DoF; In this case, the formation of sound with 3DoF includes: ignoring 6DoF related metadata; and Performing 3DoF sound generation using a 3DoF audio signal; In this case, the formation of sound with 6DoF includes: Reconstructing the original audio signals of one or more audio sources from a 3DoF audio signal using the A function -1 inverse transformation; and Performing 6DoF audio generation using reconstructed original audio signals x* and 6DoF-related metadata.

2. The method according to claim 1, characterized in that the encoded audio signal data associated with 3DoF is generated by mapping the original audio signals of one or more audio sources onto corresponding audio objects located on one or more spheres surrounding the default 3DoF listener position, using a transform function A.

3. The method according to claim 1, characterized in that the encoded data of the audio signal associated with the formation of the 3DoF audio comprises at least one of: one or more audio objects, data about the direction of one or more audio objects, and data about the distance of one or more audio objects.

4. The method according to claim 3, characterized in that one or more sound objects are located on one or more spheres surrounding the position of the listener with 3DoF by default.

5. The method according to claim 1, characterized in that the metadata associated with the formation of 6DoF sound indicates one or more default 3DoF listener positions.

6. The method of claim 1, wherein the metadata associated with generating 6DoF audio indicates at least one of: a description of a 6DoF space, directions of sound objects of one or more sound objects, a virtual reality environment, at least a parameter related to at least one of attenuation with increasing range, absorption and reverberations.

7. The method according to claim 1, characterized in that the encoded data of the audio signal associated with the formation of the 3DoF audio is determined on the basis of the original audio signals from one or more audio sources and the conversion function A.

8. The method according to paragraph 7, characterized in that encoded audio signal data associated with generating 3DoF audio is determined by transforming original audio signals from one or more audio sources into 3DoF audio signals using a transform function A, wherein the transform function A maps the original audio signals of the one or more audio sources to corresponding audio objects located on one or more spheres surrounding the default 3DoF listener position.

9. The method according to claim 1, characterized in that the bit stream is compatible with the MPEG-H 3D Audio standard.

10. The method according to paragraph 1, characterized in that the encoded audio signal data associated with the 3DoF audio generation is part of the bitstream payload and Metadata associated with the generation of 6DoF audio is part of one or more bitstream extension containers.

11. A non-transitory machine-readable storage medium storing a computer program containing instructions that, when executed by a processor, cause the processor to perform the method according to paragraph 1.

12. An audio decoding device for decoding a bit stream, comprising: a receiving module for receiving a bit stream containing encoded audio signal data associated with three degrees of freedom (3DoF) audio generation and metadata associated with six degrees of freedom (6DoF) audio generation; a decoding module for decoding the encoded audio signal data associated with 3DoF to obtain an audio signal with 3DoF; and a generating module for generating a 3DoF audio signal with generating a sound field based on one of 3DoF audio generation and 6DoF audio generation, wherein the 6DoF audio generation generates 6DoF audio signal data based on the 3DoF audio signal and metadata associated with the 6DoF; In this case, the formation of sound with 3DoF includes: ignoring 6DoF related metadata; and Performing 3DoF sound generation using a 3DoF audio signal; In this case, the formation of sound with 6DoF includes: Reconstructing the original audio signals of one or more audio sources from a 3DoF audio signal using the A function -1 inverse transformation; and Performing 6DoF audio generation using reconstructed original audio signals x* and 6DoF-related metadata.