Signaling of restrictions for immersive audio rendering and playback

The described method and apparatus address inefficiencies in immersive audio rendering by using a file format carriage structure to define relationships between audio sources and bitstreams, enabling efficient decoding and playback in augmented and virtual reality environments.

WO2026098962A1PCT designated stage Publication Date: 2026-05-15NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2025-10-22
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for signaling and rendering immersive audio, particularly in augmented and virtual reality environments, leading to inefficiencies in processing and playback of 6 degrees of freedom immersive audio.

Method used

A method and apparatus for defining a file format carriage structure that includes immersive audio bitstreams, coded audio bitstreams, and indicators to control actions associated with immersive audio players, using ISOBMFF file format aware writers to generate and transmit these structures, enabling efficient rendering and playback.

Benefits of technology

Enables efficient rendering and playback of immersive audio by defining relationships between audio sources and bitstreams, allowing for selective decoding and rendering, thereby improving the immersive audio experience in augmented and virtual reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025080451_15052026_PF_FP_ABST
    Figure EP2025080451_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A method for defining a file format carriage structure for assisting immersive audio rendering, the method comprising: obtaining at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtaining at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generating at least one file format carriage by including at least one of the obtained bitstreams; generating at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and including the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SIGNALING OF RESTRICTIONS FOR IMMERSIVE AUDIO RENDERING AND PLAYBACK

[0002] Field

[0003] The present application relates to apparatus and methods for signaling of restrictions for immersive audio rendering and playback, and not exclusively for signaling of restrictions for immersive audio rendering and playback of 6 degrees of freedom immersive audio with file format support in augmented reality and / or virtual reality apparatus.

[0004] Background

[0005] A basic building block in the ISO base media file format is called a box. Each box has a header and a payload. The box header indicates the type of the box and the size of the box in terms of bytes. A box may enclose other boxes, and the ISO file format specifies which box types are allowed within a box of a certain type. Furthermore, the presence of some boxes may be mandatory in each file, while the presence of other boxes may be optional. Additionally, for some box types, it may be allowable to have more than one box present in a file. Thus, the ISO base media file format may be considered to specify a hierarchical structure of boxes.

[0006] According to the ISO base media file format, a file includes media data and metadata that are encapsulated into boxes. Each box is identified by a four character code (400) and starts with a header which informs about the type and size of the box.

[0007] Many files formatted according to the ISO base media file format start with a file type box, also referred to as FileTypeBox or the ftyp box. The ftyp box contains information of the brands labeling the file. The ftyp box includes one major brand indication and a list of compatible brands. The major brand identifies the most suitable file format specification to be used for parsing the file. The compatible brands indicate which file format specifications and / or conformance points the file conforms to. It is possible that a file is conformant to multiple specifications. All brands indicating compatibility to these specifications should be listed, so that a reader only understanding a subset of the compatible brands can get an indication that the file can be parsed. Compatible brands also give a permission for a file parser of a particular file format specification to process a file containing the same particular file format brand in the ftyp box. A file player may check if the ftyp box of a file comprises brands it supports, and may parse and play the file only if any file format specification supported by the file player is listed among the compatible brands.

[0008] In files conforming to the ISO base media file format, the media data may be provided in one or more instances of MediaDataBox (‘mdat‘) and the MovieBox (‘moov’) may be used to enclose the metadata for timed media. In some cases, for a file to be operable, both of the ‘mdat’ and ‘moov’ boxes may be required to be present. The ‘moov’ box may include one or more tracks, and each track may reside in one corresponding TrackBox (‘trak’). Each track is associated with a handler, identified by a four-character code, specifying the track type. Video, audio, and image sequence tracks can be collectively called media tracks, and they contain an elementary media stream. Other track types comprise hint tracks and timed metadata tracks.

[0009] Tracks comprise samples, such as audio or video frames, or metadata frames. For video tracks, a media sample may correspond to a coded picture or an access unit. A media track refers to samples (which may also be referred to as media samples) formatted according to a media compression format (and its encapsulation to the ISO base media file format). A hint track refers to hint samples, containing cookbook instructions for constructing packets for transmission over an indicated communication protocol. A timed metadata track may refer to samples describing referred media and / or hint samples.

[0010] Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is MPEG-I Immersive audio (ISO / IEC 23090-4) which is being currently standardized in ISO / IEC JTC1 SC29 WG6 (MPEG Audio coding WG).

[0011] The current MPEG-I Immersive audio standard (ISO / IEC 23090-4 WD3) renderer [WD, NC318930] supports 6 degrees-of-freedom (6DoF) rendering of audio scenes comprising audio objects, channels and HOA elements in the 6DoF audio scene. The renderer is able to provide a headphone and loudspeaker output at the listener position (and orientation) in the audio scene.

[0012] Summary

[0013] There is provided according to a first aspect a method for defining a file format carriage structure for assisting immersive audio rendering, the method comprising: obtaining at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtaining at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generating at least one file format carriage by including at least one of the obtained bitstreams; generating at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and including the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

[0014] The control action associated with the immersive audio player may comprise at least one of: the at least one coded audio bitstream is to be decoded and provided to an immersive audio renderer; a skipping of rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and the at least one immersive audio bitstream is not to be employed by itself in an immersive audio renderer, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information. The method may further comprise: obtaining at least one coded audio bitstream identifier corresponding to the file format carriage; generating an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene; and including the entity grouping or track grouping structure in the file format carriage.

[0015] The at least one indicator may be explicitly signalled within the entity grouping or track grouping structure.

[0016] The at least one coded audio bitstream may be an encoded MPEG-H 3D audio signal.

[0017] Generating at least one file format carriage by including at least one of the obtained bitstreams may comprise generating at least one file format carriage by a ISOBMFF file format aware file writer.

[0018] Generating an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene may comprise obtaining the entity grouping or the track grouping structure from ISOBMFF.

[0019] The entity grouping or track grouping structure may comprise a GroupsList structure.

[0020] The at least one immersive audio bitstream may comprise: scene configuration MHAS packets; scene payload MHAS packets; and scene update MHAS packets.

[0021] Including the indicator in the at least one file format carriage may comprise adding a payload of scene configuration MHAS packet with MHASPacketType PACTYP_MPEGI_CFG and PACTYP-MPEGLPLD.

[0022] The indicator which carries the at least one immersive audio bitstream identifier may be for a corresponding audio source of the at least one audio source.

[0023] The at least one file format carriage may be a defined file format carriage structure.

[0024] The at least one immersive audio bitstream may be at least one MPEG-I immersive audio bitstream.

[0025] Generating, at least one file format carriage by including the obtained bitstreams may comprise: generating at least one file with box hierarchy and metadata; and packing the obtained bitstreams in the at least one file.

[0026] Including the at least one indicator in the file format carriage may comprise at least one of: including the at least one indicator in the immersive audio bitstream; and including the at least one indicator in the at least one coded audio bitstream.

[0027] According to a second aspect there is provided a method for assisting immersive audio rendering, the method comprising: obtaining at least one file format carriage; obtaining, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; decoding the at least one coded audio bitstream; and implementing the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0028] The control action associated with the immersive audio player may comprise at least one of: providing the decoded at least one coded audio bitstream to an immersive audio renderer; skipping rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and preventing a rendering in an immersive audio renderer employing only the at least one immersive audio bitstream, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information.

[0029] Obtaining, from the at least one file format carriage, at least one indicator identifying the control action associated with an immersive audio player may comprise: parsing the at least one file format carriage to extract the at least one indicator which is configured to describe the relationship between the at least one coded audio bitstream and the at least one immersive audio bitstream and further to configure an audio decoder for decoding the at least one coded audio bitstream and provide the decoded at least one coded audio bitstream directly to the immersive audio player or renderer.

[0030] The method may further comprise: obtaining, from the at least one file format carriage, information regarding a correspondence between the at least one coded audio bitstream and the at least one audio source; configuring the immersive audio player or renderer for immersive audio rendering based on the information regarding the correspondence between the at least one coded audio data bitstream and the at least one audio source based on the at least one file format carriage structure; and rendering at the configured immersive audio player or renderer an immersive audio signal based on the at least one immersive audio bitstream and the decoded at least one coded audio bitstream.

[0031] The at least one coded audio bitstream may be an encoded MPEG-H 3D audio signal.

[0032] The file format carriage may be a ISOBMFF file format aware file.

[0033] The information regarding a correspondence between the at least one coded audio data bitstream and the at least one audio source may be an entity grouping or a track grouping structure.

[0034] The at least one indicator may be an entity grouping or a track grouping structure.

[0035] The entity grouping or a track grouping structure may comprise an ISOBMFF entity grouping or a track grouping structure.

[0036] Implementing the control action on the decoded at least one coded audio bitstream may comprise, based on the at least one indicator identifying that the at least one coded audio bitstream is to be decoded and provided to the immersive audio player: selectively bypassing a coded audio bitstream renderer; and outputting the decoded at least one coded audio bitstream to the immersive audio player.

[0037] According to a third aspect there is provided an apparatus for defining a file format carriage structure for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to:obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generate at least one file format carriage by including at least one of the obtained bitstreams; generate at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

[0038] The control action associated with the immersive audio player may comprise at least one of: the at least one coded audio bitstream is to be decoded and provided to an immersive audio renderer; a skipping of rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and the at least one immersive audio bitstream is not to be employed by itself in an immersive audio renderer, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information.

[0039] The apparatus may be further caused to: obtain at least one coded audio bitstream identifier corresponding to the file format carriage; generate an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene; and include the entity grouping or track grouping structure in the file format carriage.

[0040] The at least one indicator may be explicitly signalled within the entity grouping or track grouping structure.

[0041] The at least one coded audio bitstream may be an encoded MPEG-H 3D audio signal.

[0042] The apparatus caused to generate at least one file format carriage by including at least one of the obtained bitstreams may be caused to generate at least one file format carriage by a ISOBMFF file format aware file writer.

[0043] The apparatus caused to generate an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene may be caused to obtain the entity grouping or the track grouping structure from ISOBMFF.

[0044] The entity grouping or track grouping structure may comprise a GroupsList structure.

[0045] The at least one immersive audio bitstream may comprise: scene configuration MHAS packets; scene payload MHAS packets; and scene update MHAS packets.

[0046] The apparatus caused to include the indicator in the at least one file format carriage may be caused to add a payload of scene configuration MHAS packet with MHASPacketType PACTYP_MPEGI_CFG and PACTYP-MPEGLPLD. The indicator which carries the at least one immersive audio bitstream identifier may be for a corresponding audio source of the at least one audio source.

[0047] The at least one file format carriage may be a defined file format carriage structure.

[0048] The at least one immersive audio bitstream may be at least one MPEG-I immersive audio bitstream.

[0049] The apparatus caused to generate, at least one file format carriage by including the obtained bitstreams may be caused to: generate at least one file with box hierarchy and metadata; and pack the obtained bitstreams in the at least one file.

[0050] The apparatus caused to include the at least one indicator in the file format carriage may be caused to at least one of: include the at least one indicator in the immersive audio bitstream; and include the at least one indicator in the at least one coded audio bitstream.

[0051] According to a fourth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain at least one file format carriage; obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; decode the at least one coded audio bitstream; and implement the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0052] The control action associated with the immersive audio player may comprise at least one of: providing the decoded at least one coded audio bitstream to an immersive audio renderer; skipping rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and preventing a rendering in an immersive audio renderer employing only the at least one immersive audio bitstream, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information.

[0053] The apparatus caused to obtain, from the at least one file format carriage, at least one indicator identifying that the at least one coded audio bitstream is to be decoded and provided to the immersive renderer may be caused to: parse the at least one file format carriage to extract the at least one indicator which is configured to describe the relationship between the at least one coded audio bitstream and the at least one immersive audio bitstream and further to configure an audio decoder for decoding the at least one coded audio bitstream and provide the decoded at least one coded audio bitstream directly to the immersive audio renderer or player.

[0054] The apparatus may further be caused to: obtain, from the at least one file format carriage, information regarding a correspondence between the at least one coded audio bitstream and the at least one audio source; configure the immersive audio renderer or player for immersive audio rendering based on the information regarding the correspondence between the at least one coded audio data bitstream and the at least one audio source based on the at least one file format carriage structure; and render at the configured immersive audio renderer or player an immersive audio signal based on the at least one immersive audio bitstream and the decoded at least one coded audio bitstream.

[0055] The at least one coded audio bitstream may be an encoded MPEG-H 3D audio signal.

[0056] The file format carriage may be a ISOBMFF file format aware file.

[0057] The information regarding a correspondence between the at least one coded audio data bitstream and the at least one audio source may be an entity grouping or a track grouping structure.

[0058] The at least one indicator may be an entity grouping or a track grouping structure.

[0059] The entity grouping or a track grouping structure may comprise an ISOBMFF entity grouping or a track grouping structure.

[0060] The apparatus caused to implement the control action on the decoded at least one coded audio bitstream may be caused to, based on the at least one indicator identifying that the at least one coded audio bitstream is to be decoded and provided to the immersive audio player: selectively bypass a coded audio bitstream renderer; and output the decoded at least one coded audio bitstream to the immersive audio player.

[0061] According to a fifth aspect there is provided an apparatus for defining a file format carriage structure for assisting immersive audio rendering, the apparatus comprising means configured to: obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generate at least one file format carriage by including at least one of the obtained bitstreams; generate at least one indicator identifying a control action associated with an immersive audio player; and include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

[0062] The control action associated with the immersive audio player may comprise at least one of: the at least one coded audio bitstream is to be decoded and provided to an immersive audio renderer; a skipping of rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and the at least one immersive audio bitstream is not to be employed by itself in an immersive audio renderer, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information.

[0063] The means may be further configured to: obtain at least one coded audio bitstream identifier corresponding to the file format carriage; generate an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene; and include the entity grouping or track grouping structure in the file format carriage.

[0064] The at least one indicator may be explicitly signalled within the entity grouping or track grouping structure. The at least one coded audio bitstream may be an encoded MPEG-H 3D audio signal.

[0065] The means configured to generate at least one file format carriage by including at least one of the obtained bitstreams may be configured to generate at least one file format carriage by a ISOBMFF file format aware file writer.

[0066] The means configured to generate an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene may be configured to obtain the entity grouping or the track grouping structure from ISOBMFF.

[0067] The entity grouping or track grouping structure may comprise a GroupsList structure.

[0068] The at least one immersive audio bitstream may comprise: scene configuration MHAS packets; scene payload MHAS packets; and scene update MHAS packets.

[0069] The means configured to include the indicator in the at least one file format carriage may be configured to add a payload of scene configuration MHAS packet with MHASPacketType PACTYP-MPEGLCFG and PACTYP_MPEGI_PLD.

[0070] The indicator which carries the at least one immersive audio bitstream identifier may be for a corresponding audio source of the at least one audio source.

[0071] The at least one file format carriage may be a defined file format carriage structure.

[0072] The at least one immersive audio bitstream may be at least one MPEG-I immersive audio bitstream.

[0073] The means configured to generate, at least one file format carriage by including the obtained bitstreams may be configured to: generate at least one file with box hierarchy and metadata; and pack the obtained bitstreams in the at least one file.

[0074] The means configured to include the at least one indicator in the file format carriage may be configured to at least one of: include the at least one indicator in the immersive audio bitstream; and include the at least one indicator in the at least one coded audio bitstream.

[0075] According to a sixth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising means configured to: obtain at least one file format carriage; obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; decode the at least one coded audio bitstream; and implement the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0076] The control action associated with the immersive audio player may comprise at least one of: providing the decoded at least one coded audio bitstream to an immersive audio renderer; skipping rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and preventing a rendering in an immersive audio Tenderer employing only the at least one immersive audio bitstream, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information.

[0077] The means configured to obtain, from the at least one file format carriage, at least one indicator identifying that the at least one coded audio bitstream is to be decoded and provided to the immersive Tenderer may be configured to: parse the at least one file format carriage to extract the at least one indicator which is configured to describe the relationship between the at least one coded audio bitstream and the at least one immersive audio bitstream and further to configure an audio decoder for decoding the at least one coded audio bitstream and provide the decoded at least one coded audio bitstream directly to the immersive audio renderer or player.

[0078] The means may further be configured to: obtain, from the at least one file format carriage, information regarding a correspondence between the at least one coded audio bitstream and the at least one audio source; configure the immersive audio renderer or player for immersive audio rendering based on the information regarding the correspondence between the at least one coded audio data bitstream and the at least one audio source based on the at least one file format carriage structure; and render at the configured immersive audio renderer or player an immersive audio signal based on the at least one immersive audio bitstream and the decoded at least one coded audio bitstream.

[0079] The at least one coded audio bitstream may be an encoded MPEG-H 3D audio signal.

[0080] The file format carriage may be a ISOBMFF file format aware file.

[0081] The information regarding a correspondence between the at least one coded audio data bitstream and the at least one audio source may be an entity grouping or a track grouping structure.

[0082] The at least one indicator may be an entity grouping or a track grouping structure.

[0083] The entity grouping or a track grouping structure may comprise an ISOBMFF entity grouping or a track grouping structure.

[0084] According to a seventh aspect there is provided an apparatus for defining a file format carriage structure for assisting immersive audio rendering, the apparatus comprising: obtaining circuitry configured to obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtaining circuitry configured to obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generating circuitry configured to generate at least one file format carriage by including at least one of the obtained bitstreams; generating circuitry configured to generate at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and including circuitry configured to include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted. According to an eighth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising: obtaining circuitry configured to obtain at least one file format carriage; obtaining circuitry configured to obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; decoding circuitry configured to decode the at least one coded audio bitstream; and implementing circuitry configured to implement the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0085] According to a ninth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus, for defining a file format carriage structure for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generate at least one file format carriage by including at least one of the obtained bitstreams; generate at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

[0086] According to a tenth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtain at least one file format carriage; obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; decode the at least one coded audio bitstream; and implement the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0087] According to an eleventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus, defining a file format carriage structure for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generate at least one file format carriage by including at least one of the obtained bitstreams; generate at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

[0088] According to a twelfth aspect there is provided a non-transitory computer readable medium comprising program instructions for assisting immersive audio rendering, the apparatus caused to perform at least the following:: obtain at least one file format carriage; obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; decode the at least one coded audio bitstream; and implement the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0089] According to a thirteenth aspect there is provided an apparatus, for defining a file format carriage structure for assisting immersive audio rendering, the apparatus comprising: means for obtaining at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; means for obtaining at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; means for generating at least one file format carriage by including at least one of the obtained bitstreams; means for generating at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and means for including the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

[0090] According to a fourteenth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising: means for obtaining at least one file format carriage; means for obtaining, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; means for decoding the at least one coded audio bitstream; and means for implementing the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0091] According to a fifteenth aspect there is provided a computer readable medium comprising instructions for causing an apparatus, for defining a file format carriage structure for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generate at least one file format carriage by including at least one of the obtained bitstreams; generate at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

[0092] According to a sixteenth aspect there is provided a computer readable medium comprising instructions for causing an apparatus, for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtain at least one file format carriage; obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; at least one indicator identifying a control action associated with an immersive audio player; decode the at least one coded audio bitstream; and implement the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

[0093] An apparatus comprising means for performing the actions of the method as described above.

[0094] An apparatus configured to perform the actions of the method as described above.

[0095] A computer program comprising program instructions for causing a computer to perform the method as described above.

[0096] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.

[0097] An electronic device may comprise apparatus as described herein.

[0098] A chipset may comprise apparatus as described herein.

[0099] Embodiments of the present application aim to address problems associated with the state of the art.

[0100] Summary of the Figures

[0101] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:

[0102] Fig.1 shows a flow diagram of an example of file format encapsulation and decapsulation for storage and streaming applications encapsulation and decapsulation operations with respect to some embodiments;

[0103] Fig.2 shows a flow diagram of example signaling restrictions in the MPEG-H 3D audio bitstream when using MPEG-H 3D audio codec for MPEG-I immersive audio rendering, according to some embodiments;.

[0104] Fig.3 shows a flow diagram of example signaling restrictions in the MPEG-I audio bitstream when using MPEG-H 3D audio codec for MPEG-I immersive audio rendering, according to some embodiments;.

[0105] Fig.4 shows an example apparatus and system configuration for implementing some embodiments; and

[0106] Fig.5 shows an example device suitable for implementing the apparatus shown in previous figures. Embodiments of the Application

[0107] The following describes in further detail suitable apparatus and possible mechanisms for the carriage of immersive audio bitstreams and in some embodiments carriage in a manner that enables storage and streaming of immersive audio bitstream such that it is possible to store the immersive audio bitstream in a file or stream it over a suitable network.

[0108] As discussed above the ‘moov’ box may include one or more tracks, and each track may reside in one corresponding TrackBox (‘trak’). The 'trak' box, in the ISO base media file format, includes in its hierarchy of boxes the SampleTableBox (also known as the sample table or the sample table box). The SampleTableBox contains the SampleDescriptionBox, which gives detailed information about the coding type used, and any initialization information needed for that coding. The SampleDescriptionBox contains an entrycount and as many sample entries as the entry-count indicates. The format of sample entries is track-type specific but derive from generic classes (e.g. VisualSampleEntry, AudioSampleEntry, VolumetricVisualSampleEntry). The type of sample entry form used for derivation the track-type specific sample entry format is determined by the media handler of the track.

[0109] A TrackTypeBox may be contained in a TrackBox. The payload of TrackTypeBox has the same syntax as the payload of FileTypeBox. The content of an instance of TrackTypeBox shall be such that it would apply as the content of FileTypeBox, if all other tracks of the file were removed and only the track containing this box remained in the file.

[0110] Movie fragments may be used, for example, when recording content to ISO files, for example, in order to avoid losing data if a recording application crashes, runs out of memory space, or some other incident occurs. Without movie fragments, data loss may occur because the file format may require that all metadata, for example, a movie box, be written in one contiguous area of the file. Furthermore, when recording a file, there may not be sufficient amount of memory space to buffer a movie box for the size of the storage available, and re-computing the contents of a movie box when the movie is closed may be too slow. Moreover, movie fragments may enable simultaneous recording and playback of a file using a regular ISO file parser. Furthermore, a smaller duration of initial buffering may be required for progressive downloading, e.g., simultaneous reception and playback of a file when movie fragments are used and the initial movie box is smaller compared to a file with the same media content but structured without movie fragments.

[0111] The movie fragment feature may enable splitting the metadata that otherwise might reside in the movie box into multiple pieces. Each piece may correspond to a certain period of time of a track. In other words, the movie fragment feature may enable interleaving file metadata and media data. Consequently, the size of the movie box may be limited and the use cases mentioned above be realized.

[0112] In some examples, the media samples for the movie fragments may reside in an mdat box. For the metadata of the movie fragments, however, a moof box may be provided. The moof box may include the information for a certain duration of playback time that would previously have been in the moov box. The moov box may still represent a valid movie on its own, but in addition, it may include an mvex box indicating that movie fragments will follow in the same file. The movie fragments may extend the presentation that is associated to the moov box in time.

[0113] Within the movie fragment there may be a set of track fragments, including anywhere from zero to a plurality per track. The track fragments may in turn include anywhere from zero to a plurality of track runs, each of which document is a contiguous run of samples for that track (and hence are similar to chunks). Within these structures, many fields are optional and can be defaulted. The metadata that may be included in the moof box may be limited to a subset of the metadata that may be included in a moov box and may be coded differently in some cases. Details regarding the boxes that can be included in a moof box may be found from the ISOBMFF specification.

[0114] A self-contained movie fragment may be defined to consist of a moof box and an mdat box that are consecutive in the file order and where the mdat box contains the samples of the movie fragment (for which the moof box provides the metadata) and does not contain samples of any other movie fragment (i.e. any other moof box). A media segment may comprise one or more self-contained movie fragments. A media segment may be used for delivery, such as streaming, e.g. in MPEG-Dynamic Adaptive Streaming over Hypertext Transfer Protocol (HTTP) (MPEG-DASH).

[0115] The track reference mechanism can be used to associate tracks with each other. The TrackReferenceBox includes box(es), each of which provides a reference from the containing track to a set of other tracks. These references are labelled through the box type (i.e. the four-character code of the box) of the contained box(es). The ISO Base Media File Format contains three mechanisms for timed metadata that can be associated with particular samples: sample groups, timed metadata tracks, and sample auxiliary information. Derived specification may provide similar functionality with one or more of these three mechanisms.

[0116] The TrackGroupBox, which is contained in TrackBox, enables indication of groups of tracks where each group shares a particular characteristic or the tracks within a group have a particular relationship. The box contains zero or more boxes, and the particular characteristic or the relationship is indicated by the box type of the contained boxes. The contained boxes include an identifier, which can be used to conclude the tracks belonging to the same track group. The tracks that contain the same type of a contained box within the TrackGroupBox and have the same identifier value within these contained boxes belong to the same track group.

[0117] The syntax of the contained boxes may be defined through TrackGroupTypeBox as follows: aligned ( 8 ) class TrackGroupTypeBox (unsigned int ( 32 ) track_group_type ) extends FullBox ( track_group_type , version = 0 , flags = 0 ) { unsigned int ( 32 ) track_group_id; / / the remaining data may be speci fied / / for a particular track_group_type }

[0118] The ISO Base Media File Format contains three mechanisms for timed metadata that can be associated with particular samples: sample groups, timed metadata tracks, and sample auxiliary information. Derived specification may provide similar functionality with one or more of these three mechanisms.

[0119] A sample grouping in the ISO base media file format and its derivatives may be defined as an assignment of each sample in a track to be a member of one sample group, based on a grouping criterion. A sample group in a sample grouping is not limited to being contiguous samples and may contain non- adjacent samples. As there may be more than one sample grouping for the samples in a track, each sample grouping may have a type field to indicate the type of grouping.

[0120] Sample groupings may be represented by two linked data structures:

[0121] (1) a SampleToGroupBox (sbgp box) represents the assignment of samples to sample groups; and

[0122] (2) a SampleGroupDescriptionBox (sgpd box) contains a sample group entry for each sample group describing the properties of the group.

[0123] There may be multiple instances of the SampleToGroupBox and SampleGroupDescriptionBox based on different grouping criteria. These may be distinguished by a type field used to indicate the type of grouping. SampleToGroupBox may comprise a grouping_type_parameter field that can be used e.g. to indicate a sub-type of the grouping.

[0124] Per-sample sample auxiliary information may be stored anywhere in the same file as the sample data itself; for self-contained media files, this is typically in a MediaDataBox or a box from a derived specification. It is stored either

[0125] (a) in multiple chunks, with the number of samples per chunk, as well as the number of chunks, matching the chunking of the primary sample data or

[0126] (b) in a single chunk for all the samples in a movie sample table (or a movie fragment). The Sample Auxiliary Information for all samples contained within a single chunk (or track run) is stored contiguously (similarly to sample data). Sample Auxiliary Information, when present, is stored in the same file as the samples to which it relates as they share the same data reference ('dref ) structure. However, this data may be located anywhere within this file, using auxiliary information offsets ('saio') to indicate the location of the data.

[0127] The restricted video ('resv') sample entry and mechanism has been specified for the ISOBMFF in order to handle situations where the file author requires certain actions on the player or renderer after decoding of a visual track.

[0128] Players not recognizing or not capable of processing the required actions are stopped from decoding or rendering the restricted video tracks. The 'resv' sample entry mechanism applies to any type of video codec. A RestrictedSchemelnfoBox is present in the sample entry of 'resv' tracks and comprises an OriginalFormatBox, SchemeTypeBox, and SchemelnformationBox. The original sample entry type that would have been unless the 'resv' sample entry type were used is contained in the OriginalFormatBox. The SchemeTypeBox provides an indication which type of processing is required in the player to process the video. The SchemelnformationBox comprises further information of the required processing. The scheme type may impose requirements on the contents of the SchemelnformationBox. For example, the stereo video scheme indicated in the SchemeTypeBox indicates that when decoded frames either contain a representation of two spatially packed constituent frames that form a stereo pair (frame packing) or only one view of a stereo pair (left and right views in different tracks). StereoVideoBox may be contained in SchemelnformationBox to provide further information e.g. on which type of frame packing arrangement has been used (e.g. side-by-side or top-bottom).

[0129] Files conforming to the ISOBMFF may contain any non-timed objects, referred to as items, meta items, or metadata items, in a meta box (fourCC: ‘meta’), which may also be called MetaBox. While the name of the meta box refers to metadata, items can generally contain metadata or media data. The meta box may reside at the top level of the file, within a movie box (fourCC: ‘moov’), and within a track box (fourCC: ‘trak’), but at most one meta box may occur at each of the file level, movie level, or track level. The meta box may be required to contain a ‘hdlr’ box indicating the structure or format of the ‘meta’ box contents. The meta box may list and characterize any number of items that can be referred and each one of them can be associated with a file name and are uniquely identified with the file by item identifier (itemjd) which is an integer value. The metadata items may be for example stored in the 'idat' box of the meta box or in an 'mdat' box or reside in a separate file. If the metadata is located external to the file then its location may be declared by the DatalnformationBox (fourCC: ‘dinf ). In the specific case that the metadata is formatted using XML syntax and is required to be stored directly in the MetaBox, the metadata may be encapsulated into either the XMLBox (fourCC: ‘xml ‘) or the BinaryXMLBox (fourcc: ‘bxml’). An item may be stored as a contiguous byte range, or it may be stored in several extents, each being a contiguous byte range. In other words, items may be stored fragmented into extents, e.g. to enable interleaving. An extent is a contiguous subset of the bytes of the resource; the resource can be formed by concatenating the extents.

[0130] The ItemPropertiesBox enables the association of any item with an ordered set of item properties. Item properties are small data records. The ItemPropertiesBox consists of two parts: ItemPropertyContainerBox that contains an implicitly indexed list of item properties, and one or more ItemPropertyAssociationBox(es) that associate items with item properties. Item property is formatted as a box.

[0131] A descriptive item property may be defined as an item property that describes rather than transforms the associated item. A transformative item property may be defined as an item property that transforms the reconstructed representation of the image item content.

[0132] An entity may be defined as a collective term of a track or an item. An entity group is a grouping of items, which may also group tracks. An entity group can be used instead of item references, when the grouped entities do not have clear dependency or directional reference relation. The entities in an entity group share a particular characteristic or have a particular relationship, as indicated by the grouping type.

[0133] An entity group is a grouping of items, which may also group tracks. The entities in an entity group share a particular characteristic or have a particular relationship, as indicated by the grouping type.

[0134] Entity groups are indicated in GroupsListBox. Entity groups specified in GroupsListBox of a file-level MetaBox refer to tracks or file-level items. Entity groups specified in GroupsListBox of a movie-level MetaBox refer to movie-level items. Entity groups specified in GroupsListBox of a track-level MetaBox refer to tracklevel items of that track.

[0135] GroupsListBox contains EntityToGroupBoxes, each specifying one entity group. The syntax of EntityToGroupBox may be specified as follows: aligned ( 8 ) class EntityToGroupBox ( grouping_type , version, flags ) extends FullBox ( grouping_type , version, flags ) { unsigned int ( 32 ) group_id; unsigned int ( 32 ) num_entities_in_group ; for ( i=0 ; i<num_entities_in_group ; i++ ) unsigned int ( 32 ) entity_id;

[0136] / / the remaining data may be speci fied

[0137] / / for a particular grouping_type } entityjd is resolved to an item, when an item with itemJD equal to entityjd is present in the hierarchy level (file, movie or track) that contains the GroupsListBox, or to a track, when a track with trackJD equal to entityjd is present and the GroupsListBox is contained in the file level.

[0138] Additionally the current MPEG-I Immersive audio standard (ISO / IEC 23090-4 WD3) defines a normative bitstream for representation of the audio scene details as well as the rendering metadata payloads that are required for generating various acoustic effects. The rendering procedure is also normative to allow a consistent rendered audio as output.

[0139] The MPEG-I immersive audio standard currently excludes discussion on audio waveform or audio data compression. Consequently, another audio compression codec is required to deliver the audio streams related to the one or more audio sources in the MPEG-I immersive audio compatible audio scene.

[0140] For example the audio data may be carried as MPEG-H 3DA coded audio which is fed to the renderer after decoding.

[0141] There is research being implemented exploring integration of Application Programming Interfaces (APIs) for the player to control the MPEG-I audio renderer such that it is aligned with other Tenderers such as, Tenderers for video or haptics media.

[0142] It is possible that the MPEG-I immersive audio bitstream and MPEG-H 3D audio bitstream is transmitted as MHAS (MPEG-H Audio Stream) packets.

[0143] MPEG-H 3DA MHAS packets are described in ISO / IEC 23090-4. Some of the notable packet types relate to Audio Codec Configuration (CFG), Audio Scene Information (ASI), and 3D Audio frame data (3DA) packets. The MPEG-H 3DA packet type can be employed for encoding and decoding audio data when using MPEG-I immersive audio. The MPEG-H 3DA codec is not used for rendering the audio data when used together with MPEG-I immersive audio. Furthermore, the MPEG-H 3DA bitstream does not carry explicit timestamps because the time can be derived based on the number of audio samples and the sampling frequency.

[0144] Carriage of MPEG-H 3D audio in ISOBMFF is defined in ISO / IEC 23008-3. The audio frames that use AudioPreRollQ following the restrictions in subclause 5.5.6 of the ISO / IEC 23008-3 are considered to be IPF (immediate playout frames) and shall be signalled as sync samples according to ISO / IEC 14496-12.

[0145] If independently decodable frames (IF) are signaled (as specified in clause 5.7 of ISO / IEC 23008-3), they shall be signalled by means of AudioPreRollEntry.

[0146] An example Box structures to carry MPEG-H 3D audio, can be as follows: aligned ( 8 ) class MHADecoderConf igurationRecord { unsigned int ( 8 ) configuration Version = 1 ; unsigned int ( 8 ) mpegh3daProf ileLevel lndication; unsigned int ( 8 ) ref erenceChannelLayout ; unsigned int ( 16 ) mpegh3daConf igLength; bit ( 8 *mpegh3daConf igLength) mpegh3daConf ig;

[0147] }

[0148] The semantics of the above function are as follows: configurationversion shall be set to 1 in this version of the specification. mpegh3daProf ileLevel lndication defined in subclause 5.2.2 of ISO / IEC 23008-3. ref erenceChannelLayout Channelconfiguration value in accordance with ISO / IEC 23091-3. mpeghCdaConf igLength length in bytes Of mpegh3daConfig. mpeghCdaConf ig the MPEG-H 3D audio configuration defined in this document.

[0149] Furthermore, as indicated from the above box the following audio sample entry types are defined for

[0150] MPEG-H 3D audio:

[0151] Box Types: ‘mhaC’, ‘mhaT, ‘mha2’

[0152] Container: Sample Table Box (‘stbl’) Mandatory: No

[0153] Quantity: One or more sample entries may be present

[0154] The MHASampieEntry shall contain a MHAConf igurationBox, as defined below. This includes the MHADecoderConfigurationRecord as defined in subclause 20.4 of ISO / IEC 23008-3. If the sample entry type is ‘mha1’, multiple streams shall not be used. If the sample entry name is ‘mha2’, multiple streams may be used.

[0155] If an ‘mha1’ or ‘mha2’ MHASampieEntry is present, each sample of the appropriate track shall contain exactly one mpeghSdaFrame as defined in this document. An optional MPEG4BitRateBox may be present in the MHASampieEntry to signal the bit rate information of the MHA stream. Extension descriptors that should be inserted into the elementary stream descriptor, when used in MPEG-4, may also be present. Other boxes may be present in the MHASampieEntry. When multiple streams are used, the MHADecoderConfigurationRecord for each track shall correspond to the appropriate mpegh3daFrame of that track. class MHAConf igurationBox ( ) extends Box ( 'mhaC' ) { MHADecoderConfigurationRecord MHAConf ig; } class MPEG4BitRateBox ( ) extends Box ( 'btrt ' ) { unsigned int ( 32 ) buf f erSi zeDB ; unsigned int ( 32 ) maxBitrate ; unsigned int ( 32 ) avgBitrate ;

[0156] } class MPEG4ExtensionDescriptorsBox ( ) extends Box ( 'm4ds ' ) {

[0157] Descriptor Descr f O . . 255 ] ; } MHASampieEntry ] ) extends AudioSampleEntry ( 'mhal ' ) {

[0158] MHAConf igurationBox config;

[0159] MPEG4BitRateBox ( ) ; / / optional

[0160] MPEG4ExtensionDescriptorsBox ( ) ; / / optional } MHASampieEntry ] ) extends AudioSampleEntry ( 'mha2 ’ ) {

[0161] MHAConf igurationBox config;

[0162] MPEG4B1 tRateBox ( ) ; / / optional MPEG4ExtensionDescriptorsBox ( ) ; / / optional

[0163] }

[0164] Semantics: channelcount inherited from AudioSampieEntry, shall besettoO (inapplicable).

[0165] The MPEG-H 3D audio decoder is capable of rendering a scene to any given loudspeaker setup. The ref erenceChannelLayout carried in theMHADecoderConf igurationRecord Shall be used to signal the preferred reproduction layout for this stream and replaces the channelcount. config MHADecoderConfigurationRecord.

[0166] Descr is a descriptor which should be placed in in the ElementaryStreamDescriptor when this stream is used in an MPEG-4 systems context. This does not include SLConfigDe scrip tor or Decoderconf igDescriptor, but includes the other descriptors in order to be placed after the SLConfigDe scrip tor. buf f erSizeDB gives the size of the decoding buffer for the elementary stream in bytes. maxBi trate gives the maximum rate in bits / second over any window of 1 second. avgBi trate gives the average rate in bits / second over any window of 1 second.

[0167] MHAS sample entry for streaming or broadcast environments:

[0168] Box Types: ‘mhml’, ‘mhm2’

[0169] Container: Sample Table Box (‘stbl’)

[0170] Mandatory: No

[0171] Quantity: One or more sample entries may be present

[0172] Especially in streaming or broadcast environments based on, e.g. MPEG-DASH or MPEG-H MMT, the MPEG-H 3D audio configuration may change at arbitrary positions of the stream and not necessarily only on fragment boundaries. To enable this use-case the ‘mhml ’ and ‘mhm2’ MHASampleEntry provides an in-band configuration mechanism for MPEG-H 3D audio files. If an ‘mhm1’ MHASampleEntry is present, each sample of the appropriate track shall contain exactly one MHAS packet with the MHASPacketType PACTYP_MPEGH3DAFRAME as defined in Clause 14.4 of ISO / IEC 23008-3. If an ‘mhm2’ MHASampleEntry is present, each sample of the appropriate track shall contain at least one MHAS packet with the MHASPacketType PACTYP_MPEGH3DAFRAME as defined in Clause 14.4 of ISO / IEC 23008-3. In addition, at most one MHAS packet with MHASPacketType PACTYP_MPEGH3DAFRAME belonging to the main stream as defined in subclause 14.6 of ISO / IEC 23008- 3 shall be contained in the sample.

[0173] A sample may contain additional MHAS Packets of other types: if present, an MHAS packet with MHASPacketType PACTYP_MPEGH3DACFG, PACTYP_AUDIOSCENEINFO or PACTYP_AUDIOTRUNCATION shall directly precede the MHAS packet of type PACTYP_MPEGH3DAFRAME.

[0174] MHAS packets with the MHASPacketType PACTYP_CRC16 and PACTYP_CRC32 shall not be present in any sample. Other MHAS packets may be present in a sample.

[0175] The first sample of the movie and the first sample of every fragment (when applicable) shall contain a MHAS packet with the type PACTYP_MPEGH3DACFG followed by an MHAS packet with the Type PACTYP-AUDIOSCENEINFO if present.

[0176] All samples of the movie that contain an MHAS packet of type PACTYP_MPEGH3DACFG shall be sync samples.

[0177] If the movie contains a configuration change, i.e. one of the samples of the movie besides the first sample contains an MHAS packet of type PACTYP_MPEGH3DACFG, all sync samples of the movie shall contain an MHAS packet of type PACTYP_MPEGH3DACFG.

[0178] If the sample entry type is ‘mhm1’, multiple streams shall not be used. If the sample entry name is ‘mhm2’, multiple streams may be used.

[0179] Optional boxes may be present in the MHASampleEntry. Optional boxes for the sample entry type ‘mhm1’ are handled according to the sample entry type is ‘mha1’, optional boxes for the sample entry type is ‘mhm2’ are handled according to the sample entry type is ‘mha2’.

[0180] In contrast to the sample entry types ‘mha1’ and ‘mha2’ the MHAConfigurationBox is optional for the sample entry types ‘mhm1’ and ‘mhm2’ and not mandatory.

[0181] Syntax

[0182] MHASampleEntry / ) extends AudioSampleEntry ( 'mhml ’ ) { MHAConfigurationBox config; / / optional

[0183] MPEG4BitRateBox ( ) ; / / optional

[0184] MPEG4ExtensionDescriptorsBox ( ) ; / / optional

[0185] } MHASampleEntry ( ) extends AudioSampleEntry ( 'mhm2 ’ ) {

[0186] MHAConf igurationBox config; / / optional

[0187] MPEG4BitRateBox ( ) ; / / optional

[0188] MPEG4ExtensionDescriptorsBox ( ) ; / / optional

[0189] Support for multiple streams is defined my MHA multi-stream signaling

[0190] Box Type: ‘maeM’

[0191] Container: MHA sample entry (‘mha1’, ‘mha2’, ‘mhm1’, ‘mhm2’)

[0192] Mandatory: No

[0193] Quantity: Zero or one

[0194] This box provides information on the location of each mae_grouplD in case of splitting the audio scene over multiple streams or files. If multiple streams are used, this box shall be present.

[0195] Syntax aligned (8) class MHAMultiStreamBox ( ) extends FullBox ( 'maeM' , version=0, 0) { unsigned int(l) isMainStream; unsigned int(7) thisStreamID; if (isMainStream) { unsigned int(l) reserved = 0; unsigned int(7) ma e_numG roups ; unsigned int(l) reserved = 0; unsigned int(7) numAuxiliaryStreams ; for (i=0; i< mae_numGroups ; i++) { unsigned int(7) mae_groupID; unsigned int(l) isInMainStream; if ( ! isInMainStream) { unsigned int(l) reserved = 0; unsigned int(7) auxiliaryStreamID;

[0196] }

[0197] } }

[0198] }

[0199] Semantics i sMa inS t re am flag indicating if this is the main stream thi s S t re ami D unique ID of the audio stream in the scope of all available streams in the MHA scene mae_numGroups total number of groups in the MHA scene. This value can have a value between 1 and 127, a minimum number of 1 and a maximum number of 127 groups. This number shall be equal to mae_numGroups in MHAGroupDe f ini t i onBox ( ) numAux i l i arys t re ams total number of auxiliary streams available mae_group i D uniquely defines the group of metadata elements i s i nMa ins t re am if this flag is set to 1 , the audio data related to the group (indicated by mae_group i D) is present in the main stream, otherwise the data is transmitted in an auxiliary stream aux i l i arys t re ami D in case the audio data identified by mae_group i D is an auxiliary stream, this integer identifies the respective auxiliary stream

[0200] The MPEG-I Audio (6DoF rendering) is specified as a set of packets to carry configuration of renderer, scene information rendering metadata, scene changes. These are defined as new MHAS packet types: Scene Config (CFG), Scene Update (UPD), Rendering payload (PLD). Furthermore, MPEG-I packets also carry timestamps (e.g., to specify when a particular update is applied). The MPEG-I immersive audio scene can be a relatively static scene with some dynamic aspects (e.g., animation which is predefined or known only during runtime). For example, a room where most of the content storyline occurs. Thus, MPEG-I immersive audio bitstream (even if huge size) is required during startup (e.g., entire scene is needed during startup). There can be a continuous set of scene changes (which can include rendering metadata changes) over a period of time, these can be delivered after the start of the playback of the scene. For example, streaming of a live performance (data is needed at startup but also over the entire duration of the scene).

[0201] The MPEG-I immersive audio standard has been envisioned to work in conjunction with the MPEG- H 3DA audio standard, where the 6DoF immersive audio rendering bitstream is coded by the former and the audio data by the latter codec. There is file format specification for the MPEG-H 3DA codec however, there is no file format support for carriage of MPEG-I immersive audio bitstream. This currently restricts delivery of the 6DoF immersive audio bitstream as elementary bitstream.

[0202] Immersive Audio Model and Formats (IAMF) is used to provide Immersive Audio content for presentation on a wide range of devices in both streaming and offline applications. These applications include internet audio streaming, multicasting / broadcasting services, file download, gaming, communication, virtual and augmented reality, and others. In these applications, audio may be played back on a wide range of devices, e.g., headphones, mobile phones, tablets, TVs, sound bars, home theater systems, and big screens.

[0203] IAMF defines a model for representing Immersive Audio contents based on Audio Substreams contributing to Audio Elements meant to be rendered and mixed to form one or more presentations

[0204] The model comprises a number of coded Audio Substreams and the metadata that describes how to decode, render, and mix the Audio Substreams for playback. The model itself is codec-agnostic; any supported audio codec may be used to code the Audio Substreams.

[0205] The model includes one or more Audio Elements, each of which consists of one or more Audio Substreams. The Audio Substreams that make up an Audio Element are grouped into one or more Channel Groups. The model further includes Mix Presentations and Parameter Substreams.

[0206] The term 3D audio signal means a representation of sound that incorporates additional information beyond traditional stereo or surround sound formats such as Ambisonics (Scene-based audio), Object-based audio and Channel-based audio (e.g., 3.1.2ch or 7.1 ,4ch).

[0207] The term channel means a component of Scene-based audio, a component of Object-based audio, or a component of Channel-based audio. When used in the context of Channel-based audio, it refers to loudspeaker-based channels.

[0208] The term Immersive Audio (IA) means the combination of 3D audio signals recreating a sound experience close to that of a natural environment.

[0209] The term Audio Substream means a sequence of audio samples, which MAY be encoded with any compatible audio codec.

[0210] The term Channel Group means a set of Audio Substream(s) which is(are) able to provide a spatial resolution of audio contents by itself or which is(are) able to provide an enhanced spatial resolution of audio contents by combining with the preceding Channel Groups.

[0211] The term Audio Element means a 3D audio signal, and is constructed from one or more Audio Substreams (grouped into one or more Channel Groups) and the metadata describing them. The Audio Substreams associated with one Audio Element use the same audio codec.

[0212] The term Mix Presentation means a series of processes to present Immersive Audio contents to endusers by using Audio Element(s). It contains metadata that describes how the Audio Element(s) is(are) rendered and mixed together for playback through physical loudspeakers or headphones, as well as loudness information.

[0213] The term Parameter Substream means a sequence of parameter values that are associated with the algorithms used for reconstructing, rendering, and mixing. It is applied to its associated Audio Element or Mix Presentation. Parameter Substreams may change their values over time and may further be animated; for example, any changes in values may be smoothed over some time duration. As such, they may be viewed as a 1 D signal with different metadata specified for different time durations.

[0214] The term Rendered Mix Presentation means a 3D audio signal after the Audio Element(s) defined in a Mix Presentation is(are) rendered and mixed together for playback through physical loudspeakers or headphones.

[0215] Based on the model, IAMF defines the Immersive Audio Model and Formats (IAMF) architecture as depicted below. For a given input 3D audio signal:

[0216] A Pre-Processor generates the Channel Group(s), Descriptors, and Parameter Substream(s);

[0217] A Codec Encoder generates the coded Audio Substream(s);

[0218] An OBU Packetizer generates an IA Sequence from the coded Audio Substream(s), Descriptors and Parameter Substream(s);

[0219] An OBU Parser outputs the coded Audio Substream(s) and the Parameter Substream(s) from the IA Sequence;

[0220] A Codec Decoder outputs decoded Channel Group(s) after decoding the coded Audio Substream(s);

[0221] An Element Reconstructor re-assembles the Audio Elements by combining the Channel Group(s) guided by Descriptors and Parameter Substream(s);

[0222] A Renderer can be used to render the Audio Elements to a multi-channel or binaural format based on Descriptors;

[0223] A Mixer sums the rendered Audio Elements and applies further mixing parameters guided by the Descriptors and the Parameter Substream(s); and

[0224] A Post-Processor outputs an Immersive Audio by using the Channel Group(s), the Descriptors, and the Parameter Substream(s).

[0225] An IA Sequence is decoded and processed to output an Immersive Audio according to a given playback layout. It includes the following steps but an IA decoder can process the steps in a different order to produce the same result:

[0226] Parsing OBUs to obtain the Descriptors and IA Data;

[0227] Selecting a Mix Presentation to use;

[0228] Decoding and reconstructing one or more Audio Elements that are referenced by the Mix Presentation, and used in the remainder of the steps below: Ambisonics decoding; and Scalable Channel Audio decoding;

[0229] Rendering each Audio Element to the playback layout;

[0230] Applying mixing parameters to the rendered Audio Element;

[0231] Synchronizing and then summing all rendered and individually processed Audio Elements;

[0232] Applying further mixing parameters to the mixed Audio Elements; and Post-processing the output mix to perform loudness normalization and peak limiting.

[0233] The IA decoder may choose to lazily parse OBUs to avoid unnecessarily parsing OBUs that are not used by the selected Mix Presentation.

[0234] An example IA decoder architecture can comprise the following modules that perform the operations described above. The decoder, for example, can comprise:

[0235] An OBU parser depacketizes the IA Sequence to output the Descriptors, Audio Substreams and Parameter Substreams;

[0236] A Codec Decoder for each Audio Substream configured to output the decoded channels;

[0237] An Audio Element Renderer configured to reconstruct the 3D audio signal from decoded channels of Codec Decoders according to Audio Element type (specified Audio Element OBU), and renders the audio channels to the playback layout;

[0238] A Synchronizer configured to synchronize all rendered and individually processed Audio Elements;

[0239] A Mixer configured to sum the synchronized Audio Elements and applies further mixing parameters; and

[0240] A Post-Processor configured to output the Immersive Audio for playback after performing loudness normalization and peak-limiting.

[0241] An IA Sequence comprises a series of OBUs in the sequence of a set of Descriptors followed by their associated IA Data.

[0242] The Descriptors may additionally be repeated redundantly and as frequently as necessary. In this case, the obu_redundant_copy field in their OBU Headers are set to 1. Within an IA Sequence, each OBU in the first Descriptors is regarded as a non-redundant OBU regardless of the value of its obu_redundant_copy.

[0243] A set of Descriptors is placed in the following order regardless of where they appear in the bitstream and it may contain one or more Reserved OBUs. The locations of Reserved OBUs complies with those specified in the section 4 Profiles:

[0244] One IA Sequence Header OBU;

[0245] All Codec Config OBUs;

[0246] All Audio Element OBUs; and

[0247] All Mix Presentation OBUs.

[0248] IA Data comprises a sequence of Audio Frame OBUs, Parameter Block OBUs and Temporal Delimiter OBUs (if they ae present), according to the rules below:

[0249] Audio Frame OBUs and Parameter Block OBUs are ordered by their implied timestamp in the timeline; If there are multiple Audio Frame OBUs that have the same implied start timestamp, they are grouped by Audio Elements;

[0250] A Temporal Delimiter OBU may be inserted at the beginning of a Temporal Unit;

[0251] If Temporal Delimiter OBUs are present, one of them is inserted at the beginning of every Temporal Unit.

[0252] Additionally, the following constraints apply to the Audio Frame OBUs and Parameter Block OBUs:

[0253] Audio Frame OBUs are provided non-redundantly (i.e., obu_redundant_copy = 0), such that for each Audio Substream, there are no two Audio Frame OBUs that are overlapping in time;

[0254] Non-redundant Parameter Block OBUs do not provide data for overlapping time regions.

[0255] If the IAMF configuration changes, a new set of Descriptors are required. In that case, a new IA Sequence of the complete set of Descriptors and their corresponding IA Data follows, in the same order as described above.

[0256] The MPEG-I immersive audio (ISO / IEC 23090-4) is related to carrying immersive audio rendering metadata which includes audio scene information and rendering specific payloads. However, there is no inherent carriage or coding of audio data currently defined in the MPEG-I standard. There are other audio coding and carriage standards such as MPEG-H 3D audio, etc. which may be used to carry the audio data related to the immersive audio scene for rendering with six degrees of freedom (6DoF).

[0257] The MPEG-H 3DA (audio coding) codec can be used for carrying the coded audio data to be used in conjunction with the MPEG-I immersive audio bitstream for rendering of audio scene with six degrees of freedom. In such a scenario, the codec is not expected to be used for rendering the audio data from the MPEG-H 3D audio bitstream. However, the MPEG-H 3D audio standard defines the normative rendering of the compressed audio bitstream. Thus, currently there is no mechanism to inform the MPEG-H 3D audio decoder to indicate that the rendering of the decoded audio should not be performed when used together with MPEG-I immersive audio bitstream. Furthermore, there is a need to inform the media player to use the decoded audio data output from the MPEG-H 3D audio decoder with the MPEG-I immersive audio bitstream.

[0258] The concept which is further detailed with respect to the following embodiments relates to carriage of restrictions for audio data decoding or control action for a immersive audio player more generally. In the following the examples discuss restrictions or restriction control actions but can indicate any suitable control action with respect to the immersive (MPEG-I) audio player or audio renderer. For example in some embodiments this can involve carriage of suitable information or indicators for signalling control actions. In some embodiments the immersive audio player comprises the immersive audio renderer. The following embodiments propose apparatus and methods for describing the conditions regarding employing an audio bitstream when used in conjunction with immersive audio rendering (with six degrees of freedom), where the restriction can be at least one of the following: perform decoding of the audio and skip rendering. In this restriction decoding can deliver un-rendered output of the decoded audio samples (e.g., PCM samples). For example, the restriction can control the renderer to decode the audio data with a MPEG-H 3D audio codec but deliver un-rendered audio channels, objects and HOA signals; and / or perform rendering of the decoded audio data (with six degrees of freedom) only when the coded audio data is associated with immersive audio rendering bitstream.

[0259] In some embodiments, in case of signaling the restrictions in the MPEG-H 3D audio bitstream, the decoder is configured to output unrendered audio data comprising audio objects, channels, HOA and associated metadata.

[0260] In some embodiments the signaling of the restrictions can be explicit, for example provided by file format structure or suitable indicator or information within the file format structure.

[0261] In some embodiments this restriction can be signalled by employing a generic sample entry ‘resa’.

[0262] In some further embodiments this restriction or control action is provided or signalled by using a TrackGroup that indicates that all contained tracks should employ a restricted rendering (in other words skipping rendering and outputting the extracted audio data).

[0263] In some embodiments this restriction (or control action) is provided or signalled using an EntrityGroup that indicates that all contained terms should employ a restricted rendering (in other words skipping the internal rendering operations)

[0264] In some embodiments the signaling of the restrictions can be explicit or implicit depending on at least one reference to the audio tracks. For example, when an audio track is referenced in or from a MPEG- I audio track then the audio track should employ a restricted rendering operation.

[0265] This enables the audio decoder to identify when to skip the audio rendering and enables the media player to determine when to perform immersive rendering with six degrees of freedom for the received audio data.

[0266] In some embodiments, in case of restrictions signaling in the carriage format encapsulation, the MPEG-H 3D audio codec skips performing any audio rendering and the decoder output is the unrendered audio data.

[0267] In some further embodiments, in scenarios implementing combined rendering of coded audio together with immersive audio bitstreams (with six degrees of freedom), the MPEG-H 3D audio decoder is configured to skip or disable rendering in order to enable an immersive audio renderer to perform the rendering. The same can also be applied in case of IAMF coded audio data where the audio streams are decoded only to obtain uncompressed audio data for the audio elements as PCM signals. In some embodiments if the coded audio data is IVAS, the IVAS data which is decoded is instantiated or otherwise labeled or indicated with an ‘external’ output format (indicating that the rendering is to be performed external to the decoder).

[0268] In some embodiments, the 6D0F rendering is not initiated or performed until the presence of MPEG- I immersive audio rendering is available in conjunction with the coded audio data.

[0269] In such embodiments audio rendering employing the MPEG-H 3D audio decoder (with associated processing and power consumption cost) can be avoided and the MPEG-H 3D audio decoder limited to decoding the coded audio data related to an immersive audio scene with six degrees of freedom (e.g., to be rendered by MPEG-I immersive audio renderer). This is can be applicable for audio from IVAS or IAMF coded audio data.

[0270] The embodiments are described initially with respect to the method steps for implementing the embodiments and the an example overall end to end system suitable for implementing the embodiments.

[0271] Following this is described example file format definitions to enable the implementation of some embodiments.

[0272] Further to this is described example file format definitions and bitstream example definitions to enable the implementation of some further embodiments.

[0273] A flow diagram showing example operations for implementing some embodiments and describing the carriage of an immersive audio rendering bitstream for an audio scene and the coded audio data bitstream is presented in Fig.1 . The encapsulation to file format enables the carriage of data because the information about the file format carriage structures is available to the media player application to determine the related audio tracks required for rendering the audio sources in an immersive audio scene.

[0274] For example in some embodiments the MPEG-I immersive audio bitstream and coded audio data bitstreams are received or otherwise obtained and furthermore made available to the file writer for encapsulation as shown in Fig.1 by 101.

[0275] Following this a ISOBMFF File Format aware file writer or suitable means can be configured to generate files with appropriate box hierarchy, metadata and packs the bitstream in the file. In other words, the operation can comprise writing file format structures for audio tracks corresponding to audio sources in the MPEG-I immersive audio bitstream as shown in Fig.1 by 103.

[0276] This can then be followed by a file writer application or suitable means being configured to extract at least one coded audio data bitstream identifier corresponding to the file format carriage structures. Having extracted these identifiers then an entity grouping or track grouping structure from ISOBMFF is used to generate the data structure which carries relationship information between the track identifier of the coded audio data with the corresponding audio sources in the immersive audio scene. This data structure is furthemore included in the appropriate file format carriage structure such as GroupsList. In other words the operation can comprise writing a file for carriage of audio tracks and include the file format structures that describe any relationships between the audio sources in the MPEG-I immersive audio bitstream as shown in Fig.1 by 105. For example, the file format structures can comprise Entity Groups, Track Groups, Track References to indicate the relations.

[0277] This can then be followed by file write application or suitable means to comprise writing a file for carriage of audio tracks and include the file format structures that describe the rendering restriction of audio tracks, as shown in Fig.1 by 106.

[0278] These above operations 101 , 103, 105, 106 can be considered to be the FF encapsulation operation as shown in Fig.1 by 100.

[0279] Subsequently the media is made available or is ready for storage and transmission as shown in Fig.1 by 107.

[0280] Following receiving the media at the media player at the media consumption device, the audio data and the immersive audio rendering data is extracted from the file format encapsulation. In other words, there are suitable means configured to parse file format boxes and extract audio data and immersive audio bitstream as shown in Fig.1 by 109.

[0281] Subsequently, information regarding the relationship between the coded audio data and the audio sources is read from the file format carriage structures to configure the decoding of the coded audio data such that it is rendered in the expected manner. For example, MPEG-H 3D audio coded shall output unrendered audio based on the restriction signaling as well as MPEG-I immersive audio bitstream is not rendered or considered as an audio that can be rendered by itself. In this operation therefore a suitable means is configured parse file format structures that describe the relationship between the coded audio data and immersive bitstream to configure the audio decoder to skip audio rendering (and provide the decoded audio data to a MPEG-I renderer) as shown in Fig.1 by 111 .

[0282] The media player subsequently configures the correct audio data to be rendered for the corresponding audio source in the audio scene. In other words, any suitable means is configured to provide the correct audio stream for the correct audio source in the audio scene as shown in Fig.1 by 113.

[0283] With respect to Fig.2 and Fig.3 there are shown example flow diagrams which describe aspects or embodiments for signaling restrictions when using MPEG-H 3D audio codec for MPEG-I immersive audio rendering.

[0284] Fig.2 for example shows an example flow diagram of some embodiments where signaling indication in the MPEG-H 3D audio bitstream is employed to output unrendered audio object, channel and HOA content.

[0285] This, in some embodiments, can be signaled as part of the PACTYP_MPEGH3DACFG MHAS packet which carries the structure mpegh3daConfig. A single flag bit can signal (for example if set to 1) that the MPEG-H 3D audio codec is configured or controlled to output unrendered audio according to the output interface (such as described in clause 17.10 of ISO / IEC 23008-3 MPEG-H 3D Audio) for audio objects, channels and HOA content.

[0286] These embodiments result in an implementation which clearly indicates in the MPEG-H 3D audio bitstream whether the audio content is expected to be used with the MPEG-I immersive audio renderer and consequently there is implicit or inbuilt support identifying the 6DoF rendering with MPEG-I decoder / renderer. In these embodiments explicit signaling is required in the 3D audio bitstream. This can in some embodiments involve a rewriting of (legacy) MPEG-H 3D audio files or be aware of the intended usage while transmitting / including the PACTYP_MPEGH3DACFG configuration MHAS packet in the bitstream.

[0287] Thus, for example the initial operation is one of receiving a MPEG-I immersive audio bitstream and coded audio data bitstream(s) (e.g., MPEG-H 3D audio) as shown in Fig.2 by 201.

[0288] Furthermore, is shown the operation of adding an indication in MPEG-H 3D audio bitstream to inform the decoder for providing unrendered audio as output from MPEG-H 3D audio decoder as shown in Fig.2 by 203.

[0289] Then is the operation of performing decoding of MPEG-H 3D audio bitstream without rendering and providing an output in accordance with the interface describing the unrendered audio output and associated metadata as shown in Fig.2 by 205.

[0290] Then any unrendered decoded audio objects, channels and HOA can be used as an input for MPEG- I immersive audio rendering as shown in Fig.2 by 207.

[0291] Fig.3 furthermore shows a further example flow diagram of alternative embodiments where signaling for output of unrendered audio from MPEG-H 3D audio decoder can be performed by signaling the condition that a media player controls a configuration of the MPEG-H 3D audio decoder to output unrendered audio (according to clause 17.10 of ISO / IEC 23008-3 MPEG-H 3D Audio) if it is grouped or associated with the MPEG-I immersive audio bitstream.

[0292] This is because the MPEG-I immersive audio decoder requires the unrendered audio data from the MPEG-H 3D audio decoder output in order to perform the rendering of the immersive audio scene with six degrees of freedom.

[0293] The advantage of such embodiments is that there is no need to modify the existing MPEG-H 3D audio bitstreams. However, a change is required in the configuration implementation for the MPEG-H 3D audio decoder and expects the player to be able to configure the MPEG-H 3D audio decoder based on the presence of the MPEG-I immersive audio renderer (or immersive audio player).

[0294] Thus, for example, the initial operation is one of receiving a MPEG-I immersive audio bitstream and coded audio data bitstream(s) (e.g., MPEG-H 3D audio) as shown in Fig.3 by 201. Furthermore, is shown the operation of adding indication in MPEG-I immersive audio bitstream to inform the player to configure the MPEG-H 3D audio decoder for providing unrendered audio as output from the MPEG-H 3D audio decoder as shown in Fig.3 by 303.

[0295] Then is the operation of performing decoding of MPEG-H 3D audio bitstream without rendering and providing an output in accordance with the interface describing the unrendered audio output and associated metadata as shown in Fig.3 by 205.

[0296] Then any unrendered decoded audio objects, channels and HOA can be used as an input for MPEG- I immersive audio rendering as shown in Fig.3 by 207.

[0297] Fig .4 shows schematically an example system where the embodiments are implemented in an encoder device 403 on a suitable sever computer 401 which performs part of the functionality; writes data into a bitstream and which can be transmitted via a network 411 , for a playback device 421 , which decodes the bitstream, performs reverberator processing according to the embodiments and outputs audio for headphone 491 listening.

[0298] The encoder side 403 of Fig.4 can be performed on content creator computers and / or network server computers 401. The output of the encoder is the Immersive audio and coded audio with File Format encapsulation for storage and streaming 410 which is made available for downloading or streaming, for example via the network 411 .

[0299] The MPEG-I immersive decoder / renderer 441 functionality (or immersive audio player) runs on an end-user-device or playback device 421 , which can be a mobile device, personal computer, sound bar, tablet computer, car media system, home HiFi or theatre system, head mounted display for AR or VR, smart watch, or any suitable system for audio consumption.

[0300] The encoder 403 is configured to receive the encoder input format or any other scene description information (for example, Scene.XML) 402 and the audio signals 404. Generally, the scene description contains an acoustically relevant description of the contents of the scene, and contains, for example, the scene geometry as a mesh or as voxels, acoustic materials, acoustic environments with reverberation parameters, positions of sound sources, and other audio element related parameters such as whether reverberation is to be rendered for an audio element or not.

[0301] The encoder 403 in some embodiments comprises an MPEG-H 3D audio encoder 407 configured to receive the audio signals 404 and is configured to determine the coding and delivery method for the scene relevant audio data and prepare coded version of the scene relevant audio data and generate the encoded audio bitstream 408. This can be passed to an elementary bitstream generator 411 .

[0302] The elementary bitstream generator 411 in some embodiments is configured to generate an elementary bitstream comprising MPEG-I audio and MPEG-H 3D audio data which is passed over the network 411 to the playback device 421 . The encoder 403 in some embodiments comprises an Immersive audio encoder (e.g., MPEG-I Immersive audio) 405 which is configured to receive the scene description (for example in a format such as EIF) 402 and the encoded audio / or audio signals 406. The immersive audio encoder (e.g., MPEG-I Immersive audio) 405 thus outputs the encoded MPEG-I audio to a File Format level mapper and restriction generator (for playback for 6DoF rendering) 409. The Immersive audio encoder (e.g., MPEG-I Immersive audio) 405 can furthermore provide the restriction generator indication at bitstream level to the encoded audio bitstream 408.

[0303] File Format level mapper and restriction generator (for playback for 6DoF rendering) 409 is configured to receive the encoded MPEG-I audio and encoded audio bitstream 406 and output an Immersive audio and coded audio with File Format encapsulation for storage and streaming 412.

[0304] In some embodiments File Format level mapper and restriction generator 409 is configured to include the bitstream signaling in a suitable MPEG-H 3D audio configuration bitstream packet (MHAS packet with MHASPacketType PACTYP_MPEGH3DACFG) to enable a configuration of the receiver decoder (and control the unrendered audio output).

[0305] The File Format level mapper and restriction generator (for playback for 6DoF rendering) 409 in some embodiments comprises a file format encapsulation module configured to write the file format boxes to enable storage and streaming of MPEG-I immersive audio bitstream and MPEG-H 3D audio bitstreams.

[0306] The file format encapsulation module furthermore can be configured to write the file format box which describes the relation between the coded audio data bitstreams and the audio sources in the audio scene.

[0307] The playback device 421 , is configured to implement a presentation engine 431 instance in some embodiments.

[0308] The presentation engine 431 is further configured to implement or comprise an extractor 435. The extractor 435 is configured to extract, from the File Format encapsulation mapping information and restriction conditions. Additionally the extractor 435 is configured to operate as a decapsulation module to extract audio data and immersive audio bitstream from the carriage format encapsulation.

[0309] The extracted Immersive audio bitstream (MPEG-I) 451 can be passed to the MPEG-I immersive renderer 441 (or immersive audio player) and the renderer pipeline 445.

[0310] The presentation engine 431 is further configured to implement or comprise a decoder 437 configured to receive and decode the Elementary bitstream with restriction conditions.

[0311] The coded audio data bitstream (MPEG-H 3DA) 453 can then be passed to a decoder / audio buffer 455 for generating decoded audio data bitstream (MPEG-H 3DA).

[0312] The presentation engine 431 further comprises a MPEG-I immersive renderer 441 (which can also be known as an interactive renderer or any suitable renderer) and the renderer pipeline 445 which is then configured to receive the decoded MPEG-I data and the MPEG-H 3DA data as well as head pose (position and orientation) generated by the head pose generator 457 and generate a renderer output audio delivered with aligned audio-visual entities and pass this to the output buffers 459 prior to output to the headset, headphones, or suitable transducer system (shown as head mounted device HMD 491) for outputting the audio scene to the listener.

[0313] In such a manner during the file writing is it possible to implement restrictions signaling. Additionally the embodiments permit the apparatus to define or describe the relationship between the audio data and the audio sources in the audio scene at the level of file format carriage information. In some embodiments the definition or describing of the relationship to enable appropriate retrieval and decoding without having to performing cascaded retrieval depending on the decoding of the immersive audio bitstream can be employed.

[0314] For example, in some implementations there is no need to download the immersive audio bitstream prior to knowing which coded audio bitstreams are related.

[0315] In some embodiments where the signalling is implemented by a file format using ISOBMFF then a MPEG-I immersive audio track is employed which carries a bitstream comprising a rendering payload and audio scene information for rendering audio with six degrees of freedom. The storage of MPEG-I immersive audio bitstream track utilizes the existing capabilities of the ISO base media file format and derived specifications.

[0316] The MPEG-I immersive audio bitstream tracks in some embodiments are represented in the file as restricted audio and can employ a generic restricted sample entry 'resa' with additional requirements that a SchemeTypeBox is present in RestrictedSchemelnfoBox and scheme_type is set to 'mpgi'. Furthermore in the track header the trackjnjnovie flag is set to 0, to indicate that this track should not be presented alone.

[0317] The signaling of the trackjnjnovie flag to zero is employed because MPEG-I immersive audio bitstream does not carry any audio data. Consequently, it cannot be rendered as any other audio track. In some embodiments, as an alternative the MPEG-I immersive audio bitstream track is marked or identified as a metadata track with a new track reference, for example called “6doa”, to further reduce the ambiguity that the contents of the bitstream are not audio data.

[0318] In another implementation embodiment, if MPEG-I immersive audio bitstream is associated with the MPEG-H 3D audio tracks via entity grouping or track grouping, the MPEG-H 3D audio decoder is configured to output unrendered audio objects, channels, HOA and associated metadata.

[0319] In embodiments where the signaling is implemented as part of the elementary bitstream then there is provided a suitable indication for using the MPEG-H 3D audio decoder such that the audio output of the decoder is the unrendered audio objects, channels, HOA and associated metadata.

[0320] This can be achieved in some embodiments by signaling the use of the decoder in a restricted manner to output unrendered audio according to clause 17.10 in ISO / IEC 23008-3. This can, for example, be implemented by signaling the information as part of the codec configuration packet. An example of a modified syntax of mpegh3daConfig() carried in PACTYP_MPEGH3DACFG cfg_reserved reserved, value shall be ignored (ISO / IEC 23008-3). cfg_unrendered_mode a value equal to 1 indicates the MPEG-H 3D audio decoder should be configured to output unrendered audio objects, channels; HOA and associated metadata in accordance with clause 17.10 of ISO / IEC 23008-3. A value equal to 0 indicates the MPEG-H 3D audio decoder shall use the decoder for rendering as specified in ISO / IEC 23008-3.

[0321] The signaling of the control action or restriction, for example a control action to skip rendering of the coded audio data after decoding can be implemented by one or more of the following methods:

[0322] Signal via the coded audio bitstream to the decoder / renderer to skip rendering of the audio elements in the bitstream;

[0323] Signal via file format encapsulation of the coded audio bitstream the decoder / renderer to skip rendering of the audio elements in the file format encapsulation;

[0324] Signal via the immersive audio bitstream to skip rendering of the coded audio bitstream after decoding. This is helpful in avoiding problems to old legacy decoder implementations of the coded audio bitstream. The drawback is that only the new decoder implementations for the coded audio bitstream will be able to perform in this manner;

[0325] Signal via file format encapsulation of immersive audio bitstream to skip rendering the coded audio bitstream after decoding. This is helpful in avoiding problems to legacy coded audio bitstream decoders which may not be designed to consider additional signaling.

[0326] In some embodiments there can be employed Restriction signaling or control action signalling for IAMF (Immersive Audio Model and Formats) coded audio data. IAMF is the immersive audio coded audio format from AOM (Alliance for Open Media) which describes a method for carriage, decoding and of immersive audio data. In some embodiments the audio data for immersive audio rendering may intend to use another renderer than the one specified by the IAMF, then the IAMF decoder should be configured to output PCM signals for the one or more audio elements.

[0327] The main decoding steps for IAMF that are described in clause 2.2. of IAMF specification (Immersive Audio Model and Formats (aomediacodec.github.io) are as follows: an OBU parser is configured to obtain the audio substreams and parameter substreams; a codec decoder outputs the channel groups after decoding the audio substreams; an audio element reconstructor is used to generate the elements for rendering; a renderer then renders the audio elements; and the output of the renderer is mixed and post processed.

[0328] In order to enable bitstream signaling for rendering skipping, in some embodiments a new profile can be defined which skips rendering steps and stops at providing PCM signals from the reconstructed audio elements or decoder output with channels as output. In some further implementation embodiments, each codec configuration OBU is configured to carry signaling to indicate whether to output unrendered audio. This avoids rendering of the audio data by the audio compression codec instead of the rendering and mixing with the IAMF codec.

[0329] Furthermore, in some embodiments, signaling to indicate is implemented employing the IA sequence header OBU to configure the subsequently used audio codecs to be used in a manner that skips the rendering but only performs audio decoding to output PCM signals.

[0330] The terms restriction and restriction signalling can be understood to be examples of control actions or signalling of control actions.

[0331] With respect to Fig.5 an example electronic device which may be used as any of the apparatus parts of the system as described above. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 2000 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. The device may for example be configured to implement the encoder or the renderer or any functional block as described above.

[0332] In some embodiments the device 2000 comprises at least one processor or central processing unit 2007. The processor 2007 can be configured to execute various program codes such as the methods described herein.

[0333] In some embodiments the device 2000 comprises a memory 2011 . In some embodiments the at least one processor 2007 is coupled to the memory 2011 . The memory 2011 can be any suitable storage means. In some embodiments the memory 2011 comprises a program code section for storing program codes implementable upon the processor 2007. Furthermore, in some embodiments the memory 2011 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 2007 whenever needed via the memory-processor coupling.

[0334] In some embodiments the device 2000 comprises a user interface 2005. The user interface 2005 can be coupled in some embodiments to the processor 2007. In some embodiments the processor 2007 can control the operation of the user interface 2005 and receive inputs from the user interface 2005. In some embodiments the user interface 2005 can enable a user to input commands to the device 2000, for example via a keypad. In some embodiments the user interface 2005 can enable the user to obtain information from the device 2000. For example, the user interface 2005 may comprise a display configured to display information from the device 2000 to the user. The user interface 2005 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 2000 and further displaying information to the user of the device 2000. In some embodiments the user interface 2005 may be the user interface for communicating. In some embodiments the device 2000 comprises an input / output port 2009. The input / output port 2009 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 2007 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.

[0335] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example, in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).

[0336] The input / output port 2009 may be configured to receive the signals.

[0337] In some embodiments the device 2000 may be employed as at least part of the renderer. The input / output port 2009 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar.

[0338] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0339] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.

[0340] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory- and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.

[0341] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0342] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.

[0343] As used in this application, the term “circuitry” may refer to one or more or all of the following:

[0344] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and

[0345] (b) combinations of hardware circuits and software, such as (as applicable):

[0346] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and

[0347] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and

[0348] I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0349] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.

Claims

43CLAIMS:

1. A method for defining a file format carriage structure for assisting immersive audio rendering, the method comprising: obtaining at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtaining at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generating at least one file format carriage by including at least one of the obtained bitstreams; generating at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and including the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

2. The method as claimed in claim 1 , wherein the control action associated with the immersive audio player comprises at least one of: the at least one coded audio bitstream is to be decoded and provided to an immersive audio renderer; a skipping of rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and the at least one immersive audio bitstream is not to be employed by itself in an immersive audio renderer, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information.

3. The method as claimed in any of claims 1 or 2, further comprising: obtaining at least one coded audio bitstream identifier corresponding to the file format carriage; generating an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene; and including the entity grouping or track grouping structure in the file format carriage.

4. The method as claimed in claim 3, wherein the at least one indicator is explicitly signalled within the entity grouping or track grouping structure.

5. The method as claimed in any of claims 1 to 4, wherein the at least one coded audio bitstream is an encoded MPEG-H 3D audio signal.

446. The method as claimed in any of claims 1 to 5, wherein generating at least one file format carriage by including at least one of the obtained bitstreams comprises generating at least one file format carriage by a ISOBMFF file format aware file writer.

7. The method as claimed in claim 6 when dependent on claim 3, wherein generating an entity grouping or a track grouping structure to define a data structure which carries relationship information between a track identifier of the at least one coded audio bitstream with a corresponding audio source in the immersive audio scene comprises obtaining the entity grouping or the track grouping structure from ISOBMFF.

8. The method as claimed in any of claims 1 to 7, wherein the at least one immersive audio bitstream comprises: scene configuration MHAS packets; scene payload MHAS packets; and scene update MHAS packets.

9. The method as claimed in any of claims 1 to 8, wherein the indicator which carries the at least one immersive audio bitstream identifier is for a corresponding audio source of the at least one audio source.

10. The method as claimed in any of claims 1 to 9, wherein the at least one file format carriage is a defined file format carriage structure.11 . The method as claimed in any of claims 1 to 10, wherein the at least one immersive audio bitstream is at least one MPEG-I immersive audio bitstream.

12. The method as claimed in any of claims 1 to 1 1 , wherein generating, at least one file format carriage by including the obtained bitstreams comprises: generating at least one file with box hierarchy and metadata; and packing the obtained bitstreams in the at least one file.

13. The method as claimed in any of claims 1 to 12, wherein including the at least one indicator in the file format carriage comprises at least one of: including the at least one indicator in the immersive audio bitstream; and including the at least one indicator in the at least one coded audio bitstream.4514. A method for assisting immersive audio rendering, the method comprising: obtaining at least one file format carriage; obtaining, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; and at least one indicator identifying a control action associated with an immersive audio player; decoding the at least one coded audio bitstream; and implementing the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

15. The method as claimed in claim 14, wherein the control action associated with the immersive audio player comprises at least one of: providing the decoded at least one coded audio bitstream to an immersive audio renderer; skipping rendering of the at least one coded audio bitstream by a coded audio bitstream renderer; and preventing a rendering in an immersive audio renderer employing only the at least one immersive audio bitstream, as the at least one immersive audio bitstream comprises only rendering related metadata or scene information.

16. The method as claimed in any of claim 14 or 15, wherein obtaining, from the at least one file format carriage, at least one indicator identifying the control action associated with an immersive audio player comprises: parsing the at least one file format carriage to extract the at least one indicator which is configured to describe the relationship between the at least one coded audio bitstream and the at least one immersive audio bitstream and further to configure an audio decoder for decoding the at least one coded audio bitstream and provide the decoded at least one coded audio bitstream directly to the immersive audio renderer.

17. The method as claimed in any of claims 14 to 16, further comprising: obtaining, from the at least one file format carriage, information regarding a correspondence between the at least one coded audio bitstream and the at least one audio source; configuring the immersive audio player for immersive audio rendering based on the information regarding the correspondence between the at least one coded audio data bitstream and the at least one audio source based on the at least one file format carriage structure; andrendering at the configured immersive audio player an immersive audio signal based on the at least one immersive audio bitstream and the decoded at least one coded audio bitstream.

18. The method as claimed in any of claims 14 to 17, wherein the at least one coded audio bitstream is an encoded MPEG-H 3D audio signal.

19. The method as claimed in claim 17, wherein the information regarding a correspondence between the at least one coded audio data bitstream and the at least one audio source is an entity grouping or a track grouping structure.

20. The method as claimed in any of claims 14 to 19, wherein at least one indicator is an entity grouping or a track grouping structure.

21. The method as claimed in any of claims 14 to 20, wherein implementing the control action on the decoded at least one coded audio bitstream comprises, based on the at least one indicator identifying that the at least one coded audio bitstream is to be decoded and provided to the immersive audio player: selectively bypassing a coded audio bitstream renderer; and outputting the decoded at least one coded audio bitstream to the immersive audio player.

22. An apparatus for defining a file format carriage structure for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generate at least one file format carriage by including at least one of the obtained bitstreams; generate at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

23. An apparatus for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: obtain at least one file format carriage; obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; and at least one indicator identifying a control action associated with an immersive audio player; decode the at least one coded audio bitstream; and implementing the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.

24. An apparatus for defining a file format carriage structure for assisting immersive audio rendering, the apparatus comprising means configured to: obtain at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; obtain at least one coded audio bitstream, the at least one coded audio bitstream being different from the at least one immersive audio bitstream related to the immersive audio scene; generate at least one file format carriage by including at least one of the obtained bitstreams; generate at least one indicator, the at least one indicator identifying a control action associated with an immersive audio player; and include the at least one indicator in the file format carriage, wherein the file format carriage is stored and / or transmitted.

25. An apparatus for assisting immersive audio rendering, the apparatus comprising means configured to: obtain at least one file format carriage; obtain, from the at least one file format carriage: at least one immersive audio bitstream, the immersive audio bitstream defining at least one audio source in an immersive audio scene; at least one coded audio bitstream; and at least one indicator identifying a control action associated with an immersive audio player; decode the at least one coded audio bitstream; and implement the control action on the immersive audio player with respect to the decoded at least one coded audio bitstream.