An apparatus and method for immersive audio rendering
By determining and encoding rendering information like pre-roll and roll-forward durations in immersive audio bitstreams, the method addresses inconsistent rendering at random access points, ensuring a consistent and deterministic audio experience.
Patent Information
- Application Number
- PCT/EP2024/085855
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-03
AI Technical Summary
Existing immersive audio rendering technologies face challenges in achieving consistent and deterministic rendering experiences across different random access points, leading to audio-visual inconsistencies and loss of acoustic effects due to unspecified rendering behaviors.
The proposed method involves determining rendering information such as pre-roll and roll-forward durations based on audio scene descriptions and environmental parameters, encoding this information in the immersive audio bitstream, and adjusting the rendering process accordingly to ensure consistent playback at any random access point.
This approach ensures a uniform and deterministic subjective experience by aligning rendered audio with the intended playback time, maintaining acoustic effects and preventing inconsistencies across different rendering devices and scenarios.
Smart Images

Figure EP2024085855_03072025_PF_FP_ABST
Abstract
Description
[0001] AN APPARATUS AND METHOD FOR IMMERSIVE AUDIO RENDERING
[0002] Field
[0003] The present application relates to apparatus and methods for immersive audio rendering, and not exclusively for immersive audio rendering in augmented reality and / or virtual reality apparatus. Thus it covers streaming as well as conversational scenarios of delivery of the media.
[0004] Background
[0005] Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is MPEG-I Immersive audio (ISO / IEC 23090-4) which is being currently standardized in ISO / IEC JTC1 SC29 WG6 (MPEG Audio coding WG). This codec is capable of delivery of immersive audio bitstream for 6DoF (six degrees of freedom) audio rendering where the audio coding technology can be, for example, MPEG-H 3DA (ISO / IEC 23008-3 codec audio signals such as audio objects, audio channels and Higher Order Ambisonics or HOA). Another example of such a codec is the future 3GPP immersive voice and audio services (IVAS) codec with 3DoF (three degrees of rotational freedom: yaw, pitch and roll) which is being designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network. Such immersive services include uses for example in immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR) as well as spatial voice communication including teleconferencing. This audio codec is expected to handle the encoding, decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions.
[0006] Summary
[0007] There is provided according to a first aspect a method for assisting immersive audio rendering, the method comprising: obtaining at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determining at least one playback time instance based on the at least one random access point; determining at least one rendering information for the at least one playback time instance; and encoding, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
[0008] The at least one rendering information for the at least one playback time instance may comprise at least one of: a pre-roll-duration parameter; a pre-roll value; pre-roll information; a roll-forward-duration parameter; and a roll-forward information.
[0009] The immersive audio rendering bitstream may be a MPEG-I immersive audio bitstream.
[0010] Determining at least one rendering information for the at least one playback time instance may comprise determining the at least one rendering information based on at least one of: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler; a maximum for at least one of the following: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler; and a user defined parameter.
[0011] Obtaining at least one random access interval associated with the audio scene content may comprise: obtaining an audio scene description; and obtaining from the audio scene description, the at least one random access interval.
[0012] The audio scene description may comprise at least one acoustic environment parameter, and obtaining from the audio scene description file, the at least one random access interval may comprise obtaining the at least one random access interval based on the at least one acoustic environment parameter.
[0013] The audio scene description may be an Encoder Input Format file.
[0014] The acoustic environment parameters may comprise at least one of: a delay line length parameter; a delay line length attenuation filter parameter; a reverberation ratio control filter parameter; a pre-delay line delay parameter; a feedback matrix coefficient parameter; and a directional configuration parameter. Encoding, the determined at least one rendering information in an immersive audio rendering bitstream may comprise encoding the at least one rendering information as part of an immersive audio configuration message.
[0015] According to a second aspect there is provided a method for immersive audio rendering, the method comprising: obtaining an immersive audio rendering bitstream; determining a presence of a rendering random access information within the immersive audio rendering bitstream; determining a type of the rendering random access information when the rendering random access information is present; modifying a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and delivering an unmodified rendered audio to the output device after the determined duration.
[0016] The rendering random access information may comprise at least one of: a pre-roll-duration parameter; a pre-roll value; pre-roll information; a roll-forward- duration parameter; and a roll-forward information.
[0017] The rendering random access information may be defined in terms of one of: a time duration; and a number of audio samples.
[0018] The immersive audio rendering bitstream may be a MPEG-I immersive audio bitstream.
[0019] Obtaining the immersive audio rendering bitstream may comprise obtaining the rendering random access information from one of: an immersive audio configuration bitstream; an immersive audio configuration packet; and an immersive audio Tenderer configuration information packet.
[0020] Modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information may comprise muting for the determined duration the rendered audio.
[0021] Modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information may comprise skipping an audio output for the determined duration.
[0022] Modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information may comprise, if the rendering random access information is a pre-roll-duration parameter, retrieving at least one of an additional immersive audio rendering bitstream and an additional audio bitstream. Modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information may comprise, if the rendering random access information is pre-roll-duration, initiating a rendering of the rendered audio to the output device at least earlier by a pre-roll- duration period prior to the playback random access point, but outputting the unmodified rendered audio to the output device at the playback random access point.
[0023] Modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information may comprise, if the rendering random access information is roll-forward-duration, initiating a rendering of the rendered audio to the output device at the playback random access point, but outputting the unmodified rendered audio to the output device after a period of roll-forward-duration after the playback random access point.
[0024] The rendered audio to an output device may comprise one of: a ODoF rendering; a 3DoF rendering; a 6DoF rendering.
[0025] According to a third aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determine at least one playback time instance based on the at least one random access point; determine at least one rendering information for the at least one playback time instance; and encode, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
[0026] The at least one rendering information for the at least one playback time instance may comprise at least one of: a pre-roll-duration parameter; a pre-roll value; pre-roll information; a roll-forward-duration parameter; and a roll-forward information.
[0027] The immersive audio rendering bitstream may be a MPEG-I immersive audio bitstream. The apparatus caused to determine at least one rendering information for the at least one playback time instance may be caused to determine the at least one rendering information based on at least one of: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler; a maximum for at least one of the following: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler; and a user defined parameter.
[0028] The apparatus caused to obtain at least one random access interval associated with the audio scene content may be caused to: obtain an audio scene description; and obtain from the audio scene description, the at least one random access interval.
[0029] The audio scene description may comprise at least one acoustic environment parameter, and the apparatus caused to obtain from the audio scene description file, the at least one random access interval may be caused to obtain the at least one random access interval based on the at least one acoustic environment parameter.
[0030] The audio scene description may be an Encoder Input Format file.
[0031] The acoustic environment parameters may comprise at least one of: a delay line length parameter; a delay line length attenuation filter parameter; a reverberation ratio control filter parameter; a pre-delay line delay parameter; a feedback matrix coefficient parameter; and a directional configuration parameter.
[0032] The apparatus caused to encode, the determined at least one rendering information in an immersive audio rendering bitstream may be caused to encode the at least one rendering information as part of an immersive audio configuration message.
[0033] According to a fourth aspect there is provided an apparatus for immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain an immersive audio rendering bitstream; determine a presence of a rendering random access information within the immersive audio rendering bitstream; determine a type of the rendering random access information when the rendering random access information is present; modify a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and deliver an unmodified rendered audio to the output device after the determined duration.
[0034] The rendering random access information may comprise at least one of: a pre-roll-duration parameter; a pre-roll value; pre-roll information; a roll-forward- duration parameter; and a roll-forward information.
[0035] The rendering random access information may be defined in terms of one of: a time duration; and a number of audio samples.
[0036] The immersive audio rendering bitstream may be a MPEG-I immersive audio bitstream.
[0037] The apparatus caused to obtain the immersive audio rendering bitstream may be caused to obtain the rendering random access information from one of: an immersive audio configuration bitstream; an immersive audio configuration packet; and an immersive audio Tenderer configuration information packet.
[0038] The apparatus caused to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be caused to mute for the determined duration the rendered audio.
[0039] The apparatus caused to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be caused to skip an audio output for the determined duration.
[0040] The apparatus caused to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be caused to, if the rendering random access information is a pre- roll-duration parameter, retrieve at least one of an additional immersive audio rendering bitstream and an additional audio bitstream.
[0041] The apparatus caused to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be caused to, if the rendering random access information is pre- roll-duration, initiate a rendering of the rendered audio to the output device at least earlier by a pre-roll-duration period prior to the playback random access point, but outputting the unmodified rendered audio to the output device at the playback random access point. The apparatus caused to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be caused to, if the rendering random access information is roll- forward-duration, initiate a rendering of the rendered audio to the output device at the playback random access point, but outputting the unmodified rendered audio to the output device after a period of roll-forward-duration after the playback random access point.
[0042] The rendered audio to an output device may comprise one of: a ODoF rendering; a 3DoF rendering; a 6DoF rendering.
[0043] According to a fifth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising means configured to: obtain at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determine at least one playback time instance based on the at least one random access point; determine at least one rendering information for the at least one playback time instance; and encode, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
[0044] The at least one rendering information for the at least one playback time instance may comprise at least one of: a pre-roll-duration parameter; a pre-roll value; pre-roll information; a roll-forward-duration parameter; and a roll-forward information.
[0045] The immersive audio rendering bitstream may be a MPEG-I immersive audio bitstream.
[0046] The means configured to determine at least one rendering information for the at least one playback time instance may be configured to determine the at least one rendering information based on at least one of: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler; a maximum for at least one of the following: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler; and a user defined parameter.
[0047] The means configured to obtain at least one random access interval associated with the audio scene content may be configured to: obtain an audio scene description; and obtain from the audio scene description, the at least one random access interval.
[0048] The audio scene description may comprise at least one acoustic environment parameter, and the means configured to obtain from the audio scene description file, the at least one random access interval may be configured to obtain the at least one random access interval based on the at least one acoustic environment parameter.
[0049] The audio scene description may be an Encoder Input Format file.
[0050] The acoustic environment parameters may comprise at least one of: a delay line length parameter; a delay line length attenuation filter parameter; a reverberation ratio control filter parameter; a pre-delay line delay parameter; a feedback matrix coefficient parameter; and a directional configuration parameter.
[0051] The means configured to encode, the determined at least one rendering information in an immersive audio rendering bitstream may be configured to encode the at least one rendering information as part of an immersive audio configuration message.
[0052] According to a sixth aspect there is provided an apparatus for immersive audio rendering, the apparatus comprising means configured to: obtain an immersive audio rendering bitstream; determine a presence of a rendering random access information within the immersive audio rendering bitstream; determine a type of the rendering random access information when the rendering random access information is present; modify a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and deliver an unmodified rendered audio to the output device after the determined duration.
[0053] The rendering random access information may comprise at least one of: a pre-roll-duration parameter; a pre-roll value; pre-roll information; a roll-forward- duration parameter; and a roll-forward information. The rendering random access information may be defined in terms of one of: a time duration; and a number of audio samples.
[0054] The immersive audio rendering bitstream may be a MPEG-I immersive audio bitstream.
[0055] The means configured to obtain the immersive audio rendering bitstream may be configured to obtain the rendering random access information from one of: an immersive audio configuration bitstream; an immersive audio configuration packet; and an immersive audio Tenderer configuration information packet.
[0056] The means configured to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be configured to mute for the determined duration the rendered audio.
[0057] The means configured to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be configured to skip an audio output for the determined duration.
[0058] The means configured to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be configured to, if the rendering random access information is a pre-roll-duration parameter, retrieve at least one of an additional immersive audio rendering bitstream and an additional audio bitstream.
[0059] The means configured to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be configured to, if the rendering random access information is pre-roll-duration, initiate a rendering of the rendered audio to the output device at least earlier by a pre-roll-duration period prior to the playback random access point, but outputting the unmodified rendered audio to the output device at the playback random access point.
[0060] The means configured to modify the rendered audio to the output device for the determined duration depending on the type of rendering random access information may be configured to, if the rendering random access information is roll-forward-duration, initiate a rendering of the rendered audio to the output device at the playback random access point, but outputting the unmodified rendered audio to the output device after a period of roll-forward-duration after the playback random access point.
[0061] The rendered audio to an output device may comprise one of: a ODoF rendering; a 3DoF rendering; a 6DoF rendering.
[0062] According to a seventh aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising: obtaining circuitry configured to obtain at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determining circuitry configured to determine at least one playback time instance based on the at least one random access point; determine at least one rendering information for the at least one playback time instance; and encoding circuitry configured to encode, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
[0063] According to an eighth aspect there is provided an apparatus for immersive audio rendering, the apparatus comprising: obtaining circuitry configured to obtain an immersive audio rendering bitstream; determining circuitry configured to determine a presence of a rendering random access information within the immersive audio rendering bitstream; determining circuitry configured to determine a type of the rendering random access information when the rendering random access information is present; modifying circuitry configured to modify a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and delivery circuitry configured to deliver an unmodified rendered audio to the output device after the determined duration.
[0064] According to a ninth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus, for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determining at least one playback time instance based on the at least one random access point; determining at least one rendering information for the at least one playback time instance; and encoding, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
[0065] According to a tenth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus for immersive audio rendering, the apparatus caused to perform at least the following: obtaining an immersive audio rendering bitstream; determining a presence of a rendering random access information within the immersive audio rendering bitstream; determining a type of the rendering random access information when the rendering random access information is present; modifying a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and delivering an unmodified rendered audio to the output device after the determined duration.
[0066] According to an eleventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus, for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determining at least one playback time instance based on the at least one random access point; determining at least one rendering information for the at least one playback time instance; and encoding, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
[0067] According to a twelfth aspect there is provided a non-transitory computer readable medium comprising program instructions for immersive audio rendering, the apparatus caused to perform at least the following: obtaining an immersive audio rendering bitstream; determining a presence of a rendering random access information within the immersive audio rendering bitstream; determining a type of the rendering random access information when the rendering random access information is present; modifying a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and delivering an unmodified rendered audio to the output device after the determined duration.
[0068] According to a thirteenth aspect there is provided an apparatus, for assisting immersive audio rendering, the apparatus comprising: means for obtaining at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; means for determining at least one playback time instance based on the at least one random access point; means for determining at least one rendering information for the at least one playback time instance; and means for encoding, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access.
[0069] According to a fourteenth aspect there is provided an apparatus for immersive audio rendering, the apparatus comprising: means for obtaining an immersive audio rendering bitstream; means for determining a presence of a rendering random access information within the immersive audio rendering bitstream; means for determining a type of the rendering random access information when the rendering random access information is present; means for modifying a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and means for delivering an unmodified rendered audio to the output device after the determined duration.
[0070] According to a fifteenth aspect there is provided a computer readable medium comprising instructions for causing an apparatus, for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determining at least one playback time instance based on the at least one random access point; determining at least one rendering information for the at least one playback time instance; and encoding, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information. According to a sixteenth aspect there is provided a computer readable medium comprising instructions for causing an apparatus, for immersive audio rendering, the apparatus caused to perform at least the following: obtaining an immersive audio rendering bitstream; determining a presence of a rendering random access information within the immersive audio rendering bitstream; determining a type of the rendering random access information when the rendering random access information is present; modifying a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and delivering an unmodified rendered audio to the output device after the determined duration.
[0071] An apparatus comprising means for performing the actions of the method as described above.
[0072] An apparatus configured to perform the actions of the method as described above.
[0073] A computer program comprising program instructions for causing a computer to perform the method as described above.
[0074] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.
[0075] An electronic device may comprise apparatus as described herein.
[0076] A chipset may comprise apparatus as described herein.
[0077] Embodiments of the present application aim to address problems associated with the state of the art.
[0078] Summary of the Figures
[0079] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:
[0080] Figures 1 a to 1 d show example bitstream random access for 6DoF rendering of an audio file with two snare hits 1 .5 seconds apart (T 1 , T2) with playback start time TO;
[0081] Figure 2 shows schematically an example encoder for implementing rendering random access information bitstream according to some embodiments; Figure 3 shows a flow diagram of the example encoder for implementing rendering random access information bitstream as shown in Figure 2 according to some embodiments;
[0082] Figure 4 shows schematically an example rendering random access information determiner as shown in Figure 2 according to some embodiments;
[0083] Figure 5 shows a flow diagram of the example rendering random access information determiner as shown in Figure 4 according to some embodiments;
[0084] Figure 6 shows schematically an example rendering random access information inserter as shown in Figure 2 according to some embodiments;
[0085] Figure 7 shows a flow diagram of the example rendering random access information inserter as shown in Figure 6 according to some embodiments;
[0086] Figures 8 and 9 show example relationships between MPEG-H 3DA bitstream, MPEG-I bitstream, bitstream random access;
[0087] Figures 10a and 10b show example rendering random access for 6DoF rendering of an audio scene according to some embodiments;
[0088] Figures 11a and 11 b show example impact of rendering with and without rendering pre-roll according to some embodiments;
[0089] Figures 12a and 12b show an example system within which some embodiments can be implemented; and
[0090] Figure 13 shows an example device suitable for implementing the apparatus shown in previous figures.
[0091] Embodiments of the
[0092] The following describes in further detail suitable apparatus and possible mechanisms for controlling the rendering of audio scenes.
[0093] As discussed above an immersive audio rendering of an audio scene implements rendering of audio objects, channels and ambisonics elements. The rendering of the audio elements is performed after decoding the received audio elements with a suitable audio codec before providing it to the Tenderer for rendering. The audio elements are rendered with one or more acoustic effects. These acoustic effects include components such as point source direct sound rendering, propagation delay, Doppler, reverberation (early reflections as well as diffuse late reverberation), diffraction. The audio scenes can be rendered such that rendered audio responds to a change of a listener position. The listener can for example be able to move with three degrees of rotational freedom (3DoF), three degrees of rotational freedom with limited translation (3DoF+) and six degrees of freedom (6DoF) within the audio scene.
[0094] The following examples are described with respect to 6DoF movement within the audio scene. These examples are furthermore described with respect to the standardization of immersive audio with six degrees of freedom ongoing in MPEG Audio coding working group (ISO / IEC JTC1 SC29 WG6). MPEG-I 6DoF audio standard (ISO / IEC 23090-4) which is currently under development. The current MPEG-I Architecture and Requirements document specifies the requirements that need to be supported by the standard. Specifically, there is a current requirement within the standard:
[0095] 8. The specification shall support random-access in time (e.g. every 0.5 seconds) and space (e.g. jump within a sub-scene or to a new subscene).
[0096] The current audio standards such as (but not limited to) ISO / IEC 23090-4 WD5 support a method to achieve random access at a bitstream level. In other words, the temporal random access can be achieved with the help of SYNC packets specified in MPEG-H 3DA specification ISO / IEC 23008-3 3rdedition MPEG-H 3D Audio.
[0097] The current working document specifies three new types of MHAS packets to enable delivery of bitstream with scene updates and random access. The three additional MHAS packet types extend the MHASPacketType values which are currently specified for MPEG-H 3DA audio (channel, objects and HOA) and metadata carriage. The three new packet types specified in MPEG-I Immersive audio standard (ISO / IEC 23090-4) are:
[0098] PACTYP_MPEGI_CFG (or MPEG-I config packet);
[0099] PACTYP_MPEGI_UPD (or MPEG-I scene update packet) and PACTYP_MPEGI_PLD (or MPEG-I payload packet).
[0100] The PACTYP_MPEGI_CFG packet type provides the necessary information to configure and start the rendering of the audio scene with the help of the metadata provided in PACTYP_MPEGI_PLD. Any updates to the scene are indicated to the renderer via the PACTYP_MPEGI_UPD. The MPEG-I Immersive audio bitstream and the MPEG-H 3DA bitstream can be synchronized with the help of the PACTYP_SYNC (Synchronization packet). A SYNC packet is inserted prior to the PACTYP_MPEGI_CFG packet, followed by one or more PACTYP_MPEGI_UPD and PACTYP_MPEGI_PLD packets.
[0101] The packets PACTYP_MPEGI_CFG (or MPEG-I config packet), PACTYP_MPEGI_UPD (or MPEG-I scene update packet) and PACTYP_MPEGI_PLD (or MPEG-I payload packet) are relevant with respect to the example implementations described herein with respect to MPEG-H 3DA coded audio and MPEG-I Immersive audio metadata for 6DoF rendering of an audio scene.
[0102] Furthermore, other coding or rendering technologies can implement the example embodiments as disclosed in further detail herein. For example, some embodiments can be implemented when the audio is coded as IVAS or EVS or AAC or any other suitable codec, and the 6DoF rendering is performed by a suitable custom Tenderer, for example implemented with Unity or another suitable rendering engine. Additionally, these other coding or rendering methods can be implemented with appropriate packet or metadata information defined that suits the implementation platform. Thus, the above listed MHAS packets are example packet names.
[0103] The embodiments as discussed herein therefore concern providing a random-access experience for a given playback time such that it is consistent with the experience of a user who started the playout time earlier.
[0104] The term ‘Random Access’ with respect to rendering refers to initializing rendering for any scene at any playback time other than the start time of the scene. The need for random access initialized rendering can occur due to the following:
[0105] • Renderer initialization information is required at the start of joining a broadcast stream.
[0106] • Start playback at any intermediate time of the content timeline (other than the start time) for a typical unicast streaming or delivery scenario.
[0107] Random access can be implemented at two levels.
[0108] • A first level implementation supporting a bitstream structure which facilitates initializing the renderer with basic bitstream support for random access requirement. Typically, this is handled by key frame or sync packets to indicate to the decoder that the bitstream decoding should be started afresh or from scratch. This is not the main focus of this invention.
[0109] • A second level implementation supporting random access with consistent playback experience compared to playback experience without random access (i.e., playback initiated from the start of the scene) or a random access which has occurred at any other earlier playback time. The second level random access can also be referred to as rendering random access.
[0110] The embodiments as discussed in further detail hereafter focus on the second level of supporting random access with consistent playback experience.
[0111] For example, in an MPEG-I implementation a consistent experience for rendering at different random access for the rendering output cannot be guaranteed when implementing only a first level random access. The first level random access is defined via inclusion of SYNC MHAS packets information comprising MPEG-I bitstream and audio data at the random-access joining time of the content timeline.
[0112] This can lead to unspecified rendering behavior at the random-access intervals resulting in the rendered output being different depending on the randomaccess time. This can furthermore cause some acoustic effects to be audible in a way different from the expected effect or be absent altogether, and also result in an inconsistent rendering experience across different Tenderer implementations, since some implementations may attempt to mitigate the acoustic effect differences with inconsistent assumptions across different implementations. The inconsistent placement or absence of an audio effect is not desirable for consistent and interoperable experience. Furthermore, different rendered outputs at different random-access times can be confusing or unpleasant for the end user. Thus, there is a risk of adverse impact on the overall subjective experience due to the different rendered output as well as audio-visual inconsistency.
[0113] Audio-visual inconsistency may happen when e.g., a cannon shot is fired away and the flames and smoke of the explosion is visible, but the associated audio is not rendered. The Audio-visual inconsistency can be caused by the difference of speed of sound (noticeable delay and thus requires modeling) and speed of light (almost no delay). Consequently, the physically modeled sound rendering needs rendering pre-roll functionality if combined with video to model the physically accurate skew in the visual event and the audible sound.
[0114] Additionally, unspecified and inconsistent rendering behavior can be detrimental because typically content creators prefer to have precise control on the expected end user experience. Consequently, the unspecified and inconsistent rendering behavior is detrimental to content creator control of the end user’s subjective experience.
[0115] The concept as discussed in further detail with respect to the following examples and embodiments are apparatus and methods for enabling random access of audio scene rendering in order to achieve a consistent rendered output at any given playback time of the audio scene playback timeline. This is performed by determining and signaling or indicating a rendering pre-roll temporal duration to substantially align the rendered audio at a given playback time compared to the rendered audio obtained by start of rendering at the beginning of the playback timeline. This aims to enable a consistent Tenderer random access experience for uniform and deterministic subjective experience of the 6DoF audio scene rendering and avoids unspecified rendering behavior.
[0116] This is achieved within a suitable encoder or encoder functionality by performing:
[0117] Obtaining an audio scene description file in a suitable format (e.g., Encoder Input Format - EIF);
[0118] Parsing acoustic environment parameters (e.g., RT60, predelay);
[0119] Parsing audio source position in the audio scene;
[0120] Obtaining user reachable region information (if the user reachable region information is specified);
[0121] Determining maximum audio source to listener distance using the audio source position and user reachable region;
[0122] Determining a rendering pre-roll value (which is a maximum of all the above delays) in milliseconds to start 6DoF audio rendering prior to the intended randomaccess consumption start point; and
[0123] Encoding the rendering pre-roll information comprising at least one of: preroll value, roll-forward value in the metadata providing instructions to the audio Tenderer. For example, in MPEG-I based implementations, in the MPEG-I bitstream, where pre-roll value indicates the start of render processing prior to delivering rendered audio output to the listener at desired random-access time and roll-forward determines and signals or indicates the rendering processing duration before delivering the rendered audio output to the listener at a roll-forward duration after the random-access time.
[0124] The encoding furthermore can be configured to signal or indicate whether a pre-roll value is suitable for client-server streaming scenarios where the content prior to the desired random-access time can be requested.
[0125] Additionally in some embodiments the encoding can determine and signal whether the roll-forward value is required for broadcast scenarios where the content prior to the tune in time cannot be requested.
[0126] In some embodiments, the roll-forward duration is used also for client-server unicast streaming on-demand streaming scenarios to simplify the random-access retrieval logic for the player.
[0127] With respect to a suitable Tenderer or Tenderer functionality this can further be achieved by:
[0128] Obtaining the rendering instructions with the audio rendering metadata. For example, in case of MPEG-I implementation, the MPEG-I immersive audio bitstream is obtained for a given random access point in time;
[0129] Determining the presence of pre-roll information comprising the roll-back value or the roll-forward value; initializing and starting rendering but deliver audio output (e.g., to headphones or loudspeakers) only after the roll-forward period, in order to avoid any additional content retrieval or in the example of broadcast scenarios; retrieving audio and MPEG-I bitstream prior to the intended random-access point in time to initialize and start rendering but delivering audio output only after the pre-roll period has elapsed in the example of unicast on demand streaming,
[0130] In some embodiments, the selected random-access point on the playback time (including pre-roll shift to the earlier time) is such that it precedes the start of the audio source emitting sound (or activation instance) in the audio scene, even if such an audio source in the audio scene is not audible to the listener. This enables inclusion of the audio from the audio source of interest in the listener’s audio experience. In some embodiments, the pre-roll information can be time varying over the playback timeline. Consequently, depending on the random-access playback time, the rendering pre-roll information can be different. Furthermore, the pre-roll or rollforward information can be determined by the content creator to consider content specific important aspects for determination of the temporal duration of pre-roll or roll-forward. For example, in some scenarios, only the reverberation parameter information is considered important and used to determine the pre-roll or rollforward duration. In some other scenarios, audio source activation time of a certain source (e.g., a very loud explosion in the scene) is considered important. Thus, the parameter used for determination of the delayed audio output after rendering for a predefined duration can be dependent on the encoder implementation and content creator preferences. The impact of different delay causing phenomena is described in the subsequent paragraphs.
[0131] In further embodiments the pre-roll information can also be spatially specified, i.e., depending on the start position of the listener in the audio scene.
[0132] In other embodiments, the decoder or player is able to select how to use the pre-roll information or skip it.
[0133] In some embodiments, the pre-roll information can be specified for each acoustic effect separately, enabling the player to make the selection of the roll-back or roll-forward period to be used.
[0134] In some embodiments, the pre-roll or roll-forward information can be specified as part of the PACTYP_MPEGI_CFG. This enables the player to perform necessary retrieval of MPEG-I Immersive audio bitstream as well as MPEG-H 3DA bitstream required for the pre-roll. In case of roll-forward information, the Tenderer can be configured to delay the audio output for the specified temporal duration.
[0135] The implementation of the example embodiments thus aims to produce the following advantages:
[0136] Enabling consistent rendering experience for the listener starting audio scene content consumption for any random-access playback time; and
[0137] Enabling consistent rendering experience for unicast streaming where the client retrieves content based on its requirements as well as for broadcast scenario where the client cannot retrieve past content. Aspects of the implementation of random access in spatial audio signal rendering and examples of unspecified and inconsistent playback with random access are first described.
[0138] A 6DoF rendering of audio signals with various acoustic effects such as reverberation (early reflections, diffuse late reverberation, distance propagation delay, etc.) is not instantaneous. The rendering can be impacted by different kinds of delays due to the acoustic modelling performed during the process of rendering the audio. Some examples of delays are:
[0139] Impact of sound propagation delay. The sound propagation delay can be significant and can be observed more prominently when loud audio sources are part of the audio scene. For example, a loud siren or explosion 400 metres away can result in a propagation delay of more than one second (as a sound wave travels approximately 343m / s). In absence of a method to account for this, in case of a random access immediately after the loud explosion, the listener will not hear the explosion. Thus, the experience is not consistent depending on the playback time when the listener starts consuming the 6DoF audio scene.
[0140] Impact of diffuse late reverberation. The reverberation delay, for example, in an acoustic environment with a large predelay or a long RT60 (e.g., a cathedral), can be such that if the listener joins immediately after somebody hits a snare drum, the snare drum hit may or may not be heard depending on the random-access playback time. In an example where a choir is singing in a cathedral, the listener will hear the choir without any reverberation of audio even in a highly reverberant cathedral at the start of the consumption. This would be different for a listener who starts consuming earlier outside the main reverberant volume of the cathedral.
[0141] The above two examples illustrate the effect of random access on rendering leading to inconsistent experience between the listeners.
[0142] Figures 1 a to 1 d illustrate an example scenario of bitstream random access for 6DoF rendering.
[0143] Figure 1 a illustrates a timeline for an example audio scene. The audio scene comprises an audio stream which has a start playback time To 100 and two snare drum hits at playback time Ti 101 and T2 3.
[0144] Figure 1 b illustrates the same timeline for the audio scene but further introduces a random access for content consumption at time TR 115. The timeline furthermore shows that the first of the two snare drum hits at playback time Ti 101 occurs 0.5 seconds 112 after the start playback time To 100, the random access for content consumption at time TR 115 occurs 0.5 seconds 111 after the first of the two snare drum hits at playback time Ti 101 , and furthermore the second of the two snare drum hits at playback time T2 103 occurs 1 .0 seconds 113 after the random access for content consumption at time TR 115. Thus, the second of the two snare drum hits at playback time T2 103 occurs 1 .5 seconds 114 after the first of the two snare drum hits at playback time Ti 101.
[0145] Figure 1 c illustrates a rendered audio output with random access at TR using a conventional Tenderer relying only on bitstream random access. In this example the Tenderer audio output for the second snare drum hit 121 occurs after an inconsistent time delay 125 after the random access for content consumption at time TR 115. In this example only the second snare drum hit effect is rendered.
[0146] Figure 1d furthermore shows the rendered audio output where the playback start time is To 100 where the first snare and second snare drum hits are rendered as shown by the Tenderer audio output for the first snare drum hit 131 and second snare drum hit 121 .
[0147] From Figure 1 the following observations can be made.
[0148] Firstly, the rendered output audio between the TR 115 and T2 103 is inconsistent if the playback start time is TR 115 instead of To 100.
[0149] Secondly, the reverberation tail clearly present if playback start time is To 100 whereas it is completely absent if playback start time is TR 115.
[0150] The above example thus shows clearly the problem as indicated above. The aim of the embodiments described herein is to make the period between TR 115 and T2 103 substantially consistent for playback start time To 100 and TR 115.
[0151] The following embodiments can be described with respect to the implementation within the following three parts of the processing of the audio signals:
[0152] I. An encoding apparatus or method to determine or derive information and subsequent bitstream that ensures consistent playback experience for different random access playback time during the scene timeline.
[0153] II. A rendering apparatus or method to alter behavior for obtaining consistent rendered output audio for the listener. III. A player apparatus or method to enable the presentation of a consistent experience.
[0154] With respect to Figure 2 is shown an example encoder 299 (and with respect to Figure 3 a flow diagram showing the example encoding method) to enable renderer random access for consistent playback.
[0155] The encoder 299 is configured to receive the audio scene description information (e.g., EIF), user reachable region (URR) (which is optional) and the associated audio files that make up the audio scene 200. Furthermore, the encoder 299 is configured to receive the bitstream random access interval 202.
[0156] In some embodiments the encoder comprises a bitstream random access timestamp determiner 201. The bitstream random access timestamp determiner 201 is configured to obtain the bitstream random access interval, which is selected based on the expected random-access requirement. These are used to determine the number of bitstream random-access points for a given audio scene (depending on the audio scene duration). The number of bitstream random access points can be the same or different from the number of Tenderer random access points. These are encoding parameters. The shorter the bitstream random access intervals the amount of excess audio data retrieval in case of pre-roll will be that much smaller. This is expected because the nearest bitstream random access point for a given pre-roll information will be closer compared to the scenario when the bitstream random access interval is large.
[0157] Furthermore, in some embodiments the encoder 299 comprises a last bitstream random access timestamp (for the scene) determiner 203. The last bitstream random access timestamp (for the scene) determiner 203 is configured to determine playback timestamps for each of these bitstream random access points. This timestamp information can be subsequently used to determine the playback time instances for which the rendering random access information should be determined.
[0158] The encoder 299 comprises a rendering random-access information determiner 205. The rendering random-access information determiner 205 is configured to determine, for each of the bitstream random access timestamps, the rendering random access information. Depending on the application scenario, the determination can comprise pre-roll information, roll-forward information or both pre-roll information, roll-forward information.
[0159] The encoder 299 comprises a rendering random access information inserter 207. The rendering random access information inserter 207 is configured to insert and encode the random access information as part of the random-access initialization bitstream for immersive audio rendering. In some embodiments, where the implementation is a MPEG-I Immersive audio codec, this information is encoded and inserted within a PACTYP_MPEGI_CFG MHAS packet.
[0160] Thus, with respect to Figure 3 is shown the obtaining of audio scene description (e.g., EIF), user reachable region (URR) and audio files or audio data 301.
[0161] Additionally, the obtaining of a bitstream random access interval is shown by 303.
[0162] Furthermore, is shown an operation of determining bitstream random access timestamps corresponding to audio scene playback timeline by 305.
[0163] A check can then be made of whether this is the last bitstream random access timestamp for the scene as shown by 307.
[0164] Where it is the last bitstream random access timestamp then there is a stop of the encoding as shown by 309.
[0165] Where it is not the last bitstream random access timestamp then there is a determination of rendering random-access information as shown by 311 .
[0166] When there is pre-roll information required then there is a determination of pre-roll information as shown by 313.
[0167] When there is roll-forward information is required then there is a determination of roll-forward information as shown by 315.
[0168] When there is both pre-roll and roll-forward information required then there is a determination of pre-roll information as shown by 313 and a determination of roll-forward information as shown by 315.
[0169] Then there is an encoding and inserting of the determined rendering random access information in the immersive audio bitstream (e.g., PACTYP_MPEGI_CFG in MPEG-I Immersive audio) 317.
[0170] With respect to Figure 4 is shown in further detail the rendering random access information determiner 204 as shown in Figure 2. As discussed above the determination of rendering random access information is based on the received audio scene information (e.g., EIF and user reachable region) and bitstream random access time.
[0171] For example, the rendering random access information determiner 204 can comprise a timestamp determiner 401. The timestamp determiner 401 in some embodiments is configured to determine at least one audio source activation timestamp and at least one associated deactivation timestamp with respect to the playback time. The audio source activation and deactivation times can be used to determine the temporal duration of relevance after the deactivation. This timestamp information can be employed in highly reverberant acoustic environments or for very loud noises (e.g., explosion, a snare drum hit, clapping, etc.) which are relevant for the storyline. This information can then be used to include the audio source prior to deactivation such that it is sufficient to be included in the subsequent audio processing.
[0172] In some embodiments the rendering random access information determiner 204 can comprise audio source duration determiner 403. The audio source duration determiner 403 can be configured to indicate the duration of the audio source.
[0173] The rendering random access information determiner 204 can furthermore comprise a loudness determiner 405. The loudness determiner 405 can be configured to determine whether the loudness of active sources is relevant for the listener (specifically for the sources that are far away and not immediately audible to the listener). In some embodiments the loudness determiner 405 determines the amount of propagation delay that should be performed in the Tenderer prior to random access to ensure an important far away sound is audible.
[0174] The rendering random access information determiner 204 can furthermore comprise a Doppler effect duration determiner 407. The Doppler effect duration determiner 407 can be configured to determine a maximum additive delay change caused by the Doppler effect for fast moving audio elements. There can in some embodiments be a speed threshold to determine if an audio source is relevant for Doppler effect.
[0175] The rendering random access information determiner 204 can furthermore comprise a farthest audio source determiner 409. The farthest audio source determiner 409 is configured to identify the farthest audio source. The rendering random access information determiner 204 can furthermore comprise a reverberation rendering duration determiner 411 which in some embodiments is configured to determine, based on the acoustic environment properties and acoustic environment dimensions / predelay, the rendering duration prior to random access required to achieve the desired diffuse late reverberation.
[0176] The rendering random access information determiner 204 can furthermore comprise a maximum duration determiner 413. The maximum duration determiner 413 is configured to determine the maximum duration considering all the above elements individually. In case all the effects are not of equal importance, a subset of parameters is considered to determine the desired consistency of the rendered audio output.
[0177] With respect to Figure 5 is shown a flow diagram showing the operations of the rendering random access information determiner 204 according to some embodiments. The first operation is obtaining a scene information, bitstream random access time as shown by 501 in Figure 5.
[0178] Then is shown an optional operation of obtaining user reachable region (if available) by 503 in Figure 5.
[0179] Furthermore, is shown determining the activation and deactivation timestamps of the audio elements with respect to playback time of the audio scene as shown in Figure 5 by 505.
[0180] Then is shown determining the duration after deactivation when the audio is perceptible (e.g., with respect to the nearest possible listener position) as shown in Figure 5 by 507.
[0181] Also is shown determining if the loudness is above a predefined threshold for the active sources (e.g., with respect to the farthest possible listener position) as shown by 509 in Figure 5.
[0182] Also is shown by 511 in Figure 5, determining the Doppler effect duration for a fast-moving audio source to determine the rendering access duration.
[0183] Then is shown determining the farthest audio source of interest to determine the propagation delay effect duration desired to be present during rendering random access, as shown by 513 in Figure 5. Additionally, as shown by 515 in Figure 5 is determining based on the acoustic environment properties and acoustic environment dimensions or predelay the rendering duration required to achieve desired diffuse late reverberation.
[0184] Also as shown in Figure 5 by 517 is determining the maximum duration of all the above or the maximum of the prioritized durations is used as the rollforwardDuration or prerollDuration.
[0185] An extension for the Tenderer random access is proposed to the mpegiSceneConfig syntax element in the MPEG-I Immersive audio standard can be as described in the following. In some embodiments the extension to the configuration packet is implemented because the configuration packet can be used for initializing the Tenderer. This places the prerollDuration or rollforwardDuration information at the beginning of the content. This simplifies the Tenderer implementation for configuring the audible audio output skipping and can result in a convenient implementation. 1. Syntax of mpegiSceneConfigO
[0186] In some embodiments the semantics for the prerollDuration and rollforwardDuration can be the same as delayBufferSize in WD3 clause 5.2.1 .3 Table 8. In such implementations the two values prerollDuration and rollforwardDuration can be specified as 3 bits uimsbf (instead of the 8 bits in the table above).
[0187] In some embodiments the playbackTimestampPresent equal to 1 indicates the presence of playbackTimestamp which is the playback time to which the configuration packet corresponds to. In absence of any other information, the immersive audio bitstream (without any timestamp) which is subsequent to the configuration packet (with timestamp) is defined or corresponds to the timestamp of the configuration packet. In some implementation embodiments, the timestamp for the configuration packet carrying the prerollDuration or rollforwardDuration is determined via higher level (e.g., system level timestamp) information.
[0188] In some embodiments the prerollDuration parameter specifies the minimum time duration (in number of samples or milliseconds or any other suitable unit) for which the Tenderer should move prior to the current random access. In some situations, the actual duration is dictated by the closest bitstream random access point for the audio bitstream (e.g., MPEG-H 3DA audio random access point) and the immersive audio bitstream random access point (e.g., MPEG-I Immersive audio). The bitstream random access point is selected as the nearest point to start prerollDuration and is required in order to achieve pre-roll rendering which is at least equal to or greater than the specified prerollDuration. In other words, a bitstream random access point is selected that is earlier than or equal to the prerollDuration prior to the rendering random access point.
[0189] Furthermore, in some embodiments the rollforwardDuration parameter specifies the minimum time duration (in number of samples or milliseconds or any other suitable unit) for which the Tenderer should perform rendering without delivering (or skipping) the audible output to the listener. The actual duration for which the audible output is skipped may be equal to or greater than the rollforwardDuration due to the bitstream random access for audio and immersive rendering metadata.
[0190] The parameters can be further assisting in their definitions due to the implication of additional audio data retrieval in case of prerollDuration and delayed audible playback output due to the rollfowardDuration.
[0191] In some embodiments the semantics indication can be as follows:
[0192] Table 1 : Value of prerollDuration or rollforwardDuration Table 2 — Value of prerollduration or rollforwardDuration
[0193] In some embodiments, the prerollDuration or rollforwardDuration is specified as the number of audio frames where each frame equals 1024 samples.
[0194] In some embodiments, the prerollDuration and rollforwardDuration are specified in units of milliseconds. In some embodiments the buffer size values (samples or any suitable unit) can be fine-tuned. In such embodiments the actual pre-roll value is therefore not directly sent to avoid bitstream bloating.
[0195] With respect to Figure 6 is shown an example Tenderer implementation according to some embodiments.
[0196] The Tenderer 600 in some embodiments comprises an MPEG-I Immersive audio Tenderer configuration packet parser 601. The MPEG-I Immersive audio Tenderer configuration packet parser 601 receives the PACTYP_MPEGI_CFG (config packet), PACTYP_MPEGI_PLD and optionally PACTYP_MPEGI_UPD to initialize the Tenderer. Additionally, the MPEG-I Immersive audio Tenderer configuration packet parser 601 is configured to parse these packets and extract the information from these packets.
[0197] The Tenderer 600 in some embodiments comprises a renderingRandomAccess determiner 603. The renderingRandomAccess determiner 603 is configured to determine the presence of renderingRandomAccess information. This is present if renderingRandomAccess flag is equal to 1 . Additionally in some embodiments the Tenderer 600 in some embodiments comprises a roll forward audio Tenderer 605 and a pre-roll Tenderer 607. The roll forward audio Tenderer 605 and a pre-roll Tenderer 607 are configured, depending on the application (or player) preference, and whether the prerollDuration or rollforwardDuration information is selected.
[0198] Without loss of generality, it is assumed in this example that both parameters are present. For scenarios where the player preferred information is not present (e.g., rollforwardDuration is preferred by the player but only the prerollDuration information is present), the Tenderer is configured to perform this operation based on available information in the bitstream.
[0199] In case of Tenderer operation with rollforwardDuration, the roll forward audio Tenderer 605 initiates rendering from the closest previous bitstream random access point. However, the output audio is skipped for the duration equal to rollforwardDuration. After the specified duration (i.e., after the skip time), the rendered audio is delivered to the listener. This rendering is particularly suitable for broadcast scenarios where retrieving content prior to the tune in time is not feasible. In some implementation embodiments, this may also be used in case of client server scenarios, if the player prefers to avoid retrieval of audio and MPEG-I bitstream prior to the random-access playback time.
[0200] In case of Tenderer operation with prerollDuration, the pre-roll audio Tenderer 607 initiates rendering from the closest bitstream random access which is at a playback time prior by the amount prerollDuration with respect to the desired random access playback time. This requires retrieval of additional immersive audio rendering bitstream (e.g., MPEG-I bitstream) and audio data (e.g., MPEG-H 3DA audio bitstream). The Tenderer skips output to the listener for at least a duration that is equal to the prerollDuration before delivering the rendered output to the listener.
[0201] In some embodiments the skip duration of audio output is obtained based on a determined difference between the closest earlier bitstream random access playback time that is greater than prerollDuration prior to the desired random access playback time.
[0202] With respect to Figure 7 is shown a flow diagram showing the operation of the Tenderer example shown in Figure 6. The initial operation is one of obtaining Tenderer initialization bitstream (e.g., comprising at least PACTYP_MPEGI_CFG and PACTYP_MPEGI_PLD and optionally PACTYP_MPEGI_UPD MHAS packets) as shown by 701.
[0203] Then is shown parsing MPEG-I Immersive audio Tenderer configuration packet (PACTYP_MPEGI_CFG MHAS packet) as shown by 703.
[0204] Also is shown determining whether renderingRandomAccess is present in PACTYP_MPEGI_CFG MHAS packet as shown by 705.
[0205] Then there is shown a check of whether the application prefers roll-forward or pre-roll as shown by 707.
[0206] For roll forward operations there is an operation of performing audio rendering according to the immersive audio bitstream information (PACTYP_MPEGI_CFG and PACTYP_MPEGI_PLD and optionally PACTYP_MPEGI_UPD MHAS packets) as shown by 709.
[0207] Then is shown the operation of skipping output audio for the temporal duration specified in rollforwardDuration parameter in the PACTYP_MPEGI_CFG MHAS packet as shown by 711 .
[0208] For pre-roll operations there is an operation of retrieving immersive audio bitstream (e.g., MPEG-I Immersive audio bitstream) and audio (e.g., MPEG-H 3DA bitstream) preceding the current random access playback time at least for the duration specified in prerollDuration parameter in the PACTYP_MPEGI_CFG MHAS packet as shown by 710.
[0209] Following this is an operation of skipping an output audio for the temporal duration until the current random access playback time as shown by 712.
[0210] Finally, there is the delivering of the rendered audio to the listener after the end of the skip time as shown by 713.
[0211] With respect to Figure 8 is shown the player operation with respect to the audio data bitstream (e.g., MPEG-H 3DA bitstream) 800, the immersive audio rendering bitstream (e.g., MPEG-I Immersive audio bitstream) 802, 804, 806 and the bitstream random access 805. The audio data bitstream 800 is shown considering MPEG-H 3DA coded audio, but can be illustrated using any other suitable codec such as AAC, EVS, IVAS, etc. The timeline is represented by axis “t”, with respect to the immersive audio bitstream 802, 804, 806. The main motivation to illustrate the audio data bitstream, the rendering metadata bitstream and the timeline together is to highlight the interplay between the playback time, the audio data and the rendering metadata. Figure 8 considers an example rollforwardDuration scenario. The audio and immersive audio rendering bitstream is represented by the MPEG-I 802, 804, 806 and MPEG-H 800 bitstream plotted with respect to the playback timeline “t”. The SYNC packets 852, 854, 856 represent the bitstream random access points (or BR points) corresponding to the different playback timestamps on t axis (BRi 812, BR2 814 and BR3816 on the playback timeline). For simplicity, without loss of generality, the MPEG-I bitstream and MPEG-H bitstream random access points are visualized as overlapping. TR 805 represents the random-access time. The rollforwardDuration 808 is obtained from the config packet of the MPEG-I Immersive audio bitstream (or any other bitstream or rendering initialization packet for any other rendering method). The skip duration 807 is derived based on the temporal duration obtained as a difference between the closest earlier bitstream random access (considering both the audio data bitstream and the rendering metadata bitstream) with respect to TR, which is BR2 and the playback time at the end of rollforwardDuration with respect to TR. Thus, the skip duration in this example amounts to TR + F - BR2. Figure 8 furthermore emphasizes the effect of bitstream random access being the necessary step before achieving the rendering random access. Thus, the rendering random access intervals specified via the rollforwardDuration in practice can be greater than or equal to the rollforwardDuration depending on the bitstream random access intervals.
[0212] Figure 9 illustrates the player operation with respect to the audio data bitstream (e.g., MPEG-H 3DA bitstream) 800, the immersive audio rendering bitstream (e.g., MPEG-I Immersive audio bitstream) 802, 804, 806, the bitstream random access 805. As described above the audio data bitstream 800 is shown considering MPEG-H 3DA coded audio, but can be illustrated using any other suitable codec such as AAC, EVS, IVAS, etc.
[0213] Additionally is shown a timeline represented by axis “t”, with respect to the immersive audio bitstream 802, 804, 806 and the audio data bitstream 800. Figure 9 considers an example prerollDuration scenario. The issue of bitstream random access becomes even more important when considering the prerollDuration because it requires retrieval of audio data and immersive audio rendering metadata prior to the current random-access point. The duration of such roll back 908 dictated by the following equation:
[0214] Roll back duration = max (prerollDuration, closest random access which is greater than or equal to the timestamp obtained by difference between random access time and prerollDuration).
[0215] The skip duration 907 is derived based on the temporal duration obtained as a difference between the closest earlier bitstream random access with respect to TR 805, which is BR2 814 and the prerollDuration with respect to TR. Thus, the skip duration amounts to TR - BR2, where TR - Ap < TR-BR2. Thus, BR2< TR - Ap < TR.
[0216] The Figure 9 highlights the effect of bitstream random access which should be considered over and above the prerollDuration in order to ensure success bitstream random access followed by rendering random access.
[0217] Figures 10a and 10b illustrate an example simulation using an example audio track with two snare drum hits 1001 , 1003 auralized in an immersive audio renderer (e.g., MPEG-I Immersive audio Tenderer) which performs diffuse late reverberation modeling. The audio track has two snare drum hits, 1 .5 seconds apart 1004, at playback time T1 and T2. The example simulates a random access at playback time TR.
[0218] Figures 11 a and 11 b illustrate an example simulation of the output with and without rendering random access with pre-roll. In case of rendering with pre-roll (Figure 11 a) which precedes 1109 the rendering initiation to TR - Ap, the rendered audio 1105, 1101 is skipped from the output to the listener until the playback time TR 1107. The reverberation effect 1105 is clearly audible after TR 1107 until the next snare drum hit. In contrast for rendering without pre-roll (Figure 11 b), there is no audible output 1101 to the listener before T2. This example clearly demonstrates the realization of a consistent rendering experience for a listener who starts consumption at a random-access time TR versus a listener who starts consumption at start of the playback timeline at To.
[0219] With respect to the play timeline determination is an important aspect of performing random access at bitstream level as well as for renderer random access. The key requirement is to determine the random-access point on the playback timeline. Due to the differences between the broadcast and client-server scenarios these can be described separately.
[0220] Broadcast scenarios:
[0221] In broadcast scenarios, there can be scenarios where the playback timestamp is not explicitly signaled. SYNC packets (containing a sync word) are also needed in the broadcast scenario to indicate and detect the correct point to begin parsing the subsequent packets in the incoming bitstream. The player tunes in at the first configuration packet which it receives and subsequently performs the playback time progression based on the audio data sampling frequency and the number of audio packets. This process assumes that the configuration packet, the payload packet carrying data for the Tenderer stage initialization and the update packets are carried together. The update packets carry information only for updates that occur after the configuration packet presence in the playback timeline. This is performed as follows:
[0222] • Starts clock after encountering the first config packet in the broadcast bitstream.
[0223] • Counts the elapsed time based on the audio frames.
[0224] • Counts the skip output duration according to the information in the config packet.
[0225] In some other scenarios, the timestamp can be useful also in case of broadcast scenarios because the payload data may not be streamed with the broadcast stream continuously but only the configuration packets are streamed with regular frequency and the payload data is retrieved separately by the player from a server. Keeping this in mind there is a provision of optionally including the playback timestamp playbackTimestamp in the configuration packet in addition to the prerollDuration and rollforwardDuration in the configuration packet.
[0226] Client-server scenarios:
[0227] The presence of playback timestamp or “scene time” relative to the scene timeline provides direct information to the player to select the random-access point in the scene timeline. The scene timeline information also enables the synchronization of the MPEG-I Immersive audio bitstream with the other bitstreams such as the MPEG-H 3DA audio bitstream by insertion of PACTYP_SYNC prior to the configuration packet. Consequently, the file or streaming segments can contain the bitstream packets in the following order: mpeghiSceneConfig (Current time=0); mpeghiScenePayload ();
[0228] PACTYP_SYNC; mpeghiSceneConfig (Current time=1 s); mpeghiScenePayload ();
[0229] PACTYP_SYNC; mpeghiSceneConfig (Current time=2s); mpeghiScenePayload (); mpeghiScenellpdate (Timestamp=Xs),
[0230] PACTYP_SYNC; mpeghiSceneConfig (Current time=Ns); mpeghiScenePayload ();
[0231] Figures 12a and 12b show schematically an example system where the embodiments are implemented in an encoder device 1901 which performs part of the functionality; writes data into a bitstream 1921 and transmits that for a Tenderer device 1941 , which decodes the bitstream, performs reverberator processing according to the embodiments and outputs audio for headphone listening.
[0232] The encoder side 1901 of Figure 12 can be performed on content creator computers and / or network server computers. The output of the encoder is the bitstream 1921 which is made available for downloading or streaming. The decoder / renderer 1941 functionality runs on an end-user-device, which can be a mobile device, personal computer, sound bar, tablet computer, car media system, home HiFi or theatre system, head mounted display for AR or VR, smart watch, or any suitable system for audio consumption.
[0233] The encoder 1901 is configured to receive the virtual scene description 1900 and the audio signals 1904. The virtual scene description 1900 can be provided in the MPEG-I encoder input format (EIF) or in another suitable format. Generally, the virtual scene description contains an acoustically relevant description of the contents of the virtual scene, and contains, for example, the scene geometry as a mesh or as voxels, acoustic materials, acoustic environments with reverberation parameters, positions of sound sources, and other audio element related parameters such as whether reverberation is to be rendered for an audio element or not.
[0234] The encoder 1901 in some embodiments comprises an Immersive audio rendering metadata determiner (e.g., MPEG-I Immersive audio) 1913 configured to use the audio scene description or EIF together with the audio data by the MPEG- I Immersive audio encoder to generate the MPEG-I Immersive audio bitstream which carries information required for the invention. The information derived is the prerollDuration and the rollforwardDuration.
[0235] The encoder 1901 further comprises a MPEG-H 3D audio encoder 1914 configured to obtain the audio signals 1904 and MPEG-H encode them and pass them to a bitstream encoder 1915.
[0236] The encoder 1901 furthermore in some embodiments comprises a bitstream encoder 1915 which is configured to receive the output of the scene and reverberation payload encoder 1913 and the encoded audio signals from the MPEG-H encoder 1914 and generate the bitstream 1921 which can be passed to the bitstream decoder 1941. The bitstream 1921 in some embodiments can be streamed to end-user devices or made available for download or stored.
[0237] The decoder 1941 in some embodiments comprises a bitstream decoder 1951 configured to decode the bitstream.
[0238] The decoder 1941 in some embodiments comprises a Tenderer random access aware output processor 1959. The Tenderer random access aware output processor 1959 determines the Tenderer random access information and passes this to a Tenderer 1999.
[0239] The Tenderer 1999 performs the rendering in accordance with the renderingRandomAccess metadata described above. The Tenderer initiates rendering according to the prerollDuration or the rollforwardDuration and the bitstream random access points. However, the rendered audio output is not delivered to the listener until after the determined skip duration. At the appropriate time, this data is delivered to the headphones or loudspeaker rendering output. The output buffer 1967 receives the rendered audio output and passes this to the output device (head mounted device) 1970.
[0240] Furthermore, the head pose generator 1957 receives information from a head mounted device 1970 or similar and generates head pose information or parameters which can be passed to the Tenderer 1999.
[0241] The decoder 1941 comprise MPEG-H 3D audio decoder 1954 which is configured to decode the audio signals and pass them to the Tenderer 1999.
[0242] In another implementation embodiment the prerollDuration and rollfowardDuration are applied for multimodal scenes whereby the skip duration of the audio output is replicated for the visual output or any other Tenderer output so that the entire multimodal scene initiates user perceptible rendering according to the prerollDuration or rollforwardDuration information. This information can be signaled to the player or presentation engine by implementing it as an additional value for the common scene description where the prerollDuration and rollforwardDuration are applied for all the modalities in the scene comprising multiple media components.
[0243] With respect to Figure 13 an example electronic device which may be used as any of the apparatus parts of the system as described above. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 2000 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. The device may for example be configured to implement the encoder or the Tenderer or any functional block as described above.
[0244] In some embodiments the device 2000 comprises at least one processor or central processing unit 2007. The processor 2007 can be configured to execute various program codes such as the methods described herein.
[0245] In some embodiments the device 2000 comprises a memory 2011. In some embodiments the at least one processor 2007 is coupled to the memory 2011 . The memory 2011 can be any suitable storage means. In some embodiments the memory 2011 comprises a program code section for storing program codes implementable upon the processor 2007. Furthermore, in some embodiments the memory 2011 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 2007 whenever needed via the memory-processor coupling.
[0246] In some embodiments the device 2000 comprises a user interface 2005. The user interface 2005 can be coupled in some embodiments to the processor 2007. In some embodiments the processor 2007 can control the operation of the user interface 2005 and receive inputs from the user interface 2005. In some embodiments the user interface 2005 can enable a user to input commands to the device 2000, for example via a keypad. In some embodiments the user interface 2005 can enable the user to obtain information from the device 2000. For example, the user interface 2005 may comprise a display configured to display information from the device 2000 to the user. The user interface 2005 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 2000 and further displaying information to the user of the device 2000. In some embodiments the user interface 2005 may be the user interface for communicating.
[0247] In some embodiments the device 2000 comprises an input / output port 2009. The input / output port 2009 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 2007 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
[0248] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example, in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802. X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).
[0249] The input / output port 2009 may be configured to receive the signals. In some embodiments the device 2000 may be employed as at least part of the Tenderer. The input / output port 2009 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar.
[0250] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0251] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
[0252] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory- and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
[0253] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0254] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0255] As used in this application, the term “circuitry” may refer to one or more or all of the following:
[0256] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and
[0257] (b) combinations of hardware circuits and software, such as (as applicable):
[0258] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and
[0259] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and
[0260] I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0261] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
[0262] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements
[0263] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.
Claims
CLAIMS:1 . A method for assisting immersive audio rendering, the method comprising: obtaining at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determining at least one playback time instance based on the at least one random access point; determining at least one rendering information for the at least one playback time instance; and encoding, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
2. The method as claimed in claim 1 , wherein the at least one rendering information for the at least one playback time instance comprises at least one of: a pre-roll-Duration parameter; a pre-roll value; pre-roll information; a roll-forward-duration parameter; and a roll-forward information.
3. The method as claimed in any of claims 1 or 2, wherein the immersive audio rendering bitstream is a MPEG-I immersive audio bitstream.
4. The method as claimed in any of claims 1 to 3, wherein determining at least one rendering information for the at least one playback time instance comprises determining the at least one rendering information based on at least one of: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler;a maximum for at least one of the following: an activation and / or a deactivation timestamp; at least one diffuse late reverberation parameter; a farthest distance audio source with loudness threshold; a speed threshold for doppler; and a user defined parameter.
5. The method as claimed in any of claims 1 to 4, wherein obtaining at least one random access interval associated with the audio scene content comprises: obtaining an audio scene description; and obtaining from the audio scene description, the at least one random access interval.
6. The method as claimed in claim 5, wherein the audio scene description comprises at least one acoustic environment parameter, and obtaining from the audio scene description file, the at least one random access interval comprises obtaining the at least one random access interval based on the at least one acoustic environment parameter.
7. The method as claimed in claim 6, wherein the audio scene description is an Encoder Input Format file.
8. The method as claimed in any of claims 6 to 7, wherein the acoustic environment parameters comprises at least one of: a delay line length parameter; a delay line length attenuation filter parameter; a reverberation ratio control filter parameter; a pre-delay line delay parameter; a feedback matrix coefficient parameter; and a directional configuration parameter.
9. The method as claimed in any of claims 1 to 8, wherein encoding, the determined at least one rendering information in an immersive audio rendering bitstream comprises encoding the at least one rendering information as part of an immersive audio configuration message.
10. A method for immersive audio rendering, the method comprising: obtaining an immersive audio rendering bitstream; determining a presence of a rendering random access information within the immersive audio rendering bitstream; determining a type of the rendering random access information when the rendering random access information is present; modifying a rendered audio to an output device for a determined duration depending on the type of rendering random access information; and delivering an unmodified rendered audio to the output device after the determined duration.
11. The method as claimed in claim 10, wherein the rendering random access information comprises at least one of: a pre-roll-Duration parameter; a pre-roll value; pre-roll information; a roll-forward-duration parameter; and a roll-forward information.
12. The method as claimed in claim 11 , wherein the rendering random access information is defined in terms of one of: a time duration; and a number of audio samples.
13. The method as claimed in any of claims 10 to 12, wherein the immersive audio rendering bitstream is a MPEG-I immersive audio bitstream.
14. The method as claimed in any of claims 10 to 13, wherein obtaining the immersive audio rendering bitstream comprises obtaining the rendering random access information from one of: an immersive audio configuration bitstream; an immersive audio configuration packet; andan immersive audio Tenderer configuration information packet.
15. The method as claimed in any of claims 10 to 14, wherein modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information comprises muting for the determined duration the rendered audio.
16. The method as claimed in any of claims 10 to 15, wherein modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information comprises skipping an audio output for the determined duration.
17. The method as claimed in any of claims 10 to 16, wherein modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information comprises, if the rendering random access information is a pre-roll-duration parameter, retrieving at least one of an additional immersive audio rendering bitstream and an additional audio bitstream.
18. The method as claimed in any of claims 10 to 17, wherein modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information comprises, if the rendering random access information is pre-roll-duration, initiating a rendering of the rendered audio to the output device at least earlier by a pre-roll-duration period prior to the playback random access point, but outputting the unmodified rendered audio to the output device at the playback random access point.
19. The method as claimed in any of claims 10 to 18, wherein modifying the rendered audio to the output device for the determined duration depending on the type of rendering random access information comprises, if the rendering random access information is roll-forward, initiating a rendering of the rendered audio to the output device at the playback random access point, but outputting the unmodified rendered audio to the output device after a period of roll-forward-duration after the playback random access point.
20. The method as claimed in any of claims 10 to 19, wherein the rendered audio to an output device comprises one of: a ODoF rendering; a 3DoF rendering; a 6DoF rendering.21 . An apparatus comprising means for performing the method of any of claims 1 to 20.
22. A computer program comprising instructions, which, when executed by an apparatus, cause the apparatus to perform the method of any of claims 1 to 20.
23. An apparatus for assisting immersive audio rendering, the apparatus comprising means configured to: obtain at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point; determine at least one playback time instance based on the at least one random access point; determine at least one rendering information for the at least one playback time instance; and encode, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
24. An apparatus for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: obtain at least one random access interval associated with the audio scene content, wherein at least the at least one random access interval is used to determine at least one random access point;determine at least one playback time instance based on the at least one random access point; determine at least one rendering information for the at least one playback time instance; and encode, the determined at least one rendering information in an immersive audio rendering bitstream, wherein the rendering information comprises at least rendering random access information.
Citation Information
Patent Citations
Audio decoder, apparatus for generating encoded audio output data and methods permitting initializing a decoder
EP2863386A1
Apparatus and method for efficient object metadata coding
US20170311106A1
Methods, apparatus and systems for generation, transportation and processing of immediate playout frames (IPFS)
US20210335376A1
Fragment-aligned audio coding
US20220167031A1
Audio decoder, audio encoder, method for decoding, method for encoding and bitstream, using a plurality of packets, the packets comprising one or more scene configuration packets defining a temporal evolution of a rendering scenario and comprising a timestamp information
WO2023083920A1