Method and apparatus for efficient delivery of 6DOF rendering
The method optimizes 6DoF audio delivery by generating scenes with multiple content sets at varying expression levels, addressing bandwidth issues by selectively delivering necessary HOA sources based on listener position, achieving up to 60% savings.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2024-04-03
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for delivering 6-degree-of-freedom (6DoF) audio scenes with Higher-Order Ambisonics (HOA) sources are bandwidth-intensive due to the need for multiple channels, leading to network congestion and latency, especially when only a subset of channels is required for rendering based on listener position.
A method and apparatus that generate audio scenes with multiple content sets at different expression levels, allowing selective delivery of full-order, reduced-order, and primary representations based on listener location, optimizing bandwidth usage by acquiring only necessary audio data for rendering.
Reduces bandwidth consumption by up to 60% while maintaining audio quality by ensuring only required HOA sources are delivered at appropriate orders, minimizing network congestion and latency.
Smart Images

Figure 2026513616000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for efficient distribution in 6-degree-of-freedom rendering, and generally, with a 6-degree-of-freedom system of audio captured by a microphone array. [Background technology]
[0002] The MPEG-I Immersive Audio standard (ISO / IEC 23090-4) has been refined and includes new features. One aspect is its function in efficiently delivering the content required for rendering. Higher-Order Ambisonics (HOA) is one of the signal types that provides a significant technology in this standard. Efficient content delivery is important for complex scenes with a large amount of HOA sources.
[0003] The HOA audio format is bandwidth-intensive due to the large number of channels required to achieve high-quality playback of audio scenes. For example, a 4th-order ambisonic representation consists of 25 channels. MPEG-H 3DA audio supports up to 6th order or 49 channels. When representing an audio scene with 6 degrees of freedom (6DoF), i.e., in the case of HOA translation, multiple HOA sources are required for good quality playback.
[0004] As a rule of thumb, the more sampling points (i.e., capture positions) there are for a given user-reachable area (or listening space size) of an audio scene, the better the playback quality while performing HOA translation compared to a scene with fewer sampling points. The number and order of HOA sources are determined by the scene creator. Some example audio scenes are as follows: An audio scene with nine sampling positions corresponding to nine HOA sources using fourth-order ambisonics results in 225 (a very large number) channels.
[0005] An audio scene representation with three HOA sources using 4th-order ambisonics results in 75 (another large number) channels.
[0006] The example above highlights the need for efficient delivery methods for audio scenes containing multiple HOA sources, emphasizing that the greater the number of HOA sources or the higher their order, the greater the need for efficient delivery. [Overview of the Initiative]
[0007] According to a first aspect, a method is provided for generating an audio scene having at least two audio content sets, comprising: obtaining a media description including information that identifies the location of an audio source in the audio scene associated with the at least two audio content sets; encoding the at least two audio content sets so that they are configured to be output at at least two expression levels; and outputting the encoded at least two audio content sets for an audio scene, wherein the output encoded at least two audio content sets comprises a first encoded at least two audio content sets encoded at a first expression level of at least two, and a second first encoded at least two audio content sets encoded at a second expression level of at least two.
[0008] At least two audio content sets may include at least two higher ambisonic audio sources.
[0009] At least two audio content sets may include metadata associated with each of at least two higher ambisonic audio sources.
[0010] The metadata associated with each of at least two higher-order ambisonic audio sources may include their location within the audio source.
[0011] Media descriptions may further include information that identifies at least two levels of representation.
[0012] At least two levels of representation may include at least two of the following: full-order representations of at least two audio content sets, reduced or lower-order representations of at least two audio content sets, and primary representations of at least two audio content sets.
[0013] The full-order representation of at least two audio content sets can be all channels of at least two higher-order ambisonic audio sources. The reduced or lower-order representation of at least two audio content sets can be a selection of channels from at least two higher-order ambisonic audio sources. The primary representation of at least two audio content sets can be a selection of the first four channels of at least two higher-order ambisonic audio sources.
[0014] The step of outputting at least two encoded audio content sets for an audio scene may include receiving at least one request relating to the audio scene, the request specifying levels of representation for at least two audio content sets, selecting at least two encoded audio content sets for the levels of representation for the at least two audio content sets specified in at least one request, and outputting the selected at least two encoded audio content sets for the levels of representation for the at least two audio content sets specified in at least one request.
[0015] Obtaining a media description that includes information identifying the location of audio sources within an audio environment associated with at least two audio content sets may further include outputting the media description to the device, and outputting the encoded at least two audio content sets for an audio scene may further include outputting them to the device.
[0016] The device could be a spatial audio player configured to generate an audio scene.
[0017] A second aspect provides a method for generating an audio scene having at least two audio content sets, comprising: obtaining a media description including information that identifies the location of an audio source in the audio scene associated with the at least two audio content sets; obtaining a listener location in the audio scene; determining a source representation level for the at least two audio content sets based on the location of an audio source associated with the at least two content sets and the listener location; obtaining at least two encoded audio content sets based on the source representation levels, wherein the at least two encoded audio content sets include a first encoded at least two audio content sets encoded at a first of at least two representation levels, and a second encoded at least two audio content sets encoded at a second of at least two representation levels; and generating a spatialized audio output based on the obtained encoded at least two audio content sets.
[0018] At least two levels of representation may include full-order representations of at least two audio content sets, reduced or lower-order representations of at least two audio content sets, and at least two primary representations of at least two audio content sets.
[0019] The media description may further specify an Ambisonics order for at least two audio content sets.
[0020] Based on the source representation level, obtaining at least two encoded audio content sets may include obtaining all the encoded channel representations for the at least two encoded audio content sets when the source representation level is a full-order representation.
[0021] Based on the source representation level, obtaining at least two encoded audio content sets may include obtaining the selected encoded channel representations for the at least two encoded audio content sets when the source representation level is a reduced or low-order representation.
[0022] Based on the source representation level, obtaining at least two encoded audio content sets may include obtaining the first four selected encoded channel representations for the at least two encoded audio content sets when the source representation level is a first-order representation.
[0023] Generating a spatialized audio output based on the obtained at least two encoded audio content sets may include performing signal interpolation by processing the channel signals of the obtained at least two encoded audio content sets based on the listener position, and generating a spatialized audio output based on the performed signal interpolation.
[0024] Based on the source representation level, obtaining at least two encoded audio content sets may include: generating a request for the representation of at least two encoded audio content sets, where the request includes the source representation level, and receiving the requested representation of at least two encoded audio content sets.
[0025] For at least two audio content sets, determining the source representation level based on the positions of the audio sources and the listener position associated with the at least two content sets may further include determining state information related to content set rendering. Obtaining at least two encoded audio content sets based on the state information related to content set rendering may include obtaining at least two encoded audio content sets based on the state information related to content set rendering.
[0026] The state information related to content set rendering is configured to specify whether the content set can be used as one of a 6 - degree - of - freedom higher - order ambisonics audio source during generating a spatialized audio output and an object rendering audio source during generating a spatialized audio output.
[0027] According to a third aspect, an apparatus is provided for generating an audio scene having at least two audio content sets, comprising: means for obtaining a media description including information that identifies the location of an audio source in an audio scene associated with at least two audio content sets; encoding the at least two audio content sets such that the at least two audio content sets are output at at least two expression levels; and outputting the encoded at least two audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first expression level of at least two, and a second first encoded at least two audio content sets encoded at a second expression level of at least two.
[0028] At least two audio content sets may include at least two higher ambisonic audio sources.
[0029] At least two audio content sets may include metadata associated with each of at least two higher ambisonic audio sources.
[0030] The metadata associated with each of at least two higher-order ambisonic audio sources may include their location within the audio source.
[0031] Media descriptions may further include information that identifies at least two levels of representation.
[0032] At least two levels of representation may include at least two of the following: full-order representations of at least two audio content sets, reduced or lower-order representations of at least two audio content sets, and primary representations of at least two audio content sets.
[0033] The full-order representation of at least two audio content sets can be all channels of at least two higher-order ambisonic audio sources. The reduced or lower-order representation of at least two audio content sets can be a selection of channels from at least two higher-order ambisonic audio sources. The primary representation of at least two audio content sets can be a selection of the first four channels of at least two higher-order ambisonic audio sources.
[0034] Means configured to output at least two encoded audio content sets for an audio scene may be configured to receive at least one request relating to an audio scene, wherein the request specifies a level of representation for at least two audio content sets, select at least two encoded audio content sets for the level of representation for the at least two audio content sets specified in the at least one request, and output at least two encoded audio content sets selected for the level of representation for the at least two audio content sets specified in the at least one request.
[0035] Means configured to obtain a media description containing information identifying the location of audio sources in an audio environment associated with at least two audio content sets may be further configured to output the media description to a device, and outputting the encoded at least two audio content sets for an audio scene may be to a further device.
[0036] Further devices could include spatial audio players configured to generate audio scenes.
[0037] According to a fourth aspect, there is a device for generating an audio scene having at least two audio content sets, which includes means for obtaining a media description including information that identifies the location of an audio source in an audio scene associated with at least two audio content sets, obtaining a listener location in the audio scene, determining a source expression level for at least two audio content sets based on the location of the audio source and the listener location associated with the at least two content sets, obtaining at least two encoded audio content sets based on the source expression levels, and the means for generating a spatialized audio output based on the obtained at least two encoded audio content sets, which includes a first encoded at least two audio content sets encoded at a first expression level and a second encoded at least two audio content sets encoded at a second expression level.
[0038] At least two levels of representation may include full-order representations of at least two audio content sets, reduced or lower-order representations of at least two audio content sets, and at least two primary representations of at least two audio content sets.
[0039] The media description may further specify the ambisonic order for at least two audio content sets.
[0040] A means configured to obtain at least two encoded audio content sets based on a source representation level may be configured to obtain all encoded channel representations for at least two encoded audio content sets when the source representation level is a full-order representation.
[0041] Means configured to obtain at least two encoded audio content sets based on the source representation level may be configured to obtain an encoded selected channel representation for the at least two encoded audio content sets when the source representation level is reduced or low-order representation.
[0042] A means configured to obtain at least two encoded audio content sets based on a source representation level may be configured to obtain the first four selected encoded channel representations for the at least two encoded audio content sets when the source representation level is a primary representation.
[0043] Means configured to generate spatialized audio output based on at least two acquired encoded audio content sets may be configured to perform signal interpolation by processing the channel signals of the acquired encoded audio content sets based on the listener position, and to generate spatialized audio output based on the signal interpolation performed.
[0044] Means configured to obtain at least two encoded audio content sets based on source representation levels may be configured to: generate requests for representations of at least two encoded audio content sets, and receive requested representations of at least two encoded audio content sets, wherein the requests include source representation levels.
[0045] Means configured to determine source representation levels based on the position and listener position of audio sources associated with at least two audio content sets may be further configured to determine content set rendering-related state information associated with at least two audio content sets, and means configured to obtain the content set rendering-related state information and at least two encoded audio content sets may be configured to obtain the at least two encoded audio content sets based on the content set rendering-related state information.
[0046] The state information related to content set rendering is configured to determine whether the content set can be used as one of the following: a 6-degree-of-freedom higher-order ambisonic audio source while generating spatialized audio output, and an object-rendered audio source while generating spatialized audio output.
[0047] According to a fifth aspect, an apparatus is provided for generating an audio scene having at least two audio content sets, the apparatus comprising at least one processor and at least one memory for storing instructions, the instructions, when executed by at least one processor, cause the apparatus to obtain a media description including information identifying the locations of audio sources in an audio scene associated with at least two audio content sets, encode the at least two audio content sets so that the at least two audio content sets are output at at least two representation levels, and output the encoded at least two audio content sets for an audio scene, the output encoded at least two audio content sets comprising a first encoded at least two audio content sets encoded at a first representation level of at least two, and a second first encoded at least two audio content sets encoded at a second representation level of at least two.
[0048] At least two audio content sets may include at least two higher ambisonic audio sources.
[0049] At least two audio content sets may include metadata associated with each of at least two higher ambisonic audio sources.
[0050] The metadata associated with each of at least two higher-order ambisonic audio sources may include their position within the audio scene.
[0051] Media descriptions may further include information that identifies at least two levels of representation.
[0052] At least two levels of representation may include at least two of the following: full-order representations of at least two audio content sets, reduced or lower-order representations of at least two audio content sets, and primary representations of at least two audio content sets.
[0053] The full-order representation of at least two audio content sets can be all channels of at least two higher-order ambisonic audio sources.
[0054] The reduction or lower-order representation of at least two audio content sets may involve the selection of channels from at least two higher-order ambisonic audio sources.
[0055] The primary representation of at least two audio content sets could be a selection of the first four channels of at least two higher-order ambisonic audio sources.
[0056] A device configured to output at least two encoded audio content sets for an audio scene may: receive at least one request relating to an audio scene, the request identifies a representation level for at least two audio content sets, select at least two encoded audio content sets that are at the representation levels for the at least two audio content sets identified within the at least one request, and output the selected at least two encoded audio content sets that are at the representation levels for the at least two audio content sets identified within the at least one request.
[0057] A device that causes a device to obtain a media description containing information that identifies the location of an audio source in an audio environment associated with at least two audio content sets may be further configured to output the media description to a further device, and the device may further configure to output at least two encoded audio content for an audio scene to a further device.
[0058] Further devices could include spatial audio players configured to generate audio scenes.
[0059] According to a sixth aspect, there is a device for generating an audio scene having at least two audio content sets, the device comprising at least one processor and at least one memory for storing instructions, the instructions, when executed by at least one processor, cause the system to: obtain a media description including information identifying the location of an audio source in an audio scene associated with at least two audio content sets; obtain a listener location in the audio scene; determine a source representation level for at least two audio content sets based on the location of an audio source and the listener location associated with at least two content sets; obtain at least two encoded audio content sets based on the source representation levels, wherein the at least two encoded audio content sets include a first encoded at least two audio content sets encoded at a first of at least two representation levels and a second encoded at least two audio content sets encoded at a second of at least two representation levels; and generate a spatialized audio output based on the obtained at least two encoded audio content sets.
[0060] At least two levels of representation may include at least two of the following: full-order representations of at least two audio content sets, reduced or lower-order representations of at least two audio content sets, and primary representations of at least two audio content sets.
[0061] The media description may further specify the ambisonic order for at least two audio content sets.
[0062] A means configured to obtain at least two encoded audio content sets based on a source representation level may be configured to obtain all encoded channel representations for at least two encoded audio content sets when the source representation level is a full-order representation.
[0063] Means configured to obtain at least two encoded audio content sets based on the source representation level may be configured to obtain an encoded selected channel representation for at least two encoded audio content sets when the source representation level is reduced or low-order representation.
[0064] A means configured to obtain at least two encoded audio content sets based on a source representation level may be configured to obtain the first four selected encoded channel representations for at least two encoded audio content sets when the source representation level is a primary representation.
[0065] Means configured to generate a spatialized audio output based on at least two acquired encoded audio content sets may be configured to perform signal interpolation by processing the channel signals of the acquired encoded audio content sets based on the listener position, and to generate a spatialized audio output based on the signal interpolation performed.
[0066] A means configured to obtain at least two encoded audio content sets based on a source representation level may further configure to generate a request for representations of the at least two encoded audio content sets, the request including a source representation level, and to receive the requested representations of the at least two encoded audio content sets.
[0067] Means configured to determine source representation levels for at least two audio content sets based on the location of the audio source associated with the at least two content sets and the listener location may be further configured to determine content set rendering-related state information associated with the at least two audio content sets, and means configured to obtain the content set rendering-related state information and the encoded at least two audio content sets may be further configured to obtain the encoded at least two audio content sets based on the content set rendering-related state information.
[0068] The state information related to content set rendering is configured to determine whether the content set can be used as one of the following: a 6-degree-of-freedom higher-order ambisonic audio source while generating spatialized audio output, and an object-rendered audio source while generating spatialized audio output.
[0069] According to a seventh aspect, an apparatus is provided for generating an audio scene having at least two audio content sets: a circuit that can be configured to obtain a media description including information that identifies the location of an audio source in an audio scene associated with at least two audio content sets; an encoding circuit configured to encode at least two audio content sets such that the at least two audio content sets are output at at least two expression levels; and an output circuit configured to output the encoded at least two audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first of at least two expression levels, and a second first encoded at least two audio content sets encoded at a second of at least two expression levels.
[0070] According to the eighth aspect, an apparatus is provided for generating an audio scene having at least two audio content sets, comprising: a circuit that can be configured to obtain a media description including information that identifies the location of an audio source in an audio scene associated with at least two audio content sets; a circuit that can be configured to obtain a listener location in an audio scene; a determination circuit configured to determine a source expression level for at least two audio content sets based on the location of an audio source associated with at least two content sets and the listener location; an acquisition circuit configured to acquire at least two encoded audio content sets based on the source expression level, wherein the at least two encoded audio content sets include a first encoded at least two audio content sets encoded at a first expression level of at least two, and a second encoded at least two audio content sets encoded at a second expression level of at least two; and generating a spatialized audio output based on the acquired encoded at least two audio content sets.
[0071] According to the ninth aspect, a computer program [or computer-readable medium including instructions] is provided which includes instructions for a device to generate an audio scene having at least two audio content sets, wherein the device is configured to: obtain a media description including information that identifies the location of an audio source in an audio scene associated with at least two audio content sets; encode at least two audio content sets such that the at least two audio content sets are output at at least two representation levels; and output the encoded at least two audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first representation level of at least two, and a second first encoded at least two audio content sets encoded at a second representation level of at least two.
[0072] According to a tenth aspect, a computer program is provided which includes instructions to cause a device to generate an audio scene having at least two audio content sets, wherein the device is configured to: obtain a media description which includes information which identifies the location of an audio source in an audio scene associated with at least two audio content sets; obtain a listener location in the audio scene; determine a source representation level for at least two audio content sets based on the location of an audio source associated with at least two content sets and the listener location; obtain at least two encoded audio content sets based on the source representation levels, wherein the at least two encoded audio content sets include a first encoded at least two audio content sets encoded at a first of at least two representation levels and a second encoded at least two audio content sets encoded at a second of at least two representation levels; and generate a spatialized audio output based on the obtained encoded at least two audio content sets.
[0073] According to the eleventh aspect, a non-temporary computer-readable medium is provided, which includes a program instruction, the program instruction causing a device for generating an audio scene having at least two audio content sets to: obtain a media description including information identifying the location of an audio source in an audio scene associated with at least two audio content sets; encode at least two audio content sets so that the at least two audio content sets are output at at least two representation levels; and output the encoded at least two audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first representation level and a second first encoded at least two audio content sets encoded at a second representation level.
[0074] According to the twelfth aspect, a non-temporary computer-readable medium including program instructions, the program instructions cause a device for generating an audio scene having at least two audio content sets to perform at least the following: obtain a media description including information identifying the location of an audio source in an audio scene associated with at least two audio content sets; obtain a listener location in an audio scene; determine a source representation level for at least two audio content sets based on the location of an audio source associated with at least two content sets and the listener location; obtain at least two encoded audio content sets based on the source representation levels, wherein the at least two encoded audio content sets include a first encoded at least two audio content sets encoded at a first of at least two representation levels, and a second encoded at least two audio content sets encoded at a second of at least two representation levels; and generate a spatialized audio output based on the obtained at least two encoded audio content sets.
[0075] According to the 13th aspect, an apparatus is provided for generating an audio scene having at least two audio content sets, comprising: means for obtaining a media description including information that identifies the location of an audio source in the audio scene associated with the at least two audio content sets; means for encoding the at least two audio content sets such that the at least two audio content sets are output at at least two expression levels; and means for outputting the encoded at least two audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first expression level of at least two, and a second first encoded at least two audio content sets encoded at a second expression level of at least two.
[0076] According to the 14th aspect, an apparatus is provided for generating an audio scene having at least two audio content sets, comprising: means for obtaining a media description including information identifying the location of an audio source in an audio scene associated with at least two audio content sets; means for obtaining a listener location in an audio scene; means for determining a source expression level for at least two audio content sets based on the location of an audio source associated with at least two content sets and the listener location; means for obtaining at least two encoded audio content sets based on the source expression level, wherein the at least two encoded audio content sets include a first encoded at least two audio content sets encoded at a first of at least two expression levels, and a second encoded at least two audio content sets encoded at a second of at least two expression levels; and generating a spatialized audio output based on the obtained encoded at least two audio content sets.
[0077] According to the 15th aspect, a computer-readable medium is provided which includes instructions, the instructions causing a device for generating an audio scene having at least two audio content sets to: obtain a media description including information identifying the location of an audio source in an audio scene associated with at least two audio content sets; encode at least two audio content sets such that the at least two audio content sets are output at at least two representation levels; and output the encoded at least two audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first representation level of at least two, and a second first encoded at least two audio content sets encoded at a second representation level of at least two.
[0078] According to the 15th aspect, a computer-readable medium is provided which includes instructions, the instructions causing a device for generating an audio scene having at least two audio content sets to perform at least the following: obtain a media description including information that identifies the location of an audio source in an audio scene associated with at least two audio content sets; obtain a listener location in an audio scene; determine a source representation level for at least two audio content sets based on the location of an audio source associated with at least two content sets and the listener location; obtain at least two encoded audio content sets based on the source representation levels, wherein the at least two encoded audio content sets include a first encoded at least two audio content sets encoded at a first representation level and a second encoded at least two audio content sets encoded at a second representation level; and generate a spatialized audio output based on the obtained at least two encoded audio content sets.
[0079] An apparatus comprising means for performing the actions of the method described above.
[0080] An apparatus configured to perform the actions of the method described above.
[0081] A computer program that contains program instructions for causing a computer to perform the actions described above.
[0082] Computer program products stored on a medium can cause the device to perform the actions described herein.
[0083] The electronic device may include an apparatus as described herein.
[0084] The chipset may include devices such as those described herein.
[0085] The embodiments of this application aim to address problems associated with the prior art.
[0086] For a better understanding of this application, the attached drawings will be used as examples from here on. [Brief explanation of the drawing]
[0087] [Figure 1] This diagram schematically illustrates an example 6DoF audio scene, showing the listener position trajectories within an audio scene containing eight HOA sources, and the changes in the HOA sources required to render the audio. [Figure 2] This figure schematically illustrates the requirements for acquiring HOA audio data that are dependent on the listener's position, based on the audio scene shown in Figure 1. [Figure 3] This figure schematically illustrates the requirements for acquiring HOA audio data that are dependent on the listener's position, based on the audio scene shown in Figure 1. [Figure 4] This diagram illustrates the selection of HOA sources for the baseline and subset channels based on the role of the HOA source, based on the listener position within the audio scene, and shows an HOA source where all channels are processed. [Figure 5] This diagram illustrates the selection of HOA sources for the baseline and subset channels based on the role of the HOA source, based on the listener's position within the audio scene, and shows an HOA source in which only the first four channels are processed. [Figure 6a] Figure 1 shows an example of channel acquisition for an HOA source during a time period ΔT1, based on the audio scene shown in Figure 1, where HOA1 is used for signal interpolation, while HOA2 and HOA3 are used for spatial metadata processing. [Figure 6b]Figure 1 shows an example of channel acquisition for an HOA source during a time period ΔT1, based on the audio scene shown in Figure 1, where HOA1 is used for signal interpolation, while HOA2 and HOA3 are used for spatial metadata processing. [Figure 6c] Figure 1 shows an example of channel acquisition for an HOA source during a time period ΔT1, based on the audio scene shown in Figure 1, where HOA1 is used for signal interpolation, while HOA2 and HOA3 are used for spatial metadata processing. [Figure 7] This figure shows an example encoding process according to several embodiments. [Figure 8] This is a flowchart illustrating the operation of an encoder as an example of several embodiments. [Figure 9] This is a flowchart illustrating the operation of a decoder or renderer as an example of several embodiments. [Figure 10] This figure shows a baseline and extended decoder as examples of several embodiments. [Figure 11] This figure shows a baseline and extended decoder as examples of several embodiments. [Figure 12] This figure shows channel selection for HOA sources based on the audio scene shown in Figure 1, according to several embodiments. [Figure 13] This figure shows an example of channel acquisition for an HOA source during a time period ΔT1 based on the audio scene shown in Figure 1. [Figure 14] This figure shows an example implementation configuration based on several embodiments. [Figure 15] This figure shows an example implementation configuration based on several embodiments. [Figure 16] This figure shows the effects of several implementation examples of different embodiments. [Figure 17] This figure shows an apparatus that is a suitable example for carrying out several embodiments. [Modes for carrying out the invention]
[0088] Concepts discussed in more detail herein with respect to the following embodiments relate to apparatus and methods for minimizing the bitrate on a player device for 6DoF HOA rendering.
[0089] Higher-order HOA sources (those with a large number of audio channels per HOA source) require higher encoding bitrates. However, higher-order HOA sources often need to represent audio scenes with more detailed spatial audio cues and higher subjective quality audio. While it is possible to reduce the bitrate by encoding all HOA signals as (first-order ambisonic) FOA audio signals, this results in poor audio quality for user spatial positioning.
[0090] Acquiring and consuming all of the HOA source signals encoded in higher order requires a much higher bitrate, as multiple higher-order HOA source signals are needed to perform 6DoF rendering. Therefore, in a certain scenario where the listener position is within a triangle formed by three HOA sources, there is at least a threefold increase in the corresponding bandwidth requirement for 6DoF rendering with increasing the encoding bitrate of one HOA source. Such a large increase in bandwidth requirements for 6DoF rendering can lead to network congestion and latency. This increases the challenges of consuming 6DoF rendering of audio scenes containing multiple higher-order HOA sources. In some less common or transient scenarios, audio data from five or six HOA sources may be required to ensure smooth listener transitions across the triangle formed by the HOA sources, or in the case of teleportation involving sudden long-distance jumps within the audio scene.
[0091] The current 6DoF HOA rendering process utilizes different HOA sources depending on the listener position: 1. All channels are required for several HOA sources.
[0092] 2. For some HOA sources, only a subset of channels is needed.
[0093] 3. For some HOA sources, no channels are required. In other words, these HOA sources are not needed to perform 6DoF HOA rendering due to the current listener position in the audio scene.
[0094] Therefore, depending on the listener location, some HOA sources are not needed (location 3), some HOA sources are needed at the highest available order (location 1), and some HOA sources are needed at a lower order or as FOA (location 2).
[0095] However, no known method addresses the delivery of a relevant subset of audio channels from a relevant subset of the HOA source. In other words, no known method delivers only the necessary data without negatively impacting audio quality. This can result in excessive and unnecessary bandwidth consumption.
[0096] Referring to Figure 1, an example 6DoF audio scene is shown. Several HOA audio sources are located within the scene, which are shown as HOA1103, HOA2105, HOA3107, HOA4109, HOA5111, HOA6113, HOA7115, and HOA8117.
[0097] Furthermore, within the audio scene, the representation of the listener or user 101 as it moves within the audio scene is shown. This movement is indicated by a line that starts at 151 and ends at 161. The starting point 151 is located between and near HOA sources HOA1103 and HOA2105. The movement then moves the listener to point 153, which is closer to HOA3107, during the first time ΔT1102. Furthermore, the movement within the scene during the next time ΔT2104 moves the listener past HOA3107, around HOA4109, to point 155. Additionally, the movement within the scene during the next time ΔT3106 moves the listener towards HOA5111, and then away from HOA5111 towards HOA7115, to point 157. Furthermore, the movement within the scene during the next time interval ΔT4108 moves the listener around HOA7115 and then to point 157 between HOA6113, HOA7115, and HOA8117. Finally, with respect to Figure 1, the movement within the scene during the next time interval ΔT5110 moves the listener towards HOA7115 and then to point 161.
[0098] Regarding Figure 2, an example of an HOA selection for rendering the motion shown in Figure 1 is presented. Therefore: Over the time interval T0 to T1, i.e., ΔT1, the rendering selects HOA audio sources HOA1103, HOA2105, and HOA3107 because they are the closest and contribute the most to the output. Over the time interval from T1 to T2, i.e., ΔT2, the rendering selects HOA audio sources HOA4109, HOA5111, and HOA6113 because they are the closest and contribute the most to the output. Over the time interval from T2 to T3, i.e., ΔT3, the rendering selects HOA audio sources HOA5111, HOA6113, and HOA7115 as significant contributors. During the time interval from T3 to T4, i.e., ΔT4, the rendering selects HOA audio sources HOA4109, HOA5111, and HOA6113 for output. From T4 to T5, i.e., over ΔT5, the rendering selects HOA audio sources HOA6113, HOA7115, and HOA8117.
[0099] Regarding Figure 3, an example system is shown where, instead of acquiring all HOA source audio signals, the selected HOA audio source is only the HOA audio signal to be acquired. Therefore, as shown in Figure 3: From T0151 to T1153, i.e., over ΔT1102, rendering retrieves HOA audio sources HOA1103, HOA2105, HOA3107, 303. From T1153 to T2155, i.e., over ΔT2104, rendering retrieves HOA audio sources HOA4109, HOA5111, HOA6113, 305. From T2155 to T3157, i.e., over ΔT3106, rendering retrieves HOA audio sources HOA5111, HOA6113, HOA7115 and 307, From T3157 to T4159, i.e., over ΔT4108, the rendering acquires HOA audio sources HOA4109, HOA5111, HOA6113, and 309, During the time from T4159 to T5161, i.e., over ΔT5110, the rendering acquires HOA audio sources HOA6113, HOA7115, and HOA8117.
[0100] Therefore, as discussed in the following example, the number of channels for selected / acquired HOA sources can vary in some situations. For example, the number of channels (or HOA degree representations) for selected HOA sources may be chosen by the player depending on the listener position.
[0101] For example, as shown in Figure 4, a first HOA-type audio signal may be acquired, where the full set 405 of channel 401 is acquired, and STFT processing is performed on the frequency bins 403 for all acquired channels.
[0102] Additionally, as shown in Figure 5, further HOA-type audio signals may be acquired, where a full set 505 of channel 401 is acquired, and STFT processing is performed on frequency bins 403 for a subset of all acquired channels (the first four or first HOA order).
[0103] Therefore, with respect to Figures 6a to 6c, examples of audio sources that are encoded / acquired when the listener is in an audio scene as shown in Figure 1 over a time period from T0151 to T1153, i.e., ΔT1102. In this example, it is shown that HOA audio source HOA1103 has the full set 601 of channels to be acquired and STFT is performed on the frequency bins, while HOA audio source HOA2105 has a subset 603 of channels and frequency bins, and HOA audio source HOA3107 has a subset 605 of channels and frequency bins.
[0104] In some situations, the channel and frequency bin subsets 603 and 605 correspond to the primary ambisonic channels (i.e., the first four channels).
[0105] Concepts discussed herein in examples and embodiments include the existence of apparatuses and methods for minimizing the bitrate on a player device for 6DoF HOA rendering. In such embodiments, this is provided by apparatuses and methods for determining the relevant HOA source representation level (full-order, first-order) to achieve a reduction in bandwidth consumption (e.g., 55% for higher-order, subject to encoding bitrate) by acquiring HOA sources used for signal interpolation in the full-order HOA representation and acquiring HOA sources used for spatial metadata processing in the first-order HOA representation. This processing includes content creation, content description for acquisition, and content acquisition enhancement.
[0106] This can be done in some embodiments by an encoder configured as follows: In addition to full-order representation, encode at least primary ambisonic representations for HOA sources including audio scenes. A media manifest file is created that declares the HOA source location, HOA source order, and all HOA source audio data, with each HOA source represented in the highest order specified by the audio scene description encoder input format, and in a primary ambisonic representation.
[0107] In some embodiments, the highest-order HOA source audio data representation is made available by the content creator.
[0108] Furthermore, the decoder / player is: Obtain the media manifest file to determine the HOA source location. Obtain the listener's position within the audio scene, Based on the listener position and HOA source position, determine the HOA source that will be acquired. Determine the HOA source representation level, ensuring that at least one HOA source is required for full-order representation and at least one HOA source is required for primary representation (considering a scenario with two HOA source scenes, one HOA source is acquired for full-order for signal interpolation, while the other HOA source is acquired for primary representation). If a full-order representation is required for rendering, retrieve the HOA source at the highest order in the manifest file, and If the primary representation is sufficient, retrieve the HOA source in the primary. It can be configured in this way.
[0109] In some embodiments, a renderer state interface is provided to the player, including an HOA source identifier and a list of HOA source roles for rendering. In another embodiment, the HOA source roles may be replaced with an HOA source order or the number of HOA channels required for 6DoF rendering.
[0110] In such a configuration, the decoder / player may be configured to acquire only the necessary HOA source audio data, aiming to achieve savings of over 60% in terms of channel data that would be acquired (for example, in the case of a scene containing a 6th-order HOA source). This could result in acquiring one HOA source for all orders and two HOA sources for first orders.
[0111] Embodiments discussed herein can be divided into two parts: An encoder configured to perform content creation and manifest creation. A decoder / player configured to provide a renderer state interface to the player and media acquisition logic.
[0112] Regarding Figure 7, an example of content creation and manifest creation is shown to achieve efficient bandwidth utilization (and to avoid acquiring HOA source channels that are not required for rendering without causing any adverse effects on subjective audio quality compared to acquiring all HOA sources in the entire sequence). The encoder is configured to properly create encoded versions of the HOA sources described in the audio scene, followed by specifying extensions to the media manifest for acquisition (e.g., for MPEG DASH, HLS, etc.) to enable content selection by the player.
[0113] In some embodiments, the encoder includes an EIF input that is passed to an ambisonic encoder 701. The (higher-order) ambisonic encoder 701 generates multiple HOA source representations, including at least a full-order representation and a first-order representation. In some scenarios, the lowest order may be higher than first-order.
[0114] The (higher-order) ambisonics encoder 701 is configured to receive an encoder input format representation (EIF) and audio data of an audio scene (e.g., audio data corresponding to HOA sources). In different implementation embodiments, any other suitable scene description format may be used. The EIF is parsed to determine the presence of an HOA group structure within the EIF, indicating the presence of HOA sources that will be used to perform 6DoF HOA rendering.
[0115] The encoded HOA source audio data is available as a full-order representation 703.
[0116] The encoder is configured to encode the HOA source at the highest required order using a suitable encoder (e.g., an MPEG-H 3DA (ISO / IEC 23008-3) encoder) to produce MPEG-H 3DA encoded HOA source audio data using the highest order (i.e., all channels specified in the EIF).
[0117] Next, the HOA source is encoded by the MPEG-H 3DA encoder as a primary representation 706 and an arbitrary lower-order representation 705. Full-order, primary, and lower-order HOA source audio data representations are generated in MPEG-H 3DA encoding. In addition, the manifest generator 707 is configured to create a manifest file to be used by the media selector 709. The manifest file can therefore be passed to the selector 709. This can also be referred to as a media description file to facilitate appropriate content selection and retrieval.
[0118] Furthermore, the full-order representation 703, lower-order representation 705, and primary representation 706 of the MPEG-H 3DA encoded dead HOA source audio data are represented in MPEG DASH MPD generated by the manifest generator 707, which will be discussed in more detail below. Note that any other variants of the media manifest representation may be generated that enable the acquisition of the HOA source in a full-order representation that holds all channels and a primary representation that holds the first four channels. The following are example operational steps for generating suitable HOA source content for 6DoF rendering according to several embodiments.
[0119] Figure 8 shows a flowchart illustrating an example of the operation of an encoder according to several embodiments.
[0120] The initial operation for receiving the Audio Scene Description (EIF) is shown by 801 in Figure 8.
[0121] Next, determine the HOA source within the EIF, as shown by 803 in Figure 8.
[0122] Next, each HOA source is encoded into a full-order MPEG-H 3DA encoded format, as shown by 805 in Figure 8.
[0123] Next, each HOA source is encoded into a primary MPEG-H 3DA encoded dead, as shown by 807 in Figure 8.
[0124] Finally, generate a media manifest to declare the available HOA source representations for content selection, as shown by 809 in Figure 8.
[0125] In some embodiments, HOA source information may be derived from an EIF or any other suitable scene description. aligned(8)HOASourceInfoStruct(){ HOASourcePositionStruct(); / / position of the HOA source unsigned int(16)hoa_source_id; / / unique identifier for each HOA source unsigned int(3)full_hoa_order; / / highest order of HOA source bit(4)reserved=0; unsigned int(1)hoa_group_id_present; / / Unique HOA group identifier if(hoa_group_id_present) unsigned int(16)hoa_group_id; / / Unique HOA group identifier } hoa_source_id is a unique identifier for each HOA source within an audio scene.
[0126] `full_hoa_order` is the highest HOA order available for an HOA source with the identifier `hoa_source_id`.
[0127] hoa_group_id corresponds to the HOAGroup ID as described in the audio scene. If hoa_group_id_present is equal to 0, the HOA source is not part of any HOA group and is not expected to be rendered with 6 degrees of freedom as a listener. aligned(8)HOASourcePositionStruct(){ signed int(32)hoa_source_pos_x; signed int(32)hoa_source_pos_y; signed int(32)hoa_source_pos_z; signed int(16) hoa_source_rot_yaw; signed int(16) hoa_source_rot_pitch; signed int(16) hoa_source_rot_roll; } The values of hoa_source_pos_x, hoa_source_pos_y, and hoa_source_pos_z define the position in 3D space in units of 1 millimeter. hoa_source_rot_yaw and hoa_source_rot_roll are from -180 * 2 16 to +180 * 2 16 -1, and in the case of hoa_source_rot_pitch from -90 * 2 16 to 90 * 2 16 -1, and are defined in units of 2 -16 degrees. In a general sense, it is 2 -16 degree step_size within the range of -180 to 180-step_size for yaw and roll, and -90 to 90-step_size for pitch.
[0128] Regarding the creation of manifests for media acquisition in DASH MPD, an HOA source element with an @schemeIdUri attribute equal to “urn:mpeg:mpegI:mia:2023:6dho” is referred to as an HOA source (6DHO) descriptor. 6DHO descriptors can exist at the adaptation set level and are not present at any other level. If an Adaptation Set within a Media Presentation does not contain a 6DHO descriptor, it is assumed that the Media Presentation does not support 6DoF rendering using HOA sources. A 6DHO descriptor indicates the HOA source to which the Adaptation Set belongs. Therefore, a 6DHO descriptor may include an @value attribute, as well as an HOASourceInfo element with its sub-elements and attributes, as specified in the table below.
[0129] [Table 1]
[0130] An additional SupplementalProperty or EssentialProperty6DoFHOAorder descriptor with a schemeIdUri equal to “urn:mpeg:mpegI:mia:2023:hord” is referred to as an HOAOrder descriptor. This 6DoFHOAorder descriptor exists for all representations in the HOA source adaptation set to indicate the HOA order of the representation.
[0131] The existence of different representations of HOA sources, including audio scenes, allows the 6DoF player to acquire the appropriate HOA source representation as needed for efficient acquisition while maintaining optimal 6DoF HOA rendering quality. <mpd> … … … / *This is HOA source ID 3000 in HOA group with ID 1* / <adaptationset id=""3000”" mimetype=""audio / mp4" profiles="’mhm2’”" codecs=""oabl”" segmentalignment=""1”"> <hoasource schemeiduri=""urn:mpeg:mpegI:mia:2023:6dho”" value=""1”"> <hoasourceinfo label=""HOA" source 1”> <position x=""1”" y=""2”" z=""3” / "> <hoagroupinfo groupid=""1” / "> < / hoagroupinfo> < / position> < / hoasourceinfo> < / hoasource> <representation id=""3001”" bandwidth=""1512000”" startwithsap=""1”"> <supplementalproperty schemeIdUri=""urn:mpeg:mpegI:mia:2023:hord”" hoa_order="4" / > <segmenttemplate media=""AudioScene1.HOASource1.Order4.$Number$.mp4”" initialization=""AudioScene1.HOASource1.Order4.init.mp4”" duration=""17”" startnumber=""1”" timescale=""30” / "> < / segmenttemplate> < / representation> <representation id=""3002”" bandwidth=""256000”" startwithsap=""1”"> <supplementalproperty schemeIdUri=""urn:mpeg:mpegI:mia:2023:hord”" hoa_order="1" / > <segmenttemplate media=""AudioScene1.HOASource1.Order4.$Number$.mp4”" initialization=""AudioScene1.HOASource1.Order4.init.mp4”" duration=""17”" startnumber=""1”" timescale=""30” / "> < / segmenttemplate> < / representation> < / adaptationset> / *This is HOA source ID4000 in HOA group with ID 1* / <adaptationset id=""4000”" mimetype=""audio / mp4" profiles="’mhm2’”" codecs=""oabl”" segmentalignment=""2”"> <hoasource schemeiduri=""urn:mpeg:mpegI:mia:2023:6dho”" value=""1”"> <hoasourceinfo label=""HOA" source 1”> <position x=""4”" y=""2”" z=""6” / "> <hoagroupinfo groupid=""1” / "> < / hoagroupinfo> < / position> < / hoasourceinfo> < / hoasource> <representation id=""4001”" bandwidth=""1512000”" startwithsap=""1”"> <supplementalproperty schemeIdUri=""urn:mpeg:mpegI:mia:2023:hord”" hoa_order="4" / > <segmenttemplate media=""AudioScene1.HOASource1.Order4.$Number$.mp4”" initialization=""AudioScene1.HOASource1.Order4.init.mp4”" duration=""17”" startnumber=""1”" timescale=""30” / "> < / segmenttemplate> < / representation> <representation id=""4002”" bandwidth=""256000”" startwithsap=""1”"> <supplementalproperty schemeIdUri=""urn:mpeg:mpegI:mia:2023:hord”" hoa_order="1" / > <segmenttemplate media=""AudioScene1.HOASource1.Order4.$Number$.mp4”" initialization=""AudioScene1.HOASource1.Order4.init.mp4”" duration=""17”" startnumber=""1”" timescale=""30” / "> < / segmenttemplate> < / representation> < / adaptationset> / *This is HOA source ID 5000 in HOA group with ID 1* / <adaptationset id=""5000”" mimetype=""audio / mp4" profiles="’mhm2’”" codecs=""oabl”" segmentalignment=""1”"> <hoasource schemeiduri=""urn:mpeg:mpegI:mia:2023:6dho”" value=""1”"> <hoasourceinfo label=""HOA" source 1”> <position x=""10”" y=""3”" z=""20” / "> <hoagroupinfo groupid=""1” / "> < / hoagroupinfo> < / position> < / hoasourceinfo> < / hoasource> <representation id=""5001”" bandwidth=""1512000”" startwithsap=""1”"> <supplementalproperty schemeIdUri=""urn:mpeg:mpegI:mia:2023:hord”" hoa_order="4" / > <segmenttemplate media=""AudioScene1.HOASource1.Order4.$Number$.mp4”" initialization=""AudioScene1.HOASource1.Order4.init.mp4”" duration=""17”" startnumber=""1”" timescale=""30” / "> < / segmenttemplate> < / representation> <representation id=""5002”" bandwidth=""256000”" startwithsap=""1”"> <supplementalproperty schemeIdUri=""urn:mpeg:mpegI:mia:2023:hord”" hoa_order="1" / > <segmenttemplate media=""AudioScene1.HOASource1.Order4.$Number$.mp4”" initialization=""AudioScene1.HOASource1.Order4.init.mp4”" duration=""17”" startnumber=""1”" timescale=""30” / "> < / segmenttemplate> < / representation> < / adaptationset> < / mpd> The DASH manifest shown above, as an example, allows the DASH player to acquire the necessary HOA source audio data in an efficient format (without the waste of acquiring excessive channels).
[0132] The decoder / player has two main components. The first component is the renderer state interface to the player. The second is the media acquisition logic for achieving efficient acquisition.
[0133] Regarding the player / decoder section as shown in Figure 7, there is a selector (for example, implemented as part of MPD) 709 configured to perform content selection. In some embodiments, content selection may be based on the listener position.
[0134] Furthermore, as shown in Figure 7, the selector 709 can receive listener position information from the player 711 and receive even more expressions that can be selected by the selector 709.
[0135] With respect to Figure 9, the operation of the selector 709 / player 711 in several embodiments follows the operation of the encoder as shown in Figure 8.
[0136] Therefore, for example, 901 indicates the operation of receiving the HOA source position within the audio scene.
[0137] Next, 903 indicates that the listener's position within the audio scene is being received.
[0138] Following this, 905 indicates that the relevant HOA source should be determined.
[0139] Next, 907 indicates that it receives identifiers of one or more HOA sources to be used for signal interpolation via the renderer interface.
[0140] What follows, as indicated by 909, is the receipt of identifiers for two or more HOA sources used for spatial metadata creation via the renderer interface.
[0141] Next, 911 indicates obtaining the HOA source from the manifest, which will be used for signal interpolation in the full-order representation.
[0142] Following this, 913 demonstrates obtaining the HOA source used for creating spatial metadata in the primary representation from the manifest.
[0143] Finally, as indicated by 915, the acquired HOA source audio data is distributed to an MPEG-H 3DA decoder, and subsequently to an MPEG-I renderer implemented in, for example, player 711.
[0144] This is further illustrated in Figures 10 and 11, which show the operation of the media acquirer according to several embodiments.
[0145] Therefore, as shown, the baseline media manifest 1000 store is configured to pass the acquired audio data based on the listener position 1001 to the decoder / player media acquirer 1003.
[0146] The media acquirer 1003 is configured to acquire the MPEG-H 3DA audio bitstream 1004 and pass it to the MPEG-H 3DA decoder 1005.
[0147] The MPEG-H 3DA decoder 1005 is configured to pass the decoder-ready MPEG-H 3DA data to the MPEG-I renderer 1007.
[0148] Additionally, the acquired MPEG-I bistream 1006, which is passed to the MPEG-I renderer 1007, is shown.
[0149] Furthermore, there is an interface 1020 from the player control 1011 to the MPEG-I relayer 1007.
[0150] The MPEG-I renderer 1007 is configured to generate an audio signal and pass it to the audio output headphones 1009.
[0151] Furthermore, the player control 1011 may include local scene updates 1010, consumption environment information (e.g., LSDF) 1012, and user location and interaction information 1014. User location 1002 may be passed to the media acquirer 1002.
[0152] Additionally, as shown in Figure 11, and unlike Figure 10, there is an enhanced media manifest 1100 that may include HOA sources in full-order and primary representations.
[0153] Furthermore, the media acquirer 1103 may be configured to receive a renderer state interface (RSI) 1103 having renderer state parameters (RSP).
[0154] The media acquirer 1103 may further be configured to acquire audio data 1101 from the media manifest 1100 based on the listener position and renderer status.
[0155] In some embodiments, the renderer interface (provided by the renderer within the MPEG-I immersive audio) is configured to retrieve, among other things, information from the bitstream, dynamic updates from the local interface, listener position, listener interaction, and an LSDF (Listener Space Description Format) file.
[0156] Media acquisition is therefore performed by the player based on a media manifest that describes the content selection options for the audio scene being consumed, as discussed above. This includes HOA source representations in all HOA dimensions (as specified in the EIF). The relevant HOA sources are acquired based on the listener position.
[0157] The renderer state interface 1102, or RSI, is configured to allow the player to obtain information about various renderer state parameters. These renderer state parameters are used by the player to determine the audio data acquisition strategy.
[0158] In some embodiments, the player does not understand mphoa rendering; in other words, the player is "ignorant" in the sense that the logic in determining which HOA sources are needed at what order is stored in the renderer.
[0159] The renderer 1007 can therefore be configured to request from a player HOA source having the order that the renderer requests from the player HOA source via the RendererStateInterface 1102. aligned(8)RendererStateInterface(){ unsigned int(8)renderType; / / Type of rendering state information if(renderType==6DoFHOARendering) unsigned int(3)num_active_triangles; / / Active triangles unsigned int(3)num_hoa_source_trr; / / TRR list for(i=0;i <num_hoa_source_trr;i++){ unsigned int(16)hoa_source_id; unsigned int(1)full_order_source; unsigned int(1)first_order_source; } } else if(renderType==ObjectRendering) / * Define for Object Rendering related state information* / for(i=0;i <num_object_sources;i++){ unsigned int(16)object_source_id; unsigned int(1)object_source_culled; unsigned int(1)interactive_object_source; unsigned int(1)primary_render_item_active; } else / * Any other renderer scenario related state information* / } }
[0160] [Table 2]
[0161] `num_active_triangles` may indicate the current number of active triangles being considered within the renderer. num_hoa_source_trr may indicate the current number of HOA sources being considered by the renderer for active processing. hoa_source_id may indicate the HOA source identifier of an HOA source in the TRR list. A full_order_source flag equal to 1 may indicate that all channels of the HOA source are required by the renderer. A first_order_source equal to 1 may indicate that the primary HOA signal is required by the renderer for rendering for this HOA source. `num_object_sources` indicates the number of object sources in the audio scene. The object_source_id indicates the object source for which the retrieval clue is provided by the interface. The `interactive_object_source` property indicates that the object source is interactive, and as a result, it may be inactive but can be activated by the user at any time. `primary_render_item_active` indicates that the object source is the first rendering item that is active and therefore should be retrieved.
[0162] If two or more flags are set to 1 by the renderer via the interface, the player implementation choice is responsible for prioritizing their retrieval.
[0163] In some embodiments, the player does not understand mphoa rendering. In other words, the player is configured to understand which order of HOA source is required based on the mphoa renderer scene state parameters. aligned(8)RendererStateInterface(){ unsigned int(8)renderType; / / Type of rendering state information if(renderType==6DoFHOARendering) unsigned int(3)num_active_triangles; / / Active triangles unsigned int(3)num_hoa_source_trr; / / TRR list for(i=0;i <num_hoa_source_trr;i++){ unsigned int(16)hoa_source_id; unsigned int(1)signal_interpolation_source; unsigned int(1)beamforming_source; unsigned int(1)spatial_metadata_source; } } else if(renderType==ObjectRendering) / * Define for Object Rendering related state information* / else / * Any other renderer scenario related state information* / } }
[0164] [Table 3]
[0165] `num_active_triangles` may indicate the current number of active triangles being considered within the renderer. num_hoa_source_trr may indicate the current number of HOA sources being considered by the renderer for active processing. hoa_source_id may indicate the HOA source identifier of an HOA source in the TRR list. A value of 1 for the signal_interpolation_source flag may indicate that the HOA source is used for signal interpolation, while a value equal to 0 indicates that it is not a signal interpolation source. Sources used for spatial metadata processing are also implicitly used for spatial metadata interpolation. A beamforming_source flag equal to 1 may indicate that the HOA source is used for beamforming, while a value equal to 0 indicates that it is not used for beamforming. A spatial_metadata_source value equal to 1 may indicate that the HOA source is used for spatial metadata processing. A value equal to 0 indicates that it is not used exclusively for spatial metadata processing. Both signal_interpolation_source and spatial_metadata_source may be equal to 1 during transitions or changes in the signal interpolation source.
[0166] As shown in Figure 12, this can result in a situation where only a subset of channels are processed (e.g., by STFT), and the unprocessed channels are skipped to avoid unnecessary acquisition of audio data.
[0167] Therefore, with respect to channel 1200 and frequency bin 1201 representing the HOA source, a subset of channels 1205 representing the FOA channels processed by STFT is shown, while the other 12 channels of a single HOA source are either not processed or are skipped.
[0168] In some embodiments, such as those described above, the decoder / player does not understand mphoa rendering, so the player receives listener positions and HOA source positions from the media manifest. Based on the listener positions, the renderer may be configured to determine the relevant HOA sources for processing. These are determined by the renderer using active triangle discrimination and triangle track records (TRRs). The player uses RSI to obtain HOA sources used for signal interpolation and HOA sources used for spatial metadata processing.
[0169] In some embodiments, as an alternative to creating a media description that includes a primary HOA representation in addition to the full-order HOA representation, the media description may hold the full-order HOA representation, in which case the player decides to retrieve audio data corresponding only to the primary HOA channels for the HOA source indicated by RendererStateInterface(). This can be implemented by making a suitable byte range request to retrieve only the audio data corresponding to the first four channels from the full-order HOA representation.
[0170] In such an implementation, generating additional primary HOA representations is avoided, and it requires additional determination of byte range information that will be obtained.
[0171] One example is shown in Figure 13, which illustrates an example of an audio source to be encoded / acquired when the listener is in an audio scene as shown in Figure 1 over a time interval of T0151 to T1153, i.e., ΔT1102. In this example, it is shown that HOA audio source HOA1103 has the full set of channels and frequency bins 1301, while HOA audio sources HOA2105 and HOA audio sources HOA3107 have subsets of channels and frequency bins 1303 and 1305, respectively. As such, the acquisition of unnecessary audio channels 1307 is avoided. In these embodiments, the player acquires a full-order HOA representation from the media manifest for the HOA source used for signal interpolation and the HOA source used for beamforming. Simultaneously, the decoder / player acquires a primary HOA representation from the media manifest for the HOA source used for spatial metadata processing. The acquired media is delivered to an MPEG-H 3DA decoder, which then passes the decoded PCM frame buffer to an MPEG-I immersive audio renderer.
[0172] In some embodiments, the decoder / player is configured to receive listener positions and HOA source positions from a media manifest. Based on the listener positions, the renderer determines the relevant HOA sources for processing. These are determined by the renderer using active triangle discrimination and triangle track records (TRRs). The player uses RSI to obtain a list of HOA sources to which HOA order representation levels will be retrieved. In one implementation embodiment, the RSI interface indicates to the player the need for full-order and primary representations. The player retrieves full-order and primary HOA representations from the media manifest. The retrieved media is delivered to an MPEG-H 3DA decoder, which passes the decoded PCM frame buffer to an MPEG-I immersive audio renderer.
[0173] Figures 14 and 15 show schematic diagrams of the current implementation (shown in Figure 14) and several embodiments of the implementation (shown in Figure 15).
[0174] Therefore, Figure 14 shows an EIF source 1401 configured to provide N HOA sources 1400 to the encoder 1405. Further shown is an MPEG-I audio source 1403 configured to also provide an MPEG-I audio signal 1402 to the encoder 1405.
[0175] The encoder 1405 comprises an MPEG-I encoder 1407 and an MPEG-H encoder 1409, which are configured to generate a bitstream 1410 for the MPEG-I content server 1411.
[0176] The MPEG-I content server 1411 is configured to store N full-order HOA sources. The decoder / player may then be configured to retrieve three HOA source full-order audio data 1412 from the MPEG-I content server 1411 to the MPEG-renderer 1413.
[0177] With respect to Figure 15, it is then additionally shown that encoder 1505 is configured to encode the FOA representation of the HOA source as shown by 1501. In these embodiments, the MPEG-I content server 1511 is configured to store both N HOA sources and new N FOA representation audio data.
[0178] In these embodiments, the MPEG-I content server 1511 is configured to indicate the availability of the FOA representation audio data of all HOA source audio data in the media manifest file 1503.
[0179] In some embodiments, a renderer state interface 1513 coupled to an MPEG-I renderer 1413 is configured to control the acquisition 1512 of 1x full-order audio data and 2x primary audio data from an MPEG-I content server 1511.
[0180] The additional cost relates to generating an additional HOA primary representation within the DASH server (MPEG-I content).
[0181] The following describes a scenario for acquiring audio data for 6DoF HOA rendering. As discussed earlier, the majority (redundant or wasteful) method is to deliver all HOA sources within the audio scene over the entire duration of the audio scene consumption. This is not efficient at all.
[0182] An improvement over the full supply example is a method for delivering only the HOA sources required for 6DoF HOA rendering based on the listener's position within the audio scene.
[0183] The following equation explains the total delivery bitrate:
[0184]
number
[0185] Total bitrate for delivering the HOA source for α=6DoF HOA rendering. β = Bitrate per HOA source n = number of HOA sources being distributed
[0186] In some embodiments, different HOA-order representations for a given HOA source in an audio scene are β j This is expressed as j, where j can range from 1 to 6 (from primary HOA source to a maximum of sixth-order HOA source). In this example, the upper limit is 6 because this is the highest order supported by MPEG-H 3DA (ISO / IEC 23080-3 third edition).
[0187] The number of active HOA sources, or the number of HOA sources required for 6DoF HOA rendering at a given listener location, can vary depending on various factors, such as whether one or more HOA sources encompassing the listener location are about to change to form a triangle, or whether the listener has performed a teleport action that results in a jump at the listener location to a new set of HOA sources required for 6DoF HOA rendering.
[0188] A typical scenario involves a listener remaining within a triangle formed by three HOA sources for a period longer than a default threshold (e.g., 30 milliseconds, derived from the duration of six 256-sample audio frames at a 48000Hz sampling rate), where the three HOA sources are considered necessary for 6DoF HOA rendering. If we consider this scenario as a benchmark for determining the required bitrate, the following equation describes the delivery bitrate using prior art or advanced technology methods: α=3 * β j (In the formula, j represents the order of the distribution HOA source). However, the proposed methods in these embodiments are governed by the following equation (in the case of equivalent scenarios): α=2 * β1+β j (In the formula, j represents the order of the distribution HOA source). This results in significant bitrate savings without any loss of quality.
[0189] Therefore, the bitrates for delivering HOA sources at different orders are illustrated in Figure 16, which shows quality 1601 for various orders. Note that the bitrate values for different orders shown in Figure 16 are examples and may be determined based on the content creator's choice (depending on the encoder implementation).
[0190] [Table 4]
[0191] When delivering HOA sources using the current implementation and proposed embodiments, it is clear that primary HOA sources 1611, i.e., FOAs, receive no benefit whatsoever. However, for higher-order sources 1613, 1615, 1617, 1619, and 1621, the benefits increase significantly. This is evident from the illustrated rate distortion curves, such as those shown in Figure 16, which demonstrate that bitrate savings increase in scenes with higher-order HOA sources.
[0192] With respect to Figure 17, an example electronic device is shown that may be used as a computer, an encoder processor, a decoder processor, or any of the functional blocks described herein. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 2000 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0193] In some embodiments, the device 2000 comprises at least one processor or central processing unit 2007. The processor 2007 may be configured to execute various program code, such as methods described herein.
[0194] In some embodiments, device 2000 includes memory 2011. In some embodiments, at least one processor 2007 is coupled to memory 2011. Memory 2011 can be any preferred storage means. In some embodiments, memory 2011 includes a program code section for storing program code that can be implemented on processor 2007. Furthermore, in some embodiments, memory 2011 may further include a storage data section for storing data, for example, data that is being processed or will be processed according to embodiments such as those described herein. The implementable program code stored in the program code section and the data stored in the storage data section can be retrieved by processor 2007 at any time as needed by the memory-processor coupling.
[0195] In some embodiments, device 2000 includes a user interface 2005. In some embodiments, the user interface 2005 may be coupled to a processor 2007. In some embodiments, the processor 2007 may control the operation of the user interface 2005 and receive input from the user interface 2005. In some embodiments, the user interface 2005 may allow a user to input commands to device 2000, for example, via a keypad. In some embodiments, the user interface 2005 may allow a user to obtain information from device 2000. For example, the user interface 2005 may include a display configured to show information from device 2000 to the user. In some embodiments, the user interface 2005 includes a touchscreen or touch interface having the ability to both input information into device 2000 and to display further information to the user of device 2000.
[0196] In some embodiments, device 2000 includes an input / output port 2009. In some embodiments, the input / output port 2009 includes a transceiver. The transceiver in such embodiments may be coupled to a processor 2007 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver or any preferred transceiver or transmitter and / or receiver means may, in some embodiments, be configured to communicate with other electronic devices or devices via wire or wired coupling.
[0197] The transceiver may communicate with further devices by any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable Short-Range Radio Frequency Communication protocol such as Bluetooth, or an Infrared Data Path (IRDA).
[0198] The transceiver input / output ports 2009 may be configured to transmit / receive audio signals and bitstreams, and in some embodiments, to perform the operations and methods described above by using a processor 2007 that executes preferred code.
[0199] In general, various embodiments of the present invention may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some embodiments may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Various embodiments of the present invention may be illustrated and described using block diagrams, flowcharts, or some other graphic representations, but it should be understood that these blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof, as non-limiting examples.
[0200] Embodiments of the present invention may be implemented by computer software executable by the data processor of a mobile device, such as within a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, it should be noted that any block of the logic flow shown in the figure may represent a program step, or an interconnected set of logic circuits, blocks, and functions, or a combination of program steps and logic circuits, blocks, and functions. The software may be stored in memory blocks implemented in such physical media as memory chips, or in processors, magnetic media, and optical media.
[0201] Memory can be any type suitable for the local technology environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. Data processors can be any type suitable for the local technology environment and, in non-limiting examples, may include one or more of general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits and processors based on multi-core processor architectures.
[0202] Embodiments of the present invention can be put into practice in various components, such as integrated circuit modules. Designing integrated circuits is generally a highly automated process. Complex and powerful software tools are available to translate logic-level designs into ready-to-etch and formed semiconductor circuit designs on semiconductor substrates.
[0203] Programs such as those offered by Synopsys, Inc. in Mountain View, California, and California and Cadence Design in San Jose, use well-established design rules and a library of pre-stored design modules to automatically route conductors and position components on a semiconductor chip. Once the semiconductor circuit design is complete, the resulting design, in a standardized electronic format (e.g., Opus, GDSII, or similar), can be sent to a semiconductor fabrication facility, or "fab," for manufacturing.
[0204] The foregoing description provides a complete and informative description of exemplary embodiments of the invention, as illustrative and non-limiting examples. However, various modifications and adaptations may become apparent to those skilled in the art when read in conjunction with the accompanying drawings and claims. Nevertheless, all such and similar modifications of the teachings of the invention shall still fall within the scope of the invention as defined in the claims.
Claims
1. A method for generating an audio scene having at least two audio content sets, To obtain a media description that includes information identifying the location of audio sources within an audio scene associated with at least two audio content sets, Encode at least two audio content sets such that at least two audio content sets are output at at least two expression levels, A method comprising outputting at least two encoded audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first level of at least two levels of representation, and a second first encoded at least two audio content sets encoded at a second level of representation.
2. The method according to claim 1, wherein at least two audio content sets include at least two higher ambisonic audio sources.
3. The method according to claim 1 or 2, wherein at least two audio content sets include metadata associated with each of at least two higher ambisonic audio sources.
4. The method according to claim 3, wherein metadata associated with each of at least two higher-order ambisonic audio sources includes a position within an audio scene.
5. The method according to any one of claims 1 to 4, wherein the media description further includes information that identifies at least two levels of representation.
6. At least two levels of expression, Full-order representation of at least two audio content sets, Reduction or lower-order representation of at least two audio content sets, and Primary representation of at least two audio content sets The method according to any one of claims 1 to 5, comprising at least two of the above.
7. The method according to claim 6, subject to claim 2, wherein the full-order representation of at least two audio content sets includes all channels of at least two higher-order ambisonic audio sources, the reduced or lower-order representation of at least two audio content sets is a selection of channels from at least two higher-order ambisonic audio sources, and the primary representation of at least two audio content sets is a selection of the first four channels from at least two higher-order ambisonic audio sources.
8. The step of outputting at least two encoded audio content sets for the audio scene is, To receive at least one request regarding an audio scene, wherein the request specifies the level of representation for at least two audio content sets, Selecting at least two audio content sets encoded at the representation level for at least two audio content sets identified within at least one request, Outputting at least two encoded audio content sets selected at the representation level for at least two audio content sets identified within at least one request, The method according to any one of claims 1 to 7, including
9. The method according to any one of claims 1 to 8, further comprising obtaining a media description including information that identifies the location of an audio source in an audio environment associated with at least two audio content sets, outputting the media description to the device, and further outputting the encoded at least two audio content sets for an audio scene to the device.
10. The method according to claim 9, wherein the device is a spatial audio player configured to generate an audio scene.
11. A method for generating an audio scene having at least two audio content sets, To obtain a media description that includes information identifying the location of audio sources within an audio scene associated with at least two audio content sets, To obtain the listener's position within the audio scene, For at least two audio content sets, the source representation level is determined based on the location of the audio source associated with at least two content sets and the listener's position. Obtaining at least two encoded audio content sets based on the source representation level, wherein the at least two encoded audio content sets include a first encoded audio content set encoded at a first of at least two representation levels, and a second encoded audio content set encoded at a second of at least two representation levels. To generate spatialized audio output based on at least two sets of encoded audio content obtained. Methods that include...
12. At least two levels of expression, Full-order representation of at least two audio content sets, Reduction or lower-order representation of at least two audio content sets, and Primary representation of at least two audio content sets The method according to claim 11, comprising at least two of the following.
13. The method according to claim 11 or 12, wherein the media description further specifies the ambisonic order for at least two audio content sets.
14. The method of claim 13, as dependent on claim 12, wherein obtaining at least two encoded audio content sets based on the source representation level includes obtaining all encoded channel representations for the at least two encoded audio content sets when the source representation level is full-order representation.
15. The method according to claim 13, as dependent on claim 12 or claim 14, wherein obtaining at least two encoded audio content sets based on the source representation level includes obtaining encoded selected channel representations for at least two encoded audio content sets when the source representation level is reduced or low-order representation.
16. The method according to claim 13, as dependent on claim 12 or claim 14 or 15, wherein obtaining at least two encoded audio content sets based on the source representation level includes obtaining the first four encoded selected channel representations for the at least two encoded audio content sets when the source representation level is primary representation.
17. It is possible to generate spatialized audio output based on at least two sets of acquired encoded audio content. Signal interpolation is performed by processing the channel signals of at least two acquired encoded audio content sets based on the listener position. To generate spatialized audio output based on the signal interpolation performed and The method according to any one of claims 11 to 16, including the method described above.
18. Based on the source representation level, it is possible to obtain at least two encoded audio content sets. To generate a request for representation of at least two encoded audio content sets, wherein the request includes a source representation level, To receive the requested representation of at least two encoded audio content sets and The method according to any one of claims 11 to 17, including the method described above.
19. The method according to any one of claims 11 to 18, wherein determining source representation levels for at least two audio content sets based on the location and listener location of the audio sources associated with at least two content sets further includes determining content set rendering-related state information associated with at least two audio content sets, and obtaining the content set rendering-related state information, the encoded at least two audio content sets, based on the content set rendering-related state information.
20. Content set rendering-related status information, the content set, A 6-degree-of-freedom higher-order ambisonic audio source during the generation of spatialized audio output, Object rendering audio source and generating spatialized audio output The method according to claim 19, configured to determine whether or not it will be used as one of the following.
21. A device for generating an audio scene having at least two audio content sets, Obtain a media description that includes information identifying the location of audio sources within an audio scene associated with at least two audio content sets, Encode at least two audio content sets so that at least two audio content sets are output at at least two expression levels, An apparatus comprising means configured to output at least two encoded audio content sets for an audio scene, wherein the output encoded at least two audio content sets include a first encoded audio content set encoded at least two levels of expression, and a second first encoded audio content set encoded at least two levels of expression.
22. A device for generating an audio scene having at least two audio content sets, Obtain a media description that includes information identifying the location of audio sources within an audio scene associated with at least two audio content sets, Obtain the listener's position within the audio scene, For at least two audio content sets, the source representation level is determined based on the position of the audio source associated with at least two content sets and the listener position. Based on the source representation level, obtain at least two encoded audio content sets, and the at least two encoded audio content sets include a first at least two encoded audio content sets encoded at a first at least two representation levels, and a second at least two encoded audio content sets encoded at a second at least two representation levels. Generate spatialized audio output based on at least two sets of acquired encoded audio content. An apparatus including means configured in such a manner.
23. An apparatus for generating an audio scene having at least two audio content sets, the apparatus comprising at least one processor and at least one memory for storing instructions, wherein when an instruction is executed by the at least one processor, the apparatus has at least Obtain a media description that includes information identifying the location of audio sources within an audio scene associated with at least two audio content sets, Encode at least two audio content sets so that at least two audio content sets are output at at least two expression levels, An apparatus that causes at least two encoded audio content sets to output for an audio scene, wherein the output encoded at least two audio content sets include a first encoded at least two audio content sets encoded at a first level of at least two levels of expression, and a second first encoded at least two audio content sets encoded at a second level of expression.
24. An apparatus for generating an audio scene having at least two audio content sets, the apparatus comprising at least one processor and at least one memory for storing instructions, wherein when an instruction is executed by the at least one processor, the apparatus has at least Obtain a media description that includes information identifying the location of audio sources within an audio scene associated with at least two audio content sets, Obtain the listener's position within the audio scene, For at least two audio content sets, the source representation level is determined based on the position of the audio source associated with at least two content sets and the listener position. Obtaining at least two encoded audio content sets based on the source representation level, wherein the at least two encoded audio content sets include a first encoded audio content set encoded at a first of at least two representation levels, and a second encoded audio content set encoded at a second of at least two representation levels. Generate spatialized audio output based on at least two sets of acquired encoded audio content. A device that causes something to happen.