Audio encoder and audio decoder
The audio decoder addresses computational constraints by mapping dynamic audio objects to static objects in predefined speaker configurations, ensuring immersive audio output even in low-computational complexity scenarios.
Patent Information
- Application Number
- JP2025186037
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-01-16
- Filing Date
- 2025-11-05
- Publication Date
- 2026-01-27
AI Technical Summary
Existing audio decoders face challenges in providing immersive audio output from dynamic audio objects due to computational constraints and bitrate limitations, especially when operating in core decode mode, where parametric reconstruction of individual dynamic audio objects is not possible.
An audio decoder with a controller that selects between different decoding modes, allowing for mapping dynamic audio objects to static audio objects in a predefined speaker configuration, even in low-computational complexity scenarios, thereby achieving immersive audio output.
Enables immersive audio output from low-bitrate bitstreams by mapping dynamic audio objects to static audio objects, reducing computational complexity and maintaining flexibility in decoding processes.
Smart Images

Figure 2026012934000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: U.S. Provisional Application No. 62 / 754,758 (Docket No. D18053USP1), filed November 2, 2018, European Patent Application No. 18204046.9 (Docket No. D18053EP), filed November 2, 2018, and U.S. Provisional Application No. 62 / 793,073 (Docket No. D18053USP2), which are incorporated herein by reference.
[0002] Technical Field This disclosure relates to the field of audio coding, and in particular to an audio decoder having at least two decoding modes, as well as related decoding methods and decoding software for such an audio decoder. This disclosure further relates to a corresponding audio encoder, and related encoding methods and encoding software for such an audio encoder. [Background technology]
[0003] An audio scene can generally contain audio objects. An audio object is an audio signal with an associated spatial location. If the spatial location of an audio object changes over time, the audio object is typically called a dynamic audio object. If the location is static, the audio object is typically called a static audio object or bed object. Bed objects are typically audio signals that directly correspond to channels of a multi-channel speaker configuration, such as a classic stereo configuration with left and right speakers, or a so-called 5.1 speaker configuration with three front speakers, two surround speakers, and a low-frequency effects speaker. A bed can contain one or many bed objects. It is thus a collection of bed objects that can match a multi-channel speaker configuration.
[0004] Because the number of audio objects can typically be very large, e.g., on the order of tens or hundreds of audio objects, an encoding method is needed that allows the audio objects to be efficiently compressed at the encoder side, for example, for transmission as a bitstream (e.g., data stream). This is especially true when low bit rates are targeted for transmission. In this case, clusters of dynamic audio objects are parametrically reconstructed into individual audio objects in certain decoding modes in the audio decoder. These are rendered into a set of output audio signals depending on the configuration of the output device (e.g., speakers, headphones, etc.) used to play the audio signals. However, in some cases, the decoder is forced to function in core mode, which means that parametric reconstruction of individual dynamic audio objects from clusters of dynamic audio objects is not possible, e.g., due to decoder processing power constraints or other reasons. This can cause problems, especially when an immersive audio experience (e.g., 3D audio) is expected from the user listening to the output audio.
[0005] Therefore, improvements are needed in this context. Summary of the Invention [Problem to be solved by the invention]
[0006] In view of the above, it is an object of the present disclosure to overcome or mitigate at least some of the problems discussed above. In particular, it is an object of the present disclosure to provide, in a decoder in a core decode mode, preferably immersive audio output from received dynamic audio objects. It is a further object of the present disclosure to provide an encoder for encoding an audio bitstream from a collection of dynamic audio objects in a manner that allows for decoding the audio bitstream into preferably immersive audio objects as described above. Further and / or alternative objects of the present disclosure will be apparent to the reader of this disclosure. [Means for solving the problem]
[0007] According to a first aspect of the present invention, there is provided an audio decoder comprising one or more buffers for storing a received audio bitstream, and a controller coupled to the one or more buffers.
[0008] The controller is configured to operate in a decoding mode selected from a plurality of different decoding modes, the plurality of different decoding modes including a first decoding mode and a second decoding mode, wherein only the first decoding mode of the first and second decoding modes allows complete decoding of one or more encoded dynamic audio objects in the bitstream into reconstructed individual audio objects.
[0009] If the selected decoding mode is a second decoding mode, the controller is configured to access the received audio bitstream, determine whether the received audio bitstream includes one or more dynamic audio objects, and in response to determining that the received audio bitstream includes one or more dynamic audio objects, map at least one of the one or more dynamic audio objects to a set of static audio objects, the set of static audio objects corresponding to a predefined speaker configuration.
[0010] By including a step of mapping at least one of the one or more dynamic audio objects to a set of static audio objects, an immersive audio output can be achieved from a low bitrate bitstream that is constrained to contain, for example, only up to 10 audio objects (dynamic and static), or up to 7, 5, etc. audio objects, even in a decoder operating in a low-computational complexity decoding mode (core decoding) where parametric reconstruction of individual dynamic audio objects from a cluster of dynamic audio objects is not possible (full decoding is not possible).
[0011] By the term "immersive audio output" in the context of this specification, a channel output configuration including channels for top speakers should be understood.
[0012] By the term "immersive speaker configuration" the same meaning should be understood, namely a speaker configuration that includes an upper speaker.
[0013] Furthermore, this embodiment provides a flexible decoding method since not all received dynamic audio objects are necessarily mapped to a set of static audio objects corresponding to a predefined speaker configuration, allowing, for example, the inclusion of additional dialogue objects in the audio bitstream that serve different purposes, e.g., dialogue and related audio.
[0014] Furthermore, this embodiment allows a flexible process of providing and later rendering a collection of static audio objects, e.g., to achieve lower computational complexity or to allow reuse of existing software code / functions used to implement the decoder, as will be discussed further below.
[0015] In general, this embodiment allows decoder-side flexibility in low bitrate, low complexity scenarios.
[0016] The step of the controller determining that the received audio bitstream contains one or more dynamic audio objects can be accomplished in various ways. According to some embodiments, this is determined from metadata in the bitstream, such as integer values or flag values. In other embodiments, this may be determined by analysis of the audio objects or associated object metadata.
[0017] The controller may select the decoding mode in various ways. For example, the selection may be made using bitstream parameters, and / or in view of the output configuration for the rendered output audio signal, and / or by checking the number of dynamic audio objects (downmix audio objects, clusters, etc.) in the audio bitstream, and / or based on user parameters, etc. It should be noted that the decision to map at least one of the one or more dynamic audio objects to a set of static audio objects may be made using more information than simply determining whether the received audio bitstream contains one or more dynamic audio objects.
[0018] According to some embodiments, the controller also bases such a decision on further data, such as bitstream parameters. For example, if it is determined that the received audio bitstream does not contain dynamic audio objects, or if it is determined for other reasons that the above-mentioned dynamic audio object mapping should not be performed, the controller may decide to directly render the received static audio objects (e.g., bed objects) into the set of output audio channels, for example, using the received rendering coefficients (e.g., downmix coefficients) applicable to the configuration of the output audio channels. In this mode of operation of the controller, the received dynamic audio objects are rendered into the output audio channels in the usual way.
[0019] According to some embodiments, if the selected decoding mode is the second decoding mode, the controller is further configured to render the set of static audio objects into the set of output audio channels, and any other static audio objects received in the audio bitstream (such as LFE) are also rendered into the set of output audio channels, advantageously in the same rendering step.
[0020] According to some embodiments, the configuration of the set of output audio channels is different from the predefined speaker configuration used to map dynamic audio objects to a set of static audio objects as described above, and increased flexibility is achieved because the predefined speaker configuration is not limited to the configuration of the output audio channels.
[0021] According to some embodiments, the audio bitstream includes a first set of downmix coefficients, and the controller is configured to utilize the first set of downmix coefficients to render the set of static audio objects into the set of output audio channels, and in the case of further received static audio objects in the bitstream, the downmix coefficients are applied to both the set of static audio objects and the further static audio objects.
[0022] In some embodiments, the controller may use the received first set of downmix coefficients directly to render a set of static audio objects into a set of output audio channels, however, in other embodiments, the first set of downmix coefficients must first be processed based on the type of encoder-side downmix operation that resulted in the one or more dynamic audio objects received in the bitstream.
[0023] In some embodiments, the controller is further configured to receive information regarding an attenuation applied to at least one of the one or more dynamic audio objects at the encoder side. The information may be received in the bitstream or may be predefined at the decoder. The controller may then be configured to modify the first set of downmix coefficients accordingly when using the first set of downmix coefficients to render the set of static audio objects into the set of output audio channels. As a result, attenuations contained in the downmix coefficients but already applied at the encoder side are not applied twice, resulting in a better listening experience.
[0024] In some embodiments, the controller is further configured to receive information regarding a downmix operation performed on the encoder side, the information defining an original channel configuration of the audio signal, the downmix operation resulting in downmixing the audio signal into the one or more dynamic audio objects. In this case, the controller may be configured to select a subset of the first set of downmix coefficients based on the information regarding the downmix information, and utilizing the first set of downmix coefficients to render the set of static audio objects into the set of output audio channels comprises utilizing the subset of the first set of downmix coefficients to render the set of static audio objects into the set of output audio channels. This may result in a more flexible decoding method that handles all types of downmix operations performed on the encoder side that result in the received one or more dynamic audio objects.
[0025] According to some embodiments, the controller is configured to perform the mapping of the at least one of the one or more dynamic audio objects and the rendering of the set of static audio objects in a combined calculation using a single matrix, which advantageously can reduce the computational complexity of rendering audio objects in the received audio bitstream.
[0026] According to some embodiments, the controller is configured to perform the mapping of at least one of the one or more dynamic audio objects and the rendering of the set of static audio objects in separate calculations using respective matrices. In this embodiment, the one or more dynamic audio objects are pre-rendered to a set of static audio objects, which defines an intermediate bed representation of the one or more dynamic audio objects. Advantageously, this allows the reuse of existing software code / functions used to implement a decoder adapted to render a bed representation of an audio scene to a set of output audio channels. Furthermore, this embodiment reduces the additional complexity of implementing the invention described herein in a decoder.
[0027] According to some embodiments, the received audio bitstream includes metadata identifying said at least one of said one or more dynamic audio objects, which allows for increased flexibility of the decoder method, since not all of the received one or more dynamic audio objects need to be mapped to a set of static audio objects, and the controller can use said metadata to easily determine which of the received one or more dynamic objects should be mapped and which should be directly forwarded to rendering in a set of output audio channels.
[0028] According to some embodiments, the metadata indicates that N of the one or more dynamic audio objects should be mapped to a set of static audio objects, and the controller is configured, in response to the metadata, to map N of the one or more dynamic audio objects selected from predefined position(s) in the received audio bitstream to the set of static audio objects. For example, the N dynamic audio objects may be the first N received dynamic audio objects or the last N received dynamic audio objects. As a result, in some embodiments, in response to the metadata, the controller is configured to map the first N of the one or more dynamic audio objects in the received audio bitstream to the set of static audio objects. This allows for less metadata, e.g., an integer value, to identify the at least one of the one or more dynamic audio objects.
[0029] According to some embodiments, the one or more dynamic audio objects included in the received audio bitstream include more than N dynamic audio objects. As mentioned above, for example, for audio containing dialogue in different languages, it may be advantageous to provide a dynamic audio object for each supported language.
[0030] According to some embodiments, the one or more dynamic audio objects included in the received audio bitstream include the N dynamic audio objects and K further dynamic audio objects, and the controller is configured to render the set of static audio objects and the K further audio objects into the set of output audio channels. Thus, for example, a selected language (i.e., a corresponding dynamic audio object) according to the above example may be rendered into the set of output audio signals together with the set of static audio objects.
[0031] According to some embodiments, the set of static audio objects consists of M static audio objects, where M>N>0. Advantageously, the number of mapped dynamic audio objects can be reduced, thus saving bitrate. Alternatively, the number of additional dynamic audio objects (K) in the audio bitstream can be increased.
[0032] According to some embodiments, the received audio bitstream further includes one or more additional static audio objects, which may include an LFE or other BED or Intermediate Spatial Format (ISF) object.
[0033] According to some embodiments, the set of output audio channels is one of: stereo output channels, 5.1 surround sound audio output channels, 5.1.2 immersive audio output channels, or 5.1.4 immersive audio output channels.
[0034] According to some embodiments, the predefined speaker configuration is a 5.0.2 speaker configuration, in which N may be equal to 5.
[0035] According to a second aspect of the present invention, the above objects are at least partly achieved by a method in a decoder comprising the following steps: receiving an audio bitstream and storing the received audio bitstream in one or more buffers; - selecting a decoding mode from a plurality of different decoding modes, the plurality of different decoding modes including a first decoding mode and a second decoding mode, wherein only the first decoding mode of the first and second decoding modes allows parametric reconstruction of individual dynamic audio objects from a cluster of dynamic audio objects; - operating a controller coupled to the one or more buffers in a selected decoding mode; If the selected decoding mode is the second decoding mode, the method further comprises the steps of: accessing, by the controller, the received audio bitstream; determining, by the controller, whether the received audio bitstream includes one or more dynamic audio objects; and in response to determining that the received audio bitstream includes one or more dynamic audio objects, mapping, by the controller, at least one of the one or more dynamic audio objects to a set of static audio objects corresponding to a predefined speaker configuration.
[0036] According to a third aspect of the present invention, at least some of the above objects are obtained by a computer program product comprising a computer readable medium having computer code instructions adapted to perform the method of the second aspect when executed by a device having processing capability.
[0037] The second and third aspects may generally have the same features and advantages as the first aspect.
[0038] According to a fourth aspect of the present invention, at least some of the above objects are obtained by an audio encoder comprising: a receiving component configured to receive a set of audio objects; a downmix component configured to downmix the set of audio objects into one or more downmixed dynamic audio objects, at least one of the one or more downmixed dynamic audio objects being intended to be mapped to a set of static audio objects in at least one of a plurality of decoding modes on a decoder side, the set of static audio objects corresponding to a predefined speaker configuration; a downmix coefficient providing component configured to determine a first set of downmix coefficients to be utilized for rendering the set of static audio objects corresponding to the predefined speaker configuration into a set of output audio channels at a decoder side; a bitstream multiplexer configured to multiplex the at least one downmixed dynamic audio object and the first set of downmix coefficients into an audio bitstream.
[0039] According to some embodiments, the downmix component is further configured to provide metadata identifying the at least one of the one or more downmixed dynamic audio objects to a bitstream multiplexer, the bitstream multiplexer being further configured to multiplex the metadata into the audio bitstream.
[0040] According to some embodiments, the encoder is further adapted to determine information regarding an attenuation to be applied to at least one of the one or more dynamic audio objects when downmixing the set of audio objects into one or more downmixed dynamic audio objects, and the bitstream multiplexer is further configured to multiplex said information regarding the attenuation into the audio bitstream.
[0041] According to some embodiments, the bitstream multiplexer is further configured to multiplex information regarding the channel configuration of the audio objects received by the receiving component.
[0042] According to a fifth aspect of the present invention, the above objects are at least partly obtained by a method in an encoder comprising the steps of: receiving a set of audio objects; - downmixing the set of audio objects into one or more downmixed dynamic audio objects, at least one of the one or more downmixed dynamic audio objects being intended to be mapped to a set of static audio objects in at least one of a plurality of decoding modes on the decoder side, the set of static audio objects corresponding to a predefined speaker configuration; - determining a first set of downmix coefficients to be used for rendering the set of static audio objects corresponding to the predefined speaker configuration into a set of output audio channels at a decoder side; - multiplexing the at least one downmixed dynamic audio object and the first set of downmix coefficients into an audio bitstream.
[0043] According to a sixth aspect of the present invention, at least some of the above objects are obtained by a computer program product comprising a computer readable medium having computer code instructions adapted to perform the method of the fifth aspect when executed by a device having processing capability.
[0044] The fifth and sixth aspects may generally have the same features and advantages as the fourth aspect. Furthermore, the fourth, fifth, and sixth aspects may generally have features corresponding to the first, second, and third aspects (but from the encoder side). For example, the encoder may be configured to include static audio objects (e.g., LFE) in the audio bitstream.
[0045] It is further noted that the present invention relates to all possible combinations of features unless expressly stated otherwise. [Brief explanation of the drawings]
[0046] The above, as well as additional objects, features, and advantages of the present invention, will be better understood from the following illustrative and non-limiting detailed description of preferred embodiments of the invention, with reference to the accompanying drawings, in which like reference numerals will be used for similar elements, in which: [Figure 1] FIG. 1 illustrates an audio decoder according to some embodiments. [Figure 2] FIG. 4 is a diagram illustrating a decoding operation according to the first embodiment. [Figure 3] FIG. 10 is a diagram illustrating a decoding operation according to the second embodiment. [Figure 4] FIG. 10 is a diagram illustrating a decoding operation according to the third embodiment. [Figure 5] FIG. 2 illustrates an encoding operation according to some embodiments. [Figure 6] An example is shown of an audio decoder unit for generating the gain matrix used to render a set of output audio channels. DETAILED DESCRIPTION OF THE INVENTION
[0047] DETAILED DESCRIPTION OF THE INVENTION The present invention now will be described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the invention are shown. The systems and devices disclosed herein will be described in operation.
[0048] In the following, the Dolby AC-4 audio format (published in document ETSI TS103 190-2 V1.2.1 (2018-02)) is used as a context for illustrating the present invention. However, it should be noted that the scope of the present invention is not limited to AC-4, and the various embodiments described herein may be used for any suitable audio format.
[0049] Due to computational constraints in some audio decoders, parametric reconstruction of individual dynamic audio objects from clusters of dynamic audio objects is not possible. Furthermore, constraints on the target bitrate for an audio bitstream may impose restrictions on the content of the audio bitstream, for example, limiting the number of transmitted audio objects / audio channels to 10. Further restrictions stem from the encoding standard used, which may restrict the use of certain coding tools in some specific cases. For example, AC-4 decoders are configured with various levels, and level 3 decoders restrict the use of coding tools such as A-JCC (Advanced Joint Channel Coding) and A-CPL (Advanced Coupling), which can be advantageously used to achieve an immersive audio experience under certain circumstances. Such situations may include mandatory channel encoding modes, where the decoder does not have the coding tools to decode such content (e.g., the use of A-JCC is not permitted). In this case, the present invention may be used to "mimic" channel-based immersion, as described below. Further possible constraints include the possibility of including both channel-based content and dynamic / static audio objects (discrete audio objects) in the same bitstream, which may not be allowed under certain circumstances.
[0050] In this document, the term "cluster" refers to an audio object that is downmixed within an encoder. This will be discussed later with reference to FIG. 5. In a non-limiting example, ten individual dynamic audio objects may be input to an encoder. In some cases, as described above, it may not be possible to encode all ten dynamic audio objects independently. For example, the target bit rate may only allow encoding five dynamic audio objects. In this case, the total number of dynamic audio objects needs to be reduced. A possible solution is to combine the ten dynamic audio objects into a smaller number, five dynamic audio objects in this example. These five dynamic audio objects derived by combining (downmixing) the ten dynamic audio objects are dynamic downmixed audio objects, referred to herein as a "cluster."
[0051] The present invention aims to avoid some of the above limitations and provide a favorable listening experience to the listener of audio output at low bitrate and decoder complexity.
[0052] 1 illustrates, by way of example, an audio decoder 100. The audio decoder includes one or more buffers 102 for storing a received audio bitstream 110. In some embodiments, the received audio bitstream includes A-JOC (Advanced Joint Object Coding) substreams, representing, for example, Music and Effects (M&E) or a combination of M&E and dialogue (D) (i.e., a full MAIN (CM)).
[0053] Advanced Joint Object Coding (A-JOC) is a parametric coding tool for efficiently coding collections of objects. A-JOC relies on a parametric model of object-based content. This coding tool determines the dependencies between audio objects and utilizes a perceptually based parametric model to achieve high coding efficiency.
[0054] The audio decoder 100 further includes a controller 104 coupled to the one or more buffers 102. Thus, the controller 104 can extract at least portions 112 of the audio bitstream 110 from the buffer 102 and decode the encoded audio bitstream into a set of audio output channels 118. The set of audio output channels 118 can then be used for playback by a set of speakers 120.
[0055] As mentioned above, the audio decoder 100, or the controller 104, can operate in different decoding modes. In the following, two decoding modes are illustrated. However, further decoding modes may also be used.
[0056] A first decoding mode (e.g., full decoding mode, complex decoding mode, etc.) allows for the parametric reconstruction of individual dynamic audio objects from clusters of dynamic audio objects. In the context of AC-4, the first decoding mode may be referred to as A-JOC full decoding. In the non-limiting example given above with 10 individual dynamic objects and 5 clusters (dynamic downmixed audio objects), the full decoding mode allows for the reconstruction of the 10 original individual dynamic objects (or approximations thereof) from the 5 clusters.
[0057] In a second decoding mode (e.g., core decoding, low-complexity decoding), such reconstruction is not performed due to constraints in the decoder 100. In the AC-4 context, the second decoding mode may be referred to as A-JOC core decoding. In the non-limiting example described above with 10 individual dynamic objects and 5 clusters (dynamic downmixed audio objects), the core decoding mode cannot reconstruct the 10 original individual dynamic objects (or an approximation thereof) from the 5 clusters.
[0058] Thus, the controller is configured to select either the first decoding mode or the second decoding mode. Such a decision can be made based on internal parameters 116 of the decoder 100, for example, stored in the memory 106 of the decoder 100. Alternatively or additionally, the decision may be based on input 114, for example, from a user. Alternatively or additionally, the decision may be based on the content of the audio bitstream 110. For example, if the received audio bitstream contains more than a threshold number of dynamically downmixed audio objects (e.g., more than 6, or more than 10, or any other suitable number depending on the context), the controller may select the second decoding mode. In some embodiments, the audio bitstream 110 may include a flag value that indicates to the controller which decoding mode to select.
[0059] For example, in the context of AC-4, according to one embodiment, the first decoding mode selection may be one or many of the following: The presentation level is less than or equal to 2 (bitstream parameters). The output stage is configured for 5.1.2 output (user parameters). An A-JOC substream contains up to 5 downmix objects (clusters) (bitstream parameters). Applications do not force core decoding via API (user parameters).
[0060] In the following, the second decoding mode (core decoding) is illustrated in connection with FIGS.
[0061] FIG. 2 shows a first embodiment 109a of the second decoding mode 109 described in connection with FIG.
[0062] The controller 104 is configured to determine whether the received audio bitstream 110 includes one or more dynamic audio objects (which in this embodiment are all mapped to a set of static audio objects) and base a decision on how to decode the received audio bitstream on that determination. According to some embodiments, the controller also bases such a decision on further data, such as bitstream parameters. For example, in AC-4, the controller may decide to decode the received audio bitstream as described in FIG. 2 according to the values of one or both of the following bitstream parameters, i.e., if one of the following is true: 1. "num_bed_obj_ajoc" is greater than 0 (for example, 1 to 7), or 2. "num_bed_obj_ajoc" is not present in the bitstream and "n_fullband_dmx_signals" is less than 6.
[0063] If the controller 104 determines that one or more dynamic audio objects 210 should be taken into account, optionally taking into account other data as described above, the controller is configured to map at least one 210 of the one or more dynamic audio objects to a set of static audio objects. In FIG. 2 , all received dynamic audio objects are mapped to a set 222 of static audio objects, which correspond to predefined speaker configurations. The mapping is performed as follows: The audio bitstream 110 includes N dynamic audio objects 210. The audio bitstream further includes N corresponding object audio metadata (OAMD) 212. Each OAMD 212 defines the attributes, such as gain and position, of each of the N dynamic audio objects 210. The N OAMDs 212 are used to calculate 206 a gain matrix 218, which is used to pre-render the N dynamic audio objects 210 to a set of static audio objects 222. The size of the collection of static audio objects is M. Thus, N dynamic audio objects 210 are transformed (rendered) into a bed 222, for example a 5.0.2 bed (M=7). Other configurations, such as a 7.0.2 (M=9), are equally possible. The bed configuration (for example, 5.0.2) is predefined in the decoder 100, which uses this knowledge to calculate 206 the gain matrix 218. In other words, the collection of static audio objects 222 corresponds to a predefined speaker configuration. Thus, the gain matrix 218 in this case has a size of M×N.
[0064] According to some embodiments, M>N>0.
[0065] An advantage of actually rendering the N dynamic audio objects 210 into a bed 222 is that the remaining operation of the decoder 100 (i.e., generating the set of output audio signals 118) can be achieved by reusing existing software code / functions used to implement a decoder adapted to render the bed 222 (and optionally further dynamic audio objects as described in FIG. 3) into the set of output audio signals 118.
[0066] The decoder generates a set of additional OAMDs 214. These OAMDs 214 define positions and gains for the intermediate rendered beds 222. Thus, the OAMDs 214 are not conveyed in the bitstream, but instead are "generated" locally within the decoder to describe the channel configuration (typically 5.0.2) that will be generated at the output of the pre-rendering 202. For example, if the intermediate bed 222 is configured as 5.0.2, the OAMDs 214 define the positions (L, R, C, Ls, Rs, Ltm, Rtm) and gains for the 5.0.2 bed 222. If another configuration of the intermediate bed is used, such as 3.0.0, the positions would be L, R, C. Thus, the number of OAMDs 214 in this embodiment corresponds to the number of still audio objects 222, e.g., 7 for the 5.0.2 bed 222. In some embodiments, the gain of each OAMD 214 is 1. Thus, the OAMD 214 includes attributes for the collection of static audio objects 222, such as the gain and position for each static audio object 222. In other words, the OAMD 214 represents a predefined configuration of the beds 222.
[0067] The audio bitstream 110 further includes downmix coefficients 216. Depending on the configuration of the set of output channels 118, the controller selects the corresponding downmix coefficients 216 to be utilized when calculating the second gain matrix 220. By way of example, the set of output audio channels may be stereo output channels; 5.1 surround audio output channels; 5.1.2 immersive audio output channels (immersive audio output configuration); 5.1.4 immersive audio output channels; 7.1 surround audio output channels; or 9.1 surround audio output channels. Thus, the resulting gain matrix has a size of Ch (the number of output channels) × M. The selected downmix coefficients may be used as is when calculating the second gain matrix 220. However, as will be further described below in connection with FIG. 6, the selected downmix coefficients may need to be modified to compensate for attenuations performed at the encoder side when downmixing the original audio signal to achieve the N dynamic audio objects 210. Furthermore, in some embodiments, the selection process of which of the received downmix coefficients 216 should be used to calculate the second gain matrix 220 can be based on the downmix operation performed on the encoder side in addition to the configuration of the set of output channels 118, as will be further described below in connection with FIG.
[0068] The second gain matrix is used in the rendering stage 204 of the decoder 100 to render the set of static audio objects 222 into the set of output audio channels 118 .
[0069] Note that the LFE is not shown in Figure 2. In this context, the LFE should be transmitted directly to the final rendering stage 204 to be included in (or mixed into) the set of output audio channels 118.
[0070] In Figure 3, a second embodiment 109b of the second decoding mode 109 is shown. Similar to the embodiment shown in Figure 2, this embodiment shows a low-rate transmission (low-bitrate audio bitstream) decoded in core decoding mode. The difference in Figure 3 is that the received audio bitstream 110 carries further audio objects 302 in addition to the N dynamic audio objects 210 that are mapped to static audio objects 222. Such additional audio objects may include discrete congruent (A-JOC) dynamic audio objects and / or static audio objects (bed objects) or ISFs. For example, the additional audio objects 302 may include: LFE (zero to many) Other bed objects Other dynamic objects ·ISF.
[0071] Thus, in some embodiments, the dynamic audio objects included in the received audio bitstream are more than N dynamic audio objects 210. For example, the dynamic audio objects included in the received audio bitstream include N dynamic audio objects and K additional dynamic audio objects. According to some embodiments, the received audio bitstream includes M&E+D. In that case, if separate dialogue is added when rendering the set of output channels 118, this may cause problems in low-rate cases where only 10 audio objects may be included in the received audio bitstream 110. If the set of output channels 118 is in a 5.1.2 configuration and bed objects are used (i.e., a legacy solution), eight bed objects would need to be transmitted. This leaves only two possible audio objects representing dialogue, which may be too few if, for example, five different dialogue objects are to be supported. Using the present invention, immersive output audio can be achieved in this case by transmitting, for example, four (N) dynamic audio objects for M&E that are mapped 202 to a set of static audio objects 222, one additional static object 302 for LFE, and five (K) additional dynamic objects for dialogue.
[0072] In the embodiment of FIG. 3, N dynamic audio objects 210 are pre-rendered into M static audio objects 222 as described above in connection with FIG.
[0073] A set of OAMDs 214 is used for rendering 204. The received audio bitstream includes six OAMDs 214, in this example, one for each additional audio object 302. These six OAMDs are then included in the audio bitstream at the encoder side and used in the decoder 100 for the decoding process described herein. Additionally, as described above in connection with FIG. 2, the decoder generates a further set of OAMDs 214 that define positions and gains for the intermediate rendered beds 222. In this example, there are a total of 13 OAMDs 214. The OAMDs 214 include attributes for the set of static audio objects 222, e.g., the gain (i.e., 1) and position for each static audio object 222, and attributes for the additional audio objects 302, e.g., the gain and position for each additional audio object 302.
[0074] The audio bitstream 110 further includes downmix coefficients 216, which are utilized to render a set of output channels 118 similar to those described above in connection with FIG. 2 and below in connection with FIG. 6.
[0075] The second gain matrix 220 is used in the rendering stage 204 of the decoder 100 to render the set of static audio objects 222 and the set of further audio objects 302 (which may include dynamic audio objects and / or static audio objects and / or ISF objects as defined above) into the set of output audio channels 118.
[0076] In the case described in FIG. 3, the controller needs to know which received dynamic audio objects should be mapped to the collection of static audio objects 222 and which should be passed directly to the final rendering stage 204. This can be achieved in several different ways. For example, each received audio object may include a flag value that informs the controller whether the audio object is to be mapped (pre-rendered). In another example, the received audio bitstream includes metadata that identifies the dynamic audio object(s) to be mapped. It should be noted that in the AC-4 context, the subset to be sent to the pre-renderer 202 needs to be found, e.g., using flag values or metadata as described above, only if the additional dynamic objects are part of the same A-JOC substream as the N dynamic audio objects.
[0077] In one embodiment, the metadata indicates that N of the one or more dynamic audio objects should be mapped to a set of static audio objects, thereby informing the controller that these N dynamic audio objects should be selected from predefined positions or positions within the received audio bitstream. The mapped dynamic audio objects 210 may be, for example, the first or last N audio objects within the audio bitstream 110. The number of mapped audio objects may be indicated by the flag values Num_bed_obj_ajoc (also referred to as num_obj_with_bed_render_info) and / or n_fullband_dmx_signals in the AC-4 standard (published in document ETSI TS103 190-2 V1.2.1(2018-02)). Other standards may use other names for the flag values. It should also be noted that the flag values may be renamed for newer versions of the aforementioned AC-4 standard. According to some embodiments, if num_bed_obj_ajoc is greater than zero, this means that num_bed_obj_ajoc dynamic objects are mapped to a set of static audio objects. According to some embodiments, if num_bed_obj_ajoc is not present and n_fullband_dmx_signals is less than 6, this means that all dynamic objects are mapped to a set of static audio objects.
[0078] In some embodiments, the dynamic audio object is received before any static audio objects in the received bitstream 110. In other embodiments, the LFE is received first in the bitstream 110, before the dynamic audio object and any further static audio objects.
[0079] FIG. 4 illustrates, by way of example, a third embodiment 109c of the second decoding mode 109. The dual rendering stages 202, 204 of the embodiments of FIGS. 2-3 may be considered inefficient in some cases due to computational complexity. As a result, in some embodiments, the two gain matrices 218, 220 are combined into a single matrix 404 before rendering 204 the audio objects 210, 302 of the received audio bitstream 110 to the set of output channels 118. In this embodiment, a single rendering stage 204 is used. The setup of FIG. 4 is applicable both to the case described in FIG. 2, i.e., when the received audio bitstream 110 contains only the dynamic object 210 that is mapped to the set of static audio objects 222, and to the case described in FIG. 3, i.e., when the received audio bitstream 110 further contains additional audio objects 302. It should be noted that in the case of FIG. 3, in case matrix multiplication according to FIG. 4 is to be used, matrix 218 needs to be augmented with additional columns and / or rows to handle the "through" of additional objects 302.
[0080] FIG. 5 illustrates, by way of example, an encoder 500 for encoding an audio bitstream 110 to be decoded according to any of the above-described embodiments. In general terms, the encoder 500 includes components corresponding to the content of the audio bitstream 110 to achieve such a bitstream 110, as will be understood by readers of this disclosure. Typically, the encoder 500 includes a receiving component (not shown) configured to receive a collection of audio objects (dynamic and / or static). The encoder 500 further includes a downmix component 502 configured to downmix the collection of audio objects 508 into one or more downmixed dynamic audio objects 510, where at least one of the downmixed dynamic audio objects 510 is intended to be mapped to a collection of static audio objects corresponding to a predefined speaker configuration at the decoder side in at least one of a plurality of decoding modes. The downmix component 502 may attenuate some of the audio objects, as will be described below in connection with FIG. 6. In this case, the attenuation performed needs to be compensated for on the decoder side. As a result, information about the attenuation performed and / or the configuration of the audio object 508 is included in some embodiments in the bitstream 110. In other embodiments, the decoder is pre-configured with all or part of this information, and as a result, such information may be omitted from the bitstream 110. In other words, in some embodiments, the bitstream multiplexer 506 is further configured to multiplex information about the channel configuration of the audio object 508 received by the receiving component into said audio bitstream. The original channel configuration (format of the original audio signal) may be any suitable configuration, such as 7.1.4, 5.1.4, etc.In some embodiments, the encoder (e.g., downmix component 502) is further adapted to determine, when downmixing the collection of audio objects 508 into one or more downmixed dynamic audio objects 510, information regarding the attenuation to be applied to at least one of the one or more dynamic audio objects 510. This information (not shown in FIG. 5) is then transmitted to a bitstream multiplexer 506 configured to multiplex information regarding the attenuation into the audio bitstream 110.
[0081] The encoder 500 further includes a downmix coefficient providing component 504 configured to determine a first set of downmix coefficients 516 to be utilized for rendering a set of static audio objects corresponding to a predefined speaker configuration into a set of output audio channels at the decoder side. As will be described below in connection with Figure 6, depending on, for example, the downmix operation performed by the downmix component (attenuation and / or what type of downmixing was performed, from what configuration to what configuration), the decoder may need to perform a further selection process and / or adjustment among the first set of downmix coefficients 516 before actually using the resulting downmix coefficients for rendering.
[0082] The encoder further includes a bitstream multiplexer 506 configured to multiplex the at least one downmixed dynamic audio object 510 and the first set of downmix coefficients 516 into an audio bitstream 110 .
[0083] In some embodiments, the downmix component 502 also provides metadata 514 that identifies the at least one downmixed audio object 510 of the one or more downmixed dynamic audio objects to the bitstream multiplexer 506. In this case, the bitstream multiplexer 506 is further configured to multiplex the metadata 514 into the audio bitstream 110.
[0084] In some embodiments, the downmix component 502 receives a target bitrate 509 to determine the details of the downmix operation, for example, how many downmixed audio objects should be calculated from the set of dynamic audio objects 508. In other words, the target bitrate can determine the clustering parameters for the downmix operation.
[0085] As will be appreciated, if the one or more downmixed dynamic audio objects 510 include more dynamic audio objects than those intended to be mapped to the set of static audio objects at the decoder, downmix coefficients also need to be calculated for them. Furthermore, static audio objects (e.g., LFE, etc.) may be sent by the bitstream multiplexer 506 for inclusion in the audio bitstream 110 along with their corresponding downmix coefficients. Furthermore, each audio object included in the audio bitstream 110 has an associated OAMD, e.g., OAMDs 512 associated with all dynamic audio objects 510 intended to be mapped to the set of static audio objects at the decoder, which are multiplexed into the audio bitstream 110.
[0086] FIG. 6 shows, by way of example, further details of how the second gain matrix 220 of FIGS. 2-4 may be determined using the gain matrix calculation unit 208. As described above, the gain matrix calculation unit 208 receives the downmix coefficients 216 from the bitstream. In this embodiment, the gain matrix calculation unit 208 also receives data 612 regarding the type of downmix of the audio signal performed on the encoder side. Thus, the data 612 includes information about the downmix operation performed on the encoder side that resulted in the N dynamic audio objects 210. The data 612 may define / indicate the original channel configuration of the audio signal being downmixed into the N dynamic audio objects 210. Based on the received data 612 and the received downmix coefficients 216, the downmix coefficient (DC) selection and modification unit 606 determines downmix coefficients 608, which are then used in the gain matrix calculation unit 610 to form the second gain matrix 220 using the above-mentioned OAMD 214 and output channel 118 configuration, e.g., 5.1. Thus, the gain matrix calculation unit 610 selects those coefficients from the downmix coefficients 608 that are suitable for the requested configuration of output channels 118 and determines the second gain matrix 220 to be used for this particular audio rendering setup. In some embodiments, the DC selection and modification unit 606 may directly select the set of downmix coefficients 608 from the received downmix coefficients 216. In other embodiments, the DC selection and modification unit 606 may need to first select the downmix coefficients and then modify them to derive the downmix coefficients 608 that are used in the gain matrix calculation unit 610 to calculate the second gain matrix 220.
[0087] The functionality of the DC selection and modification unit 606 will now be illustrated for a specific setup of encoded and decoded audio.
[0088] In some embodiments, the encoder applies attenuation to / on some of the transmitted audio objects 210. Such attenuation is the result of a downmix process within the encoder of an original audio signal to a downmix audio signal. For example, if the original audio signal is in the format 7.1.4 (L, R, C, LFE, Ls, Rs, Lb, Rb, Tfl, Tfr, Tbl, Tbr) and is downmixed in the encoder to a 5.1.2 (Ld, Rd, Cd, LFE, Lsd, Rsd, Tld, Trd) format, the Lsd signal is downmixed in the encoder to: N dB(Ls+Lb) and the Tld signal is determined in the encoder as: M dB(Tfl+Tbl) is determined as follows.
[0089] Typically, N=M=3, although other attenuation levels may be applied.
[0090] In this setup, an attenuation of 3 dB has thus already been applied in Lsd and Tld. In these examples, only the left channel is described, but the right channel is treated correspondingly.
[0091] It should be noted that to further reduce the bitrate, the downmix (e.g., 5.1.2 channel audio) is then further reduced in the encoder to, e.g., five dynamic audio objects (210 in Figures 2 and 3).
[0092] In this case, the associated downmix coefficients 216 transmitted in the bitstream are: gain_tfb_to_tm: Gain from top front and / or top back to top center gain_t2a, gain_t2b: Gain of the upper front channels to the front and surround channels, respectively Typical / Default: gain_t2a is mapped to -Inf dB and gain_t2b is mapped to -3dB, which means downmix to surround channels at -3dB. gain_t2d, gain_t2e: Gain of the upper rear channels to the front or surround channels Typical / Default: gain_t2d is mapped to -Inf dB, gain_t2e is mapped to -3dB, which means downmix to surround channels at -3dB. gain_b4_to_b2: Rear and surround channels to surround channels Typical / Default: Mapped to -3dB.
[0093] However, if the above downmix coefficients are applied directly when the audio format of the output channels 118 is 5.1, the upper channels Tfl and Tbl will be attenuated by 6 dB in the surround output, i.e., M=3 dB already applied in the encoder and 3 dB of the gain_t2b downmix coefficient received in the bitstream. The same applies to the lower channels Ls and Lb, which will also be attenuated by 6 dB in the surround output, i.e., N=3 dB already applied in the encoder and 3 dB of the gain_b4_to_b2 downmix coefficient received in the bitstream. In order to compensate for the attenuation already made on the encoder side, the DC selection and modification unit 606 is configured in this case to determine the downmix coefficients 608 such that the output channels are rendered as follows: L out =L d +(+M dB+gain_t2a)Tl d =L+gain_t2a(Tfl+Tbl) Ls out =(+N dB+gain_b4_to_b2)Ls d +(+M dB+gain_t2b)Tl d=gain_b4_to_b2(Ls+Lb)+gain_t2b(Tfl+Tbl)
[0094] In this embodiment, the decoder selects the gains for the upper front channels, gain_t2a and gain_t2b, to the front and surround channels, respectively. These are therefore preferable to the gains for the upper rear channels, gain_t2d and gain_t2e. It should also be noted that the above equations are intended to convey the idea of compensation in the decoder for the attenuation made by the encoder; in practice, the equations that achieve this would be designed to ensure that, for example, the conversion from gain / attenuation in the log-dB domain to linear gain is handled correctly.
[0095] To achieve this, the decoder needs to know the attenuation applied by the encoder. In some embodiments, the values of N(dB) and M(dB) are indicated in the bitstream as additional metadata 602. The additional metadata 602 thus defines information about the attenuation applied to at least one of the one or more dynamic audio objects at the encoder side. In other embodiments, the decoder is pre-configured (in memory 604) with the attenuation 603 to be applied at the encoder side. For example, in the case of a downmix from 7.1.4 (or 5.1.4) to 5.1.2 at the encoder, the decoder may know that an attenuation of 3 dB is always performed. In some embodiments, the decoder has received information 602, 603 about the attenuation applied to at least one of the one or more dynamic audio objects at the encoder side. This information 602, 603, in conjunction with received data 612 indicating what type of downmix was performed at the encoder, may be used to select and / or adjust the downmix coefficients 216 in the DC selection and modification unit 606. The selected and / or adjusted coefficients 608 are used by the gain matrix calculation unit 610 in conjunction with the OAMD 214 and the configuration of the output audio signal 118 to form the second gain matrix 220, as described above.
[0096] In another exemplary setup, the original audio signal at the encoder is 5.1.2 with top front channels (L, R, C, LFE, Ls, Rs, Tfl, Tfr), which is downmixed to a 5.1.2 format with top center channels (Ld, Rd, Cd, LFE, Lsd, Rsd, Tld, Trd) instead. In this embodiment, no attenuation is performed at the encoder. However, in this case, the DC selection and modification unit 606 needs to know what the original signal configuration was at the encoder side in order to select appropriate downmix coefficients for the 5.1 output signal 118. In this case, the relevant downmix coefficients 216 transmitted in the bitstream are: gain_t2a, gain_t2b, which are gains for the top front channels, front and surround channels, respectively. The DC selection and modification unit 606 is configured in this case to determine the downmix coefficients 608 such that the output channels 118 are rendered as follows: L out =L d +gain_t2a(Tld)=L+gain_t2a(Tfl) Ls out =Ls d +gain_t2b(Tld)=Ls+gain_t2b(Tfl)
[0097] Further embodiments of the present disclosure will be apparent to those skilled in the art after reviewing the above description. While the description and drawings disclose embodiments and examples, the present disclosure is not limited to such specific examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure, which is defined by the appended claims. Any reference signs appearing in the claims should not be construed as limiting the scope thereof.
[0098] Moreover, variations to the disclosed embodiments can be understood and implemented by those skilled in the art in practicing the present disclosure, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprises" does not exclude other elements or steps, and the word "a" or "an" does not exclude a plurality. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
[0099] The systems and methods disclosed above may be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks among functional units referred to in the above description does not necessarily correspond to a division into physical units. Conversely, a single physical component may have multiple functions, and a single task may be performed by several cooperating physical components. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or application-specific integrated circuits. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, those skilled in the art will appreciate that communication media typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0100] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). [EEE1] one or more buffers for storing received audio bitstreams; a controller coupled to the one or more buffers, the controller comprising: operating in a decoding mode selected from a plurality of different decoding modes, the plurality of different decoding modes including a first decoding mode and a second decoding mode, wherein only the first decoding mode of the first and second decoding modes allows parametric reconstruction of individual audio objects from a cluster of dynamic audio objects; If the selected decoding mode is the second decoding mode: Accessing said received audio bitstream; determining whether the received audio bitstream includes one or more dynamic audio objects; and in response to determining that the received audio bitstream includes one or more dynamic audio objects, mapping at least one of the one or more dynamic audio objects to a set of static audio objects, the set of static audio objects corresponding to a predefined speaker configuration. Audio decoder. [EEE2] The audio decoder of EEE1, wherein if the selected decoding mode is the second decoding mode, the controller is further configured to render the set of static audio objects into a set of output audio channels. [EEE3] The audio decoder of EEE2, wherein the audio bitstream includes a first set of downmix coefficients, and wherein the controller is configured to utilize the first set of downmix coefficients to render the set of static audio objects into the set of output audio channels. [EEE4] The audio decoder of EEE3, wherein the controller is further configured to receive information regarding an attenuation applied to at least one of the one or more dynamic audio objects at the encoder side, and wherein the controller is configured to modify the first set of downmix coefficients accordingly when using the first set of downmix coefficients to render the set of static audio objects into the set of output audio channels. [EEE5] 5. The audio decoder of claim 3, wherein the controller is further configured to receive information regarding a downmix operation performed at an encoder side, the information defining an original channel configuration of the audio signal, the downmix operation resulting in downmixing the audio signal into the one or more dynamic audio objects, and the controller is configured to select a subset of the first set of downmix coefficients based on the information regarding the downmix information, and wherein utilizing the first set of downmix coefficients for rendering the set of static audio objects into a set of output audio channels comprises utilizing the subset of the first set of downmix coefficients for rendering the set of static audio objects into a set of output audio channels. [EEE6] 6. The audio decoder of any one of claims 8 to 10, wherein the controller is configured to perform the mapping of at least one of the one or more dynamic audio objects and the rendering of the set of static audio objects in a combined calculation using a single matrix. [EEE7] 6. The audio decoder of any one of claims EEE2 to 5, wherein the controller is configured to perform the mapping of at least one of the one or more dynamic audio objects and the rendering of the set of static audio objects in separate calculations using respective matrices. [EEE8] 8. An audio decoder according to any one of claims 8 to 7, wherein the received audio bitstream includes metadata identifying at least one of the one or more dynamic audio objects. [EEE9] the metadata indicates that N of the one or more dynamic audio objects should be mapped to the set of static audio objects; In response to the metadata, the controller is configured to map N of the one or more dynamic audio objects selected from predefined position(s) within the received audio bitstream to the set of static audio objects. Audio decoder according to EEE8. [EEE10] 8. The audio decoder according to claim 6, wherein the one or more dynamic audio objects included in the received audio bitstream include more than N dynamic audio objects. [EEE11] 8. The audio decoder of claim 6, wherein the one or more dynamic audio objects included in the received audio bitstream include the N dynamic audio objects and K further dynamic audio objects, and the controller is configured to render the set of static audio objects and the K further audio objects onto a set of output audio channels. [EEE12] 12. The audio decoder of any one of claims 8 to 11, wherein in response to the metadata, the controller is configured to map a first N of the one or more dynamic audio objects in the received audio bitstream to the set of static audio objects. [EEE13] 13. An audio decoder according to any one of claims EEE9 to 12, wherein the set of static audio objects consists of M static audio objects, where M>N>0. [EEE14] 14. An audio decoder according to any one of claims 1 to 13, wherein the received audio bitstream further comprises one or more further static audio objects. [EEE15] 5.1 surround sound audio output channels; 5.1.2 immersive audio output channels; or 5.1.4 immersive audio output channels. [EEE16] 16. The audio decoder of any one of claims 1 to 15, wherein the predefined speaker configuration is a 5.0.2 speaker configuration. [EEE17] 1. A method in a decoder comprising: receiving an audio bitstream and storing the received audio bitstream in one or more buffers; selecting a decoding mode from a plurality of different decoding modes, the plurality of different decoding modes including a first decoding mode and a second decoding mode, wherein only the first decoding mode of the first and second decoding modes allows parametric reconstruction of individual dynamic audio objects from a cluster of dynamic audio objects; operating a controller coupled to the one or more buffers in a selected decoding mode; If the selected decoding mode is the second decoding mode, the method further comprises: accessing, by the controller, the received audio bitstream; determining, by the controller, whether the received audio bitstream includes one or more dynamic audio objects; and in response to determining that the received audio bitstream includes one or more dynamic audio objects, mapping, by the controller, at least one of the one or more dynamic audio objects to a set of static audio objects corresponding to a predefined speaker configuration. method. [EEE18] 1. An audio encoder, comprising: a receiving component configured to receive a set of audio objects; a downmix component configured to downmix the set of audio objects into one or more downmixed dynamic audio objects, at least one of the one or more downmixed dynamic audio objects being intended to be mapped to a set of static audio objects in at least one of a plurality of decoding modes on a decoder side, the set of static audio objects corresponding to a predefined speaker configuration; a downmix coefficient providing component configured to determine a first set of downmix coefficients to be utilized for rendering the set of static audio objects corresponding to the predefined speaker configuration into a set of output audio channels at a decoder; a bitstream multiplexer configured to multiplex the at least one downmixed dynamic audio object and the first set of downmix coefficients into an audio bitstream. Encoder. [EEE19] the downmix component is further configured to provide metadata to the bitstream multiplexer identifying the at least one of the one or more downmixed dynamic audio objects; the bitstream multiplexer is further configured to multiplex the metadata into the audio bitstream. Encoder according to EEE18. [EEE20] the encoder is further adapted to determine, when downmixing the set of audio objects into one or more downmixed dynamic audio objects, information regarding an attenuation to be applied to at least one of the one or more dynamic audio objects; the bitstream multiplexer is further configured to multiplex the information regarding attenuation into the audio bitstream. 1. An encoder as defined in EEE18 or 19. [EEE21] 21. The encoder of any one of EEE18 to 20, wherein the bitstream multiplexer is further configured to multiplex information about a channel configuration of the audio objects received by the receiving component into the audio bitstream. [EEE22] 1. A method in an encoder comprising: receiving a set of audio objects; downmixing the set of audio objects into one or more downmixed dynamic audio objects, at least one of the one or more downmixed dynamic audio objects being intended to be mapped to a set of static audio objects in at least one of a plurality of decoding modes on a decoder side, the set of static audio objects corresponding to a predefined speaker configuration; determining a first set of downmix coefficients to be used for rendering the set of static audio objects corresponding to the predefined speaker configuration into a set of output audio channels at a decoder side; and multiplexing the at least one downmixed dynamic audio object and the first set of downmix coefficients into an audio bitstream. method. [EEE23] A computer program product comprising a computer readable medium having instructions adapted to perform the method of any one of claims EEE17 to EEE22 when executed by a device having processing capability.
Claims
1. one or more buffers for storing the received audio bitstreams; a controller coupled to the one or more buffers, the controller comprising: operating in a decoding mode selected from a plurality of different decoding modes for decoding the received audio bitstream into one or more dynamic or static audio objects, the dynamic or static audio objects comprising audio signals associated with time-varying or static spatial locations, the plurality of different decoding modes comprising a first decoding mode and a second decoding mode, wherein only the first decoding mode of the first and second decoding modes allows full decoding of one or more encoded dynamic audio objects in the bitstream into reconstructed individual audio objects; If the selected decoding mode is the first decoding mode: and in response to determining that the received audio bitstream includes one or more dynamic audio objects, mapping at least one of the one or more dynamic audio objects to a set of static audio objects; Audio decoder.
2. 2. The audio decoder of claim 1, wherein if the selected decoding mode is the first decoding mode, the controller is further configured to render the set of static audio objects to a set of output audio channels.
3. 3. The audio decoder of claim 2, wherein the audio bitstream includes a first set of downmix coefficients, and the controller is configured to utilize the first set of downmix coefficients to render the set of static audio objects into the set of output audio channels.
4. 4. An audio decoder according to claim 1, wherein the received audio bitstream includes metadata identifying the at least one of the one or more dynamic audio objects.
5. 5. An audio decoder according to claim 1, wherein the received audio bitstream further comprises one or more further static audio objects.
6. 6. An audio decoder according to any one of claims 1 to 5, with reference to claim 2, wherein the set of output audio channels is one of: stereo output channels; 5.1 surround sound audio output channels; 5.1.2 immersive audio output channels; or 5.1.4 immersive audio output channels.
7. 7. An audio decoder according to claim 1, wherein the set of static audio objects corresponds to a predefined immersive speaker configuration, the predefined immersive speaker configuration being a 5.0.2 speaker configuration.
8. 1. A method in a decoder comprising: receiving an audio bitstream; selecting a decoding mode from a plurality of different decoding modes for decoding the received audio bitstream into one or more dynamic or static audio objects, the dynamic or static audio objects comprising audio signals associated with time-varying or static spatial locations, the plurality of different decoding modes comprising a first decoding mode and a second decoding mode, wherein only the first decoding mode of the first and second decoding modes allows full decoding of one or more encoded dynamic audio objects in the bitstream into reconstructed individual audio objects; operating a controller coupled to the one or more buffers in a selected decoding mode; If the selected decoding mode is the first decoding mode, the method further comprises: and in response to determining that the received audio bitstream includes one or more dynamic audio objects, mapping, by the controller, at least one of the one or more dynamic audio objects to a set of static audio objects. method.