Program loudness of presentation base having no correlation to transmission
By incorporating loudness data and DRC for each presentation, the solution ensures accurate and flexible loudness consistency across different audio content sub-streams, addressing inaccuracies in existing technologies and maintaining regulatory compliance.
Patent Information
- Application Number
- JP2025072821
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-10-10
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2035-10-06
AI Technical Summary
Existing audio encoding and decoding technologies struggle to achieve accurate loudness consistency across different audio content sub-streams, leading to potential inaccuracies greater than the specified tolerances, especially when switching languages or adding commentary tracks.
The proposed solution involves providing loudness data for each presentation data structure, using mixing coefficients and dynamic range compression (DRC) data to ensure consistent loudness levels, and allowing for flexible content selection while maintaining compliance with regulatory standards.
This approach achieves consistent loudness across various audio presentations, reducing computational load and ensuring compliance with loudness regulations, even when user inputs deviate from default settings.
Smart Images

Figure 2025107210000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 062,479, filed Oct. 10, 2014. The content of the application is incorporated herein by reference in its entirety.
[0002] Technical Field The present invention relates to audio signal processing, and more particularly, to encoding and decoding of audio data bitstreams to achieve a desired loudness level of an output audio signal.
Background Art
[0003] Dolby AC - 4 is an audio format for efficiently distributing rich media content. AC - 4 provides a flexible framework for broadcasters and content producers to distribute and encode content in an efficient manner. The content can be distributed through several sub - streams. For example, one sub - stream contains M&E (music and effects), and a second sub - stream contains dialog. For some audio content, it may be advantageous to be able to switch, for example, the language of the dialog from one language to another, or to be able to add additional sub - streams, such as a commentary sub - stream to the content or an explanation for visually impaired persons.
[0004] To ensure proper leveling of the content presented to the consumer, the loudness of the content needs to be known with a certain degree of accuracy. Current loudness requirements have tolerances of 2 dB (ATSC A / 85) and 0.5 dB (EBU R128), while some specifications have tolerances as low as around 0.1 dB. That is, the loudness of an output audio signal with a commentary track and dialog in a first language should be substantially the same as that of an output audio signal without a commentary track and with dialog in a second language.
Brief Description of the Drawings
[0005] Here, exemplary embodiments will be described with reference to the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0006] In view of the above, an object is to provide an encoder, a decoder, and an associated method, which aim to provide a desired loudness level for an output audio signal regardless of what content sub-streams are mixed in the output audio signal.
[0007] 〈I. Overview - Decoder〉 According to a first aspect, an exemplary embodiment proposes a decoding method, a decoder, and a computer program product for decoding. The proposed method, decoder, and computer program product may generally have the same features and advantages.
[0008] According to an exemplary embodiment, a method for processing a bitstream including a plurality of content sub-streams each representing an audio signal is provided. The method includes: extracting, from the bitstream, one or more presentation data structures, each presentation data structure including a reference to at least one of the content sub-streams, and each presentation data structure further including a reference to a metadata sub-stream representing loudness data describing a combination of one or more of the referenced content sub-streams; receiving data indicating a selected presentation data structure among the one or more presentation data structures and a desired loudness level; decoding one or more content sub-streams referenced by the selected presentation data structure; forming an output audio signal based on the decoded content sub-streams, and the method further includes processing the decoded one or more content sub-streams or the output audio signal based on the loudness data referenced by the selected presentation data structure to achieve the desired loudness level.
[0009] Data indicating the selected presentation data structure and the desired loudness level are typically user settings available in the decoder. The user may, for example, use a remote control to select a presentation data structure where the dialog is in French and / or increase or decrease the desired output loudness level. In many embodiments, the output loudness level is related to the capacity of the playback device. According to some embodiments, the output loudness level is controlled by volume. As a result, data indicating the selected presentation data structure and the desired loudness level are typically not included in the bitstream received by the decoder.
[0010] In the usage in this document, "loudness" represents a modeled psychoacoustic measurement of the intensity of sound. In other words, loudness represents an approximation of the volume of the sound(s) perceived by an average user.
[0011] In the usage in this document, "loudness data" refers to data resulting from the measurement of the loudness level of a particular presentation data structure by a function that models psychoacoustic loudness perception. In other words, it is a collection of values indicating the loudness attributes of a combination of one or more content sub-streams being referenced. According to embodiments, the average loudness level of the combination of the one or more content sub-streams referenced by a particular presentation data structure can be measured. For example, the loudness data may refer to the dialnorm value (based on ITU-R BS.1770) of the one or more content sub-streams referenced by a particular presentation data structure. Other suitable loudness measurement standards may be used, such as the Glasberg and Moore loudness models that provide modifications and extensions to the Zwicker loudness model.
[0012] In the usage in this document, "presentation data structure" refers to metadata related to the content of the output audio signal. The output audio signal is also referred to as a "program". The presentation data structure is also referred to as a "presentation".
[0013] Audio content can be distributed through several sub-streams. In the usage in this document, "content sub-stream" refers to such a sub-stream. For example, a content sub-stream may include music of the audio content, dialog of the audio content, or a commentary track to be included in the output audio signal. The content sub-stream may be channel-based or object-based. In the latter case, time-dependent spatial position data is included in the content sub-stream. The content sub-stream may be included in a bitstream or may be part of an audio signal (i.e., as a channel group or an object group).
[0014] In the usage in this document, "output audio signal" refers to the actually output audio signal, which is rendered to the user.
[0015] The inventor has recognized that by providing loudness data, such as a dialnorm value, for each presentation, specific loudness data indicating exactly how much loudness is for at least one content sub-stream referred to when decoding that specific presentation becomes available to the decoder.
[0016] In the prior art, loudness data may be provided for each content sub-stream. The problem with providing loudness data for each content sub-stream is that in that case, it is left to the decoder to combine the various loudness data into a presentation loudness. It may not be accurate to arrive at a loudness value for a presentation by adding up the individual loudness data values of the sub-streams representing the average loudnesses of the sub-streams, and in many cases, it does not result in the actual average loudness value of the combined sub-streams. Adding up the loudness data for each of the referenced content sub-streams may be mathematically impossible due to signal attributes, loudness algorithms, and the nature of loudness perception, which is typically not additive, and can lead to potential inaccuracies greater than the above tolerances.
[0017] Using this embodiment, the difference between the average loudness level of the selected presentation provided by the loudness data for the selected presentation and the desired loudness level can thus be used to control the playback gain of the output audio signal.
[0018] By providing and using loudness data as described above, a consistent loudness, i.e., a loudness close to the desired loudness level, can be achieved between various presentations. Furthermore, consistent loudness can be achieved between different programs on a television channel, for example, between a television program and its commercial, or across television channels.
[0019] According to an exemplary embodiment, the selected presentation data structure references two or more content sub-streams and further references at least two mixing coefficients to be applied thereto, and the forming of the output signal further includes additively mixing one or more decoded content sub-streams by applying the mixing coefficient(s).
[0020] By providing at least two mixing coefficients, an increased flexibility of the content of the output audio signal is achieved.
[0021] For example, the selected presentation data structure may refer to one mixing coefficient to be applied to each sub-stream of the two or more content sub-streams. According to this embodiment, the relative loudness levels between the content sub-streams can be changed. For example, cultural preferences may require different balances between different content sub-streams. Consider a situation where the Spanish-speaking region does not desire as much attention to music. Thus, the music sub-stream can be attenuated by 3 dB. According to other embodiments, a signal mixing coefficient may be applied to a subset of the two or more content sub-streams.
[0022] According to an exemplary embodiment, the bitstream includes a plurality of time frames, and the mixing coefficients referenced by the selected presentation data structure can be assigned independently for each time frame. The effect of providing time-varying mixing coefficients is that ducking can be achieved. For example, the loudness level over a certain time segment of a content sub-stream may be reduced by the increased loudness in the same time segment of another content sub-stream.
[0023] According to an exemplary embodiment, the loudness data represents the value of the loudness function regarding the application of gating to the audio input signal.
[0024] The audio input signal is a signal to which a loudness function (e.g., the dialnorm function) has been applied on the encoder side. Then, the resulting loudness data is transmitted to the decoder in the bitstream. A noise gate (also referred to as a mute gate) is an electronic device or software used to control the volume of an audio signal. Gating is the use of such a gate. A noise gate attenuates signals that indicate values below a threshold. The noise gate may attenuate the signal by a fixed amount known as a range. In its simplest form, the noise gate allows a signal to pass only when it is above a set threshold.
[0025] Gating may also be based on the presence of dialogue in the audio input signal. As a result, according to an exemplary embodiment, the loudness data represents values related to a time segment of the dialogue of the audio input signal for the loudness function. According to other embodiments, gating is based on a minimum loudness level. Such a minimum loudness level may be an absolute threshold or a relative threshold. The relative threshold may be based on the loudness level measured using the absolute threshold.
[0026] According to an exemplary embodiment, the presentation data structure further includes a reference to dynamic range compression (DRC) data for one or more content substreams being referenced, and the method further includes processing the decoded one or more content substreams or the output audio signal based on the DRC data. Here, the processing includes applying one or more DRC gains to the decoded one or more content substreams or the output audio signal.
[0027] Dynamic range compression reduces the volume of loud sounds or amplifies quiet sounds, thereby narrowing or "compressing" the dynamic range of an audio signal. By providing DRC data uniquely for each presentation, an improved user experience of the output audio signal can be achieved, regardless of the presentation chosen. Further, by providing DRC data for each presentation, a consistent user experience of the audio output signal can be achieved across each of a plurality of presentations, between programs as described above, and across television channels.
[0028] The DRC gain always varies with time. In each time segment, the DRC gain may be a single gain for the audio output signal or multiple DRC gains that vary for each sub-stream. The DRC gain may be applied to groups of channels and / or may be frequency-dependent. Additionally, the DRC gain included in the DRC data may represent the DRC gain for two or more DRC time segments. For example, a sub-frame of a time frame defined by an encoder.
[0029] According to an exemplary embodiment, the DRC data includes at least one set of the one or more DRC gains. Thus, the DRC data may include a plurality of DRC profiles corresponding to DRC modes, each providing a different user experience of the audio output signal. By including the DRC gain directly in the DRC data, a reduced computational load on the decoder can be achieved.
[0030] According to an exemplary embodiment, the DRC data includes at least one compression curve, and the one or more DRC gains are obtained by calculating one or more loudness values of the one or more content sub-streams or the audio output signal using a predefined loudness function and mapping the one or more loudness values to DRC gains using the compression curve. By providing compression curves in the DRC data and calculating DRC gains based on those curves, the bit rate required to transmit the DRC data to an encoder can be reduced. The predefined loudness function may be taken from, for example, the ITU-R BS.1770 recommendation document, but any suitable loudness function may be used.
[0031] According to an exemplary embodiment, the mapping of the loudness values includes a smoothing operation of the DRC gain. The effect of this can be a more perceptible output audio signal. The time constant for smoothing the DRC gain may be transmitted as part of the DRC data. Such a time constant may vary depending on the signal attributes. For example, in some embodiments, the time constant may be smaller when the loudness value is greater than the corresponding previous loudness value compared to when the loudness value is smaller than the corresponding previous loudness value.
[0032] According to an exemplary embodiment, the referenced DRC data is included in a metadata sub-stream. This can reduce the complexity of decoding the bitstream.
[0033] According to an exemplary embodiment, each of the decoded one or more content sub-streams includes sub-stream level loudness data that describes the loudness level of that content sub-stream, and the processing of the decoded one or more content sub-streams or the output audio signal further includes ensuring that loudness consistency is provided based on the loudness level of the content sub-stream.
[0034] In the usage in this document, "loudness consistency" means that the loudness is consistent among presentations with different loudness, that is, it is consistent across a plurality of output audio signals formed based on different content sub-streams. Further, this term means that the loudness is consistent among programs with different loudness, that is, among completely different output audio signals such as the audio signal of a TV program and the audio signal of a commercial. Further, this term means that the loudness is consistent across different TV channels.
[0035] Providing loudness data that describes the loudness level of a content sub-stream may, in some cases, help a decoder provide loudness consistency. For example, the formation of the output audio signal includes combining two or more decoded content sub-streams using alternative mixing coefficients, and the sub-stream level loudness data is used to compensate the loudness data to provide loudness consistency. These alternative mixing coefficients may be derived from user input, for example, when a user decides to deviate from the default presentation (such as with dialog enhancement, dialog attenuation, scene personalization, etc.). This may jeopardize loudness compliance. The influence by the user may cause the loudness of the audio output signal to deviate from the compliance regulation. To assist loudness consistency in such a case, this embodiment provides an option to transmit sub-stream level loudness data.
[0036] According to some embodiments, a reference to at least one of the content sub-streams is a reference to at least one content sub-stream group consisting of one or more of the content sub-streams. Since multiple presentations can share a content sub-stream group (e.g., a sub-stream group consisting of a content sub-stream related to music and a content sub-stream related to effects), this can reduce the complexity of the decoder. This can also reduce the required bit rate for transmitting the bitstream.
[0037] According to some embodiments, a selected presentation data structure refers to a single mixing coefficient applied to each of the one or more of the content sub-streams that make up a content sub-stream group for that content sub-stream group.
[0038] This can be advantageous when the relative ratio of the loudness levels of the content sub-streams in the content sub-stream group is okay, but the overall loudness level of the content sub-streams in that content sub-stream group should be increased or decreased compared to other content sub-stream(s) or content sub-stream group(s) referred to by the selected presentation data structure.
[0039] In some embodiments, the bitstream includes a plurality of time frames, and data indicating the selected presentation data structure among the one or more presentation data structures can be independently assigned for each time frame. As a result, when multiple presentation data structures are received for a program, the selected presentation data structure may be changed during the progress of the program, for example, by the user. As a result, this embodiment provides a more flexible way to select the content of the output audio while at the same time providing loudness consistency of the output audio signal.
[0040] According to some embodiments, the method further comprises: extracting, from the bitstream, for a first one of the plurality of time frames, one or more presentation data structures; and extracting, from the bitstream, for a second one of the plurality of time frames, one or more presentation data structures different from the one or more presentation data structures extracted from the first one of the plurality of time frames, wherein data indicating the selected presentation data structure indicates the selected presentation data structure for the time frame to which it is assigned. As a result, a plurality of presentation data structures may be received in the bitstream, some of those presentation data structures being related to a first set of time frames and some of those presentation data structures being related to a second set of time frames. For example, a commentary track may be available only for a certain time segment of the program. Further, during program progression, currently applicable presentation data structures at a particular point in time may be used to select the selected presentation data structure. As a result, this embodiment provides a more flexible way of selecting the content of the output audio while at the same time providing loudness consistency of the output audio signal.
[0041] According to some embodiments, only the one or more content sub-streams referred to by the selected presentation data structure are decoded from the plurality of content sub-streams included in the bitstream. This embodiment may provide an efficient decoder with reduced computational load.
[0042] According to some embodiments, the bitstream includes two or more distinct bitstreams each including at least one of the plurality of content bitstreams, and the step of decoding the one or more content substreams referenced by the selected presentation data structure comprises: for each particular bitstream of the two or more distinct bitstreams, separately decoding content substream(s) from the referenced content substream(s) included in that particular bitstream. According to this embodiment, each distinct bitstream may be received by a distinct decoder. The decoder decodes the content substream(s) required based on the selected presentation data structure provided in that distinct bitstream. Since the distinct decoders can function in parallel, this can improve the decoding speed. As a result, the decoding performed by the distinct decoders may at least partially overlap. However, it should be noted that it is not essential for the decoding performed by the distinct decoders to overlap.
[0043] Furthermore, by splitting the content substreams among several bitstreams, this embodiment allows the at least two distinct bitstreams to be received through different infrastructures as described below. As a result, this exemplary embodiment provides a more flexible way for a decoder to receive the plurality of content substreams.
[0044] Each decoder may process the decoded substream(s) based on loudness data referenced by the selected presentation data structure, and / or apply a DRC gain, and / or apply a mixing coefficient to the decoded substream(s). Then, the processed or unprocessed content substream(s) may be provided from all of the at least two decoders to a mixing component for forming an output audio signal. Alternatively, the mixing component may perform loudness processing, and / or apply a DRC gain, and / or apply a mixing coefficient. In some embodiments, a first decoder may receive a first bitstream of the two or more separate bitstreams through a first infrastructure (e.g., cable television broadcast), while a second decoder may receive a second bitstream of the two or more separate bitstreams through a second infrastructure (e.g., through the Internet). According to some embodiments, the one or more presentation data structures are present in all of the two or more separate bitstreams. In this case, the presentation definition and loudness data are present in all of the separate decoders. This allows for independent operation of their decoding up to the mixing component. References to substreams that are not present in the corresponding bitstream may be shown as being provided externally.
[0045] According to an exemplary embodiment, a decoder is provided for processing a bitstream including a plurality of content sub-streams each representing an audio signal. The decoder includes: a receiving component configured to receive the bitstream; a demultiplexer configured to extract from the bitstream one or more presentation data structures, each presentation data structure including a reference to at least one of the content sub-streams, and further including a reference to a metadata sub-stream representing loudness data describing a combination of one or more of the content sub-streams being referenced; a playback state component configured to receive data indicating a selected presentation data structure among the one or more presentation data structures and a desired loudness level; and a mixing component configured to decode the one or more content sub-streams referenced by the selected presentation data structure and form an output audio signal based on the decoded content sub-streams, the mixing component further being configured to process the decoded one or more content sub-streams or the output audio signal based on the loudness data referenced by the selected presentation data structure so as to achieve the desired loudness level.
[0046] 〈II. Overview - Encoder〉 According to a second aspect, the exemplary embodiment proposes an encoding method, an encoder, and a computer program product for encoding. The proposed method, encoder, and computer program product may generally have the same features and advantages. Generally, the features of the second aspect may have the same advantages as the corresponding features of the first aspect.
[0047] According to an exemplary embodiment, an audio encoding method is provided. The method includes: receiving a plurality of content sub-streams representing respective audio signals; defining one or more presentation data structures each referring to at least one of the plurality of content sub-streams; for each of the one or more presentation data structures, applying a predefined loudness function to obtain loudness data describing a combination of one or more content sub-streams being referred to, including a reference from the presentation data structure to the loudness data; and forming a bitstream including the plurality of content sub-streams, the one or more presentation data structures, and the loudness data referred to by those presentation data structures.
[0048] As described above, the term "content sub-stream" encompasses sub-streams both within a bitstream and within an audio signal. An audio encoder typically receives audio signals which are then encoded into bitstreams. Those audio signals may be grouped, and each group can be characterized as an individual encoder input audio signal. Each group may then be encoded into a sub-stream.
[0049] According to some embodiments, the method further includes: for each of the one or more presentation data structures, determining dynamic range compression (DRC) data for one or more content sub-streams being referred to, the DRC data quantifying at least one desired compression curve or at least a set of DRC gains; and including the DRC data in the bitstream.
[0050] According to some embodiments, the method further comprises: for each of the plurality of content sub-streams, applying the predefined loudness function to obtain loudness data at the sub-stream level of the content sub-stream; and including the loudness data at the sub-stream level in the bitstream.
[0051] According to some embodiments, the predefined loudness function is related to the application of gating of the audio signal.
[0052] According to some embodiments, the predefined loudness function is related only to time segments of the audio signal that represent dialog.
[0053] According to some embodiments, the predefined loudness function includes at least one of: frequency-dependent weighting of the audio signal, channel-dependent weighting of the audio signal, ignoring segments of the audio signal with signal power below a threshold, and calculation of an energy measure of the audio signal.
[0054] According to an exemplary embodiment, an audio encoder is provided. The encoder includes: a loudness component configured to obtain loudness data that describes a combination of one or more content sub-streams representing respective audio signals by applying a predefined loudness function; a presentation data component configured to define one or more presentation data structures, each presentation data structure including a reference to one or more content sub-streams among the plurality of content sub-streams and a reference to loudness data that describes a combination of the content sub-streams being referenced; and a multiplexing component configured to form a bitstream including the plurality of content sub-streams, the one or more presentation data structures, and the loudness data referenced by those presentation data structures.
Example
[0055] 〈III. Exemplary Embodiments〉 FIG. 1 shows, by way of example, a generalized block diagram of a decoder 100 for processing a bitstream P to achieve a desired loudness level of an output audio signal 114.
[0056] The decoder 100 has a receiving component (not shown) configured to receive a bitstream P including a plurality of content sub-streams each representing an audio signal.
[0057] Decoder 100 further includes a demultiplexer 102 configured to extract one or more presentation data structures 104 from bitstream P. Each presentation data structure includes a reference to at least one of the content sub-streams. In other words, a presentation data structure or a presentation is a description of which content sub-streams should be combined. As described above, content sub-streams encoded in two or more separate sub-streams may be combined into one presentation.
[0058] Each presentation data structure further includes a reference to a metadata sub-stream representing loudness data that describes a combination of one or more of the referenced content sub-streams.
[0059] The content of the presentation data structure and its various references will be described here in connection with FIG. 4.
[0060] In FIG. 4, various sub-streams 412, 205 that can be referenced by one or more of the extracted presentation data structures 104 are shown. Of the three presentation data structures 104, a selected presentation data structure 110 is selected. As is clear from FIG. 4, bitstream P includes content sub-stream 412, metadata sub-stream 205, and the one or more presentation data structures 104. The content sub-stream 412 may include a sub-stream for music, a sub-stream for effects, a sub-stream for ambience, a sub-stream for English dialog, a sub-stream for Spanish dialog, associated audio (AA) in English, for example, a sub-stream for an English commentary track, and AA in Spanish, for example, a sub-stream for a Spanish commentary track.
[0061] In FIG. 4, all content sub-streams 412 are encoded in the same bit stream P, but as described above, this is not always necessary. The broadcaster of the audio content may use a single bit stream configuration, such as a single packet identifier (PID) configuration in the MPEG standard, or a multiple bit stream configuration, such as a two-PID configuration, to send the audio content to the client, i.e., the decoder.
[0062] The present disclosure introduces an intermediate level in the form of a sub-stream group that exists between the presentation layer and the sub-stream layer. The content sub-stream group may group or reference one or more content sub-streams. Then, the presentation may be able to reference the content sub-stream group. In FIG. 4, the content sub-streams of music, effects, and ambient sound are grouped to form a content sub-stream group 410. This is referenced (404) by the selected presentation data structure 110.
[0063] The content sub-stream group provides additional flexibility in combining content sub-streams. In particular, the sub-stream group level provides a means to group or aggregate several content sub-streams into a unique group, such as a group 410 that includes music, effects, and ambient sound.
[0064] This can be advantageous because a content sub-stream group (e.g., for music and effects, or for music, effects, and ambient sound) can be used for more than one presentation, e.g., in connection with a dialog in English or Spanish. Similarly, a certain content sub-stream can also be used in more than one content sub-stream group.
[0065] Furthermore, depending on the syntax of the presentation data structure, using content sub-stream groups may provide the possibility of mixing a greater number of content sub-streams for presentation.
[0066] According to some embodiments, presentations 104, 110 always consist of one or more sub-stream groups.
[0067] The selected presentation data structure 110 in FIG. 4 includes a reference 404 to a content sub-stream group 410 composed of one or more of the content sub-streams. The selected presentation data structure 110 further includes references to content sub-streams for a Spanish dialog and references to content sub-streams for AA in Spanish. Furthermore, the selected presentation data structure 110 includes a reference 406 to a metadata sub-stream 205 representing loudness data 408 that describes a combination of one or more of the referenced content sub-streams. Clearly, the other two presentation data structures of the plurality of presentation data structures 104 may include data similar to the selected presentation data structure 110. According to other embodiments, the bitstream P may include additional metadata sub-streams similar to the metadata sub-stream 205. Here, the additional metadata sub-streams are referenced from other presentation data structures. In other words, each presentation data structure of the plurality of presentation data structures 104 may reference dedicated loudness data.
[0068] The selected presentation data structure may change over time, i.e., when the user decides to turn off the Spanish commentary track AA (ES). In other words, the bitstream P includes a plurality of time frames, and the data (reference numeral 108 in FIG. 1) indicating the selected presentation data structure among the one or more presentation data structures 104 can be assigned independently for each time frame.
[0069] As described above, the bitstream P includes a plurality of time frames. According to some embodiments, the one or more presentation data structures 104 may be related to different time segments of the bitstream P. In other words, the demultiplexer (reference numeral 102 in FIG. 1) is configured to extract one or more presentation data structures from the bitstream P for the first of the plurality of time frames, and further, from the bitstream P, for the second of the plurality of time frames, one or more presentation data structures different from the one or more presentation data structures extracted from the first of the plurality of time frames may be configured to be extracted. In this case, the data (reference numeral 108 in FIG. 1) indicating the selected presentation data structure indicates the selected presentation data structure for the time frame to which it is assigned.
[0070] Here, referring to FIG. 1, the decoder 100 further has a playback state component 106. The playback state component 106 is configured to receive data 108 indicating a selected presentation data structure 110 among the one or more presentation data structures 104. The data 108 also includes a desired loudness level. As described above, the data 108 may be provided by a consumer of the audio content decoded by the decoder 100. The desired loudness value may be a decoder-specific setting depending on the playback equipment used for playback of the output audio signal. The consumer may choose, for example, as understood from the above, that the audio content should include a Spanish dialogue.
[0071] Decoder 100 further has a hybrid component that receives the selected presentation data structure 110 from the playback state component 106 and decodes the one or more content sub-streams referred to by the selected presentation data structure 110 from the bitstream P. According to some embodiments, only the one or more content sub-streams referred to by the selected presentation data structure 110 are decoded by the hybrid component. As a result, if a consumer selects a presentation with a Spanish dialog, for example, any content sub-stream representing an English dialog is not decoded. This reduces the computational load of decoder 100.
[0072] Hybrid component 112 is configured to form an output audio signal based on the decoded content sub-stream.
[0073] Furthermore, hybrid component 112 is configured to process the decoded one or more content sub-streams or the output audio signal based on the loudness data referred to by the selected presentation data structure 110 to achieve the desired dialog loudness level.
[0074] Figures 2 and 3 depict different embodiments of hybrid component 112.
[0075] In FIG. 2, the bit stream P is received by the sub-stream decode component 202, and the sub-stream decode component 202 decodes, from the bit stream P, the one or more content sub-streams 204 referred to by the selected presentation data structure 110 based on the selected presentation data structure 110. Then, the one or more decoded content sub-streams 204 are transmitted to a component 206 that forms the output audio signal 114 based on the decoded content sub-streams 204 and the metadata sub-stream 205. When forming the audio output signal, the component 206 may take into account, for example, the time-dependent spatial position data included in the content sub-stream(s) 204 if any. The component 206 may further take into account the DRC data included in the metadata sub-stream 205. Alternatively, the loudness component 210 (described later) processes the output audio signal 114 based on the DRC data. In some embodiments, the component 206 receives (not shown in FIG. 2) mixing coefficients (described later) from the presentation data structure 110 and applies them to the corresponding content sub-streams 204. Then, the output audio signal 114* is transmitted to the loudness component 210, and the loudness component 210 processes the output audio signal 114* to achieve the desired loudness level based on the loudness data (included in the metadata sub-stream 205) referred to by the selected presentation data structure 110 and the desired loudness level included in the data 108, and thus outputs the loudness-processed output audio signal 114.
[0076] In FIG. 3, a similar mixing component 112 is shown. The difference from the mixing component 112 described in FIG. 2 is that the component 206 forming the output audio signal and the loudness component 210 have exchanged positions with each other. As a result, the loudness component 210 processes the one or more decoded content sub-streams 204 so as to achieve the desired loudness level (based on the loudness data included in the metadata sub-stream 205), and outputs one or more loudness-processed content sub-streams 204*. Then these are transmitted to the component 206 for forming the output audio signal, and the component 206 outputs the loudness-processed output audio signal 114. As described in connection with FIG. 2, the DRC data (included in the metadata sub-stream 205) can be applied either in the component 206 or in the loudness component 210. Further, in some embodiments, the component 206 receives mixing coefficients (described later) from the presentation data structure 110 (not shown in FIG. 3) and applies these coefficients to the corresponding content sub-streams 204*.
[0077] Each of the one or more presentation data structures 104 includes dedicated loudness data indicating what the loudness of the content sub-stream referred to by the presentation data structure will actually be when decoded. According to some embodiments, the loudness data represents a value for applying gating of the loudness function to its audio input signal. For example, if the loudness data is based on a band-limiting loudness function, frequency bands containing only noise can be ignored, so the background noise of the audio input signal is not taken into account when calculating the loudness data.
[0078] Furthermore, the loudness data may represent values of the loudness function that relate to a time segment of the audio input signal that represents a dialog. This is in accordance with the ATSC A / 85 standard, in which dialnorm is explicitly defined with respect to the loudness of dialog (anchor element): "The value of the dialnorm parameter indicates the loudness of the anchor element of the content."
[0079] Processing of the decoded one or more content sub-streams or the output audio signal, or leveling of the output audio signal, to achieve the desired loudness level ORL based on the loudness data referenced by the selected presentation data structure L can thus be performed using the presentation dialnorm, DN(pres), calculated according to the above: g L = ORL - DN(pres) where DN(pres) and ORL are typically both values expressed in dB FS (dB relative to a full-scale 1 kHz sine wave (or rectangular wave)).
[0080] According to some embodiments, the selected presentation data structure references two or more content sub-streams, and the selected presentation data structure further references at least one mixing coefficient to be applied to the two or more content sub-streams. The mixing coefficient(s) can be used to provide a modified relative loudness level between the content sub-streams referenced by the selected presentation. These mixing coefficients may be applied as broadband gains to channels / objects within a content sub-stream before mixing channels / objects within that content sub-stream with channels / objects within other content sub-stream(s).
[0081] At least one mixing coefficient is typically static, but may be assigned independently for each time frame of the bitstream. For example, to achieve dithering.
[0082] As a result, the mixing coefficients do not need to be transmitted for each time frame in the bitstream. They can continue to be valid until overwritten.
[0083] The mixing coefficients may be defined for each content substream. In other words, the selected presentation data structure may refer to one mixing coefficient to be applied to the corresponding substream for each of the two or more substreams.
[0084] According to other embodiments, the mixing coefficients are defined for each content substream group and may be applied to all content substreams within the content substream group. In other words, the selected presentation data structure refers to a single mixing coefficient to be applied to each of the one or more of the content substreams that make up the substream group for the content substream group.
[0085] According to yet another embodiment, the selected presentation data structure may refer to a single mixing coefficient to be applied to each of the two or more content substreams.
[0086] Table 1 below shows an example of object transmission. The objects are clustered into categories that are distributed across several sub-streams. All presentation data structures combine music and effects that include the main part of the audio content without dialog. Thus, this combination is a content sub-stream group. Depending on the selected presentation data structure, a certain language is selected. For example, English (D#1) or Spanish D#2. Further, the content sub-stream includes one associated audio sub-stream in English (Desc#1) and one associated audio sub-stream in Spanish (Desc#2). The associated audio may include enhancement audio such as audio description, a narrator for the hard-of-hearing, a narrator for the visually impaired, a commentary track, etc.
[0087] [Table 1] In Presentation 1, there is no mixing gain via the mixing coefficient to be applied. Thus, Presentation 1 does not refer to any mixing coefficient at all.
[0088] Due to cultural preferences, different balances between categories may be required. This is illustrated in Presentation 2. Consider a situation where the Spanish-speaking region does not want as much attention paid to music. Thus, the music sub-stream is attenuated by 3 dB. In this example, Presentation 2 refers to one mixing coefficient to be applied to each of the two or more sub-streams for each respective sub-stream.
[0089] Presentation 3 includes a Spanish language description stream for visually impaired persons. This stream was recorded at the booth and is too loud to be mixed into the presentation as is, so it is attenuated by 6 dB. In this example, Presentation 3 refers to one mixing coefficient to be applied to each of the two or more sub-streams, for each such sub-stream.
[0090] In Presentation 4, both the music sub-stream and the effects sub-stream are attenuated by 3 dB. In this case, Presentation 4 refers to a single mixing coefficient to be applied to each of the one or more content sub-streams that make up the M&E sub-stream group, for the M&E sub-stream group.
[0091] According to some embodiments, a user or consumer of the audio content can provide user input such that the output audio signal deviates from the selected presentation data structure. For example, the user may request dialog boost or dialog attenuation, or the user may wish to perform some kind of scene personalization, such as increasing the volume of sound effects. In other words, alternative mixing coefficients may be provided when combining two or more decoded content sub-streams to form the output audio signal. This may affect the loudness level of the audio output signal. In order to provide loudness consistency in this case, each of the one or more decoded content sub-streams may include loudness data at the sub-stream level that describes the loudness level of that content sub-stream. The sub-stream level loudness data may then be used to compensate the loudness data to provide loudness consistency.
[0092] Loudness data at the substream level may be the same as the loudness data referenced by the presentation data structure, and advantageously, optionally using a larger range to cover generally quieter signals in the content substream, may represent the values of the loudness function.
[0093] There are many ways to use this data to achieve loudness consistency. The following algorithm is shown as an example.
[0094] Let DN(P) be the presentation dialnorm and DN(S i ) be the substream loudness of substream i.
[0095] The decoder forms an audio output signal based on a presentation that references the music content substream S M and the effect content substream S E as one content substream group S M&E , and further the dialog content substream S D . When wanting to maintain consistent loudness while applying a 9 dB dialog enhancement DE, the decoder adds the content substream loudness values:
Equation
[0096] As described above, performing such an addition of substream loudness when approximating presentation loudness may result in a loudness that is very different from the actual loudness. Thus, an alternative is to calculate the approximation without DE and find the offset from the actual loudness.
[0097]
Equation
[0098]
Number
[0099] According to some embodiments, the DRC data referenced by the presentation data structure corresponds to multiple DRC profiles. These DRC profiles are customized for the specific audio signal to which they are applied. These profiles can range from no compression (“none”) to fairly light compression (e.g., “Music Light”) to very aggressive compression (e.g., “Speech”). As a result, the DRC data may include multiple sets of DRC gains or multiple compression curves from which the multiple sets of DRC gains are derived.
[0100] According to embodiments, the referenced DRC data may be included in the metadata substream 205 of FIG. 4.
[0101] The bitstream P may, according to some embodiments, include two or more separate bitstreams, and it should be noted that in this case, the respective content substreams may be encoded in different bitstreams. The one or more presentation data structures are, in this case, advantageously included in all of the separate bitstreams, i.e., for each separate bitstream, one or several decoders (given to each separate decoder) can function separately and completely independently to decode the content substream referenced by the selected presentation data structure. According to some embodiments, these decoders can function in parallel. Each separate decoder decodes the substream present in the separate bitstream it receives. According to embodiments, to achieve the desired loudness level, each separate decoder performs processing of the content substream it decodes. The processed content substream is then provided to a further mixing component that forms an output audio signal having the desired loudness level.
[0102] According to other embodiments, each separate decoder provides its decoded, unprocessed sub-stream to the further mixing component, which performs loudness processing and then forms an output audio signal from all of the one or more content sub-streams referred to by the selected presentation data structure, or first mixes the one or more content sub-streams and performs loudness processing on the mixed signal. According to other embodiments, each separate decoder performs a mixing operation on two or more of its decoded sub-streams. Then, a further mixing component mixes the pre-mixed contributions of the separate decoders.
[0103] FIG. 5 shows, by way of example, an audio encoder 500 in the context of FIG. 6. The encoder 500 has a presentation data component 504 configured to define one or more presentation data structures 506, each presentation data structure including references 604, 605 to one or more of the plurality of content sub-streams 502 and a reference 608 to loudness data 510 that describes a combination of the content sub-streams 612 being referenced. The encoder 500 further has a loudness component 508 configured to take the loudness data 510 that describes a combination of one or more content sub-streams representing respective audio signals by applying a predefined loudness function 514. The encoder further has a multiplexing component 512 configured to form a bitstream P including the plurality of content sub-streams, the one or more presentation data structures 506, and the loudness data 510 referenced by the one or more presentation data structures 506. The loudness data 510 typically includes several loudness data instances, one instance for each of the one or more presentation data structures 506.
[0104] Encoder 500 may further be adapted to determine, for each of the one or more presentation data structures 506, dynamic range compression (DRC) data for the one or more referenced content sub-streams. The DRC data quantifies at least one desired compression curve or at least a set of DRC gains. The DRC data is included in bitstream P. According to various embodiments, the DRC data and the loudness data 510 may be included in metadata sub-stream 614. As discussed above, loudness data is typically presentation-dependent. Further, the DRC data may also be presentation-dependent. In these cases, the loudness data and, if applicable, the DRC data for a particular presentation data structure are included in a dedicated metadata sub-stream 614 for that particular presentation data structure.
[0105] The encoder may further be adapted to apply the predefined loudness function to each of the plurality of content sub-streams 502 to obtain loudness data at the sub-stream level of that content sub-stream and include the loudness data at the sub-stream level in the bitstream. The predefined loudness function may be related to gating of the audio signal. According to other embodiments, the predefined loudness function may be related only to time segments of the audio signal that represent dialogue. The predefined loudness function, according to some embodiments: · includes frequency-dependent weighting of the audio signal, · includes channel-dependent weighting of the audio signal, · ignores segments of the audio signal having signal power below a threshold, · ignores segments of the audio signal not detected as speech, · may include at least one of the calculation of measures of energy / power / root mean square of the audio signal.
[0106] As can be understood from the above, the loudness function is non-linear. That is, if the loudness data were only calculated from different content sub-streams, the loudness for a given presentation could not be calculated by simply adding up the loudness data of the content sub-streams being referenced. Further, when different audio tracks, i.e., content sub-streams, are combined together for simultaneous playback, there may be an apparent combined effect between the coherent / incoherent portions of the different audio tracks or in different frequency regions, which further makes the addition of loudness data for the audio tracks mathematically impossible.
[0107] 〈IV. Equivalents, Extensions, Alternatives, etc.〉 Upon consideration of the above description, further embodiments of the present disclosure will be apparent to those skilled in the art. Although the description and drawings disclose embodiments and examples, the present disclosure is not limited to such specific examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure, which is only defined by the appended claims. Even if reference numerals appear in the claims, they are not to be construed as limiting the scope.
[0108] Furthermore, from a consideration of the drawings, the present disclosure, and the appended claims, variations to the disclosed embodiments that would be apparent to and can be implemented by those skilled in the art when practicing the present disclosure will be understood. In the claims, the word "comprising" does not exclude other elements or steps, and the singular form does not exclude a plurality. The mere fact that certain measures are described in mutually different dependent claims does not indicate that a combination of those measures cannot be used advantageously.
[0109] The apparatus and method disclosed above can be implemented as software, firmware, hardware, or a combination thereof. In a hardware implementation, the partitioning of tasks among the functional units mentioned in the above description does not necessarily correspond to the partitioning into physical units. Rather, one physical component may have multiple functions, or one task may be executed by several cooperating physical components. Some or all of the components may be implemented as software executed by a digital signal processor or a microprocessor, or alternatively as hardware or as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium that may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it is well known to those skilled in the art that communication media typically embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and include any information delivery media.
[0110] Some aspects will be described. 〔Aspect 1〕 A method for processing a bitstream including a plurality of content sub-streams each representing an audio signal, comprising: Extracting one or more presentation data structures from the bitstream, each presentation data structure including a reference to one or more of the content sub-streams, and each presentation data structure further including a reference to a metadata sub-stream representing loudness data describing a combination of one or more of the referenced content sub-streams; Receiving data indicating a selected presentation data structure of the one or more presentation data structures and a desired loudness level; Decoding the one or more content sub-streams referenced by the selected presentation data structure; Forming an output audio signal based on the decoded content sub-streams, and the method further includes processing the decoded one or more content sub-streams or the output audio signal to achieve the desired loudness level based on the loudness data referenced by the selected presentation data structure. Method. 〔Aspect 2〕 The selected presentation data structure references two or more content sub-streams and further references at least two mixing coefficients to be applied thereto, and the forming of the output audio signal further includes additively mixing the decoded one or more content sub-streams by applying the mixing coefficient(s). The method according to aspect 1. 〔Aspect 3〕 The method according to aspect 2, wherein the bitstream includes a plurality of time frames, and the mixing coefficient(s) referenced by the selected presentation data structure are assignable independently for each time frame. 〔Aspect 4〕 The method according to aspect 2 or 3, wherein the selected presentation data structure refers to one mixing coefficient to be applied to each of the two or more sub-streams, respectively, for each sub-stream of the two or more sub-streams. [Aspect 5] The method according to any one of aspects 1 to 4, wherein the loudness data represents a value related to the application of gating of a loudness function to its audio input signal. [Aspect 6] The method according to aspect 5, wherein the loudness data represents a value related to a time segment representing a dialog of its audio input signal of a loudness function. [Aspect 7] The presentation data structure further includes a reference to dynamic range compression (DRC) data for one or more content sub-streams to be referenced. The method further includes processing the decoded one or more content sub-streams or the output audio signal based on the DRC data, the processing including applying one or more DRC gains to the decoded one or more content sub-streams or the output audio signal. The method according to any one of aspects 1 to 6. [Aspect 8] The method according to aspect 7, wherein the DRC data includes at least one set of the one or more DRC gains. [Aspect 9] The DRC data includes at least one compression curve, and the one or more DRC gains are: Calculating one or more loudness values of the one or more content sub-streams or the audio output signal to be referenced using a predefined loudness function. Obtained by mapping the one or more loudness values to DRC gains using the compression curve. The method according to aspect 7. [Aspect 10] The mapping of the loudness value is the method according to aspect 9, including the smoothing operation of the DRC gain. [Aspect 11] The referenced DRC data is the method according to any one of aspects 7 to 10, included in the metadata substream. [Aspect 12] Each of the decoded one or more content substreams includes loudness data at the substream level that describes the loudness level of that content substream, and the processing of the decoded one or more content substreams or the output audio signal further includes providing loudness consistency based on the loudness level of the content substream, which is the method according to any one of aspects 1 to 11. [Aspect 13] The formation of the output audio signal includes combining two or more decoded content substreams using alternative mixing coefficients, and the substream-level loudness data is used to compensate the loudness data to provide loudness consistency, which is the method according to aspect 12. [Aspect 14] The alternative mixing coefficients are related to either dialog boost or dialog attenuation, which is the method according to aspect 13. [Aspect 15] The reference to at least one of the content substreams is a reference to at least one content substream group consisting of one or more of the content substreams, which is the method according to any one of aspects 1 to 14. [Aspect 16] When aspect 15 quotes aspect 2, the selected presentation data structure refers to a single mixing coefficient applied to each of the one or more of the content substreams that make up the substream group for a certain content substream group, which is the method according to aspect 15. [Aspect 17] The method according to any one of aspects 1 to 16, wherein the bitstream includes a plurality of time frames, and data indicating the selected presentation data structure among the one or more presentation data structures is assignable independently for each time frame. [Aspect 18] extracting, from the bitstream, one or more presentation data structures for a first one of the plurality of time frames; extracting, from the bitstream, one or more presentation data structures different from the one or more presentation data structures extracted from the first one of the plurality of time frames for a second one of the plurality of time frames; wherein data indicating the selected presentation data structure indicates the selected presentation data structure for the time frame to which it is assigned. The method according to aspect 17. [Aspect 19] The method according to any one of aspects 1 to 18, wherein only one or more content sub-streams referred to by the selected presentation data structure are decoded from the plurality of content sub-streams included in the bitstream. [Aspect 20] The bitstream includes two or more separate bitstreams each including at least one of the plurality of content sub-streams, and the step of decoding the one or more content sub-streams referred to by the selected presentation data structure includes: for each particular bitstream of the two or more separate bitstreams, separately decoding content sub-stream(s) from the referenced content sub-stream(s) included in that particular bitstream. The method according to any one of aspects 1 to 19. [Aspect 21] A decoder for processing a bitstream including a plurality of content sub-streams each representing an audio signal, the decoder comprising: a receiving component configured to receive the bitstream; a demultiplexer configured to extract one or more presentation data structures from the bitstream, each presentation data structure including a reference to at least one of the content sub-streams, and further including a reference to a metadata sub-stream representing loudness data describing a combination of one or more of the referenced content sub-streams; a playback state component configured to receive data indicating a selected presentation data structure of the one or more presentation data structures and a desired loudness level; a mixing component configured to decode the one or more content sub-streams referenced by the selected presentation data structure and to form an output audio signal based on the decoded content sub-streams; wherein the mixing component is further configured to process the decoded one or more content sub-streams or the output audio signal to achieve the desired loudness level based on the loudness data referenced by the selected presentation data structure; a decoder. 〔Aspect 22〕 An audio encoding method comprising: receiving a plurality of content sub-streams representing respective audio signals; defining one or more presentation data structures each referencing at least one of the plurality of content sub-streams; for each of the one or more presentation data structures, applying a predefined loudness function to obtain loudness data describing a combination of one or more of the referenced content sub-streams, and including a reference (608) from the presentation data structure to the loudness data; Forming a bitstream that includes the plurality of content sub-streams, the one or more presentation data structures, and the loudness data referenced by those presentation data structures. Method. 〔Aspect 23〕 For each of the one or more presentation data structures, determining dynamic range compression (DRC) data for one or more content sub-streams that are referenced, the DRC data quantifying at least one desired compression curve or at least one set of DRC gains. Further including the step of including the DRC data in the bitstream. The method according to aspect 22. 〔Aspect 24〕 For each of the plurality of content sub-streams, applying the predefined loudness function to obtain loudness data at the sub-stream level of that content sub-stream; Further including the step of including the loudness data at the sub-stream level in the bitstream. The method according to aspect 22 or 23. 〔Aspect 25〕 The method according to any one of aspects 22 to 24, wherein the predefined loudness function is related to gating of the audio signal. 〔Aspect 26〕 The method according to aspect 25, wherein the predefined loudness function is related only to time segments of the audio signal that represent dialog. 〔Aspect 27〕 The predefined loudness function is: Frequency-dependent weighting of the audio signal, Channel-dependent weighting of the audio signal, Ignoring segments of the audio signal with signal power below a threshold, Including at least one of calculating an energy measure of the audio signal. The method according to any one of aspects 22 to 26. [Aspect 28] A loudness component configured to obtain loudness data describing a combination of one or more content sub-streams representing respective audio signals by applying a predefined loudness function; A presentation data component configured to define one or more presentation data structures, each presentation data structure including a reference to one or more content sub-streams among the plurality of content sub-streams and a reference to loudness data describing a combination of the content sub-streams referred to, the presentation data component; A multiplexing component configured to form a bitstream including the plurality of content sub-streams, the one or more presentation data structures, and the loudness data referred to by the one or more presentation data structures, An audio encoder. [Aspect 29] A computer program product having a computer-readable medium having instructions for performing the method according to any one of aspects 1 to 20 and 22 to 27.
Claims
Claim 1 A method for processing a bitstream (P) including a plurality of content sub-streams (412) each representing an audio signal, comprising: extracting from the bitstream one or more presentation data structures (104), each presentation data structure including references (404, 405) to a plurality of the content sub-streams, each presentation data structure further including a reference (406) to loudness data (408) included in a metadata sub-stream (205) and dynamic range compression (DRC) data, the loudness data being dedicated to the presentation data structure and indicating what the loudness of a combination (204) of the plurality of content sub-streams to be referred to when decoded will be, the DRC data including a set of frequency-dependent DRC gains; receiving data (108) indicating a selected presentation data structure among the one or more presentation data structures (104) and a desired loudness level; decoding the plurality of content sub-streams (204) referred to by the selected presentation data structure (110); forming an output audio signal (114) based on the decoded content sub-streams (204), the method further including processing the decoded plurality of content sub-streams (204) or the output audio signal (114) to achieve the desired loudness level based on the loudness data referred to by the selected presentation data structure and the set of frequency-dependent DRC gains. Method. Claim 2 The selected presentation data structure further references at least two mixing coefficients to be applied to the plurality of content sub-streams, and the forming of the output audio signal further includes additively mixing the decoded plurality of content sub-streams by applying the mixing coefficients, The method according to claim 1. Claim 3 The bitstream includes a plurality of time frames, and the mixing coefficients referred to by the selected presentation data structure are assignable independently for each time frame, and / or The selected presentation data structure references, for each sub-stream of the plurality of sub-streams, one mixing coefficient to be applied to the respective sub-stream. The method according to claim 2.
4. The method according to any one of claims 1 to 3, wherein the bitstream includes a plurality of time frames, and data indicating the selected presentation data structure among the one or more presentation data structures is assignable independently for each time frame.
5. Extracting one or more presentation data structures from the bitstream for a first one of the plurality of time frames. Extracting, from the bitstream, one or more presentation data structures different from the one or more presentation data structures extracted from the first one of the plurality of time frames for a second one of the plurality of time frames. The data indicating the selected presentation data structure indicates the selected presentation data structure for the time frame to which it is assigned. The method according to claim 4.
6. A decoder for processing a bitstream including a plurality of content sub-streams each representing an audio signal, comprising: A decoder having one or more components configured to execute the method according to any one of claims 1 to 5.
7. A computer program product having instructions for executing the method according to any one of claims 1 to 5 when executed by a computing device or system.
Citation Information
Patent Citations
How to modify metadata that affects playback volume and dynamic range of audio information
JP2008505586A
Loudness range control system, transmitting device, receiving device, transmitting program and receiving program
JP2013157659A
Audio encoder and decoder with program loudness and boundary metadata
JP2018022180A