Transmission-independent presentation-based program loudness
By incorporating loudness and DRC data in audio encoding and decoding, the solution addresses inconsistent loudness issues across substreams, ensuring accurate and flexible loudness levels in audio presentations.
Patent Information
- Application Number
- JP2025140913
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-10-10
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-03
AI Technical Summary
Existing audio encoding and decoding technologies struggle to achieve consistent loudness levels across different content substreams, leading to inaccuracies greater than the tolerances specified by standards like ATSC A/85 and EBU R128, especially when switching languages or adding commentary tracks.
The proposed solution involves providing loudness data and dynamic range compression (DRC) data at the encoder level, allowing decoders to adjust playback gain and mix coefficients to achieve a desired loudness level, using presentation data structures to combine substreams accurately.
This approach ensures consistent loudness across different presentations, programs, and channels, providing a more accurate and flexible audio experience by compensating for cultural preferences and user inputs while adhering to compliance constraints.
Smart Images

Figure 2025176056000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 062,479, filed October 10, 2014, the contents of which are incorporated herein by reference in their entirety.
[0002] Technical Field The present invention relates to audio signal processing, and more particularly to encoding and decoding audio data bitstreams to achieve a desired loudness level in an output audio signal. [Background technology]
[0003] Dolby AC-4 is an audio format for efficient distribution of rich media content. AC-4 provides a flexible framework for broadcasters and content creators to distribute and encode content in an efficient manner. Content can be distributed through several substreams; for example, M&E (music and effects) in one substream and dialogue in a second substream. For some audio content, it can be advantageous to be able to switch the language of the dialogue from one language to another, or to add additional substreams containing, for example, a commentary substream to the content or descriptions for the visually impaired.
[0004] To ensure proper leveling of content presented to consumers, the loudness of the content needs to be known with some accuracy. Current loudness requirements have tolerances of 2 dB (ATSC A / 85) and 0.5 dB (EBU R128), while some specifications have tolerances as low as 0.1 dB. This means that the loudness of an output audio signal with dialogue in a first language and a commentary track should have substantially the same loudness as an output audio signal with dialogue in a second language but without a commentary track. [Brief explanation of the drawings]
[0005] Exemplary embodiments will now be described with reference to the accompanying drawings. [Figure 1] FIG. 1 is a generalized block diagram showing, by way of example, a decoder for processing a bitstream and achieving a desired loudness level in the output audio signal. [Figure 2] FIG. 2 is a generalized block diagram of a first embodiment of the mixing component of the decoder of FIG. 1. [Figure 3] FIG. 2 is a generalized block diagram of a second embodiment of the mixing component of the decoder of FIG. 1. [Figure 4] FIG. 1 describes a presentation data structure according to embodiments. [Figure 5] FIG. 1 is a generalized block diagram of an audio encoder according to various embodiments. [Figure 6] FIG. 6 illustrates a bitstream formed by the audio encoder of FIG. 5. All drawings are schematic and generally show only parts necessary for clarity of the disclosure, while other parts may be omitted or merely suggested. Unless otherwise noted, like reference numerals refer to like parts in different figures. DETAILED DESCRIPTION OF THE INVENTION
[0006] In view of the above, it is an object to provide an encoder and decoder and associated methods that aim to provide a desired loudness level for an output audio signal regardless of what content substreams are mixed into the output audio signal.
[0007] I. Overview - Decoder According to a first aspect, exemplary embodiments propose a decoding method, a decoder and a computer program product for decoding. The proposed method, decoder and computer program product may generally have the same features and advantages.
[0008] According to an exemplary embodiment, a method for processing a bitstream including multiple content substreams, each representing an audio signal, is provided, the method including: extracting from the bitstream one or more presentation data structures, each presentation data structure including a reference to at least one of the content substreams, each presentation data structure further including a reference to a metadata substream representing loudness data describing a combination of the referenced content substreams; receiving a selected one of the one or more presentation data structures and data indicating a desired loudness level; decoding one or more content substreams referenced by the selected presentation data structure; and forming an output audio signal based on the decoded content substreams, the method further including processing the decoded content substream(s) or the output audio signal based on the loudness data referenced by the selected presentation data structure to achieve the desired loudness level.
[0009] The data indicating the selected presentation data structure and the desired loudness level are typically user settings available at the decoder. A user may, for example, use a remote control to select a presentation data structure in which the dialogue is in French and / or to increase or decrease the desired output loudness level. In many embodiments, the output loudness level is related to the capacity of the playback device. According to some embodiments, the output loudness level is controlled by the volume. As a result, the data indicating the selected presentation data structure and the desired loudness level are typically not included in the bitstream received by the decoder.
[0010] As used herein, "loudness" refers to a modeled psychoacoustic measurement of sound intensity. In other words, loudness represents an approximation of the volume of a sound or sounds as perceived by an average user.
[0011] As used herein, "loudness data" refers to data resulting from measuring the loudness level of a particular presentation data structure using a function that models psychoacoustic loudness perception. In other words, it is a collection of values that indicate the loudness attributes of a combination of one or more referenced content substreams. According to various embodiments, the average loudness level of the combination of one or more content substreams referenced by a particular presentation data structure may be measured. For example, loudness data may refer to the dialnorm values (based on ITU-R BS.1770) of the one or more content substreams referenced by a particular presentation data structure. Other suitable loudness measurement standards, such as the Glasberg and Moore loudness models, which provide modifications and extensions to the Zwicker loudness model, may also be used.
[0012] As used herein, a "presentation data structure" refers to metadata related to the content of an output audio signal. The output audio signal is also referred to as a "program." A presentation data structure is also referred to as a "presentation."
[0013] Audio content can be distributed through several substreams. As used herein, "content substream" refers to such a substream. For example, a content substream may contain the music of the audio content, the dialogue of the audio content, or a commentary track to be included in the output audio signal. Content substreams may be channel-based or object-based. In the latter case, time-dependent spatial position data is included in the content substream. Content substreams may be included in the bitstream or may be part of the audio signal (i.e., as a channel group or object group).
[0014] As used herein, "output audio signal" refers to the audio signal that is actually output and rendered to the user.
[0015] The inventors have come to realise that by providing loudness data, e.g., dialnorm values, for each presentation, specific loudness data will be available to the decoder indicating exactly what the loudness will be for at least one referenced content substream when decoding that particular presentation.
[0016] In prior art, loudness data may be provided for each content substream. The problem with providing loudness data for each content substream is that it is then up to the decoder to combine the various loudness data into a presentation loudness. Adding the individual loudness data values of the substreams, representing their average loudness, to arrive at a loudness value for a presentation may be inaccurate and often does not result in the actual average loudness value of the combined substreams. Adding the loudness data for each referenced content substream may be mathematically impossible due to signal attributes, loudness algorithms, and the typically non-additive nature of loudness perception, leading to potential inaccuracies greater than the tolerances noted above.
[0017] With this embodiment, the difference between the average loudness level of the selected presentation provided by the loudness data for the selected presentation and the desired loudness level can thus be used to control the playback gain of the output audio signal.
[0018] By providing and using loudness data as described above, consistent loudness, i.e., loudness that approaches a desired loudness level, can be achieved between different presentations. Furthermore, consistent loudness can be achieved between different programs on a television channel, for example, between a television program and its commercials, or across television channels.
[0019] According to an exemplary embodiment, the selected presentation data structure references two or more content substreams and further references at least two mixing coefficients to be applied thereto, and said forming of the output signal further includes additively mixing the decoded one or more content substreams by applying said mixing coefficient(s).
[0020] By providing at least two mix coefficients, increased flexibility in the content of the output audio signal is achieved.
[0021] For example, the selected presentation data structure may reference, for each of the two or more content substreams, one mix coefficient to be applied to each substream. According to this embodiment, the relative loudness levels between the content substreams may be altered. For example, cultural preferences may require a different balance between different content substreams. Consider a situation where a Spanish-speaking region desires less attention to music; therefore, the music substream may be attenuated by 3 dB. According to other embodiments, signal mix coefficients may be applied to a subset of the two or more content substreams.
[0022] According to an exemplary embodiment, the bitstream includes multiple time frames, and the mix coefficients referenced by the selected presentation data structure can be assigned independently for each time frame. An advantage of providing time-varying mix coefficients is that ducking can be achieved. For example, the loudness level over a time segment of one content substream may be reduced due to increased loudness in the same time segment of another content substream.
[0023] According to an exemplary embodiment, the loudness data represents values related to the application of gating of a loudness function to the audio input signal.
[0024] The audio input signal is subjected to a loudness function (e.g., a dialnorm function) at the encoder side. The resulting loudness data is then transmitted in the bitstream to the decoder. A noise gate (also called a silence gate) is an electronic device or software used to control the volume of an audio signal. Gating is the use of such a gate. A noise gate attenuates signals that exhibit values below a threshold. A noise gate may attenuate a signal by a fixed amount, known as its range. In its simplest form, a noise gate allows signals to pass only if they are above a set threshold.
[0025] Gating may also be based on the presence of dialogue in the audio input signal. As a result, according to an exemplary embodiment, the loudness data represents values of a loudness function that relate to time segments of the audio input signal that represent dialogue. According to other embodiments, gating is based on a minimum loudness level. Such a minimum loudness level may be an absolute threshold or a relative threshold. A relative threshold may be based on a loudness level measured using an absolute threshold.
[0026] According to an exemplary embodiment, the presentation data structure further includes a reference to dynamic range compression (DRC) data for the referenced content substream or substreams, and the method further includes processing the decoded content substream or substreams or the output audio signal based on the DRC data, where the processing includes applying one or more DRC gains to the decoded content substream or substreams or the output audio signal.
[0027] Dynamic range compression reduces the volume of loud sounds and amplifies quiet sounds, thereby narrowing or "compressing" the dynamic range of an audio signal. By providing DRC data uniquely for each presentation, an improved user experience of the output audio signal may be achieved, regardless of the presentation chosen. Furthermore, by providing DRC data for each presentation, a consistent user experience of the audio output signal may be achieved across each of multiple presentations, and as noted above, between programs and across television channels.
[0028] The DRC gain is always time-varying. In each time segment, the DRC gain may be a single gain for the audio output signal or multiple DRC gains that vary per substream. The DRC gains may be applied to groups of channels and / or may be frequency-dependent. In addition, the DRC gains included in the DRC data may represent DRC gains for more than one DRC time segment, for example, subframes of a time frame defined by the encoder.
[0029] According to an exemplary embodiment, the DRC data includes at least one set of the one or more DRC gains. Thus, the DRC data may include multiple DRC profiles corresponding to DRC modes, each of which provides a different user experience of the audio output signal. By including the DRC gains directly in the DRC data, reduced computational complexity of the decoder may be achieved.
[0030] According to an exemplary embodiment, the DRC data includes at least one compression curve, and the one or more DRC gains are obtained by: calculating one or more loudness values of the one or more content substreams or the audio output signal using a predefined loudness function, and mapping the one or more loudness values to a DRC gain using the compression curve. By providing compression curves in the DRC data and calculating DRC gains based on these curves, the required bit rate for transmitting the DRC data to an encoder may be reduced. The predefined loudness function may be taken, for example, from the ITU-R BS.1770 recommendation document, although any suitable loudness function may be used.
[0031] According to example embodiments, the mapping of loudness values includes a smoothing operation of the DRC gain. The effect of this may be a better perceived output audio signal. A time constant for smoothing the DRC gain may be transmitted as part of the DRC data. Such a time constant may vary depending on signal attributes. For example, in some embodiments, the time constant may be smaller when a loudness value is greater than the immediately preceding corresponding loudness value than when the loudness value is less than the immediately preceding corresponding loudness value.
[0032] According to an exemplary embodiment, the referenced DRC data is included in the metadata substream, which may reduce the complexity of decoding the bitstream.
[0033] According to an exemplary embodiment, each of the one or more decoded content substreams includes loudness data at a substream level describing a loudness level of that content substream, and the processing of the decoded one or more content substreams or the output audio signal further includes ensuring loudness consistency based on the loudness levels of the content substreams.
[0034] As used in this document, "loudness consistency" refers to loudness that is consistent across different presentations, i.e., across multiple output audio signals formed based on different content substreams. Furthermore, the term refers to loudness that is consistent across different programs, i.e., between an audio signal for a television program and an audio signal for a commercial break. Furthermore, the term refers to loudness that is consistent across different television channels.
[0035] Providing loudness data describing the loudness levels of content substreams can help decoders provide loudness consistency in some cases. For example, when the formation of an output audio signal involves combining two or more decoded content substreams using alternative mix coefficients, and the substream-level loudness data is used to compensate for the loudness data to provide loudness consistency. These alternative mix coefficients may be derived from user input, for example, if the user decides to deviate from the default presentation (e.g., with dialogue enhancement, dialogue attenuation, scene personalization, etc.). This can jeopardize loudness compliance, as user influence can cause the loudness of the audio output signal to deviate from compliance constraints. To support loudness consistency in such cases, this embodiment provides the option to transmit substream-level loudness data.
[0036] According to some embodiments, the reference to at least one of the content substreams is a reference to at least one content substream group consisting of one or more of the content substreams. This may reduce decoder complexity, as multiple presentations can share a content substream group (e.g., a substream group consisting of a music-related content substream and an effects-related content substream). This may also reduce the required bitrate for transmitting the bitstream.
[0037] According to some embodiments, the selected presentation data structure references, for a content substream group, a single mixing coefficient that is applied to each of the one or more of the content substreams that make up the substream group.
[0038] This can be advantageous when the relative loudness levels of the content substreams in a content substream group are OK, but the overall loudness level of the content substreams in that content substream group should be increased or decreased relative to the other content substream(s) or content substream group(s) referenced by the selected presentation data structure.
[0039] In some embodiments, the bitstream includes multiple time frames, and data indicating the selected one of the one or more presentation data structures is independently assignable for each time frame. As a result, when multiple presentation data structures are received for a program, the selected presentation data structure may be changed during the course of the program, e.g., by a user. As a result, this embodiment provides a more flexible way of selecting output audio content, while at the same time providing loudness consistency in the output audio signal.
[0040] According to some embodiments, the method further includes: extracting from the bitstream one or more presentation data structures for a first of the plurality of time frames; and extracting from the bitstream one or more presentation data structures for a second of the plurality of time frames that differ from the one or more presentation data structures extracted from the first of the plurality of time frames, wherein the data indicative of the selected presentation data structure indicates the selected presentation data structure for the time frame to which it is assigned. As a result, multiple presentation data structures may be received in the bitstream, some of which relate to a first set of time frames and some of which relate to a second set of time frames. For example, a commentary track may be available only for certain time segments of the program. Furthermore, as the program progresses, the currently applicable presentation data structures at a particular time point may be used to select a selected presentation data structure. As a result, this embodiment provides a more flexible way of selecting output audio content while simultaneously providing loudness consistency in the output audio signal.
[0041] According to some embodiments, from the plurality of content substreams included in the bitstream, only the one or more content substreams referenced by the selected presentation data structure are decoded, which may provide an efficient decoder with reduced computational complexity.
[0042] According to some embodiments, the bitstream includes two or more separate bitstreams, each including at least one of the plurality of content bitstreams, and decoding the one or more content substreams referenced by the selected presentation data structure includes: for each particular bitstream of the two or more separate bitstreams, separately decoding a content substream or substreams from the referenced content substreams included in that particular bitstream. According to this embodiment, each separate bitstream may be received by a separate decoder, which decodes the required content substream or substreams based on the selected presentation data structure provided in the separate bitstream. This may improve decoding speed, as separate decoders can function in parallel. As a result, the decoding performed by the separate decoders may at least partially overlap. However, it should be noted that it is not necessary for the decoding performed by the separate decoders to overlap.
[0043] Furthermore, by splitting the content substreams into several bitstreams, this embodiment allows the at least two separate bitstreams to be received through different infrastructures, as described below. As a result, this exemplary embodiment provides a more flexible way to receive the multiple content substreams at a decoder.
[0044] Each decoder may process the decoded substream(s) based on loudness data referenced by the selected presentation data structure, and / or apply DRC gains and / or apply blending factors to the decoded substream(s). The processed or unprocessed content substreams may then be provided from all of the at least two decoders to a blending component for forming an output audio signal. Alternatively, the blending component may perform loudness processing and / or apply DRC gains and / or apply blending factors. In some embodiments, a first decoder may receive a first bitstream of the two or more separate bitstreams through a first infrastructure (e.g., cable television broadcast), while a second decoder may receive a second bitstream of the two or more separate bitstreams through a second infrastructure (e.g., over the Internet). According to some embodiments, the one or more presentation data structures are present in all of the two or more separate bitstreams. In this case, presentation definitions and loudness data are present in all separate decoders. This allows for independent operation of their decoding up to the mixed components. References to substreams that do not exist in the corresponding bitstream may be indicated as being externally provided.
[0045] According to an exemplary embodiment, a decoder for processing a bitstream including multiple content substreams, each representing an audio signal, is provided, the decoder including: a receiving component configured to receive the bitstream; a demultiplexer configured to extract one or more presentation data structures from the bitstream, each presentation data structure including a reference to at least one of the content substreams and further including a reference to a metadata substream representing loudness data describing a combination of the referenced content substreams; a playback state component configured to receive a selected one of the one or more presentation data structures and data indicative of a desired loudness level; and a mixing component configured to decode the one or more content substreams referenced by the selected presentation data structure and form an output audio signal based on the decoded content substreams, the mixing component further configured to process the decoded content substream or the output audio signal based on the loudness data referenced by the selected presentation data structure to achieve the desired loudness level.
[0046] II. Overview - Encoders According to a second aspect, exemplary embodiments propose an encoding method, an encoder, and a computer program product for encoding. The proposed method, encoder, and computer program product may generally have the same features and advantages. In general, the features of the second aspect may have the same advantages as the corresponding features of the first aspect.
[0047] According to an exemplary embodiment, an audio encoding method is provided, the method including: receiving a plurality of content substreams representing respective audio signals; defining one or more presentation data structures, each of which references at least one of the plurality of content substreams; applying a predefined loudness function to each of the one or more presentation data structures to obtain loudness data describing the combination of the referenced one or more content substreams, including a reference to the loudness data from the presentation data structures; and forming a bitstream including the plurality of content substreams, the one or more presentation data structures, and the loudness data referenced by the presentation data structures.
[0048] As noted above, the term "content substream" encompasses both substreams within a bitstream and within an audio signal. An audio encoder typically receives audio signals, which are then encoded into bitstreams. The audio signals may be grouped, with each group being characterized as an individual encoder input audio signal. Each group may then be encoded into a substream.
[0049] According to some embodiments, the method further includes: for each of the one or more presentation data structures, determining dynamic range compression (DRC) data for one or more referenced content substreams, the DRC data quantifying at least one desired compression curve or at least one set of DRC gains; and including the DRC data in the bitstream.
[0050] According to some embodiments, the method further includes: for each of the plurality of content substreams, applying the predefined loudness function to obtain substream-level loudness data for that content substream; and including the substream-level loudness data in the bitstream.
[0051] According to some embodiments, the predefined loudness function relates to the application of gating to the audio signal.
[0052] According to some embodiments, the predefined loudness function relates only to time segments of the audio signal that represent dialogue.
[0053] According to some embodiments, the predefined loudness function comprises at least one of: a frequency-dependent weighting of the audio signal, a channel-dependent weighting of the audio signal, ignoring segments of the audio signal with signal power below a threshold, and calculating an energy measure of the audio signal.
[0054] According to an exemplary embodiment, an audio encoder is provided, comprising: a loudness component configured to apply a predefined loudness function to obtain loudness data describing a combination of one or more content substreams representing a respective audio signal; a presentation data component configured to define one or more presentation data structures, each presentation data structure including references to one or more content substreams of a plurality of content substreams and to loudness data describing the combination of the referenced content substreams; and a multiplexing component configured to form a bitstream including the plurality of content substreams, the one or more presentation data structures, and the loudness data referenced by the presentation data structures. [Example]
[0055] III. ILLUSTRATIVE EMBODIMENTS FIG. 1 shows, by way of example, a generalized block diagram of a decoder 100 for processing a bitstream P to achieve a desired loudness level in an output audio signal 114.
[0056] The decoder 100 comprises a receiving component (not shown) configured to receive a bitstream P that includes multiple content substreams, each representing an audio signal.
[0057] The decoder 100 further comprises a demultiplexer 102 configured to extract one or more presentation data structures 104 from the bitstream P. Each presentation data structure contains a reference to at least one of the content substreams. In other words, a presentation data structure or presentation is a description of which content substreams should be combined. As noted above, content substreams that are encoded in two or more separate substreams may be combined into one presentation.
[0058] Each presentation data structure further includes a reference to a metadata substream representing loudness data that describes the combination of one or more referenced content substreams.
[0059] The contents of the presentation data structure and its various references will now be described in conjunction with FIG.
[0060] 4 illustrates various substreams 412, 205 that may be referenced by the extracted presentation data structure(s) 104. Of the three presentation data structures 104, a selected presentation data structure 110 is chosen. As can be seen from FIG. 4, the bitstream P includes a content substream 412, a metadata substream 205, and the one or more presentation data structures 104. The content substreams 412 may include a substream for music, a substream for effects, a substream for ambience, a substream for English dialogue, a substream for Spanish dialogue, associated audio (AA) in English, e.g., a substream for an English commentary track, and AA in Spanish, e.g., a substream for a Spanish commentary track.
[0061] 4, all content substreams 412 are encoded in the same bitstream P, but as noted above, this need not always be the case. Broadcasters of audio content may use a single bitstream configuration, such as a single packet identifier (PID) configuration in the MPEG standard, or a multiple bitstream configuration, such as a two-PID configuration, to transmit audio content to clients, i.e., decoders.
[0062] This disclosure introduces an intermediate level in the form of substream groups that reside between the presentation layer and the substream layer. A content substream group may group or reference one or more content substreams. A presentation may then reference a content substream group. In Figure 4, the music, effects, and ambient content substreams are grouped to form a content substream group 410, which is referenced 404 by the selected presentation data structure 110.
[0063] Content substream groups provide additional flexibility in combining content substreams. In particular, the substream group level provides a means to organize or group several content substreams into a unique group, for example, a group 410 containing music, effects, and ambient sounds.
[0064] This can be advantageous because a content substream group (e.g., for music and effects, or for music, effects, and ambient sounds) can be used for more than one presentation, e.g., in conjunction with English or Spanish dialogue. Similarly, a content substream can also be used in more than one content substream group.
[0065] Additionally, depending on the syntax of the presentation data structure, using content substream groups may offer the possibility to mix a larger number of content substreams for presentation.
[0066] According to some embodiments, a presentation 104, 110 always consists of one or more sub-stream groups.
[0067] The selected presentation data structure 110 in FIG. 4 includes a reference 404 to a content substream group 410 consisting of one or more of the content substreams. The selected presentation data structure 110 further includes a reference to a content substream for Spanish dialogue and a reference to a content substream for AA in Spanish. Furthermore, the selected presentation data structure 110 includes a reference 406 to a metadata substream 205 representing loudness data 408 describing the combination of the referenced content substream(s). Obviously, two other presentation data structures of the plurality of presentation data structures 104 may contain data similar to the selected presentation data structure 110. According to other embodiments, the bitstream P may include an additional metadata substream similar to the metadata substream 205, where the additional metadata substream is referenced by the other presentation data structure. In other words, each presentation data structure of the plurality of presentation data structures 104 may reference its own loudness data.
[0068] The selected presentation data structure may change over time, i.e., if the user decides to turn off the Spanish commentary track AA(ES). In other words, the bitstream P includes multiple time frames, and data indicating the selected presentation data structure (reference numeral 108 in FIG. 1) of the one or more presentation data structures 104 is assignable independently for each time frame.
[0069] As mentioned above, bitstream P includes a plurality of time frames. According to some embodiments, the one or more presentation data structures 104 may relate to different time segments of bitstream P. In other words, the demultiplexer (reference number 102 in FIG. 1 ) may be configured to extract, for a first one of the plurality of time frames, one or more presentation data structures from bitstream P, and further configured to extract, for a second one of the plurality of time frames, one or more presentation data structures from bitstream P that differ from the one or more presentation data structures extracted from the first one of the plurality of time frames. In this case, the data indicative of the selected presentation data structure (reference number 108 in FIG. 1 ) indicates the selected presentation data structure for the time frame to which it is assigned.
[0070] 1, the decoder 100 further includes a playback state component 106. The playback state component 106 is configured to receive data 108 indicating a selected presentation data structure 110 from the one or more presentation data structures 104. The data 108 also includes a desired loudness level. As noted above, the data 108 may be provided by a consumer of the audio content to be decoded by the decoder 100. The desired loudness value may be a decoder-specific setting, depending on the playback equipment used for playback of the output audio signal. The consumer may, for example, select that the audio content should include Spanish dialogue, as understood from the above.
[0071] The decoder 100 further includes a mixing component that receives the selected presentation data structure 110 from the playback state component 106 and decodes the one or more content substreams referenced by the selected presentation data structure 110 from the bitstream P. According to some embodiments, only the one or more content substreams referenced by the selected presentation data structure 110 are decoded by the mixing component. As a result, if a consumer selects a presentation with, for example, Spanish dialogue, any content substreams representing English dialogue are not decoded. This reduces the computational complexity of the decoder 100.
[0072] The mixing component 112 is configured to form an output audio signal based on the decoded content substreams.
[0073] Further, mixing component 112 is configured to process the decoded content substream(s) or the output audio signal based on loudness data referenced by the selected presentation data structure 110 to achieve the desired dialogue loudness level.
[0074] 2 and 3 describe different embodiments of the mixing component 112. FIG.
[0075] 2 , bitstream P is received by substream decode component 202, which decodes the one or more content substreams 204 referenced by the selected presentation data structure 110 from bitstream P based on the selected presentation data structure 110. The one or more decoded content substreams 204 are then transmitted to component 206, which forms output audio signal 114 based on the decoded content substreams 204 and metadata substream 205. When forming the audio output signal, component 206 may, for example, take into account time-dependent spatial position data, if any, included in content substream(s) 204. Component 206 may also take into account DRC data included in metadata substream 205. Alternatively, loudness component 210 (described below) processes output audio signal 114 based on the DRC data. In some embodiments, component 206 receives mixing coefficients (described below) from presentation data structure 110 (not shown in FIG. 2 ) and applies them to the corresponding content substreams 204. The output audio signal 114* is then transmitted to loudness component 210, which processes the output audio signal 114* to achieve the desired loudness level based on the loudness data (contained in metadata substream 205) referenced by the selected presentation data structure 110 and the desired loudness level contained in data 108, thereby outputting a loudness-processed output audio signal 114.
[0076] In Figure 3, a similar mixing component 112 is shown. It differs from the mixing component 112 described in Figure 2 in that the component 206 for forming the output audio signal and the loudness component 210 have swapped positions. As a result, the loudness component 210 processes the decoded content substream(s) 204 to achieve the desired loudness level (based on the loudness data contained in the metadata substream 205) and outputs one or more loudness-processed content substreams 204*. These are then transmitted to the component 206 for forming the output audio signal, which outputs the loudness-processed output audio signal 114. As described in connection with Figure 2, the DRC data (contained in the metadata substream 205) can be applied either in the component 206 or in the loudness component 210. Additionally, in some embodiments, component 206 receives mixing coefficients (described below) from presentation data structure 110 (not shown in FIG. 3) and applies these coefficients to the corresponding content substreams 204*.
[0077] Each of the one or more presentation data structures 104 includes dedicated loudness data that indicates what the loudness of the content substream referenced by the presentation data structure will actually be when decoded. According to some embodiments, the loudness data represents values of a loudness function that apply gating to the audio input signal. For example, if the loudness data is based on a band-limiting loudness function, frequency bands containing only noise may be ignored, so that background noise in the audio input signal is not taken into account when calculating the loudness data.
[0078] Additionally, the loudness data may represent values of loudness functions relating to time segments of the audio input signal that represent dialogue. This is in line with the ATSC A / 85 standard, where dialnorm is explicitly defined with respect to the loudness of dialogue (anchor elements): "The value of the dialnorm parameter indicates the loudness of the anchor element of the content."
[0079] processing the decoded content substream(s) or the output audio signal to achieve the desired loudness level ORL based on loudness data referenced by the selected presentation data structure, or leveling the output audio signal. L can thus be performed with the dialnorm of the presentation, DN(pres), calculated according to above: g L =ORL-DN(pres) where DN(pres) and ORL are typically both in dB. FS This is a value expressed in dB (based on a full-scale 1 kHz sine wave (or square wave)).
[0080] According to some embodiments, the selected presentation data structure references two or more content substreams, and the selected presentation data structure further references at least one mix coefficient to be applied to the two or more content substreams. The mix coefficient(s) may be used to provide a modified relative loudness level between the content substreams referenced by the selected presentation. These mix coefficients may be applied as wideband gain to channels / objects in a content substream before mixing the channels / objects in that content substream with channels / objects in other content substream(s).
[0081] At least one blending coefficient is typically static, but may be independently assignable for each time frame of the bitstream, for example to achieve ducking.
[0082] As a result, the mixing coefficients do not need to be transmitted for each time frame in the bitstream, but can remain valid until overwritten.
[0083] A mixing coefficient may be defined for each content substream, i.e., the selected presentation data structure may reference, for each substream of the two or more substreams, one mixing coefficient to be applied to the corresponding substream.
[0084] According to other embodiments, a mixing factor may be defined per content substream group and applied to all content substreams within the content substream group, i.e., the selected presentation data structure references, for a content substream group, a single mixing factor that is applied to each of the one or more content substreams that make up that substream group.
[0085] According to yet another embodiment, the selected presentation data structure may reference a single mix coefficient that is applied to each of the two or more content substreams.
[0086] Table 1 below shows an example of object transmission. The objects are clustered into categories distributed across several substreams. Every presentation data structure combines music and effects, which contain the main part of the dialogue-free audio content. This combination is thus a content substream group. Depending on the selected presentation data structure, a language is chosen, for example, English (D#1) or Spanish (D#2). Furthermore, the content substream contains one accompanying audio substream in English (Desc#1) and one accompanying audio substream in Spanish (Desc#2). The associated audio may also contain enhancement audio, such as audio description, a narrator for the hearing impaired, a narrator for the visually impaired, a commentary track, etc.
[0087] [Table 1] In Exhibit 1, there is no mix gain via mix factors to be applied, so Exhibit 1 does not refer to mix factors at all.
[0088] Cultural preferences may dictate different balances between categories. This is exemplified in Exhibit 2. Consider a situation where a Spanish-speaking region desires less attention to music. Therefore, the music substream is attenuated by 3 dB. In this example, Exhibit 2 refers to one blending coefficient to be applied to each substream of the two or more substreams.
[0089] Exhibit 3 includes a Spanish description stream for the visually impaired. This stream was recorded in a booth and is attenuated by 6 dB because it is too loud to be mixed directly into the presentation. In this example, Exhibit 3 references one mix coefficient to be applied to each substream of the two or more substreams.
[0090] In Exhibit 4, both the music substream and the effects substream are attenuated by 3 dB, where Exhibit 4 refers to a single mix factor for an M&E substream group that should be applied to each of the one or more content substreams that make up the M&E substream group.
[0091] According to some embodiments, a user or consumer of audio content can provide user input to cause the output audio signal to deviate from the selected presentation data structure. For example, a user may request dialogue enhancement or dialogue attenuation, or the user may wish to perform some kind of scene personalization, such as increasing the volume of sound effects. In other words, alternative mix coefficients may be provided to be used when combining two or more decoded content substreams to form the output audio signal. This may affect the loudness level of the audio output signal. To provide loudness consistency in this case, each of the decoded content substreams may include substream-level loudness data describing the loudness level of that content substream. The substream-level loudness data may then be used to compensate the loudness data to provide loudness consistency.
[0092] The loudness data at the substream level may be similar to the loudness data referenced by the presentation data structure, and may advantageously represent loudness function values, optionally using a larger range to cover the generally quieter signals in the content substreams.
[0093] There are many ways to use this data to achieve loudness consistency. The algorithm below is given as an example.
[0094] DN(P) is the presentation dialnorm, and DN(S i ) is the substream loudness of substream i.
[0095] The decoder detects the music content substream S M and effect content substreams E One content sub-stream group M&E and even dialogue content sub-streams D and if it wants to maintain consistent loudness while applying 9 dB dialogue enhancement DE, the decoder shall add the content substream loudness values:
number
[0096] As mentioned above, performing such addition of substream loudness when approximating the presented loudness may result in a loudness that is very different from the actual loudness, so an alternative is to compute the approximation without DE and find the offset from the actual loudness.
[0097]
number
[0098]
number
[0099] According to some embodiments, the DRC data referenced by the presentation data structure corresponds to multiple DRC profiles. These DRC profiles are custom-tailored to the particular audio signal to which they are applied. These profiles can range from no compression ("None"), to fairly light compression (e.g., "Music Light"), to very aggressive compression (e.g., "Speech"). As a result, the DRC data may include multiple sets of DRC gains or multiple compression curves from which the multiple sets of DRC gains are derived.
[0100] The referenced DRC data may, according to various embodiments, be included in the metadata substream 205 of FIG.
[0101] It should be noted that, according to some embodiments, the bitstream P may include two or more separate bitstreams, and the content substreams may in this case be encoded in different bitstreams. The one or more presentation data structures are in this case advantageously included in all of the separate bitstreams, i.e., several decoders, one for each separate bitstream, can function separately and entirely independently to decode the content substreams referenced by the selected presentation data structure (and are provided to each separate decoder). According to some embodiments, the decoders can function in parallel. Each separate decoder decodes a substream present in the separate bitstream it receives. According to some embodiments, each separate decoder performs processing of the content substream it decodes to achieve a desired loudness level. The processed content substreams are then provided to a further mixing component, which forms an output audio signal with the desired loudness level.
[0102] According to other embodiments, each separate decoder provides its decoded, unprocessed substream to the further mixing component, which performs loudness processing and then forms an output audio signal from all of the one or more content substreams referenced by the selected presentation data structure, or first mixes the one or more content substreams and performs loudness processing on the mixed signal. According to other embodiments, each separate decoder performs a mixing operation on two or more of its decoded substreams. A further mixing component then mixes the pre-mixed contributions of the separate decoders.
[0103] Figure 5, in conjunction with Figure 6, illustrates an example audio encoder 500. The encoder 500 includes a presentation data component 504 configured to define one or more presentation data structures 506, each including references 604, 605 to one or more content substreams 612 of a plurality of content substreams 502 and a reference 608 to loudness data 510 describing the combination of the referenced content substreams 612. The encoder 500 also includes a loudness component 508 configured to apply a predefined loudness function 514 to the loudness data 510 describing the combination of one or more content substreams representing a respective audio signal. The encoder also includes a multiplexing component 512 configured to form a bitstream P including the plurality of content substreams, the one or more presentation data structures 506, and the loudness data 510 referenced by the one or more presentation data structures 506. The loudness data 510 typically includes several loudness data instances, one for each of the one or more presentation data structures 506 .
[0104] The encoder 500 may further be adapted to determine, for each of the one or more presentation data structures 506, dynamic range compression DRC data for the referenced content substream or substreams. The DRC data quantifies at least one desired compression curve or at least one set of DRC gains. The DRC data is included in the bitstream P. The DRC data and loudness data 510 may, according to various embodiments, be included in a metadata substream 614. As discussed above, loudness data is typically presentation-dependent. Furthermore, the DRC data may also be presentation-dependent. In these cases, the loudness data and, if applicable, the DRC data for a particular presentation data structure are included in a dedicated metadata substream 614 for that particular presentation data structure.
[0105] The encoder may further be adapted to apply the predefined loudness function to each of the plurality of content substreams 502 to obtain substream-level loudness data for that content substream; and to include the substream-level loudness data in the bitstream. The predefined loudness function may relate to gating of the audio signal. According to other embodiments, the predefined loudness function may relate only to time segments of the audio signal that represent dialogue. According to some embodiments, the predefined loudness function may be: a frequency-dependent weighting of said audio signal; a channel-dependent weighting of said audio signal; ignoring segments of said audio signal having a signal power below a threshold; disregarding segments of said audio signal that are not detected as speech; · Calculation of at least one of energy / power / root mean square measures of said audio signal.
[0106] As can be seen from the above, the loudness function is non-linear. That is, if loudness data were only calculated from different content substreams, the loudness for a presentation could not be calculated by adding up the loudness data of the referenced content substreams. Furthermore, when different audio tracks, i.e. content substreams, are combined together for simultaneous playback, there may be combining effects between coherent / incoherent parts of different audio tracks or in different frequency domains, which further makes the addition of loudness data for audio tracks mathematically impossible.
[0107] IV. Equivalents, Extensions, Substitutions, and More Further embodiments of the present disclosure will be apparent to those skilled in the art after reviewing the above description. While the description and drawings disclose embodiments and examples, the present disclosure is not limited to such specific examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure, which is defined solely by the appended claims. Any reference signs appearing in the claims should not be construed as limiting the scope thereof.
[0108] Furthermore, variations to the disclosed embodiments can be understood and implemented by those skilled in the art in practicing the present disclosure, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprises" does not exclude other elements or steps, and the word "a" or "an" does not exclude a plurality. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
[0109] The above-disclosed apparatus and methods may be implemented as software, firmware, hardware, or a combination thereof. In a hardware implementation, the division of tasks among functional units referred to in the above description does not necessarily correspond to a division into physical units. Rather, a single physical component may have multiple functions, and a single task may be performed by several cooperating physical components. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, those skilled in the art will recognize that communication media typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0110] Several aspects will be described. [Aspect 1] 1. A method of processing a bitstream including multiple content substreams, each of the content substreams representing an audio signal, comprising: extracting one or more presentation data structures from the bitstream, each presentation data structure including a reference to one or more of the content substreams, each presentation data structure further including a reference to a metadata substream representing loudness data describing a combination of the referenced one or more content substreams; receiving data indicative of a selected one of the one or more presentation data structures and a desired loudness level; decoding the one or more content substreams referenced by the selected presentation data structure; forming an output audio signal based on the decoded content sub-stream; The method further includes processing the decoded one or more content substreams or the output audio signal to achieve the desired loudness level based on loudness data referenced by the selected presentation data structure. method. [Aspect 2] the selected presentation data structure references two or more content substreams and further references at least two mixing coefficients to be applied thereto; said forming an output audio signal further comprising additively mixing said decoded one or more content sub-streams by applying said mixing coefficient(s). The method of embodiment 1. Aspect 3 3. The method of embodiment 2, wherein the bitstream includes multiple time frames, and wherein the mixing coefficient(s) referenced by the selected presentation data structure are independently assignable for each time frame. Aspect 4 4. The method of aspect 2 or 3, wherein the selected presentation data structure references, for each substream of the two or more substreams, one mixing coefficient to be applied to the respective substream. Aspect 5 5. The method of any one of aspects 1 to 4, wherein the loudness data represents a value of a loudness function relating to the application of gating to the audio input signal. Aspect 6 6. The method of claim 5, wherein the loudness data represents values of a loudness function relating to time segments of the audio input signal that represent dialogue. Aspect 7 the presentation data structure further includes a reference to dynamic range compression (DRC) data for one or more referenced content substreams; The method further includes processing the decoded one or more content substreams or the output audio signal based on the DRC data, the processing including applying one or more DRC gains to the decoded one or more content substreams or the output audio signal. 7. The method of any one of embodiments 1 to 6. Aspect 8 8. The method of claim 7, wherein the DRC data includes at least one set of the one or more DRC gains. Aspect 9 The DRC data includes at least one compression curve, and the one or more DRC gains are: calculating one or more loudness values of the referenced one or more content substreams or the audio output signal using a predefined loudness function; obtained by mapping the one or more loudness values to a DRC gain using the compression curve. The method of embodiment 7. Aspect 10 10. The method of embodiment 9, wherein the mapping of loudness values includes a smoothing operation of the DRC gain. Aspect 11 11. The method of any one of aspects 7 to 10, wherein the referenced DRC data is included in the metadata substream. Aspect 12 12. The method of any one of aspects 1 to 11, wherein each of the decoded one or more content substreams includes loudness data at a substream level describing a loudness level of that content substream, and wherein the processing of the decoded one or more content substreams or the output audio signal further includes providing loudness consistency based on the loudness levels of the content substreams. Aspect 13 13. The method of claim 12, wherein the forming of the output audio signal includes combining two or more decoded content substreams using alternative mixing coefficients, and wherein the substream-level loudness data is used to compensate the loudness data to provide loudness consistency. Aspect 14 14. The method of embodiment 13, wherein the alternative mix coefficients relate to one of: dialogue enhancement and dialogue attenuation. Aspect 15 15. The method of any one of aspects 1 to 14, wherein the reference to at least one of the content substreams is a reference to at least one content substream group consisting of one or more of the content substreams. Aspect 16 The method of aspect 15, when aspect 15 cites aspect 2, wherein the selected presentation data structure, for a content substream group, references a single mixing coefficient to be applied to each of one or more of the content substreams that make up the substream group. Aspect 17 17. A method according to any one of aspects 1 to 16, wherein the bitstream includes multiple time frames, and data indicating the selected presentation data structure from the one or more presentation data structures is independently assignable for each time frame. Aspect 18 extracting from the bitstream one or more presentation data structures for a first of the plurality of time frames; extracting from the bitstream, for a second of the plurality of time frames, one or more presentation data structures that differ from the one or more presentation data structures extracted from the first of the plurality of time frames; the data indicating the selected presentation data structure indicates the selected presentation data structure for the time frame to which it is assigned; The method of embodiment 17. Aspect 19 19. The method of any one of aspects 1 to 18, wherein from the plurality of content substreams included in the bitstream, only the one or more content substreams referenced by the selected presentation data structure are decoded. Aspect 20 the bitstream includes two or more separate bitstreams, each including at least one of the plurality of content substreams, and decoding the one or more content substreams referenced by the selected presentation data structure includes: for each particular bitstream of the two or more separate bitstreams, separately decoding content substream(s) from referenced content substreams included in that particular bitstream; 20. The method of any one of embodiments 1 to 19. Aspect 21 1. A decoder for processing a bitstream including a plurality of content substreams, each representing an audio signal, comprising: a receiving component configured to receive the bitstream; a demultiplexer configured to extract one or more presentation data structures from the bitstream, each presentation data structure including a reference to at least one of the content substreams and further including a reference to a metadata substream representing loudness data describing a combination of the referenced one or more content substreams; a playback state component configured to receive data indicating a selected one of the one or more presentation data structures and a desired loudness level; a mixing component configured to decode the one or more content substreams referenced by the selected presentation data structure and to form an output audio signal based on the decoded content substreams; the mixing component is further configured to process the decoded one or more content substreams or the output audio signal to achieve the desired loudness level based on loudness data referenced by the selected presentation data structure. decoder. Aspect 22 An audio encoding method comprising: receiving a plurality of content substreams each representing an audio signal; defining one or more presentation data structures, each of which references at least one of the plurality of content substreams; applying a predefined loudness function to each of the one or more presentation data structures to obtain loudness data describing the combination of the referenced one or more content substreams, and including a reference (608) to the loudness data from the presentation data structure; forming a bitstream including the plurality of content substreams, the one or more presentation data structures, and the loudness data referenced by those presentation data structures; method. Aspect 23 determining, for each of the one or more presentation data structures, dynamic range compression (DRC) data for one or more referenced content substreams, the DRC data quantifying at least one desired compression curve or at least one set of DRC gains; and including the DRC data in the bitstream. 23. The method of embodiment 22. Aspect 24 applying the predefined loudness function to each of the plurality of content substreams to obtain substream-level loudness data for that content substream; and including loudness data at the substream level in the bitstream. 24. The method of embodiment 22 or 23. Aspect 25 25. The method of any one of aspects 22-24, wherein the predefined loudness function relates to gating of the audio signal. Aspect 26 26. The method of embodiment 25, wherein the predefined loudness function pertains only to time segments of the audio signal that represent dialogue. Aspect 27 The predefined loudness functions are: a frequency-dependent weighting of said audio signal; a channel-dependent weighting of said audio signal; ignoring segments of the audio signal having a signal power below a threshold; calculating an energy measure of the audio signal. 27. The method of any one of embodiments 22 to 26. Aspect 28 a loudness component configured to apply a predefined loudness function to obtain loudness data describing a combination of one or more content substreams representing a respective audio signal; a presentation data component configured to define one or more presentation data structures, each presentation data structure including references to one or more content substreams of a plurality of content substreams and references to loudness data describing a combination of the referenced content substreams; a multiplexing component configured to form a bitstream including the plurality of content substreams, the one or more presentation data structures, and the loudness data referenced by the one or more presentation data structures. Audio encoder. Aspect 29 28. A computer program product having a computer-readable medium having instructions for performing the method of any one of aspects 1 to 20 and 22 to 27.
Claims
1. obtaining the encoded bitstream by a decoding device; extracting, by the decoding device, an audio signal and metadata from the encoded bitstream, the metadata including compression curve data and loudness data, the loudness data indicating a loudness level of the audio signal; generating, by said decoding device, one or more loudness values using said loudness data; mapping, by the decoding device, the one or more loudness values to a dynamic range compression (DRC) gain using the compression curve data; smoothing the DRC gain by the decoding device; applying, by the decoding device, a smoothed DRC gain to the audio signal. method.
2. 2. The method of claim 1, wherein the encoded bitstream includes DRC data, the DRC data including a plurality of DRC profiles corresponding to DRC modes, each DRC profile tailored to a particular audio signal to which the DRC gain is applicable.
3. The method of claim 1 , wherein the loudness data comprises a loudness function comprising a channel-dependent weighting of the audio signal.
4. The method of claim 1 , wherein mapping the loudness values to DRC gains comprises discarding segments of the audio signal that are not detected as being speech.
5. one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations being: obtaining an encoded bitstream; extracting an audio signal and metadata from the encoded bitstream, the metadata including compression curve data and loudness data, the loudness data indicating a loudness level of the audio signal; generating one or more loudness values using the loudness data; mapping the one or more loudness values to a dynamic range compression (DRC) gain using the compression curve data; smoothing the DRC gain; applying a smoothed DRC gain to the audio signal. Decoding device.
6. 6. The decoding device of claim 5, wherein the encoded bitstream includes DRC data, the DRC data including a plurality of DRC profiles corresponding to DRC modes, each DRC profile tailored to a particular audio signal to which the DRC gain is applicable.
7. 6. Decoding device according to claim 5, wherein said loudness data comprises a loudness function comprising a channel-dependent weighting of said audio signal.
8. 6. The decoding apparatus of claim 5, wherein mapping the loudness values to DRC gains comprises discarding segments of the audio signal that are not detected as being speech.
9. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 4.
10. A computer program product for causing a computer to carry out the method according to any one of claims 1 to 4.