Audio encoder and decoder with program information or substream structure metadata

By embedding SSM and PIM in audio bitstreams, the patent addresses the issue of redundant processing in distributed audio systems, ensuring adaptive and high-quality audio processing and compliance with regulations.

JP2026021554APending Publication Date: 2026-02-10DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025191628
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-06-19
Filing Date
2025-11-12
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing audio data processing systems fail to account for the processing history of audio data across distributed networks, leading to unnecessary processing and degradation of audio content, particularly when multiple audio processing units are cascaded.

Method used

Incorporating substream structure metadata (SSM) and program information metadata (PIM) into audio bitstreams, allowing audio processing units to adapt their operations based on the processing history, thereby avoiding redundant processing and enhancing audio quality.

Benefits of technology

Enables efficient and adaptive audio processing across distributed systems, ensuring consistent audio quality and compliance with regulatory standards by utilizing metadata for error detection and correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021554000001_ABST
    Figure 2026021554000001_ABST
Patent Text Reader

Abstract

To provide an audio encoder and decoder with program information or substream structure metadata.SOLUTION: The decoder receives an encoded audio bitstream including substream structure metadata (SSM) and / or program information metadata (PIM) and audio data, stores at least one frame of the received audio bitstream, and extracts and decodes the SSM and / or PIM and the audio data from each frame.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 61 / 836,865, filed June 19, 2013, the contents of which are incorporated herein by reference in their entirety.

[0002] Technical Field The present invention relates to audio signal processing, and more particularly to encoding and decoding audio data bitstreams with metadata indicating substream structure and / or program information about the audio content represented by the bitstream. Some embodiments of the present invention generate or decode audio data in one of the formats known as Dolby Digital (AC-3), Dolby Digital Plus (Enhanced AC-3 or E-AC-3), or Dolby E. [Background technology]

[0003] Dolby, Dolby Digital, Dolby Digital Plus, and Dolby E are trademarks of Dolby Laboratories Licensing Corporation. Dolby Laboratories offers proprietary implementations of AC-3 and E-AC-3 known as Dolby Digital and Dolby Digital Plus, respectively.

[0004] Audio data processing units typically operate in a blind manner, paying no attention to the processing history of audio data that occurred before the data was received. This may work in a processing paradigm in which a single entity performs all audio data processing and encoding for various target media rendering devices, which in turn perform all decoding and rendering of the encoded audio data. However, this blind processing does not work well (or at all) in situations in which multiple audio processing units are distributed across various networks or arranged in cascade (i.e., chained) and are expected to optimally perform each type of audio processing. For example, some audio data may be encoded for a high-performance media system and may need to be converted along the media processing chain to a reduced form suitable for mobile devices. Thus, an audio processing unit may unnecessarily perform a type of processing on the audio data that has already been performed. For example, a volume leveling unit may perform processing on an input audio clip regardless of whether the same or similar volume leveling has previously been performed on the input audio clip. As a result, the volume leveling unit may perform leveling even when it is not necessary, and this unnecessary processing may result in the degradation and / or removal of certain characteristics when rendering the audio data content. Summary of the Invention [Means for solving the problem]

[0005] In one class of embodiments, the invention is an audio processing unit capable of decoding an encoded bitstream that includes substream structure metadata and / or program information metadata (and optionally other metadata, e.g., loudness processing state metadata) in at least one segment of at least one frame of the bitstream, and audio data in at least one other segment of said frame. As used herein, substream structure metadata (or "SSM") refers to metadata of an encoded bitstream (or collection of encoded bitstreams) that indicates the substream structure of the audio content of the encoded bitstreams, and "program information metadata" (or "PIM") refers to metadata of an encoded audio bitstream that indicates at least one audio program (e.g., two or more audio programs) and that indicates at least one attribute or characteristic of the audio content of at least one of said programs (e.g., metadata indicating the type or parameters of processing performed on the audio data of a program, or metadata indicating which channel of a program is the active channel).

[0006] Typically (e.g., when the encoded bitstream is an AC-3 or E-AC-3 bitstream), the program information metadata (PIM) indicates program information that cannot be practically carried in other parts of the bitstream. For example, the PIM may indicate the processing that was applied to the PCM audio prior to encoding (e.g., AC-3 or E-AC-3 encoding), which frequency bands of the audio program were encoded using a particular audio coding technique, and the compression profile used to generate the dynamic range compression (DRC) data in the bitstream.

[0007] In another class of embodiments, the method includes multiplexing encoded audio data in each frame (or each of at least some frames) of the bitstream with an SSM and / or PIM. In typical decoding, the decoder extracts the SSM and / or PIM from the bitstream (including by parsing and demultiplexing the SSM and / or PIM from the audio data) and processes the audio data to generate a stream of decoded audio data (possibly also performing adaptive processing of the audio data). In some embodiments, the decoded audio data and SSM and / or PIM are forwarded from the decoder to a post-processor configured to perform adaptive processing on the decoded audio data using the SSM and / or PIM.

[0008] In one class of embodiments, the encoding method of the present invention generates an encoded audio bitstream (e.g., an AC-3 or E-AC-3 bitstream) that includes an audio data segment containing encoded audio data (e.g., all or part of the AB0-AB5 segments of the frame shown in FIG. 4 or the AB0-AB5 segments of the frame shown in FIG. 7) and a metadata segment (containing an SSM and / or a PIM and, optionally, other metadata) time-division multiplexed with the audio data segment. In some embodiments, each metadata segment (sometimes referred to herein as a "container") has a format that includes a metadata segment header (and, optionally, other required or "core" elements) and one or more metadata payloads following the metadata segment header. The SIM, if present, is included in one of the metadata payloads (identified by a payload header and typically having a first type of format). The PIM, if present, is included in another of the metadata payloads (identified by a payload header and typically having a second type of format). Similarly, each other type of metadata, if present, is included in a separate one of the metadata payloads (identified by a payload header and typically having a format specific to that type of metadata). This exemplary format allows convenient access to the SSM, PIM, and other metadata outside of decoding (e.g., by a post-processor following decoding or by a processor configured to recognize metadata without performing a full decode on the encoded bitstream), and allows convenient and efficient error detection and correction (e.g., of substream identification) during bitstream decoding. For example, without access to the SSM in the exemplary format above, a decoder may incorrectly identify the correct number of substreams associated with a program.One metadata payload in a metadata segment may contain an SSM, another metadata payload in the metadata segment may contain a PIM, and optionally, at least one other metadata payload in the metadata segment may also contain other metadata (e.g., loudness processing state metadata or "LPSM"). [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram of an embodiment of a system that may be configured to perform an embodiment of the method of the present invention. [Figure 2] FIG. 2 is a block diagram of an encoder that is an embodiment of the audio processing unit of the present invention. [Figure 3] 2 is a block diagram of a decoder that is an embodiment of an audio processing unit of the present invention and a post-processor that is another embodiment of an audio processing unit of the present invention coupled thereto; [Figure 4] 1 is a diagram depicting an AC-3 frame, including the segments into which it is divided. [Figure 5] 1 illustrates the synchronization information (SI) segment of an AC-3 frame, including the segments into which it is divided. [Figure 6] 1 illustrates the bitstream information (BSI) segment of an AC-3 frame, including the segments into which it is divided. [Figure 7] 1 is a diagram depicting an E-AC-3 frame, including the segments into which it is divided. [Figure 8] 8 is a diagram of a metadata segment of an encoded bitstream generated in accordance with one embodiment of the present invention, including a metadata segment header containing a container sync word (identified in FIG. 8 as "container sync") and version and key ID values, followed by multiple metadata payloads and protection bits. DETAILED DESCRIPTION OF THE INVENTION

[0010] Notation and Nomenclature Throughout this disclosure, including the claims, the expression performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation either directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performing the operation).

[0011] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0012] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0013] Throughout this disclosure, including the claims, the terms "audio processor" and "audio processing unit" are used interchangeably and broadly to refer to a system configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools).

[0014] Throughout this disclosure, including the claims, the expression "metadata" (of an encoded audio bitstream) refers to data that is separate and distinct from the corresponding audio data in the bitstream.

[0015] Throughout this disclosure, including the claims, the expression "substream structure metadata" (or "SSM") refers to metadata of an encoded audio bitstream (or a collection of encoded audio bitstreams) that indicates the substream structure of the audio content of the encoded bitstream.

[0016] Throughout this disclosure, including the claims, the expression "program information metadata" (or "PIM") refers to metadata in an encoded audio bitstream that represents at least one audio program (e.g., two or more audio programs) and that indicates at least one attribute or characteristic of the audio content of at least one of said programs (e.g., metadata that indicates the type or parameters of processing that has been performed on the audio data of a program or metadata that indicates which channel of a program is the active channel).

[0017] Throughout this disclosure, including the claims, the term "processing state metadata" (e.g., as in the expression "loudness processing state metadata") refers to metadata (of an encoded audio bitstream) associated with audio data in the bitstream, indicating the processing state of the corresponding (associated) audio data (e.g., what type(s) of processing have already been performed on that audio data), and typically also indicating at least one feature or characteristic of that audio data. The association of processing state metadata with audio data is time-synchronous. In this manner, current (most recently received or updated) processing state metadata indicates that the corresponding audio data concurrently contains the results of the indicated type(s) of audio data processing. In some cases, the processing state metadata may include some or all of the processing history and / or parameters used in and / or derived from the indicated type of processing. Furthermore, the processing state metadata may include at least one feature or characteristic of the corresponding audio data calculated or extracted from the audio data. The processing state metadata may also include other metadata not related to or derived from any processing of the corresponding audio data, such as third-party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, etc., that may be added by a particular audio processing unit and passed to other audio processing units.

[0018] Throughout this disclosure, including the claims, the expression "loudness processing state metadata" (or "LPSM") refers to processing state metadata that indicates the loudness processing state of corresponding audio data (e.g., what type(s) of loudness processing have already been performed on the audio data), and typically also at least one feature or characteristic (e.g., loudness) of the corresponding audio data. The loudness processing state metadata may include data (e.g., other metadata) that is not (considered in isolation) loudness processing state metadata.

[0019] Throughout this disclosure, including the claims, the term "channel" (or "audio channel") refers to a monophonic audio signal.

[0020] Throughout this disclosure, including the claims, the term "audio program" refers to a collection of one or more audio channels and optionally associated metadata (e.g., metadata describing a desired spatial audio presentation and / or PIM and / or SSM and / or LPSM and / or program boundary metadata).

[0021] Throughout this disclosure, including the claims, the expression "program boundary metadata" refers to metadata of an encoded audio bitstream that indicates at least one audio program (e.g., two or more audio programs), where the program boundary metadata indicates the location in the bitstream of at least one boundary (beginning and / or end) of at least one of said audio programs. For example, program boundary metadata (of an encoded audio bitstream that indicates an audio program) may include metadata that indicates the location of the beginning of the program (e.g., the beginning of the Nth frame of the bitstream or the Mth sample position of the Nth frame of the bitstream) and additional metadata that indicates the location of the end of the program (e.g., the beginning of the Jth frame of the bitstream or the Kth sample position of the Jth frame of the bitstream).

[0022] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0023] Detailed Description of the Embodiments of the Invention A typical stream of audio data contains both audio content (e.g., one or more channels of audio content) and metadata that describes at least one characteristic of the audio content. For example, in an AC-3 bitstream, there are several audio metadata parameters that are specifically intended for use in altering the sound of the program delivered to a listening environment. One such metadata parameter is the DIALNORM parameter, which is intended to indicate the average level of dialogue in an audio program and is used to determine the audio playback signal level.

[0024] During playback of a bitstream containing a sequence of different audio program segments (each with a different DIALNORM parameter), an AC-3 decoder uses the DIALNORM parameter of each segment to perform some type of loudness processing to modify the playback level or loudness so that the perceived loudness of the dialogue for that sequence of segments is at a consistent level. Each encoded audio segment (item) in a sequence of encoded audio items will (generally) have a different DIALNORM parameter, and the decoder will scale the level of each item so that the playback level or loudness of the dialogue for each item is the same or very similar. However, this may require applying different amounts of gain to different items during playback.

[0025] DIALNORM is typically set by the user and is not automatically generated, although there is a default DIALNORM value if one is not set by the user. For example, a content creator may perform loudness measurements using equipment external to the AC-3 encoder and then forward the results (which indicate the loudness of the spoken dialogue in an audio program) to the encoder to set the DIALNORM value. Thus, we rely on the content creator to set the DIALNORM parameter correctly.

[0026] There are several different reasons why the DIALNORM parameter in an AC-3 bitstream may be incorrect. First, each AC-3 encoder has a default DIALNORM value that is used during bitstream generation if no DIALNORM value is set by the content creator. This default value may differ substantially from the actual dialogue loudness level of the audio. Second, even if the content creator measures loudness and sets the DIALNORM value accordingly, a loudness measurement algorithm or meter that does not follow the recommended AC-3 loudness measurement method may have been used, leading to an incorrect DIALNORM value. Third, even if an AC-3 bitstream was generated with a DIALNORM value that was correctly measured and set by the content creator, it may have been changed to an incorrect value during transmission and / or storage of the bitstream. For example, in television broadcast applications, it is not uncommon for an AC-3 bitstream to be decoded, modified, and then re-encoded using incorrect DIALNORM metadata information. Thus, the DIALNORM values ​​contained in the AC-3 bitstream may be incorrect or inaccurate, and thus may have a negative impact on the quality of the listening experience.

[0027] Furthermore, the DIALNORM parameter does not indicate the loudness processing state of the corresponding audio data (e.g., what type(s) of loudness processing have been performed on that audio data). Loudness processing state metadata (in the format provided in some embodiments of the present invention) is useful for facilitating adaptive loudness processing of audio bitstreams and / or verification of the loudness processing state and loudness effectiveness of audio content in a particularly efficient manner.

[0028] Although the present invention is not limited to use with AC-3 bitstreams, E-AC-3 bitstreams, or Dolby E bitstreams, for convenience it will be described in embodiments that generate, decode, or otherwise process such bitstreams.

[0029] An AC-3 encoded bitstream contains metadata and one to six channels of audio content. The audio content is audio data compressed using perceptual audio coding. The metadata includes several audio metadata parameters intended for use in altering the sound of the program delivered to the listening environment.

[0030] Each frame in an AC-3 encoded audio bitstream contains audio content and metadata for 1536 samples of digital audio. For a 48 kHz sampling rate, this represents 32 milliseconds of digital audio or an audio rate of 31.25 frames per second.

[0031] Each frame in an E-AC-3 encoded audio bitstream contains audio content and metadata for 256, 512, 768, or 1536 samples of digital audio, depending on whether the frame contains one, two, three, or six blocks of audio data. For a 48 kHz sampling rate, this represents 5.333, 10.667, 16, or 32 milliseconds of digital audio, respectively, or audio rates of 189.9, 93.75, 62.5, or 31.25 frames per second, respectively.

[0032] As shown in Figure 4, each AC-3 frame is divided into sections (segments): a synchronization information (SI) section containing a synchronization word (SW) and the first of two error correction words (CRC1); a bitstream information (BSI) section containing most of the metadata; six audio blocks (AB0 through AB5) containing the data-compressed audio content (and possibly metadata); a waste bit segment (W) (also known as the "skip field") containing any unused bits left after the audio content is compressed; an auxiliary (AUX) information section which may contain further metadata; and the second of two error correction words (CRC2).

[0033] As shown in Figure 7, each E-AC-3 frame is divided into sections (segments): a synchronization information (SI) section containing the synchronization word (SW) (as shown in Figure 5); a bitstream information (BSI) section containing most of the metadata; between one and six audio blocks (AB0 through AB5) containing the data-compressed audio content (and possibly metadata); a waste bit segment (W) (also known as a "skip field") containing any unused bits left after the audio content is compressed (although only one waste bit segment is shown, each audio block is typically followed by a different waste bit or skip field segment); an auxiliary (AUX) information section which may contain further metadata; and an error correction word (CRC).

[0034] In AC-3 (or E-AC-3) bitstreams, there are several audio metadata parameters that are specifically intended for use in altering the sound of the program delivered to the listening environment. One such metadata parameter is the DIALNORM parameter, which is contained in the BSI segment.

[0035] As shown in Figure 6, the BSI segment of an AC-3 frame includes a five-bit parameter ("DIALNORM") indicating the DIALNORM value for that program. If the audio coding mode ("acmod") of that AC-3 frame is "0," indicating a dual mono or "1+1" channel configuration, it also includes a five-bit parameter ("DIALNORM2") indicating the DIALNORM value for a second audio program carried in the same AC-3 frame.

[0036] The BSI segment includes a flag ("addbsie") indicating the presence (or absence) of additional bitstream information following the "addbsie" bit, a parameter ("addbsil") indicating the length of additional bitstream information, if any, following the "addbsil" value, and up to 64 bits of additional bitstream information ("addbsi") following the "addbsil" value.

[0037] The BSI segment may include other metadata values ​​not specifically shown in FIG.

[0038] According to one class of embodiments, an encoded audio bitstream represents multiple substreams of audio content. In some cases, the substreams represent the audio content of a multi-channel program, with each substream representing one or more channels of that program. In other cases, the multiple substreams of the encoded audio bitstream represent the audio content of several audio programs, typically a "main" audio program (which may be a multi-channel program) and at least one other audio program (e.g., a program that is commentary for the main audio program).

[0039] An encoded audio bitstream representing at least one audio program necessarily includes at least one "independent" substream of audio content, which represents at least one channel of the audio program (e.g., the independent substream may represent the five full-range channels of a typical 5.1 audio program), which audio program is herein referred to as the "main" program.

[0040] In some classes of embodiments, the encoded audio bitstream represents two or more audio programs (a "main" program and at least one other audio program). In such cases, the bitstream includes two or more independent substreams: a first independent substream representing at least one channel of the main program, and at least one other independent substream representing at least one channel of another audio program (a program different from the main program). Each independent substream is independently decodable, and a decoder is operable to decode only a subset (but not all) of the independent substreams of the encoded bitstream.

[0041] In a typical example of an encoded audio bitstream showing two independent substreams, one independent substream shows standard format speaker channels of a multichannel main program (e.g., the left, right, center, left surround, and right surround full-range speaker channels of a 5.1 channel main program), while the other independent substream shows monophonic audio commentary for the main program (e.g., director's commentary for a film if the main program is a film soundtrack). In another example of an encoded audio bitstream showing multiple independent substreams, one independent substream shows standard format speaker channels of a multichannel main program (e.g., a 5.1 channel main program) containing dialogue in a first language (e.g., one of the speaker channels of the main program may show the dialogue), while each of the other independent substreams shows a monophonic translation (into another language) of that dialogue.

[0042] Optionally, an encoded bitstream representing a main program (and optionally at least one other audio program) includes at least one "dependent" substream of audio content. Each dependent substream is associated with one independent substream of the bitstream and represents at least one additional channel of a program (e.g., the main program) whose content is represented by the associated independent substream. (That is, a dependent substream represents at least one channel of a program not represented by its associated independent substream, and the associated independent substream represents at least one channel of that program.) In an example of an encoded bitstream that includes an independent substream (representing at least one channel of a main program), the bitstream also includes a subordinate substream (associated with the independent bitstream) that represents one or more additional speaker channels of the main program. Such additional speaker channels are additional to the main program channel(s) represented by the independent substream. For example, if an independent substream represents the left, right, center, left surround, and right surround full-range speaker channels of a standard format for a 7.1-channel main program, a subordinate substream may represent two other full-range speaker channels of the main program.

[0043] According to the E-AC-3 standard, an E-AC-3 bitstream must represent at least one independent substream (e.g., a single AC-3 bitstream) and may represent up to eight independent substreams, each of which may be associated with up to eight dependent substreams.

[0044] An E-AC-3 bitstream includes metadata that indicates the substream structure of the bitstream. For example, the "chanmap" field in the Bitstream Information (BSI) section of an E-AC-3 bitstream determines the channel map for the program channels represented by the bitstream's subordinate substreams. However, the substream structure metadata is typically included in an E-AC-3 bitstream in a format that is convenient only for access and use by an E-AC-3 decoder (during decoding of the encoded E-AC-3 bitstream), and is not convenient for access and use after decoding (e.g., by a post-processor) or before decoding (e.g., by a processor configured to recognize such metadata). Furthermore, there is a risk that a decoder might misidentify substreams in a typical E-AC-3 encoded bitstream using such conventionally included metadata. Until the present invention, it was unknown how to include substream structure metadata in an encoded bitstream (e.g., an encoded E-AC-3 bitstream) in a format that allows convenient and efficient detection and correction of errors in substream identification during bitstream decoding.

[0045] An E-AC-3 bitstream may also include metadata about the audio content of an audio program. For example, an E-AC-3 bitstream representing an audio program may include metadata indicating the minimum and maximum frequencies at which spectral spreading processing (and channel-joint encoding) was used to encode the program's content. However, such metadata is typically included in the E-AC-3 bitstream in a format convenient only for access and use by an E-AC-3 decoder (during decoding of the encoded E-AC-3 bitstream), and not for access and use after decoding (e.g., by a post-processor) or before decoding (e.g., by a processor configured to recognize such metadata). Also, such metadata is not included in the E-AC-3 bitstream in a format that allows convenient and efficient error detection and correction of the identification of such metadata during decoding of the bitstream.

[0046] According to an exemplary embodiment of the present invention, the PIM and / or SSM (and optionally other metadata, e.g., loudness processing state metadata or "LPSM") are embedded in one or more reserved fields (or slots) of a metadata segment of an audio bitstream that also contains audio data in other segments (audio data segments). Typically, at least one segment of each frame of the bitstream contains a PIM or SSM, and at least one other segment of the frame contains corresponding audio data (i.e., audio data whose substream structure is indicated by the SSM and / or has at least one characteristic or attribute indicated by the PIM).

[0047] In one class of embodiments, each metadata segment is a data structure (sometimes referred to herein as a container) that may contain one or more metadata payloads. Each payload contains a header that contains a specific payload identifier (and payload configuration data) to provide an unambiguous indication of the type of metadata present in the payload. The order of payloads within a container is undefined; therefore, payloads can be stored in any order; a parser must be able to parse the entire container to extract meaningful payloads and ignore non-meaningful or unsupported payloads. Figure 8 (described below) illustrates the structure of such a container and the payloads within it.

[0048] Communicating metadata (e.g., SSM and / or PIM and / or LPSM) in an audio data processing chain is particularly useful when two or more audio processing units need to function cascaded with each other throughout the processing chain (or content lifecycle). Without including metadata in the audio bitstream, serious media processing problems such as quality, level, and spatial degradation can occur, for example, when two or more audio codecs are utilized in the chain and single-ended volume leveling is applied more than once along the bitstream's path to the media consumption device (or rendering point of the bitstream's audio content).

[0049] Loudness Processing State Metadata (LPSM) embedded in an audio bitstream according to some embodiments of the present invention may be authenticated and validated, for example, to allow a loudness regulatory entity to verify whether the loudness of a particular program is already within a specified range and that the corresponding audio data itself has not been modified (thereby ensuring compliance with applicable regulations). To verify this, instead of recalculating the loudness, the loudness value included in the data block containing the loudness processing state metadata may be read. In response to the LPSM, a regulatory authority may determine that the corresponding audio content (as indicated by the LPSM) complies with loudness legislation and / or regulatory requirements (e.g., regulations promulgated under the Commercial Advertisement Loudness Mitigation Act, also known as the "CALM Act") without having to calculate the loudness of the audio content.

[0050] 1 is a block diagram of an exemplary audio processing chain (audio data processing system), one or more of the system's elements may be configured in accordance with an embodiment of the present invention. The system includes the following elements coupled together as shown: a pre-processing unit, an encoder, a signal analysis and metadata correction unit, a transcoder, a decoder, and a pre-processing unit. Variations on the illustrated system may omit one or more of the elements or may include additional audio data processing units.

[0051] In some implementations, the preprocessing unit of FIG. 1 is configured to accept as input PCM (time-domain) samples containing audio content and to output processed PCM samples. An encoder may be configured to accept as input the PCM samples and to output an encoded (e.g., compressed) audio bitstream representing the audio content. The bitstream data representing the audio content is sometimes referred to herein as "audio data." When an encoder is configured according to exemplary embodiments of the present invention, the audio bitstream output from the encoder includes PIM and / or SSM (and optionally loudness processing state metadata and / or other metadata) in addition to the audio data.

[0052] The signal analysis and metadata correction unit of Figure 1 may accept one or more encoded audio bitstreams as input and determine (e.g., validate) whether the metadata (e.g., processing state metadata) in each encoded audio bitstream is correct by performing signal analysis (e.g., using program boundary metadata in the encoded audio bitstreams). If the signal analysis and metadata correction unit finds that the included metadata is invalid, it typically replaces the incorrect value(s) with the correct value(s) obtained from the signal analysis. In this manner, each encoded audio bitstream output from the signal analysis and metadata correction unit may include corrected (or uncorrected) processing state metadata in addition to the encoded audio data.

[0053] 1 may accept an encoded audio bitstream as input and, in response, output a modified (e.g., differently encoded) audio bitstream (e.g., by decoding the input stream and re-encoding the decoded stream in a different encoding format). When the transcoder is configured according to an exemplary embodiment of the present invention, the audio bitstream output from the transcoder includes the encoded audio data as well as an SSM and / or PIM (and typically other metadata), which may have been included in the input bitstream.

[0054] 1 may accept as input an encoded (e.g., compressed) bitstream and (in response) output a stream of decoded PCM audio samples. When the decoder is configured according to an exemplary embodiment of the present invention, the output of the decoder in exemplary operation may be or include any of the following: a stream of audio samples and at least one corresponding stream of SIM and / or PIM (and typically other metadata as well) extracted from the input encoded bitstream; or a stream of audio samples and a corresponding stream of control bits determined from the SSM and / or PIM extracted from the input encoded bitstream (and typically other metadata, e.g., LPSM as well); or a stream of audio samples without a corresponding stream of metadata or control bits determined from the metadata. In this last case, a decoder may extract metadata from the input encoded bitstream and perform at least one operation on the extracted metadata (e.g., validation) without outputting the extracted metadata or control bits determined from it.

[0055] 1 in accordance with an exemplary embodiment of the present invention, the post-processing unit is configured to accept a stream of decoded PCM audio samples and perform post-processing thereon (e.g., volume leveling of the audio content) using the SSM and / or PIM (and typically other metadata, e.g., LPSM) received with the samples or control bits determined by the decoder from the metadata received with the samples. The post-processing unit is also typically configured to render the post-processed audio content for playback over one or more speakers.

[0056] An exemplary embodiment of the present invention provides an improved audio processing chain in which audio processing units (e.g., encoders, decoders, transcoders, and pre-processing and post-processing units) adapt their respective processing applied to audio data according to the concurrent state of the media data as indicated by metadata received by each audio processing unit.

[0057] Audio data input to any audio processing unit of the system of Figure 1 (e.g., the encoder or transcoder of Figure 1) may include an SSM and / or a PIM (and optionally other metadata) in addition to the audio data (e.g., encoded audio data). According to some embodiments of the invention, this metadata may have been included in the input audio by other elements of the system of Figure 1 (or other sources not shown in Figure 1). This processing unit that receives the input audio (with metadata) may perform at least one action on (e.g., validity checking) or in response to (e.g., adaptive processing of the input audio) the metadata, and may also typically be configured to include in its output audio the metadata, a processed version of the metadata, or control bits determined from the metadata.

[0058] Exemplary embodiments of the audio processing unit (or audio processor) of the present invention are configured to perform adaptive processing of the audio data based on the state of the audio data indicated by metadata corresponding to the audio data. In some embodiments, the adaptive processing is (or includes) a loudness process (if the metadata indicates that no loudness or similar processing has already been performed on the audio data), but is not (or does not include) a loudness process (if the metadata indicates that such a loudness or similar processing has already been performed on the audio data). In some embodiments, the adaptive processing is or includes a metadata validation (e.g., performed in a metadata validation subunit) to ensure that the audio processing unit performs other adaptive processing of the audio data based on the state of the audio data indicated by the metadata. In some embodiments, the validation determines the reliability of metadata associated with the audio data (e.g., included in a bitstream together with the audio data). For example, if the metadata is validated as reliable, results from a previously performed audio processing of a certain type may be reused, and new execution of the same type of audio processing may be avoided. On the other hand, if the metadata is found to be tampered with (or otherwise untrustworthy), the type of media processing that was purportedly performed previously (as indicated by the untrustworthy metadata) may be repeated by the audio processing unit, and / or other processing may be performed by the audio processing unit on the metadata and / or audio data. If the audio processing unit determines that the metadata is valid (e.g., based on a match between the extracted cryptographic value and the reference cryptographic value), it may be configured to signal to other downstream audio processing units in the enhanced media processing chain that the metadata (e.g., present in the media bitstream) is valid.

[0059] 2 is a block diagram of an encoder (100) embodying an audio processing unit of the present invention. Any of the components or elements of encoder 100 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits) in hardware, software, or a combination of hardware and software. Encoder 100 includes, connected as shown, a frame buffer 110, a parser 111, a decoder 101, an audio state validator 102, a loudness processing stage 103, an audio stream selection stage 104, an encoder 105, a stuffer / formatter stage 107, a metadata generation stage 106, a dialogue loudness measurement subsystem 108, and a frame buffer 109. Typically, encoder 100 also includes other processing elements (not shown).

[0060] Encoder 100 (which is a transcoder) is configured to convert an input audio bitstream (which may be, for example, an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream) into an encoded output audio bitstream (which may be, for example, another AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream), including by performing adaptive and automated loudness processing using loudness processing state metadata included in the input bitstream. For example, encoder 100 may be configured to convert an input Dolby E bitstream (a format typically used in production and broadcast facilities, but not in consumer devices that receive broadcast audio programs) into an encoded output audio bitstream in AC-3 or E-AC-3 format (suitable for broadcast to consumer devices).

[0061] 2 also includes an encoded audio delivery subsystem 150 (which stores and / or delivers the encoded bitstream output from encoder 100) and a decoder 152. The encoded audio bitstream output from encoder 100 may be stored by subsystem 150 (e.g., in the form of a DVD or Blu-ray disc), transmitted by subsystem 150 (which may implement a transmission link or network), or both stored and transmitted by subsystem 150. Decoder 152 is configured to decode the encoded audio bitstream (produced by encoder 100) received via subsystem 150, including by extracting metadata (PIM and / or SSM and optionally loudness processing state metadata and / or other metadata) from each frame of the bitstream (optionally also extracting program boundary metadata from the bitstream) and generating decoded audio data. Typically, decoder 152 is configured to perform adaptive processing on the decoded audio data using the PIM and / or SSM and / or LPSM (and optionally also program boundary metadata) and / or forward the decoded audio data and metadata to a post-processor configured to perform adaptive processing on the decoded audio data using the metadata. Typically, decoder 152 includes a buffer that stores (e.g., non-temporarily) the encoded audio bitstream received from subsystem 150.

[0062] Various implementations of the encoder 100 and decoder 152 are configured to perform various embodiments of the method of the present invention.

[0063] Frame buffer 110 is a buffer memory coupled to receive the input encoded audio bitstream. In operation, buffer 110 stores (e.g., non-temporarily) at least one frame of the encoded audio bitstream, and a sequence of frames of the encoded audio bitstream is presented from buffer 110 to parser 111.

[0064] Parser 111 is coupled to and configured to extract PIM and / or SSM and loudness processing metadata (LPSM), and optionally also program boundary metadata (and / or other metadata), from each frame of encoded input audio containing such metadata, present at least the LPSM (and optionally also the program boundary metadata and / or other metadata) to audio state validity checker 102, loudness processing stage 103, stage 106 and subsystem 108, extract audio data from the encoded input audio, and present the audio data to decoder 101. Decoder 101 of encoder 100 is configured to decode the audio data to generate decoded audio data, and present the decoded audio data to loudness processing stage 103, audio stream selection stage 104, subsystem 108 and typically also to state validity checker 102.

[0065] The state validator 102 is configured to authenticate and validate the LPSM (and optionally other metadata) presented to it. In some embodiments, the LPSM is (or is included in) a data block that was included in the input bitstream (e.g., in accordance with an embodiment of the present invention). The block may include a cryptographic hash (a hash-based message authentication code or "HMAC") for processing the LPSM (and optionally other metadata) and / or the underlying audio data (provided to the validator 102 from the decoder 101). The data block may, in these embodiments, be digitally signed, allowing downstream audio processing units to relatively easily authenticate and validate the processing state metadata.

[0066] For example, an HMAC may be used to generate a digest, which may be included in the protection value(s) included in the bitstream of the present invention. For an AC-3 frame, the digest may be generated as follows: 1. After the AC-3 data and LPSM are encoded, the frame data bytes (concatenated Frame Data #1 and Frame Data #2) and the LPSM data bytes are used as input for the hash function HMAC. Other data that may be present in the ancillary data field are not taken into account for calculating this digest. Such other data may be bytes that do not belong to either AC-3 data or LSPSM data. Protection bits contained in the LPSM may not be taken into account for calculating the HMAC digest. 2. After the digest is calculated, it is written to the bitstream in the field reserved for the protection bits. 3. The final step in generating a complete AC-3 frame is the calculation of the CRC check, which is written at the very end of the frame and takes into account all data belonging to this frame, including the LPSM bits.

[0067] Other cryptographic methods, including, but not limited to, any one or more non-HMAC cryptographic methods, may be used to validate the LPSM and / or other metadata (e.g., in validator 102) to ensure secure transmission and receipt of the metadata and / or underlying audio data. For example, validation (using such cryptographic methods) may be performed at each audio processing unit receiving an audio bitstream embodiment of the present invention to determine whether the metadata and corresponding audio data included in the bitstream have been subjected to particular processing (and / or result from particular loudness processing) (as indicated by the metadata) and have not been modified since such particular processing was performed.

[0068] The state validator 102 presents control data to the audio stream selection stage 104, the metadata generator 106 and the dialogue loudness measurement subsystem 108 to indicate the result of the validation operation. In response to the control data, stage 104 can select (and communicate up to the encoder 105) either: the adaptively processed output of loudness processing stage 103 (e.g., when LPSM indicates that the audio data output from decoder 101 has not undergone a particular type of loudness processing and the control bit from validator 102 indicates that LPSM is enabled); or (For example, when LPSM indicates that the audio data output from decoder 101 has already undergone a particular type of loudness processing to be performed by stage 103 and the control bit from validity checker 102 indicates that LPSM is valid) said audio data output from decoder 101.

[0069] Stage 103 of encoder 100 is configured to perform adaptive loudness processing on the decoded audio data output from decoder 101 based on one or more audio data characteristics indicated by the LPSM extracted by decoder 101. Stage 103 may be an adaptive transform-domain real-time loudness and dynamic range control processor. Stage 103 may receive user input (e.g., user target loudness / dynamic range values ​​or dialnorm values) or other metadata input (e.g., one or more types of third-party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, etc.) and / or other input (e.g., from a fingerprinting process) and use such input to process the decoded audio data output from decoder 101. Stage 103 may perform adaptive loudness processing on decoded audio data (output from decoder 101) that indicates a single audio program (indicated by the program boundary metadata extracted by parser 111), and may reset the loudness processing in response to receiving decoded audio data (output from decoder 101) that indicates a different audio program as indicated by the program boundary metadata extracted by parser 111.

[0070] Dialogue loudness measurement subsystem 108 may operate to determine the loudness of segments of decoded audio (from decoder 101) indicating dialogue (or other speech) using, for example, the LPSM (and / or other metadata) extracted by decoder 101 if the control bit from validity checker 102 indicates that LPSM is disabled. If the control bit from validity checker 102 indicates that LPSM is enabled, operation of dialogue loudness measurement subsystem 108 may be disabled when the LPSM indicates a previously determined loudness of a dialogue (or other speech) segment of decoded audio (from decoder 101). Subsystem 108 may perform loudness measurements on decoded audio data indicating a single audio program (indicated by program boundary metadata extracted by parser 111) and may reset the measurements in response to receiving decoded audio data indicating a different audio program as indicated by such program boundary metadata.

[0071] There are useful tools (e.g., a Dolby LM100 loudness meter) for conveniently and easily measuring the level of dialogue in audio content. Some embodiments of the APU (e.g., stage 108 of encoder 100) of the present invention are implemented to include such a tool (or to perform the functionality of such a tool) to measure the average dialogue loudness of the audio content of an audio bitstream (e.g., a decoded AC-3 bitstream presented to stage 108 from decoder 101 of encoder 100).

[0072] If stage 108 is implemented to measure the true average dialogue loudness of the audio data, the measurement may include isolating segments of the audio content that contain primarily speech. The primarily speech audio segments are then processed according to a loudness measurement algorithm. For audio data decoded from an AC-3 bitstream, this algorithm may be the standard K-weighted loudness measure (in accordance with international standard ITU-R BS.1770). Alternatively, other loudness measures (e.g., based on psychoacoustic models of loudness) may be used.

[0073] Isolating speech segments is not essential for measuring the average dialogue loudness of audio data, but it improves the accuracy of the metric and typically gives more satisfying results from the listener's perspective. Because not all audio content contains dialogue, a loudness metric for the entire audio content may provide a good approximation of the dialogue level of that audio if speech were present.

[0074] Metadata generator 106 generates (and / or passes to stage 107) metadata to be included by stage 107 in the encoded bitstream output from encoder 100. Metadata generator 106 may pass to stage 107 LPSMs (and optionally LIMs and / or PIMs and / or program boundary metadata and / or other metadata) extracted by encoder 101 and / or parser 111 (e.g., if a control bit from validity checker 102 indicates that the LPSMs and / or other metadata are valid), or may generate new LIMs and / or PIMs and / or LPSMs and / or program boundary metadata and / or other metadata and present the new metadata to stage 107 (e.g., if a control bit from validity checker 102 indicates that the metadata extracted by decoder 101 is invalid). Alternatively, metadata generator 106 may present to stage 107 a combination of the metadata extracted by decoder 101 and / or parser 111 and the newly generated metadata. Metadata generator 106 may include the loudness data generated by subsystem 108 and at least one value indicating the type of loudness processing performed by subsystem 108 in an LPSM that is submitted to stage 107 for inclusion in the encoded bitstream output from encoder 100.

[0075] Metadata generator 106 may generate protection bits (which may consist of or include a hash-based message authentication code or "HMAC") useful for at least one of decrypting, authenticating, or validating the LPSM (and optionally other metadata) to be included in the encoded bitstream and / or the underlying audio data to be included in the encoded bitstream. Metadata generator 106 may provide such protection bits to stage 107 for inclusion in the encoded bitstream.

[0076] In typical operation, the dialogue loudness measurement subsystem 108 processes the audio data output from the decoder 101 and, in response, generates loudness values ​​(e.g., gated and ungated dialogue loudness values) and dynamic range values. In response to these values, the metadata generator 106 may generate loudness processing state metadata (LPSM) for inclusion (by the stuffer / formatter 107) in the encoded bitstream output from the encoder 100.

[0077] Additionally, optionally, or alternatively, subsystems 106 and / or 108 of encoder 100 may perform additional analysis of the audio data to generate metadata indicative of at least one characteristic of the audio data for inclusion in the encoded bitstream output from stage 107.

[0078] Encoder 105 encodes the audio data output from selection stage 104 (e.g., by performing compression on it) and presents the encoded audio to stage 107 for inclusion in an encoded bitstream output from stage 107.

[0079] Stage 107 multiplexes the encoded audio from encoder 105 and the metadata (including PIM and / or SSM) from generator 106 to produce an encoded bitstream that is output from stage 107. Preferably, the encoded bitstream has a format specified by a preferred embodiment of the present invention.

[0080] Frame buffer 109 is a buffer memory that stores (e.g., non-temporarily) at least one frame of the encoded audio bitstream output from stage 107. The sequence of those frames of the encoded audio bitstream is then presented from buffer 109 to delivery system 150 as output from encoder 100.

[0081] The LPSM generated by metadata generator 106 and included in the encoded bitstream by stage 107 typically indicates the loudness processing state of the corresponding audio data (e.g., what type(s) of loudness processing have been performed on the audio data) and the loudness of the corresponding audio data (e.g., measured dialogue loudness, gated and / or ungated loudness and / or dynamic range).

[0082] In this document, "gating" loudness and / or level measurements performed on audio data refers to a specific level or loudness threshold, whereby calculated value(s) above the threshold are included in the final measurement (e.g., ignoring short-term loudness values ​​below -60 dBFS in the final measured value). Gating on an absolute value refers to a fixed level or loudness, while gating on a relative value refers to a value that is dependent on the current "ungated" measurement.

[0083] In some implementations of encoder 100, the encoded bitstream buffered in memory 109 (and output to delivery system 150) is an AC-3 or E-AC-3 bitstream and includes audio data segments (e.g., segments AB0-AB5 of the frame shown in FIG. 4) and metadata segments, where the audio data segments represent audio data and at least some of the metadata segments each include a PIM and / or SSM (and optionally other metadata). Stage 107 inserts the metadata segments (including the metadata) into the bitstream in the following format: Each metadata segment including a PIM and / or SSM is included in an extra bits segment of the bitstream (e.g., extra bits segment "W" shown in FIG. 4 or FIG. 7), or in the "addbsi" field of a bitstream information ("BSI") segment of a frame of the bitstream, or in an auxiliary data field at the end of a frame of the bitstream (e.g., the AUX segment shown in FIG. 4 or FIG. 7). A frame of the bitstream may contain one or two metadata segments, each containing metadata; if a frame contains two metadata segments, one may be present in the addbsi field of the frame and the other in the AUX field of the frame.

[0084] In some embodiments, each metadata segment (sometimes referred to herein as a "container") inserted by stage 107 has a format that includes a metadata segment header (optionally as well as other required or "core" elements) and one or more metadata payloads following the metadata segment header. The SIM, if present, is included in one of the metadata payloads (identified by a payload header and typically having a first type of format). The PIM, if present, is included in another one of the metadata payloads (identified by a payload header and typically having a second type of format). Similarly, each other type of metadata (if present) is included in another one of the metadata payloads (identified by a payload header and typically having a format specific to the metadata type). This exemplary format allows convenient access to the SSM, PIM, and other metadata at times other than during decoding (e.g., after decoding by a post-processor or by a processor configured to recognize that metadata, without performing a full decode on the encoded bitstream), allowing convenient and efficient error detection and correction (e.g., of substream identification) during bitstream decoding. For example, without access to the SSMs in this exemplary format, a decoder may incorrectly identify the correct number of substreams associated with a program. One metadata payload in a metadata segment may contain the SSM, another metadata payload in the metadata segment may contain the PIM, and optionally, at least one other metadata payload in the metadata segment may also contain other metadata (e.g., loudness processing state metadata or "LPSM").

[0085] In some embodiments, a Substream Structure Metadata (SSM) payload included (by stage 107) within a frame of an encoded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) includes an SSM in the following format: a payload header, which typically includes at least one identification value (e.g., a 2-bit value indicating the SSM format version and, optionally, length, period, count, and substream association values); After the header, Independent substream metadata indicating the number of independent substreams of the program represented by the bitstream; and Dependent substream metadata indicating whether each independent substream of a program has at least one associated dependent substream (i.e., whether each independent substream has at least one associated dependent substream), and if so, the number of dependent substreams associated with each independent substream of the program.

[0086] It is contemplated that an independent substream of an encoded bitstream may represent a set of speaker channels of an audio program (e.g., the speaker channels of a 5.1 speaker channel audio program), and that one or more dependent substreams (associated with said independent substream as indicated by dependent substream metadata) may each represent an object channel of the program. Typically, however, an independent substream of an encoded bitstream represents a set of speaker channels of a program, and each dependent substream associated with that independent substream (as indicated by dependent substream metadata) represents at least one additional speaker channel of the program.

[0087] In some embodiments, the program information metadata (PIM) payload included (by stage 107) within frames of the encoded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) has the following format: a payload header, which typically includes at least one identifying value (e.g., a value indicating the PIM format version and, optionally, length, period, count, and substream association values); and After the header, a PIM in the following format: Active channel metadata indicating each silent and non-silent channel of an audio program (i.e., which channels of the program contain audio information and which channels, if any, contain only silence (typically for the duration of that frame)). In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the active channel metadata in a frame of the bitstream may be used in conjunction with additional metadata of the bitstream (e.g., the audio coding mode ("acmod") field of that frame and the chanmap field, if present, in that frame or associated dependent substream frame(s)) to determine which channels of the program contain audio information and which channels contain silence. The "acmod" field in an AC-3 or E-AC-3 frame indicates the number of full-range channels of the audio program represented by the audio content of that frame (e.g., whether the program is a 1.0-channel monophonic program, a 2.0-channel stereo program, or a program containing L, R, C, Ls, and Rs full-range channels), or indicates that the frame represents two independent 1.0-channel monophonic programs. The "chanmap" field in an E-AC-3 bitstream indicates the channel map for dependent substreams represented by the bitstream. Active channel metadata can be useful for implementing upmixing downstream of the decoder (in a post-processor), for example, to add audio to channels containing silence at the decoder's output;

[0088] Down-mixing processing state metadata indicating whether a program was down-mixed (before or during encoding) and, if so, the type of down-mixing that was applied. The down-mixing processing state metadata may be useful for implementing up-mixing downstream of the decoder (in a post-processor), for example, to up-mix the audio content of the program using parameters that best match the type of down-mixing that was applied. In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the down-mixing processing state metadata may be used in conjunction with the audio coding mode ("acmod") field of a frame to determine the type of down-mixing, if any, that was applied to the channels of the program;

[0089] Upmixing process state metadata indicating whether a program was upmixed (e.g., from fewer channels) before or during encoding and, if so, the type of upmixing that was applied. The upmixing process state metadata may be useful for implementing downmixing downstream of the decoder (in a post-processor), for example, to downmix the audio content of a program in a manner that is compatible with the type of upmixing applied to the program (e.g., Dolby Pro Logic or Dolby Pro Logic II Movie mode or Dolby Pro Logic II Music mode or Dolby Professional Upmixer). In embodiments where the encoded bitstream is an E-AC-3 bitstream, the upmixing process state metadata may be used in conjunction with other metadata (e.g., the value of the 'strmtyp' field of the frame) to determine the type of upmixing (if any) that was applied to the channels of the program. The value of the "strmtyp" field (in the BSI segment of a frame of an E-AC-3 bitstream) indicates whether the audio content of the frame belongs to an independent stream (which determines a program) or an independent substream (of a program that contains or is associated with multiple substreams) and therefore can be decoded independently of any other substream represented by that E-AC-3 bitstream, or whether the audio content of the frame belongs to a dependent substream (of a program that contains or is associated with multiple substreams) and therefore needs to be decoded in conjunction with the associated independent substream;

[0090] Preprocessing state metadata indicating whether preprocessing was performed on the audio content of this frame (prior to encoding the audio content to produce the encoded bitstream) and, if so, the type of preprocessing that was performed.

[0091] In some implementations, the preprocessing state metadata indicates: whether surround attenuation was applied (for example, whether the surround channels of an audio program were attenuated by 3 dB prior to encoding); whether a 90 degree phase shift was applied (for example, to the surround Ls and Rs channels of an audio program prior to encoding), Whether a low-pass filter was applied to the LFE channel of the audio program prior to encoding, Whether the level of the LFE channel of the Program was monitored during production and, if so, the monitored level of the LFE channel relative to the levels of the Program's full-range audio channels.

[0092] Whether dynamic range compression should be performed (e.g., in a decoder) for each block of the decoded audio content of the program, and if so, the type (and / or parameters) of dynamic range compression to be performed (e.g., preprocessing state metadata of this type may indicate which of the following compression profile types was assumed by the encoder to generate the dynamic range compression control values ​​to be included in the encoded bitstream: film standard, film light, music standard, music light, or speech. Alternatively, preprocessing state metadata of this type may indicate that heavy dynamic range compression ('compr' compression) should be performed for each frame of the decoded audio content of the program, in a manner determined by the dynamic range compression control values ​​to be included in the encoded bitstream).

[0093] Whether spectral extension processing and / or channel-combining encoding was used to encode a particular frequency range of the program content, and if so, the minimum and maximum frequencies of the frequency components of the content that underwent spectral extension encoding and the minimum and maximum frequencies of the frequency components of the content that underwent channel-combining encoding. This type of preprocessing state metadata information may be useful for performing equalization downstream of the decoder (in a post-processor). Both channel-combining and spectral extension information are also useful for optimizing quality during transcoding operations and applications. For example, an encoder may optimize its behavior (including adaptation of preprocessing stages such as headphone virtualization, upmixing, etc.) based on the state of parameters such as spectral extension and channel-combining information. Furthermore, an encoder may dynamically adapt its combining and spectral extension parameters to match and / or to optimal values ​​based on the state of the incoming (and authenticated) metadata.

[0094] Whether dialogue enhancement adjustment range data is included in the encoded bitstream and, if so, the range of adjustment available during dialogue enhancement processing (e.g., in a post-processor downstream of the decoder) to adjust the level of dialogue content relative to the level of non-dialogue content in the audio program.

[0095] In some implementations, additional pre-processing state metadata (e.g., metadata indicating headphone-related parameters) is included (by stage 107) in the PIM payload of the encoded bitstream output from encoder 100.

[0096] In some embodiments, the LPSM payload included (by stage 107) in a frame of an encoded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) includes an LPSM in the following format: a header (typically containing a synchronization word identifying the beginning of the LPSM payload, followed by at least one identification value, such as the LPSM format version, length, period, count, and sub-stream association values ​​shown in Table 2 below); After the header, at least one dialogue indication value (e.g., parameter "Dialogue Channel" in Table 2) indicating whether the corresponding audio data indicates dialogue or not (e.g., which channel of the corresponding audio data indicates dialogue); at least one loudness regulation compliance value indicating whether the corresponding audio data conforms to an indicated set of loudness regulations (e.g., the parameter "Loudness Regulation Type" in Table 2); At least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data (e.g., one or more of the parameters "Dialogue-Gated Loudness Compensation Flag" and "Loudness Compensation Type" in Table 2); and At least one loudness value indicating at least one loudness (e.g., peak or average loudness) characteristic of the corresponding audio data (e.g., one or more of the parameters "ITU relative gated loudness", "ITU speech-gated loudness", "ITU (EBU3341) short-term 3s loudness", and "true peak").

[0097] In some embodiments, each metadata segment that includes a PIM and / or SSM (and optionally other metadata) includes a metadata segment header (and optionally additional core elements) and, following the metadata segment header (or following the metadata segment header and other core elements), at least one metadata payload segment having the following format: A payload header, which typically contains at least one identifying information value (e.g., SSM or PIM format version, length, period, count, and substream association values); After the payload header, there is the SSM or PIM (or other type of metadata).

[0098] In some implementations, each of the metadata segments inserted by stage 107 into the extra bits / skip field segments (or "addbsi" fields or ancillary data fields) of frames of the bitstream has the following format: a metadata segment header (typically containing a sync word identifying the start of the metadata segment, followed by identifying information values ​​such as version, length, period, extension element count, and substream association values ​​as shown in Table 1 below); and after the metadata segment header, at least one protection value useful for at least one of decryption, authentication, or validation of at least one of the metadata of the metadata segment or the corresponding audio data (e.g., the HMAC digest and audio fingerprint values ​​of Table 1); and Also following the metadata segment header are metadata payload identification information ("ID") and payload configuration values ​​that identify the type of metadata in each subsequent metadata payload and indicate at least one aspect of the configuration of each such payload (e.g., size).

[0099] Each metadata payload is followed by a corresponding payload ID and payload configuration value.

[0100] In some embodiments, each metadata segment in the redundant bits segment (or ancillary data field or "addbsi" field) of a frame has a three-level structure: a high-level structure (e.g., a metadata segment header), which contains a flag indicating whether the extra bits (or ancillary data or addbsi) field contains metadata, at least one ID value indicating what type(s) of metadata are present, and typically also a value indicating how many bits of metadata (e.g., of each type) are present (if metadata is present). One type of metadata that may be present is PIM, another type of metadata that may be present is SSM, and other types of metadata that may be present are LPSM and / or program boundary metadata and / or media research metadata; a mid-level structure, which contains data related to each identified type of metadata (e.g., metadata payload header, protection value, and payload ID and payload configuration value for each identified type of metadata); and A low-level structure, which is a metadata payload for each identified type of metadata (e.g., a sequence of PIM values ​​if PIM is identified as present and / or other types of metadata values ​​(e.g., SSM or LPSM) if other types of metadata are identified as present).

[0101] The data values ​​in such a three-level structure may be nested. For example, the protection value(s) for each payload (e.g., each PIM or SSM or other metadata payload) identified by the high-level and mid-level structures may be included after the payload (and thus after the metadata payload header for that payload), and the protection value(s) for all metadata payloads identified by the high-level and mid-level structures may be included after the final metadata payload in the metadata segment (and thus after the metadata payload headers for all payloads in that metadata segment).

[0102] In one example (described below with reference to the metadata segment or "container" in Figure 8), the metadata segment header identifies four metadata payloads. As shown in Figure 8, the metadata segment header includes a container sync word (identified as "container sync") and version and key ID values. Following the metadata segment header are the four metadata payloads and protection bits. The metadata segment header is followed by a payload ID and payload configuration (e.g., payload size) values ​​for a first payload (e.g., a PIM payload), followed by the first payload itself, followed by a payload ID and payload configuration (e.g., payload size) values ​​for a second payload (e.g., an SSM payload), followed by the first payload, followed by the second payload itself, followed by a payload ID and payload configuration (e.g., payload size) values ​​for a third payload (e.g., an LPSM payload), followed by the second payload, followed by the third payload itself, followed by these IDs and configuration values, followed by a payload ID and payload configuration (e.g., payload size) values ​​for a fourth payload, followed by the third payload, followed by the fourth payload itself, followed by these IDs and configuration values, and finally by a protection value(s) (identified in FIG. 8 as "protected data") for all or part of the payload (or for all or part of the payload for high-level and mid-level structures).

[0103] In some embodiments, when decoder 101 receives an audio bitstream generated in accordance with an embodiment of the present invention that has a cryptographic hash, the decoder is configured to parse and extract the cryptographic hash from a determined data block from the bitstream, the block including metadata. Validator 102 may use the cryptographic hash to validate the received bitstream and / or the associated metadata. For example, if validator 102 finds the metadata to be valid based on a match between a reference cryptographic hash and the cryptographic hash extracted from the data block, validator 102 may disable processor 103's operation on the corresponding audio data and may pass the audio data through selection stage 104 (unaltered). Additionally, optionally, or alternatively, other types of cryptographic techniques may be used instead of cryptographic hash-based methods.

[0104] 2 may determine (in response to the LPSM extracted by decoder 101, and optionally also to program boundary metadata) that a post / pre-processing unit has performed some type of loudness processing on the audio data to be encoded (in elements 105, 106, and 107), and may therefore generate (in generator 106) loudness processing state metadata that includes particular parameters used in and / or derived from the previously performed loudness processing. In some implementations, encoder 100 may generate (and include in the encoded bitstream output therefrom) metadata indicative of the processing history on the audio content, so long as the encoder recognizes the type of processing that has been performed on the audio content.

[0105] FIG. 3 is a block diagram of a decoder (200) and a post-processor (300) coupled thereto, which are an embodiment of an audio processing unit of the present invention. The post-processor (300) is also an embodiment of an audio processing unit of the present invention. Any of the components or elements of the decoder 200 and the post-processor 300 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. The decoder 200 includes a frame buffer 201, a parser 205, an audio decoder 202, an audio state validity check stage (validator) 203, and a control bit generation stage 204, connected as shown. Typically, the decoder 200 also includes other processing elements (not shown).

[0106] The frame buffer 201 (buffer memory) stores (e.g., non-temporarily) at least one frame of the encoded audio bitstream received by the decoder 200. A sequence of frames of the encoded audio bitstream is presented from the buffer 201 to a parser 205.

[0107] Parser 205 is coupled and configured to extract the PIM and / or SSM (and optionally other metadata, e.g., LPSM) from each frame of the encoded input audio, provide at least a portion of the metadata (e.g., LPSM and program boundary metadata (if extracted) and / or PIM and / or SSM) to audio state validator 203 and stage 204, provide the extracted metadata as output (e.g., to post-processor 300), extract audio data from the encoded input audio, and provide the extracted audio data to decoder 202.

[0108] The encoded audio bitstream input to the decoder 200 may be one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream.

[0109] The system of Figure 3 also includes a post-processor 300. Post-processor 300 has a frame buffer 301 and other processing elements (not shown), including at least one processing element coupled to buffer 301. Frame buffer 301 stores (e.g., non-temporarily) at least one frame of the decoded audio bitstream received by post-processor 300 from decoder 200. The processing elements of post-processor 300 are coupled and configured to receive the sequence of frames of the decoded audio bitstream output from buffer 301 and to adaptively process them using the metadata output from decoder 200 and / or the control bits output from stage 204 of decoder 200. Typically, the post-processor 300 is configured to perform adaptive loudness processing on the decoded audio data using metadata from the decoder 200 (e.g., adaptive loudness processing on the encoded audio data using LPSM values ​​and optionally program boundary metadata, where the adaptive processing may be based on loudness processing states and / or one or more audio characteristics indicated by the LPSM for audio data representing a single audio program).

[0110] Various implementations of the decoder 200 and post-processor 300 are configured to perform various embodiments of the method of the present invention.

[0111] The audio decoder 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 to generate decoded audio data and to provide the decoded audio data as output (e.g., to a post-processor 300).

[0112] State validator 203 is configured to authenticate and validate metadata presented to it. In some embodiments, the metadata is (or is contained in) a data block included in the input bitstream (e.g., in accordance with an embodiment of the present invention). The block may include a cryptographic hash (a hash-based message authentication code or "HMAC") for processing the metadata and / or the underlying audio data (provided to validator 203 from parser 205 and / or decoder 202). The data block may, in these embodiments, be digitally signed, allowing downstream audio processing units to relatively easily authenticate and validate the processing state metadata.

[0113] Other cryptographic methods, including but not limited to any one or more non-HMAC cryptographic methods, may be used to validate the metadata (e.g., in validator 203) to ensure secure transmission and reception of the metadata and / or underlying audio data. For example, validation (using such cryptographic methods) may be performed at each audio processing unit receiving an embodiment of an audio bitstream of the present invention to determine whether the loudness processing state metadata and corresponding audio data included in the bitstream have been subjected to (and / or result from) a particular loudness processing (as indicated by the metadata) and have not been modified since such particular loudness processing was performed.

[0114] State validator 203 provides control data to control bit generator 204 and / or provides the control data as output (e.g., to post-processor 300) to indicate the result of the validation operation. In response to the control data (and optionally other metadata extracted from the input bitstream), stage 204 may generate (and provide to post-processor 300) any of the following: a control bit indicating that the decoded audio data output from decoder 202 has undergone a particular type of loudness processing (e.g., when LPSM indicates that the audio data output from decoder 202 has undergone a particular type of loudness processing and the control bit from validator 203 indicates that LPSM is enabled); or A control bit indicating that the decoded audio data output from decoder 202 should be subjected to a particular type of loudness processing (for example, when LPSM indicates that the audio data output from decoder 202 has not undergone a particular type of loudness processing, or when LPSM indicates that the audio data output from decoder 202 has undergone a particular type of loudness processing but the control bit from validity checker 203 indicates that LPSM is not valid).

[0115] Alternatively, the decoder 200 presents the metadata extracted by the decoder 202 from the input bitstream and the metadata extracted by the parser 205 from the input bitstream to the post-processor 300, which uses the metadata to perform adaptive processing on the decoded audio data or performs a validation check of the metadata and then, if the validation check indicates that the LPSM is valid, uses the metadata to perform adaptive processing on the decoded audio data.

[0116] In some embodiments, when decoder 200 receives an audio bitstream generated in accordance with an embodiment of the present invention that has a cryptographic hash, the decoder is configured to parse and extract the cryptographic hash from a determined data block from the bitstream. The block includes loudness processing state metadata (LPSM). Validator 203 may use the cryptographic hash to validate the received bitstream and / or associated metadata. For example, if validator 203 finds the LPSM to be valid based on a match between a reference cryptographic hash and the cryptographic hash extracted from the data block, validator 203 may signal a downstream audio processing unit (e.g., post-processor 300, which may be or include a volume leveling unit) to pass the audio data of the bitstream through (without modification). Additionally, optionally, or alternatively, other types of cryptographic techniques may be used instead of cryptographic hash-based methods.

[0117] In some implementations of decoder 200, the received encoded bitstream (and buffered in memory 201) is an AC-3 or E-AC-3 bitstream, and includes audio data segments (e.g., segments AB0-AB5 of the frame shown in FIG. 4) and metadata segments, where the audio data segments represent audio data, and at least some of the metadata segments each include a PIM or SSM (or other metadata). Decoder stage 202 (and / or parser 205) is configured to extract the metadata from the bitstream. Each metadata segment containing a PIM and / or SSM (and optionally other metadata) is included in an extra bits segment of a frame of the bitstream, an "addbsi" field of a bitstream information ("BSI") segment of a frame of the bitstream, or an auxiliary data field at the end of a frame of the bitstream (e.g., the AUX segment shown in FIG. 4). A frame of the bitstream may contain one or two metadata segments, each containing metadata; if a frame contains two metadata segments, one may be present in the addbsi field of the frame and the other may be present in the AUX field of the frame.

[0118] In some embodiments, each metadata segment (sometimes referred to herein as a "container") of a bitstream buffered in buffer 201 has a format that includes a metadata segment header (and optionally other required or "core" elements) followed by one or more metadata payloads. The SIM, if present, is included in one of the metadata payloads (identified by a payload header and typically having a first type of format). The PIM, if present, is included in another of the metadata payloads (identified by a payload header and typically having a second type of format). Similarly, each of the other types of metadata, if present, is included in another of the metadata payloads (identified by a payload header and typically having a format specific to the metadata type). This exemplary format allows convenient access to the SSM, PIM, and other metadata at times other than during decoding (e.g., by post-processor 300 following decoding or by a processor configured to recognize metadata without performing a full decode on the encoded bitstream), allowing convenient and efficient error detection and correction (e.g., of substream identification) during bitstream decoding. For example, without access to the SSM in the exemplary format, decoder 200 may incorrectly identify the correct number of substreams associated with a program. One metadata payload in a metadata segment may include an SSM, another metadata payload in the metadata segment may include a PIM, and optionally, at least one other metadata payload in the metadata segment may also include other metadata (e.g., loudness processing state metadata or "LPSM").

[0119] In some embodiments, a substream structure metadata (SSM) payload included within a frame of an encoded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) buffered in buffer 201 includes an SSM in the following format: a payload header, which typically includes at least one identification value (e.g., a 2-bit value indicating the SSM format version and, optionally, length, period, count, and substream association values); After the header, Independent substream metadata indicating the number of independent substreams of the program represented by the bitstream; and Dependent substream metadata indicating whether each independent substream of a program has at least one dependent substream associated with it, and if so, the number of dependent substreams associated with each independent substream of the program.

[0120] In some embodiments, the program information metadata (PIM) payload included within frames of an encoded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) buffered in buffer 201 has the following format: a payload header, which typically includes at least one identifying value (e.g., a value indicating the PIM format version and, optionally, length, period, count, and substream association values); and After the header, a PIM in the following format: Active channel metadata indicating each silent and non-silent channel of an audio program (i.e., which channels of the program contain audio information and which channels, if any, contain only silence (typically for the duration of that frame)). In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the active channel metadata in a frame of the bitstream may be used in conjunction with additional metadata of the bitstream (e.g., the audio coding mode ("acmod") field of that frame and the chanmap field, if present, in that frame or associated dependent substream frame(s)) to determine which channels of the program contain audio information and which channels contain silence;

[0121] Down-mixing processing state metadata indicating whether a program was down-mixed (before or during encoding) and, if so, the type of down-mixing that was applied. The down-mixing processing state metadata may be useful for implementing up-mixing downstream of the decoder (e.g., in post-processor 300), for example, to up-mix the audio content of the program using parameters that best match the type of down-mixing that was applied. In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the down-mixing processing state metadata may be used in conjunction with the audio coding mode ("acmod") field of a frame to determine the type of down-mixing, if any, that was applied to the channels of the program;

[0122] Upmixing process state metadata indicating whether a program was upmixed (e.g., from fewer channels) before or during encoding and, if so, the type of upmixing that was applied. The upmixing process state metadata may be useful for implementing downmixing downstream of the decoder (in a post-processor), for example, to downmix the audio content of a program in a manner that is compatible with the type of upmixing applied to the program (e.g., Dolby Pro Logic or Dolby Pro Logic II Movie mode or Dolby Pro Logic II Music mode or Dolby Professional Upmixer). In embodiments where the encoded bitstream is an E-AC-3 bitstream, the upmixing process state metadata may be used in conjunction with other metadata (e.g., the value of the 'strmtyp' field of the frame) to determine the type of upmixing (if any) that was applied to the channels of the program. The value of the "strmtyp" field (in the BSI segment of a frame of an E-AC-3 bitstream) indicates whether the audio content of the frame belongs to an independent stream (which determines a program) or an independent substream (of a program that contains or is associated with multiple substreams) and therefore can be decoded independently of any other substream represented by that E-AC-3 bitstream, or whether the audio content of the frame belongs to a dependent substream (of a program that contains or is associated with multiple substreams) and therefore needs to be decoded in conjunction with the associated independent substream;

[0123] Preprocessing state metadata indicating whether preprocessing was performed on the audio content of this frame (prior to encoding the audio content to produce the encoded bitstream) and, if so, the type of preprocessing that was performed.

[0124] In some implementations, the preprocessing state metadata indicates: whether surround attenuation was applied (for example, whether the surround channels of an audio program were attenuated by 3 dB prior to encoding); whether a 90 degree phase shift was applied (for example, to the surround Ls and Rs channels of an audio program prior to encoding), Whether a low-pass filter was applied to the LFE channel of the audio program prior to encoding, Whether the level of the LFE channel of the Program was monitored during production and, if so, the monitored level of the LFE channel relative to the levels of the Program's full-range audio channels.

[0125] Whether dynamic range compression should be performed (e.g., in a decoder) for each block of the decoded audio content of the program, and if so, the type (and / or parameters) of dynamic range compression to be performed (e.g., preprocessing state metadata of this type may indicate which of the following compression profile types was assumed by the encoder to generate the dynamic range compression control values ​​to be included in the encoded bitstream: film standard, film light, music standard, music light, or speech. Alternatively, preprocessing state metadata of this type may indicate that heavy dynamic range compression ('compr' compression) should be performed for each frame of the decoded audio content of the program, in a manner determined by the dynamic range compression control values ​​to be included in the encoded bitstream).

[0126] Whether spectral extension processing and / or channel-combining encoding was used to encode a particular frequency range of the program content, and if so, the minimum and maximum frequencies of the frequency components of the content that underwent spectral extension encoding and the minimum and maximum frequencies of the frequency components of the content that underwent channel-combining encoding. This type of preprocessing state metadata information may be useful for performing equalization downstream of the decoder (in a post-processor). Both channel-combining and spectral extension information are also useful for optimizing quality during transcoding operations and applications. For example, an encoder may optimize its behavior (including adaptation of preprocessing stages such as headphone virtualization, upmixing, etc.) based on the state of parameters such as spectral extension and channel-combining information. Furthermore, an encoder may dynamically adapt its combining and spectral extension parameters to match and / or to optimal values ​​based on the state of the incoming (and authenticated) metadata.

[0127] Whether dialogue enhancement adjustment range data is included in the encoded bitstream and, if so, the range of adjustment available during dialogue enhancement processing (e.g., in a post-processor downstream of the decoder) to adjust the level of dialogue content relative to the level of non-dialogue content in the audio program.

[0128] In some embodiments, the LPSM payload included in a frame of an encoded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) buffered in buffer 201 includes an LPSM in the following format: a header (typically containing a synchronization word identifying the beginning of the LPSM payload, followed by at least one identification value, such as the LPSM format version, length, period, count, and sub-stream association values ​​shown in Table 2 below); After the header, at least one dialogue indication value (e.g., parameter "Dialogue Channel" in Table 2) indicating whether the corresponding audio data indicates dialogue or not (e.g., which channel of the corresponding audio data indicates dialogue); at least one loudness regulation compliance value indicating whether the corresponding audio data conforms to an indicated set of loudness regulations (e.g., the parameter "Loudness Regulation Type" in Table 2); At least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data (e.g., one or more of the parameters "Dialogue-Gated Loudness Compensation Flag" and "Loudness Compensation Type" in Table 2); and At least one loudness value indicating at least one loudness (e.g., peak or average loudness) characteristic of the corresponding audio data (e.g., one or more of the parameters "ITU relative gated loudness", "ITU speech-gated loudness", "ITU (EBU3341) short-term 3s loudness", and "true peak").

[0129] In some implementations, the parser 205 (and / or the decoder stage 202) is configured to extract each metadata segment from the extra bits segment or the "addbsi" field or the ancillary data field of a frame of the bitstream, having the following format: a metadata segment header (typically containing a synchronization word identifying the start of the metadata segment, followed by at least one identifying value, such as version, length, period, extension element count, and substream association value); and after the metadata segment header, at least one protection value useful for at least one of decryption, authentication, or validation of at least one of the metadata of the metadata segment or the corresponding audio data (e.g., the HMAC digest and audio fingerprint values ​​of Table 1); and Also following the metadata segment header are metadata payload identification information ("ID") and payload configuration values ​​that identify the type of each subsequent metadata payload and at least one aspect of its configuration (e.g., size).

[0130] Each metadata payload (preferably in the format specified above) is followed by a corresponding metadata payload ID and payload configuration value.

[0131] More generally, encoded audio bitstreams produced by preferred embodiments of the present invention have a structure that provides a mechanism for labeling metadata elements and sub-elements as core (mandatory) or extension (optional) elements or sub-elements. This allows the data rate of the bitstream (including its metadata) to scale across many applications. Core (mandatory) elements of the preferred bitstream syntax should also be able to signal the presence (in-band) and / or remote location (out of band) of extension (optional) elements associated with the audio content.

[0132] A core element or elements are required to be present in all frames of the bitstream. Some sub-elements of a core element are optional and may be present in any combination. Extension elements are not required to be present in all frames (to limit bitrate overhead). Thus, an extension element may be present in some frames and absent in others. Some sub-elements of an extension element are optional and may be present in any combination, but some sub-elements of an extension element may be mandatory (i.e., mandatory if the extension element is present in a frame of the bitstream).

[0133] In one class of embodiments, an encoded audio bitstream is generated (e.g., by an audio processing unit embodying the present invention) that includes a sequence of audio data segments and metadata segments. The audio data segments represent audio data, and at least some of the metadata segments each include a PIM and / or SSM (and optionally at least one other type of metadata), and the audio data segments are time-division multiplexed with the metadata segments. In preferred embodiments of this class, each of the metadata segments has a preferred format described herein.

[0134] In one preferred format, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each of the metadata segments containing the SSM and / or PIM is included as additional bitstream information in an "addbsi" field (shown in FIG. 6) of a bitstream information ("BSI") segment of a frame of the bitstream, or in an auxiliary data field of a frame of the bitstream, or in an extra bits segment of a frame of the bitstream (e.g., by stage 107 of a preferred implementation of encoder 100).

[0135] In the preferred format described above, each frame includes a metadata segment (also referred to herein as a metadata container or container) in the frame's Extra Bits segment (or addbsi field). The metadata segment has required elements (collectively referred to as "core elements") with the format shown in Table 1 below (and may include optional elements shown in Table 1). At least some of the required elements shown in Table 1 are included in the metadata segment header of the metadata segment, but may be included elsewhere in the metadata segment.

[0136] [Table 1] In the preferred format, each metadata segment (within the redundant bits segment or addbsi or ancillary data field of a frame of the encoded bitstream) containing an SSM, PIM, or LPSM contains a metadata segment header (and optionally additional core elements) and, following the metadata segment header (or following the metadata segment header and other core elements), one or more metadata payloads. Each metadata payload contains a metadata payload header (indicating the particular type of metadata contained in the payload (e.g., SSM, PIM, or LPSM)) followed by metadata of that particular type. Typically, the metadata payload header contains the following values ​​(parameters): A payload ID (identifying the type of metadata, e.g., SSM, PIM, or LPSM), which follows the metadata segment header (which may contain values ​​specified in Table 1, for example); A payload configuration value (typically indicating the size of the payload), which follows the payload ID; Optionally, additional payload configuration values ​​(e.g., an offset value indicating the number of audio samples from the beginning of the frame to the first audio sample for the payload, as well as a payload priority value, e.g., indicating the conditions under which the payload may be discarded).

[0137] Typically, the metadata in the payload has one of the following formats:

[0138] payload metadata is an SSM, which includes independent substream metadata indicating the number of independent substreams of the program represented by the bitstream, and dependent substream metadata indicating whether each independent substream of the program has at least one dependent substream associated with it, and if so, the number of dependent substreams associated with each independent substream of the program; The payload metadata is PIM. active channel metadata indicating which channels of the audio program contain audio information and which channels (if any) contain only silence (typically for the duration of the frame); downmixing state metadata indicating whether the program was downmixed (before or during encoding) and, if so, the type of downmixing that was applied; upmixing state metadata indicating whether the program was upmixed (e.g., from fewer channels) before or during encoding and, if so, the type of upmixing that was applied; and preprocessing state metadata indicating whether preprocessing was performed on the audio content of the frame (before encoding the audio content to generate the encoded bitstream) and, if so, the type of preprocessing that was performed; The payload metadata is LPSM data and has the format shown in the following table (Table 2).

[0139] [Table 2-1] [Table 2-2] In another preferred format of an encoded bitstream generated in accordance with the present invention, the bitstream is an AC-3 or E-AC-3 bitstream, and each of the metadata segments containing a PIM and / or SSM (and optionally at least one other type of metadata) is included (e.g., by stage 107 in a preferred implementation of encoder 100) in any of: an Extra Bits segment of a frame of the bitstream; or an "addbsi" field of a Bitstream Information ("BSI") segment of a frame of the bitstream (as shown in FIG. 6); or an auxiliary data field at the end of a frame of the bitstream (e.g., an AUX segment shown in FIG. 4). A frame may include one or two metadata segments, each containing a PIM and / or SSM; and (in some embodiments) if a frame includes two metadata segments, one may be in the addbsi field of the frame and the other may be in the AUX field of the frame. Each metadata segment preferably has the format specified above with reference to Table 1 above (i.e., it contains the core elements specified in Table 1, followed by a payload ID (which identifies the type of metadata in each payload of the metadata segment) and a payload configuration value, and each metadata payload). Each metadata segment containing an LPSM preferably has the format specified above with reference to Tables 1 and 2 above (i.e., it contains the core elements specified in Table 1, followed by a payload ID (which identifies the metadata as an LPSM) and a payload configuration value, followed by the payload (LPSM data having the format shown in Table 2).

[0140] In another preferred format, the encoded bitstream is a Dolby E bitstream, and each of the metadata segments containing PIM and / or SSM (and optionally other metadata) is the first N sample positions of a Dolby E guard band interval. Dolby E bitstreams containing such metadata segments containing LPSM preferably include a value indicating the LPSM payload length signaled in the Pd word of the SMPTE 337M preamble (the SMPTE 337M Pa word repetition rate preferably remains the same as the associated video frame rate).

[0141] In a preferred format in which the encoded bitstream is an E-AC-3 bitstream, each of the metadata segments containing a PIM and / or SSM (and optionally an LPSM and / or other metadata) is included as additional bitstream information (e.g., by stage 107 of a preferred implementation of encoder 100) in an extra bits segment or in an "addbsi" field of a bitstream information ("BSI") segment of a frame of the bitstream. Further aspects of encoding an E-AC-3 bitstream with LPSMs in this preferred format are now described.

[0142] 1. During the generation of an E-AC-3 bitstream, while the E-AC-3 encoder (inserting LPSM values ​​into the bitstream) is "active", for every frame (sync frame) generated, the bitstream should contain a metadata block (containing LPSM) carried in the frame's addbsi field (or extra bits segment). The bits required to carry the metadata block should not increase the encoder bitrate (frame length).

[0143] 2. All metadata blocks (including LPSM) should contain the following information: loudness_correction_type_flag: where "1" indicates that the loudness of the corresponding audio data has been corrected upstream of the encoder, and "0" indicates that the loudness has been corrected by a loudness corrector built into the encoder (e.g., the loudness processor 103 of the encoder 100 of FIG. 2); speech_channel: indicates which source channel(s) contain speech (during the previous 0.5 seconds). If no speech is detected, this is indicated; speech_loudness: indicates the integrated speech loudness (during the previous 0.5 seconds) of each corresponding audio channel that contains speech; ITU_loudness: indicates the combined ITU BS.1770-3 loudness of each corresponding audio channel; Gain: The loudness complex gain(s) to invert in the decoder (to demonstrate reversibility).

[0144] 3. While an E-AC-3 encoder (which inserts LPSM values ​​into the bitstream) is "active" and receives an AC-3 frame with the "trusted" flag, the loudness controller in that encoder (e.g., loudness processor 103 of encoder 100 in FIG. 2) should be bypassed. The "trusted" source dialnorm and DRC values ​​should be passed (e.g., by generator 106 of encoder 100) to the E-AC-3 encoder component (e.g., stage 107 of encoder 100). LPSM block generation continues, with loudness_correction_type_flag set to "1". The loudness controller bypass sequence should be synchronized to the beginning of the decoded AC-3 frame in which the "trusted" flag appears. The loudness controller bypass sequence should be implemented as follows: The leveler_amount control is decremented from value 9 to value 0 over 10 audio block periods (i.e., 53.3 msec), and the leveler_back_end_meter control is put into bypass mode (this action should give a seamless transition). The term "trusted" bypass of the leveler implies that the dialnorm value of the source bitstream is also reused at the encoder's output (e.g., if a "trusted" source bitstream has a dialnorm value of -30, the encoder's output should use -30 for the outgoing dialnorm value). While an E-AC-3 encoder (which inserts LPSM values ​​into the bitstream) is "active" and receives AC-3 frames without the "confidence" flag, the loudness controller built into the encoder (e.g., loudness processor 103 of encoder 100 in Figure 2) should be active. LPSM block generation continues and loudness_correction_type_flag is set to "0". The loudness controller activation sequence should be synchronized to the beginning of the decoded AC-3 frame where the "confidence" flag disappears. The loudness controller activation sequence should be implemented as follows: the leveler_amount control is incremented from value 0 to value 9 over one audio block period (i.e., 5.3 msec), and the leveler_back_end_meter control is put into "active" mode (this action should give a seamless transition and include an integrated reset of the back_end_meter).

[0145] 5. During encoding, the graphical user interface (GUI) should show the user the following parameters: "Input Audio Program [Trusted / Not Trusted]" - the state of this parameter is based on the presence of the "trust" flag in the input signal; and "Real-time Loudness Correction: [Enabled / Disabled]" - the state of this parameter is based on whether the loudness controller built into the encoder is active or not.

[0146] When decoding an AC-3 or E-AC-3 bitstream that has LPSMs included in the extra bits or skip fields segment or the "addbsi" field of the bitstream information ("BSI") segment of each frame of the bitstream (in the preferred format described above), the decoder SHOULD parse the LPSM block data (in the extra bits segment or addbsi field) and pass all of the extracted LPSM values ​​to the graphical user interface (GUI). The set of extracted LPSM values ​​is refreshed every frame.

[0147] In another preferred format of an encoded bitstream generated in accordance with the present invention, the encoded bitstream is an AC-3 or E-AC-3 bitstream, and each of the metadata segments containing a PIM and / or SSM (and optionally also an LPSM and / or other metadata) is included (e.g., by stage 107 of a preferred implementation of encoder 100) in an extra bits segment, or in an Aux segment, or as additional bitstream information in an "addbsi" field (shown in FIG. 6) of a Bitstream Information ("BSI") segment of a frame of the bitstream. In this format (which is a variation on the format described above with reference to Tables 1 and 2), each of the addbsi (or Aux or Extra Bits) fields containing an LPSM contains the following LPSM value:

[0148] The core elements are as specified in Table 1, followed by a payload ID (which identifies the metadata as an LPSM) and payload configuration values, followed by the payload (LPSM data). The LPSM data has the following format (similar to the required elements shown in Table 2 above):

[0149] LPSM Payload Version: A 2-bit field indicating the version of the LPSM payload.

[0150] dialchan: A 3-bit field indicating whether the left, right, and / or center channels of the corresponding audio data contain spoken dialogue. The bit assignment of the dialchan field may be as follows: bit 0, indicating the presence of dialogue in the left channel, is stored in the most significant bit of the dialchan field, and bit 2, indicating the presence of dialogue in the center channel, is stored in the least significant bit of the dialchan field. Each bit of the dialchan field is set to "1" if the corresponding channel contains dialogue spoken during the preceding 0.5 seconds of the program.

[0151] loudregtyp: A 4-bit field that indicates which loudness regulatory standard the program loudness complies with. Setting the "loudregtyp" field to "000" indicates that the LPSM does not indicate loudness regulatory compliance. For example, one value of this field (e.g., 0000) may indicate that compliance with a loudness regulatory standard is not indicated, another value of this field (e.g., 0001) may indicate that the audio data of the program complies with the ATSC A / 85 standard, and another value of this field (e.g., 0010) may indicate that the audio data of the program complies with the EBU R128 standard. In this example, if this field is set to any value other than "0000", the loudcorrdialgat and loudcorrtyp fields should follow the payload.

[0152] loudcorrdialgat: A 1-bit field indicating whether dialogue-gated loudness correction has been applied. If the program's loudness has been corrected using dialogue gating, the value of the loudcorrdialgat field is set to "1". Otherwise, it is set to "0".

[0153] loudcorrtyp: A 1-bit field indicating the type of loudness correction applied to the program. If the program's loudness has been corrected using an infinite look-ahead (file-based) loudness correction process, the value of the loudcorrtyp field is set to "0". If the program's loudness has been corrected using a combination of real-time loudness measurement and dynamic range control, the value of this field is set to "1".

[0154] loudrelgate: A 1-bit field indicating whether relative gated loudness data (ITU) is present. If the loudrelgate field is set to "1", it should be followed in the payload by the 7-bit ituloudrelgat field.

[0155] loudrelgat: A 7-bit field indicating the relative gated program loudness (ITU). This field indicates the integrated loudness of the audio program, measured according to ITU-R BS.1770-3, without any gain adjustments due to dialnorm and dynamic range compression (DRC) applied. Values ​​from 0 to 127 are interpreted as -58 LKFS to +5.5 LKFS in 0.5 LKFS steps.

[0156] loudspchgate: A 1-bit field indicating whether speech-gated loudness data (ITU) is present. If the loudspchgate field is set to "1", it should be followed in the payload by a 7-bit loudspchgat field.

[0157] loudspchgat: A 7-bit field indicating the speech-gated program loudness. This field indicates the integrated loudness of the entire corresponding audio program, measured according to formula (2) of ITU-R BS.1770-3, without any gain adjustments due to dialnorm and dynamic range compression applied. Values ​​from 0 to 127 are interpreted as -58 LKFS to +5.5 LKFS in steps of 0.5 LKFS.

[0158] loudstrm3se: A 1-bit field indicating whether short-term (3-second) loudness data is present. If this field is set to "1", it should be followed by a 7-bit loudstrm3s field in the payload.

[0159] loudstrm3s: A 7-bit field indicating the ungated loudness of the preceding 3 seconds of the corresponding audio program, measured according to ITU-R BS.1770-1, without any gain adjustments due to dialnorm and dynamic range compression applied. Values ​​0 through 256 are interpreted as -116 LKFS to +5.5 LKFS, in steps of 0.5 LKFS.

[0160] truepke: A 1-bit field that indicates whether true peak loudness data is present. If the truepke field is set to "1", it should be followed by an 8-bit truepk field in the payload.

[0161] truepk: An 8-bit field indicating the true peak sample value of the program, measured according to Annex 2 of ITU-R BS.1770-3, without any gain adjustments due to dialnorm and dynamic range compression applied. Values ​​0 to 256 are interpreted as -116 LKFS to +11.5 LKFS, in steps of 0.5 LKFS.

[0162] In some embodiments, a metadata segment core element in the extra bits segment or ancillary data (or "addbsi") field of a frame of an AC-3 or E-AC-3 bitstream includes a metadata segment header (typically including an identification value, e.g., a version), followed by: a value indicating whether fingerprint data (or other protection value) is included for the metadata of the metadata segment; a value indicating whether external data (related to the audio data corresponding to the metadata of the metadata segment) is present; a payload ID and payload configuration value for each type of metadata identified by the core element (e.g., PIM and / or SSM and / or LPSM and / or certain types of metadata); and a protection value for at least one type of metadata identified by the metadata segment header (or other core elements of the metadata segment). The metadata payload(s) of the metadata segment follow the metadata segment header and are (possibly) nested within the metadata segment core element.

[0163] Embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination of both (e.g., as a programmable logic array). Unless otherwise specified, the algorithms or processes included as part of the present invention are not inherently related to any particular computer or other apparatus. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus (e.g., an integrated circuit) to perform the required method steps. Thus, the present invention may be implemented in one or more computer programs running on one or more programmable computer systems (e.g., an implementation of any of the elements of FIG. 1 or the encoder 100 (or element thereof) of FIG. 2 or the decoder 200 (or element thereof) of FIG. 3 or the post-processor (or element thereof) of FIG. 3). Each computer system has at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.

[0164] Each such program may be implemented in any desired computer language to communicate with a computer system, including machine, assembly, or high-level procedural, logical, or object-oriented programming languages, and in any case, the language may be a compiled or interpreted language.

[0165] For example, when implemented by a sequence of computer software instructions, various functions and steps of embodiments of the present invention may be implemented by a multi-threaded sequence of software instructions executed on suitable digital signal processing hardware, in which case various units, steps and functions of the embodiments may correspond to portions of the software instructions.

[0166] Each such computer program is preferably stored on or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., semiconductor memory or media or magnetic or optical media) that, when read by a computer system, configures or operates the computer to perform the procedures described herein. The system of the present invention may also be implemented as a computer-readable storage medium configured with (i.e., having stored thereon) a computer program, which causes the computer system to operate in a specific, predefined manner to perform the functions described herein.

[0167] While several embodiments of the present invention have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the invention. Numerous modifications and variations of the present invention are possible in light of the above teachings. It will be understood that, within the scope of the appended claims, the invention may be practiced otherwise than as specifically described herein.

[0168] Several aspects will be described. [Aspect 1] 1. An audio processing unit comprising: a buffer memory; and at least one processing subsystem coupled to the buffer memory, the buffer memory stores at least one frame of an encoded audio bitstream, the frame including program information metadata or substream structure metadata in at least one metadata segment of at least one skip field of the frame and audio data in at least one other segment of the frame; the processing subsystem is coupled and configured to perform at least one of generating the bitstream, decoding the bitstream, or adaptively processing audio data of the bitstream using metadata of the bitstream, or authenticating or verifying at least one of audio data or metadata of the bitstream using metadata of the bitstream; The metadata segment includes at least one metadata payload, the metadata payload comprising: Header and; after the header, including at least a portion of the program information metadata or at least a portion of the substream structure metadata; Audio processing unit. [Aspect 2] the encoded audio bitstream indicates at least one audio program, and the metadata segment includes a program information metadata payload, the program information metadata payload comprising: Program information metadata header and; after the program information metadata header, program information metadata indicating at least one attribute or characteristic of the audio content of the program; the program information metadata includes active channel metadata indicating each non-silence channel and each silence channel of the program; 2. The audio processing unit of embodiment 1. Aspect 3 The program information metadata: downmix processing state metadata indicating whether the program has been downmixed and, if so, the type of downmix applied to the program; upmixing process state metadata indicating whether the program has been upmixed and, if so, the type of upmixing applied to the program; preprocessing state metadata indicating whether preprocessing was performed on the audio content of said frame and, if so, the type of preprocessing that was performed on said audio content; or spectral spreading or channel combining metadata indicating whether spectral spreading or channel combining has been applied to the program and, if so, the frequency range to which the spectral spreading or channel combining has been applied; 3. The audio processing unit of claim 2, further comprising at least one of: Aspect 4 The encoded audio bitstream indicates at least one audio program having at least one independent substream of audio content, and the metadata segment includes a substream structure metadata payload, the substream structure metadata payload comprising: Substream structure metadata payload header and; the substream structure metadata payload header is followed by independent substream metadata indicating the number of independent substreams of the program and dependent substream metadata indicating whether each independent substream of the program has at least one associated dependent substream; 2. The audio processing unit of embodiment 1. Aspect 5 The metadata segment: Metadata segment headers and; the metadata segment header is followed by at least one protection value useful for at least one of decryption, authentication, or validation of the program information metadata or the substream structure metadata or at least one of the audio data corresponding to the program information metadata or the substream structure metadata; the metadata segment header is followed by a metadata payload identification and a payload configuration value, and the metadata payload follows the metadata payload identification and the payload configuration value; 2. The audio processing unit of embodiment 1. Aspect 6 An audio processing unit as described in aspect 5, wherein the metadata segment includes a synchronization word that identifies the beginning of the metadata segment and at least one identification value following the synchronization word, and the header of the metadata payload includes the at least one identification value. Aspect 7 2. The audio processing unit of embodiment 1, wherein the encoded audio bitstream is an AC-3 bitstream or an E-AC-3 bitstream. Aspect 8 2. The audio processing unit of claim 1, wherein the buffer memory stores the frames in a non-temporary manner. Aspect 9 2. The audio processing unit of embodiment 1, wherein the audio processing unit is an encoder. Aspect 10 the processing subsystem: a decoding subsystem configured to receive an input audio bitstream and extract input metadata and input audio data from the input audio bitstream; an adaptive processing subsystem coupled and configured to perform adaptive processing on the input audio data using the input metadata, thereby generating processed audio data; an encoding subsystem coupled and configured to generate the encoded audio bitstream in response to the processed audio data, including by including the program information metadata or the substream structure metadata in the encoded audio bitstream, and to present the encoded audio bitstream to the buffer memory; 10. The audio processing unit of embodiment 9. Aspect 11 2. The audio processing unit of embodiment 1, wherein the audio processing unit is a decoder. Aspect 12 12. The audio processing unit of claim 11, wherein the processing subsystem is a decoding subsystem coupled to the buffer memory and configured to extract the program information metadata or the substream structure metadata from the encoded audio bitstream. Aspect 13 a subsystem coupled to the buffer memory and configured to extract the program information metadata or the substream structure metadata from the encoded audio bitstream and to extract the audio data from the encoded audio bitstream; a post-processor coupled to the subsystem and configured to perform adaptive processing on the audio data using at least one of the program information metadata or the sub-stream structure metadata extracted from the encoded audio bitstream. 2. The audio processing unit of embodiment 1. Aspect 14 2. The audio processing unit of embodiment 1, wherein the audio processing unit is a digital signal processor. Aspect 15 2. The audio processing unit of claim 1, wherein the audio processing unit is a preprocessor configured to extract the program information metadata or the substream structure metadata and the audio data from the encoded audio bitstream, and to perform adaptive processing on the audio data using at least one of the program information metadata or the substream structure metadata extracted from the encoded audio bitstream. Aspect 16 1. A method of decoding an encoded bitstream, comprising: receiving an encoded audio bitstream; extracting metadata and audio data from the encoded audio bitstream, wherein the metadata is or includes program information metadata and substream structure metadata; the encoded audio bitstream comprises a sequence of frames and indicates at least one audio program, the program information metadata and the substream structure metadata indicate the program, each frame comprises at least one audio data segment, each of the audio data segments comprises at least a portion of the audio data, and each frame of at least a subset of the frames comprises a metadata segment, each of the metadata segments comprises at least a portion of the program information metadata and at least a portion of the substream structure metadata; method. Aspect 17 The metadata segment includes a program information metadata payload, the program information metadata payload comprising: Program information metadata header and; after the program information metadata header, program information metadata indicating at least one attribute or characteristic of the audio content of the program; the program information metadata includes active channel metadata indicating each non-silence channel and each silence channel of the program; The method of embodiment 16. Aspect 18 The program information metadata: downmix processing state metadata indicating whether the program has been downmixed and, if so, the type of downmix applied to the program; Upmixing process state metadata indicating whether the program has been upmixed and, if so, the type of upmixing that has been applied to the program; or Preprocessing state metadata indicating whether preprocessing was performed on the audio content of the frame and, if so, the type of preprocessing that was performed on the audio content. 20. The method of embodiment 17, further comprising at least one of: Aspect 19 The encoded audio bitstream indicates at least one audio program having at least one independent substream of audio content, and the metadata segment includes a substream structure metadata payload, the substream structure metadata payload comprising: Substream structure metadata payload header and; the substream structure metadata payload header is followed by independent substream metadata indicating the number of independent substreams of the program and dependent substream metadata indicating whether each independent substream of the program has at least one associated dependent substream; The method of embodiment 16. Aspect 20 The metadata segment: Metadata segment headers and; the metadata segment header is followed by at least one protection value useful for at least one of decrypting, authenticating, or validating the program information metadata or the substream structure metadata or at least one of the audio data corresponding to the program information metadata and the substream structure metadata; a metadata payload after the metadata segment header, the metadata payload including the at least a portion of the program information metadata and the at least a portion of the substream structure metadata; The method of embodiment 16. Aspect 21 17. The method of embodiment 16, wherein the encoded audio bitstream is an AC-3 bitstream or an E-AC-3 bitstream. Aspect 22 performing adaptive processing on the audio data using at least one of the program information metadata or the substream structure metadata extracted from the encoded audio bitstream. The method of embodiment 16.

Claims

1. one or more processors; a memory coupled to the one or more processors configured to store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations; an audio processing unit comprising: receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data for a set of one or more audio channels and metadata associated with the set of audio channels, the metadata including dynamic range control (DRC) metadata, loudness metadata, and metadata indicating a number of channels in the set of audio channels, the DRC metadata including a DRC value and a DRC profile metadata indicating a DRC profile used to generate the DRC value, and the loudness metadata including metadata indicating a dialogue loudness of the audio program, the dialogue loudness of the audio program being measured in accordance with ITU-R BS.1770; decoding the encoded audio data to obtain decoded audio data for the set of audio channels; obtaining, from the metadata of the encoded audio data, the DRC value and metadata indicative of the dialogue loudness of the audio program; modifying the decoded audio data of the set of audio channels in response to the DRC value and metadata indicative of the dialogue loudness of the audio program, wherein modifying the decoded audio data includes performing loudness control of the decoded audio data using the dialogue loudness of the audio program; Including, Audio processing unit.

2. 2. The audio processing unit of claim 1, wherein the encoded audio bitstream includes a metadata container, the metadata container including a header followed by one or more metadata payloads, the one or more metadata payloads including the DRC metadata.

3. A method performed by an audio processing unit, comprising: receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data for a set of one or more audio channels and metadata associated with the set of audio channels, the metadata including dynamic range control (DRC) metadata, loudness metadata, and metadata indicating a number of channels in the set of audio channels, the DRC metadata including a DRC value and a DRC profile metadata indicating a DRC profile used to generate the DRC value, and the loudness metadata including metadata indicating a dialogue loudness of the audio program, the dialogue loudness of the audio program being measured in accordance with ITU-R BS.1770; decoding the encoded audio data to obtain decoded audio data for the set of audio channels; obtaining, from the metadata of the encoded audio data, the DRC value and metadata indicative of the dialogue loudness of the audio program; modifying the decoded audio data of the set of audio channels in response to the DRC value and metadata indicative of the dialogue loudness of the audio program, wherein modifying the decoded audio data includes performing loudness control of the decoded audio data using the dialogue loudness of the audio program; Including, method.

4. 1. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data for a set of one or more audio channels and metadata associated with the set of audio channels, the metadata including dynamic range control (DRC) metadata, loudness metadata, and metadata indicating a number of channels in the set of audio channels, the DRC metadata including a DRC value and a DRC profile metadata indicating a DRC profile used to generate the DRC value, and the loudness metadata including metadata indicating a dialogue loudness of the audio program, the dialogue loudness of the audio program being measured in accordance with ITU-R BS.1770; decoding the encoded audio data to obtain decoded audio data for the set of audio channels; obtaining, from the metadata of the encoded audio data, the DRC value and metadata indicative of the dialogue loudness of the audio program; modifying the decoded audio data of the set of audio channels in response to the DRC value and metadata indicative of the dialogue loudness of the audio program, wherein modifying the decoded audio data includes performing loudness control of the decoded audio data using the dialogue loudness of the audio program; Including, A non-transitory computer-readable storage medium.