Audio encoder and decoder with program loudness and boundary metadata

The audio processing unit addresses the challenge of inefficient loudness processing and boundary determination in existing audio technologies by embedding metadata within audio bitstreams, enabling adaptive processing and maintaining audio quality.

JP2025072557AActive Publication Date: 2025-05-09DOLBY LABORATORIES LICENSING CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025018892
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-05-16
Filing Date
2025-02-07
Publication Date
2025-05-09
Estimated Expiration
2034-01-15

AI Technical Summary

Technical Problem

Existing audio signal processing technologies, such as AC-3 and E-AC-3, lack efficient mechanisms for encoding and decoding audio data bitstreams with metadata indicating the loudness processing state and program boundary locations, leading to unnecessary processing and deterioration of audio quality.

Method used

An audio processing unit that includes a buffer memory, an audio decoder, and a parser, capable of extracting and verifying metadata from encoded audio bitstreams. This unit embeds loudness processing state metadata (LPSM) and program boundary metadata within the bitstream, allowing for adaptive loudness processing and accurate program boundary determination.

Benefits of technology

The proposed solution enables efficient adaptive loudness processing, reduces unnecessary audio processing, and maintains audio quality by accurately determining program boundaries within audio bitstreams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072557000001_ABST
    Figure 2025072557000001_ABST
Patent Text Reader

Abstract

To provide an audio processing unit, an audio processing method and a non-temporary medium, which decode a bitstream.SOLUTION: A method by a decoder for decoding an encoded audio bitstream including program loudness metadata and audio data includes adaptive loudness processing of audio data of an audio program indicated by the bitstream or processing for performing authentication and / or validation confirmation of metadata and / or audio data of such audio program. The decoder includes a buffer memory for storing at least one frame of the audio bitstream encoded based on any embodiment of the method.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 61 / 754,882, filed January 21, 2013, and U.S. Provisional Patent Application No. 61 / 824,010, filed May 16, 2013, each of which is incorporated herein by reference in its entirety.

[0002] Technical Field The present invention relates to audio signal processing, and more particularly to encoding and decoding audio data bitstreams having metadata indicating the loudness processing state of the audio content and the locations of audio program boundaries indicated by the bitstream. Some embodiments of the present invention generate or decode audio data in one of the formats known as AC-3, Enhanced AC-3 or E-AC-3 or Dolby E. [Background technology]

[0003] Dolby, Dolby Digital, Dolby Digital Plus, and Dolby E are trademarks of Dolby Laboratories Licensing Corporation. Dolby Laboratories offers proprietary implementations of AC-3 and E-AC-3 known as Dolby Digital and Dolby Digital Plus, respectively.

[0004] Audio data processing units typically operate in a blind manner, paying no attention to the processing history of audio data that has been performed before the data is received. This may work in a processing framework where a single entity performs all audio data processing and encoding for various target media rendering devices, which perform all decoding and rendering of the encoded audio data. However, this blind processing does not work well (or at all) in situations where multiple audio processing units are distributed or arranged in cascade (chain) across a diverse network, and are expected to perform each type of audio processing optimally. For example, some audio data may have been encoded for a high-performance media system, and may need to be converted along the media processing chain to a reduced form suitable for mobile devices. Thus, audio processing units may unnecessarily perform a type of processing on the audio data that has already been performed. For example, a volume leveling unit may perform processing on an input audio clip regardless of whether the same or similar volume leveling has been performed on the input audio clip previously. As a result, the volume leveling unit may perform leveling when it is not necessary, and this unnecessary processing may result in the degradation and / or removal of certain characteristics when rendering the content in the audio data.

[0005] A typical stream of audio data contains both audio content (e.g., one or more channels of audio content) and metadata that describes at least one characteristic of the audio content. For example, in an AC-3 bitstream, there are several audio metadata parameters that are specifically intended for use in altering the sound of the program delivered to a listening environment. One of the metadata parameters is the DIALNORM parameter, which is intended to indicate the average level of dialogue appearing in an audio program and is used to determine the audio playback signal level.

[0006] During playback of a bitstream having a sequence of various audio program segments (each with a different DIALNORM parameter), an AC-3 decoder uses the DIALNORM parameter of each segment to perform some type of loudness processing that modifies the playback level or loudness so that the perceived loudness of the dialogue of the sequence of segments is at a consistent level. Each encoded audio segment (item) in a sequence of encoded audio items will (generally) have a different DIALNORM parameter, and the decoder will scale the level of each item so that the playback level or loudness of the dialogue for each item is the same or very similar. However, this may require applying different amounts of gain to different items during playback.

[0007] DIALNORM is typically set by the user and is not generated automatically, although there is a default DIALNORM value if no value is set by the user. For example, a content creator may make loudness measurements using equipment external to the AC-3 encoder and then forward the results (indicating the loudness of the spoken dialogue of the audio program) to the encoder to set the DIALNORM value. Thus, it is up to the content creator to set the DIALNORM parameter correctly.

[0008] There are several different reasons why the DIALNORM parameter in an AC-3 bitstream may be incorrect. First, each AC-3 encoder has a default DIALNORM value that is used during bitstream generation if the DIALNORM value is not set by the content creator. This default value may differ substantially from the actual dialogue loudness level of the audio. Second, even if the content creator measures the loudness and sets the DIALNORM value accordingly, a loudness measurement algorithm or meter that does not comply with the recommended AC-3 loudness measurement method may have been used, resulting in an incorrect DIALNORM value. Third, even if an AC-3 bitstream was generated with a DIALNORM value that was measured and correctly set by the content creator, it may have been changed to an incorrect value during transmission and / or storage of the bitstream. For example, in television broadcast applications, it is not uncommon for AC-3 bitstreams to be decoded, modified, and then encoded using incorrect DIALNORM metadata information. Thus, the DIALNORM values ​​included in the AC-3 bitstream may be incorrect or inaccurate, and therefore may have a negative impact on the quality of the listening experience.

[0009] Furthermore, the DIALNORM parameter does not indicate the loudness processing state of the corresponding audio data (e.g., what type(s) of loudness processing have been performed on that audio data). Until the present invention, audio bitstreams have not included metadata that indicates the loudness processing state of the audio content of the audio bitstream (e.g., the type(s) of loudness processing that have been applied thereto) or the loudness processing state and loudness of the audio content of the bitstream in a format of the type described in this disclosure. Loudness processing metadata in such a format is useful for facilitating adaptive loudness processing of an audio bitstream and / or verification of the effectiveness of the loudness processing state and loudness of the audio content in a particularly efficient manner.

[0010] Although the present invention is not limited to use with AC-3, E-AC-3, or Dolby E bitstreams, for convenience it will be described in embodiments that generate, decode, or otherwise process such bitstreams that include loudness processing state metadata.

[0011] An AC-3 encoded bitstream contains metadata and one to six channels of audio content. The audio content is audio data compressed using perceptual audio coding. The metadata includes several audio metadata parameters that are intended for use in altering the sound of the program delivered to the listening environment.

[0012] Details of AC-3 (also known as Dolby Digital) encoding are well known and described in many publications, including non-patent document 1, and patent documents 1, 2, 3, 4, and 5, all of which are incorporated herein by reference in their entireties.

[0013] Details of Dolby Digital Plus (E-AC-3) are described in Non-Patent Document 2.

[0014] Details of Dolby E encoding are described in Non-Patent Document 3 and Non-Patent Document 4.

[0015] Each frame in an AC-3 encoded audio bitstream contains audio content and metadata for 1536 samples of digital audio. For a sampling rate of 48 kHz, this represents 32 milliseconds of digital audio or a rate of 31.25 frames of audio per second.

[0016] Each frame in an E-AC-3 encoded audio bitstream contains audio content and metadata for 256, 512, 768 or 1536 samples of digital audio, respectively, depending on whether the frame contains one, two, three or six blocks of audio data. For a sampling rate of 48 kHz, this represents 5.333, 10.667, 16 or 32 milliseconds of digital audio, respectively, or a rate of 189.9, 93.75, 62.5 or 31.25 frames per second of audio, respectively.

[0017] As shown in Figure 4, each AC-3 frame is divided into sections (segments). The sections include (as shown in Figure 5) a synchronization information (SI) section, which contains the synchronization word (SW) and the first of two error correction words (CRC1); a bitstream information (BSI) section, which contains most of the metadata; six audio blocks (AB0 through AB5), which contain the data-compressed audio content (and may also contain metadata); a waste bits segment (W), which contains any unused bits left after the audio content is compressed; an auxiliary (AUX) information section, which may contain further metadata; and the second of two error correction words (CRC2). The waste bits segment (W) is sometimes referred to as the "skip field."

[0018] As shown in Figure 7, each E-AC-3 frame is divided into sections (segments). The sections include a synchronization information (SI) section, which contains the synchronization word (SW) (as shown in Figure 5); a bitstream information (BSI) section, which contains most of the metadata; between one and six audio blocks (AB0 to AB5) which contain the data compressed audio content (and may also contain metadata); a waste bits segment (W) which contains any unused bits left after the audio content is compressed (although only one waste bits segment is shown, there is typically a different waste bits segment following each audio block); an auxiliary (AUX) information section which may contain further metadata; and an error correction word (CRC). The waste bits segment (W) is sometimes referred to as the "skip field".

[0019] In an AC-3 (or E-AC-3) bitstream, there are several audio metadata parameters that are specifically intended for use in altering the sound of the program delivered to the listening environment. One such metadata parameter is the DIALNORM parameter, which is contained in the BSI segment.

[0020] As shown in Figure 6, the BSI segment of an AC-3 frame includes a five-bit parameter ("DIALNORM") that indicates the DIALNORM value for the program, and a five-bit parameter ("DIALNORM2") that indicates the DIALNORM value for a second audio program carried in the same AC-3 frame if the audio coding mode ("acmod") for the AC-3 frame is "0", indicating that a dual mono or "1+1" channel configuration is being used.

[0021] The BSI segment includes a flag ("addbsie") indicating the presence (or absence) of additional bitstream information following the "addbsie" bit, a parameter ("addbsil") indicating the length of the additional bitstream information, if any, following the "addbsil" value, and up to 64 bits of additional bitstream information ("addbsi") following the "addbsil" value.

[0022] The BSI segment includes other metadata values ​​not specifically shown in FIG. [Prior art documents] [Patent documents]

[0023] [Patent Document 1] U.S. Patent No. 5,583,962 [Patent Document 2] U.S. Patent No. 5,632,005 [Patent Document 3] U.S. Patent No. 5,633,981 [Patent Document 4] U.S. Patent No. 5,727,119 [Patent Document 5] U.S. Patent No. 6,021,386 [Non-patent literature]

[0024] [Non-Patent Document 1] ATSC Standard A52 / A: Digital Audio Compression Standard (AC-3), Revision A, Advanced Television Systems Committee, 20 Aug. 2001 [Non-Patent Document 2] Introduction to Dolby Digital Plus, an Enhancement to the Dolby Digital Coding System, AES Convention Paper 6196, 117th AES Convention, October 28, 2004 [Non-Patent Document 3] "Efficient Bit Allocation, Quantization, and Coding in an Audio Distribution System", AES Preprint 5068, 107th AES Conference, August 1999 [Non-Patent Document 4] "Professional Audio Coder Optimized for Use with Video", AES Preprint 5033, 107th AES Conference August 1999 Summary of the Invention [Means for solving the problem]

[0025] In one class of embodiments, the invention is an audio processing unit including a buffer memory, an audio decoder, and a parser. The buffer memory stores at least one frame of an encoded audio bitstream. The encoded audio bitstream includes audio data and a metadata container. The metadata container includes a header, one or more metadata payloads, and protection data. The header includes a synchronization word that identifies the beginning of the container. The one or more metadata payloads describe an audio program associated with the audio data. The protection data follows the one or more metadata payloads. The protection data may also be used to verify the integrity of the metadata container and the one or more payloads within the metadata container. The audio decoder is coupled to the buffer memory and is capable of decoding the audio data. The parser is coupled to or integrated with the audio decoder and is capable of parsing the metadata container.

[0026] In an exemplary embodiment, the method includes receiving an encoded audio bitstream, the audio data being segmented into one or more frames. The audio data is extracted from the encoded audio bitstream along with a metadata container. The metadata container includes a header followed by one or more metadata payloads followed by protection data. Finally, the integrity of the container and the one or more metadata payloads is verified through use of the protection data. The one or more metadata payloads may include a program loudness payload including data indicative of a measured loudness of an audio program associated with the audio data.

[0027] A payload of program loudness metadata embedded in an audio bitstream according to an exemplary embodiment of the present invention, referred to as loudness processing state metadata (LPSM), may be authenticated and validated, for example to allow a loudness regulatory entity to verify whether the loudness of a particular program is already within a specified range and that the corresponding audio data itself has not been modified (thereby ensuring compliance with the applicable regulations). To verify this, instead of recalculating the loudness, the loudness value contained in the data block containing the loudness processing state metadata may be read. In response to the LPSM, a regulatory authority may determine that the corresponding audio content (as indicated by the LPSM) complies with loudness legislation and / or regulatory requirements (e.g. the Commercial Advertisement Loudness Mitigation Act, also known as the "CALM Act") without the need to calculate the loudness of the audio content.

[0028] Loudness measurements required for compliance with some loudness legislation and / or regulatory requirements (e.g., rules issued under the CALM Act) are based on integrated program loudness. Integrated program loudness requires that loudness measurements, either at dialogue level or full mix level, be made for the entire audio program. Thus, in order to make program loudness measurements (e.g., at various stages in the broadcast chain) to verify compliance with typical legal requirements, it is essential that the measurements are made with knowledge of which audio data (and metadata) define the entire audio program, which typically requires knowledge of the beginning and end locations of that program (e.g., during processing of a bitstream representing a sequence of audio programs).

[0029] According to an exemplary embodiment of the present invention, an encoded audio bitstream indicates at least one audio program (e.g., a sequence of audio programs), and program boundary metadata and LPSM included in the bitstream allow resetting of program loudness measurements at the end of a program, thus providing an automated method of measuring integrated program loudness. Exemplary embodiments of the present invention include program boundary metadata in an encoded audio bitstream in an efficient manner, which allows accurate and robust determination of at least one boundary between consecutive audio programs indicated by the bitstreams. Exemplary embodiments allow accurate and robust determination of program boundaries, in the sense that they allow accurate program boundary determination even when bitstreams indicating different programs are spliced ​​together (to generate a bitstream of the present invention) in a manner that truncates one or both of the spliced ​​bitstreams (thus discarding program boundary metadata included in at least one of the bitstreams prior to splicing).

[0030] In a typical embodiment, the program boundary metadata for a frame of a bitstream of the present invention is a program boundary flag indicating a frame count. Typically, this flag indicates the number of frames between the current frame (the frame containing the flag) and a program boundary (the start or end of the current audio program). In some preferred embodiments, program boundary flags are inserted in a symmetric and efficient manner at the start and end of each bitstream segment representing a single program (i.e., in frames occurring within some predetermined number of frames after the start of the segment and in frames occurring within some predetermined number of frames before the end of the segment). Thus, when two such bitstreams are concatenated (thereby representing a sequence of two programs), program boundary metadata can be present (e.g., symmetrically) on both sides of the boundary between the two programs.

[0031] To limit the data rate increase resulting from including program boundary metadata in an encoded audio bitstream (which may represent an audio program or a sequence of audio programs), in an exemplary embodiment, program boundary flags are inserted into only a subset of the frames of the bitstream. Typically, the boundary flag insertion rate is a non-increasing function of the increasing distance of each frame of the bitstream (in which the flag is inserted) from its nearest program boundary. Here, the "boundary flag insertion rate" represents the average ratio of the number of frames (indicating a program) that contain a program boundary flag to the number of frames (indicating the program) that do not contain a program boundary flag, where the average is a running average over a number (e.g., a relatively small number) of consecutive frames of the encoded audio bitstream. In one class of embodiments, the boundary flag insertion rate is a logarithmically decreasing function of increasing distance (of each flag insertion location) from the nearest program boundary, and for each flag-containing frame that contains one of the flags, the size of the flag in the flag-containing frame is greater than or equal to the size of each flag in a frame that is located closer to the nearest program boundary than the flag-containing frame (i.e., the size of the program boundary flag in each flag-containing frame is a non-decreasing function of the flag-containing frame's increasing distance from the nearest program boundary).

[0032] Another aspect of the invention is an audio processing unit (APU) configured to perform any of the embodiments of the method of the invention. In another class of embodiments, the invention is an APU that includes a buffer memory (buffer) storing (e.g., in a non-temporary manner) at least one frame of an encoded audio bitstream generated by any of the embodiments of the method of the invention. Examples of APUs include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems (preprocessors), post-processing systems (postprocessors), audio bitstream processing systems, and combinations of such elements.

[0033] In another class of embodiments, the invention is an audio processing unit (APU) configured to generate an encoded audio bitstream including audio data segments and metadata segments. The audio data segments are indicative of audio data, and each of at least some of the metadata segments includes loudness processing state metadata (LPSM), and optionally also program boundary metadata. Typically, at least one such metadata segment in a frame of the bitstream includes at least one segment of LPSM indicating whether a first type of loudness processing has been performed on the audio data of that frame (i.e., audio data in at least one audio data segment of that frame), and at least one other segment of LPSM indicating loudness of at least a portion of the audio data of that frame (e.g., dialogue loudness of at least a portion of data of the audio data of that frame that is indicative of dialogue). In an embodiment of this class, the APU is an encoder configured to encode input audio to generate encoded audio, and the audio data segments include the encoded audio. In an exemplary embodiment of this class, each of the metadata segments has a preferred format described herein.

[0034] In some embodiments, each of the metadata segments of an encoded bitstream (in some embodiments, an AC-3 bitstream or an E-AC-3 bitstream) that includes LPSMs (e.g., LPSMs and program boundary metadata) is included in the extra bits of a skip field segment of a frame of the bitstream (e.g., an extra bits segment W of the type shown in FIG. 4 or FIG. 7). In other embodiments, each of the metadata segments of an encoded bitstream (in some embodiments, an AC-3 bitstream or an E-AC-3 bitstream) that includes LPSMs (e.g., LPSMs and program boundary metadata) is included as additional bitstream information in an "addbsi" field of a bitstream information ("BSI") segment of a frame of the bitstream, or in an auxiliary data field at the end of a frame of the bitstream (e.g., an AUX segment of the type shown in FIG. 4 or FIG. 7). Each metadata segment that includes an LPSM may have a format as described herein with reference to Tables 1 and 2 below (i.e., includes the core elements described in Table 1 or a variation thereof, followed by a payload ID (which identifies the metadata as an LPSM) and a payload size value, followed by the payload (LPSM data having a format as shown in Table 2 or a variation on Table 2 described herein). In some embodiments, a frame may include one or more metadata segments, and if a frame includes two metadata segments, one may be present in the addbsi field of the frame and the other may be present in the AUX field of the frame.

[0035] In one class of embodiments, the invention is a method that includes encoding audio data to generate an AC-3 or E-AC-3 encoded audio stream by including LPSM and program boundary metadata, and optionally other metadata about the audio program to which the frame belongs, in a metadata segment (of at least one frame of the bitstream). In some embodiments, each such metadata segment is included in an addbsi field of the frame or an ancillary data field of the frame. In other embodiments, each such metadata is included in an extra bits segment of the frame. In some embodiments, each metadata segment that includes LPSM and program boundary metadata includes a core header (and optionally additional core elements) and, after the core header (or after the core header and other core elements), an LPSM payload (or container) segment having the following format:

[0036] A header, which typically contains at least one identification value (e.g., LPSM format version, length, period, count and substream association values, as described in Table 2 of this document); After the header comes the LPSM and program boundary metadata. The program boundary metadata may include a program boundary frame count, a code value (e.g., an "offset_exist" value) indicating whether the frame contains only the program boundary frame count or both the program boundary frame count and an offset value, and (optionally) an offset value.

[0037] The LPSM may include: at least one dialogue indication value indicating whether the corresponding audio data indicates dialogue or not (e.g., which channels of the corresponding audio data indicate dialogue), the dialogue indication value(s) may indicate whether dialogue is present in any combination or all of the channels of the corresponding audio data; at least one loudness regulations compliance value indicating whether the corresponding audio data complies with an indicated set of loudness regulations; at least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data; and At least one loudness value indicating at least one loudness characteristic of the corresponding audio data (e.g., peak or average loudness).

[0038] In other embodiments, the encoded bitstream is a bitstream that is not an AC-3 or E-AC-3 bitstream, and each of the metadata segments that contain LPSMs (and optionally also program boundary metadata) is included in a segment (or field or slot) of the bitstream that is reserved for the storage of additional data. Each metadata segment that contains an LPSM may have a format similar or identical to that described herein with reference to Tables 1 and 2 below (i.e., including core elements similar or identical to those described in Table 1, followed by a payload ID and payload size value (that identifies the metadata as an LPSM), followed by a payload (LPSM data having a format similar or identical to that shown in Table 2 or a variation of Table 2 described herein).

[0039] In some embodiments, the encoded bitstream includes a sequence of frames, each frame including a bitstream information ("BSI") segment including an "addbsi" field (sometimes referred to as a segment or slot) and an ancillary data field or slot (e.g., the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream). The bitstream includes audio data segments (e.g., the AB0-AB5 segments of the frames shown in FIG. 4) and metadata segments, where the audio data segments represent audio data and where each of at least some of the metadata segments includes loudness processing state metadata (LPSM) and optionally program boundary metadata. The LPSMs are present in the bitstream in the following format: Each metadata segment that includes an LPSM is included in the "addbsi" field of a BSI segment of a frame of the bitstream, or in an ancillary data field of a frame of the bitstream, or in a redundant bits segment of a frame of the bitstream. Each metadata segment that includes an LPSM includes an LPSM payload (or container) segment having the following format:

[0040] A header, which typically contains at least one identifying information value, such as the LPSM format version, length, period, count and sub-stream association values, as shown in Table 2 below; After the header is the LPSM and optionally program boundary metadata, which may include a program boundary frame count, a code value (e.g., an "offset_exist" value) indicating whether the frame contains only the program boundary frame count or both the program boundary frame count and an offset value, and (possibly) an offset value.

[0041] The LPSM may include: at least one dialogue indication value (e.g., parameter "Dialogue Channels" in Table 2) that indicates whether the corresponding audio data indicates dialogue or not (e.g., which channels of the corresponding audio data indicate dialogue); the dialogue indication value(s) may indicate whether dialogue is present in any combination or all of the channels of the corresponding audio data; at least one loudness regulation compliance value indicating whether the corresponding audio data complies with an indicated set of loudness regulations (e.g. the parameter "Loudness regulation type" in Table 2); At least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data (e.g., one or more of the parameters "Dialogue-Gated Loudness Compensation Flag" and "Loudness Compensation Type" in Table 2); and At least one loudness value indicating at least one loudness (e.g. peak or average loudness) characteristic of the corresponding audio data (e.g. one or more of the parameters "ITU relative gated loudness", "ITU speech gated loudness", "ITU (EBU3341) short-term 3s loudness" and "true peak" in Table 2).

[0042] In any embodiment of the present invention that considers, uses or generates at least one loudness value indicative of corresponding audio data, the loudness value(s) may be indicative of at least one loudness measurement characteristic that is utilized to process the loudness and / or dynamic range of the audio data.

[0043] In some implementations, each metadata segment in the "addbsi" field or ancillary data field or redundant bits segment of a frame of the bitstream has the following format:

[0044] A core header (typically a synchronization word identifying the start of a metadata segment, followed by identifying information values, such as the core element version, length and period, extension element count and sub-stream association values ​​shown in Table 1 below); and after the core header, at least one protection value useful for at least one of decryption, authentication, and validation of the loudness processing state metadata and / or the corresponding audio data (e.g., an HMAC digest and an audio fingerprint value, where the HMAC digest may be a 256-bit HMAC digest calculated (using the SHA-2 algorithm) over the audio data of the entire frame, the core element, and all expanded elements as shown in Table 1); and Also after the core header, if the metadata segment contains an LPSM, an LPSM payload identification ("ID") and an LPSM payload size value that identifies the metadata that follows as an LPSM payload and indicates the size of the LPSM payload. An LPSM payload segment (preferably having the format described above) follows the LPSM payload ID and LPSM payload size values.

[0045] In some embodiments of the type described in the previous paragraph, each metadata segment in the ancillary data field (or "addbsi" field or extra bits segment) of a frame has a three-level structure: A high-level structure that contains a flag indicating whether the ancillary data (or addbsi) field contains metadata, at least one ID value indicating what type(s) of metadata are present, and typically also a value indicating how many bits of metadata (e.g. of each type) are present (if metadata is present). One type of metadata that can be present is LPSM, another type of metadata that can be present is program boundary metadata, and another type of metadata that can be present is media research metadata; a mid-level structure, which contains a core element for each identified type of metadata (e.g., a core header, protection value, and payload ID and payload size values ​​of the types described above for each identified type of metadata); and A low-level structure that includes each payload for a core element (e.g., an LPSM payload if the core element identifies that an LPSM payload is present and / or another type of metadata payload if the core element identifies that a metadata payload is present).

[0046] The data values ​​in such a three-level structure may be nested. For example, a protection value(s) for the LPSM payload and / or another metadata payload identified by the core element may be included after each payload identified by the core element (and thus after the core header of the core element). In one example, the core header may identify the LPSM payload and another metadata payload, a payload ID and payload size value for a first payload (e.g., an LPSM payload) may follow the core header, the first payload itself may follow the ID and size values, a payload ID and payload size value for a second payload may follow the first payload, the second payload itself may follow these ID and size values, and a protection value(s) for one or both of the payloads (or for the core element value and one or both of the payloads) may follow the final payload.

[0047] In some embodiments, a core element of a metadata segment in the auxiliary field (or "addbsi" field or redundancy bit segment) of a frame includes a core header (which typically includes an identification value, e.g., the core element version), followed by: a value indicating whether fingerprint data is included for the metadata of the metadata segment, a value indicating whether external data is present (related to the audio data corresponding to the metadata of the metadata segment), a payload ID and payload size value for each type of metadata identified by the core element (e.g., LPSM and / or non-LPSM types of metadata), and a protection value for at least one type of metadata identified by the core element. The metadata payload(s) of the metadata segment follow the core header and are (possibly) nested within the value of the core element.

[0048] In another preferred format, the encoded bitstream is a Dolby E bitstream, and each of the metadata segments containing LPSM (and optionally program boundary metadata) is contained within the first N sample positions of the Dolby E guard band interval.

[0049] In another class of embodiments, the invention is an APU (e.g., a decoder) coupled to and configured to receive an encoded audio bitstream including audio data segments and metadata segments, where the audio data segments represent audio data, and where each of at least some of the metadata segments includes loudness processing state metadata (LPSM) and optionally program boundary metadata. The APU is also coupled to and configured to, in response to the audio data, extract the LPSM from the bitstream to generate decoded audio data, and to perform at least one adaptive loudness processing operation on the audio data using the LPSM. Some embodiments of this class also include a post-processor coupled to the APU, where the post-processor is coupled to and configured to perform at least one adaptive loudness processing operation on the audio data using the LPSM.

[0050] In another class of embodiments, the invention is an audio processing unit (APU) including a buffer memory (buffer) and a processing subsystem coupled to the buffer. The APU is coupled to receive an encoded audio bitstream including audio data segments and metadata segments. The audio data segments represent audio data, and each of at least some of the metadata segments includes loudness processing state metadata (LPSM) and optionally program boundary metadata. The buffer stores at least one frame of the encoded audio bitstream (e.g., in a non-temporal manner), and the processing subsystem is configured to extract the LPSM from the bitstream and perform at least one adaptive loudness processing operation on the audio data using the LPSM. In an exemplary embodiment of this class, the APU is one of an encoder, a decoder, and a post-processor.

[0051] In some implementations of the method, the generated audio bitstream is one of an AC-3 encoded bitstream, an E-AC-3 bitstream, or a Dolby E bitstream and includes loudness processing state metadata and other metadata (e.g., DIALNORM metadata parameters, dynamic range control metadata parameters, and other metadata parameters). In other implementations of the method, the generated audio bitstream is another type of encoded bitstream.

[0052] Aspects of the invention include systems or devices configured (e.g., programmed) to perform any embodiment of the inventive method, as well as computer readable media (e.g., disks) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the inventive method, or steps thereof. For example, a system of the invention may be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed by software or firmware and / or otherwise configured to perform any of a variety of operations, including embodiments of the inventive method, or steps thereof, on data. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and processing circuitry that is programmed (and / or otherwise configured) to perform embodiments of the inventive method (or steps thereof) in response to data presented to it. [Brief description of the drawings]

[0053] [Figure 1] FIG. 1 is a block diagram of an embodiment of a system that may be configured to perform an embodiment of the method of the present invention. [Diagram 2] FIG. 2 is a block diagram of an encoder that is an embodiment of the audio processing unit of the present invention. [Diagram 3] 2 is a block diagram of a decoder which is an embodiment of an audio processing unit of the present invention, and a post-processor which is another embodiment of an audio processing unit of the present invention, coupled thereto; [Figure 4] 1 illustrates an AC-3 frame including the segments into which it is divided. [Diagram 5] 1 illustrates the synchronization information (SI) segment of an AC-3 frame, including the segments into which it is divided. [Figure 6] 1 illustrates a Bitstream Information (BSI) segment of an AC-3 frame, including the segments into which it is divided. [Figure 7]A diagram depicting an E-AC-3 frame, including the segments into which it is divided. [Figure 8] A diagram of frames of an encoded audio bitstream including program boundary metadata formatted according to an embodiment of the present invention. [Figure 9] 10 is a diagram of other frames of the encoded audio bitstream of FIG. 9, some of which include program boundary metadata having a format according to an embodiment of the present invention. [Figure 10] The diagram depicts two encoded audio bitstreams: one (IEB) whose program boundary (labeled "Boundary") is aligned with the transition between two frames of the bitstream, and another (TB) whose program boundary (labeled "True Boundary") is offset by 512 samples from the transition between two frames of the bitstream. [Figure 11]FIG. 11 is a set of diagrams showing four encoded audio bitstreams. The top bitstream in FIG. 11 (labeled "Scenario 1") shows a first audio program (P1) with program boundary metadata followed by a second audio program (P2) that also includes program boundary metadata; the second bitstream (labeled "Scenario 2") shows a first audio program (P1) with program boundary metadata followed by a second audio program (P2) that does not include program boundary metadata; the third bitstream (labeled "Scenario 3") shows a truncated first audio program (P1) with program boundary metadata spliced ​​with the entire second audio program (P2) with program boundary metadata; and the fourth bitstream (labeled "Scenario 4") shows a truncated first audio program (P1) with program boundary metadata and a truncated second audio program (P2) with program boundary metadata spliced ​​with a portion of the first audio program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0054] Notation and Nomenclature Throughout this disclosure, including the claims, the expression performing an operation "on" a signal or data (e.g., filtering, scaling, transforming or applying a gain to the signal or data) is used broadly to refer to performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing prior to performing the operation).

[0055] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to an apparatus, a system, or a subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0056] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors that are programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0057] Throughout this disclosure, including the claims, the terms "audio processor" and "audio processing unit" are used interchangeably and broadly to refer to a system configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools).

[0058] Throughout this disclosure, including the claims, the expression "processing state metadata" (e.g. as in the expression "loudness processing state metadata") refers to data that is separate and distinct from the corresponding audio data (the audio content of an audio data stream that also includes the processing state metadata). The processing state metadata is associated with the audio data and indicates the loudness processing state of the corresponding audio data (e.g. what type(s) of processing have already been performed on the audio data), and typically also indicates at least one feature or characteristic of the audio data. The association of the processing state metadata with the audio data is time-synchronous. In this way, the current (most recently received or updated) processing state metadata indicates that the corresponding audio data contemporaneously comprises the results of the indicated type(s) of audio data processing. In some cases, the processing state metadata may include some or all of the processing history and / or parameters used in and / or derived from the indicated type of processing. Furthermore, the processing state metadata may include at least one feature or characteristic of the corresponding audio data calculated or extracted from the audio data. The processing state metadata may also include other metadata not related to or derived from any processing of the corresponding audio data, such as third party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, and the like, which may be added by a particular audio processing unit and passed to other audio processing units.

[0059] Throughout this disclosure, including the claims, the expression "loudness processing state metadata" (or "LPSM") refers to processing state metadata that indicates the loudness processing state of the corresponding audio data (e.g., what type(s) of loudness processing have already been performed on the audio data), and typically also at least one feature or characteristic (e.g., loudness) of the corresponding audio data. The loudness processing state metadata may include data (e.g., other metadata) that is not (considered in isolation) loudness processing state metadata.

[0060] Throughout this disclosure, including the claims, the term "channel" (or "audio channel") refers to a monophonic audio signal.

[0061] Throughout this disclosure, including the claims, the term "audio program" refers to a collection of one or more audio channels and optionally associated metadata (e.g., metadata describing a desired spatial audio presentation and / or LPSM and / or program boundary metadata).

[0062] Throughout this disclosure, including the claims, the expression "program boundary metadata" refers to metadata of an encoded audio bitstream that indicates at least one audio program (e.g., two or more audio programs), where the program boundary metadata indicates the location in the bitstream of at least one boundary (beginning and / or end) of at least one of said audio programs. For example, program boundary metadata (of an encoded audio bitstream that indicates an audio program) may include metadata that indicates the location of the start of a program (e.g., the beginning of the Nth frame of the bitstream or the Mth sample position of the Nth frame of the bitstream) and additional metadata that indicates the location of the end of the program (e.g., the beginning of the Jth frame of the bitstream or the Kth sample position of the Jth frame of the bitstream).

[0063] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0064] Detailed Description of the Preferred Embodiments of the Invention According to an exemplary embodiment of the present invention, a payload of program loudness metadata, referred to as loudness processing state metadata ("LPSM"), and optionally also program boundary metadata, is embedded in one or more reserved fields (or slots) of a metadata segment of an audio bitstream, which also contains audio data in other segments (audio data segments). Typically, at least one segment of each frame of the bitstream contains the LPSM and at least one other segment of the frame contains corresponding audio data (i.e., audio data whose loudness processing state and loudness are indicated by said LPSM). In some embodiments, the data volume of the LPSM may be small enough to be carried without affecting the bitrate allocated to carry the audio data.

[0065] Communicating loudness processing state metadata in an audio data processing chain is particularly useful when two or more audio processing units need to function cascaded with each other throughout the processing chain (or content lifecycle). Without including loudness processing state metadata in the audio bitstream, serious media processing problems such as quality, level and spatial degradation can occur when, for example, two or more audio codecs are utilized in the chain and single-ended volume leveling is applied more than once during the bitstream's journey to the media consumption device (or rendering point of the bitstream's audio content).

[0066] 1 is a block diagram of an exemplary audio processing chain (audio data processing system), where one or more of the system's elements may be configured in accordance with an embodiment of the present invention. The system includes the following elements coupled together as shown: a pre-processing unit, an encoder, a signal analysis and metadata correction unit, a transcoder, a decoder, and a pre-processing unit. Variations of the illustrated system omit one or more of the elements or include additional audio data processing units.

[0067] In some implementations, the pre-processing unit of FIG. 1 is configured to accept as input PCM (time domain) samples comprising audio content and to output processed PCM samples. The encoder is configured to accept as input the PCM samples and to output an encoded (e.g., compressed) audio bitstream representative of the audio content. The bitstream data representative of the audio content is sometimes referred to herein as "audio data." When an encoder is configured according to an exemplary embodiment of the present invention, the audio bitstream output from the encoder includes loudness processing state metadata (and typically other metadata, optionally including program boundary metadata) in addition to the audio data.

[0068] The signal analysis and metadata correction unit of Figure 1 may accept as input one or more encoded audio bitstreams and determine (e.g., validate) whether the processing state metadata in each encoded audio bitstream is correct by performing signal analysis (e.g., using program boundary metadata in the encoded audio bitstreams). If the signal analysis and metadata correction unit finds that the included metadata is invalid, it typically replaces the incorrect value(s) with the correct value(s) obtained from the signal analysis. In this way, each encoded audio bitstream output from the signal analysis and metadata correction unit may include corrected (or uncorrected) processing state metadata in addition to the encoded audio data.

[0069] 1 may accept as input an encoded audio bitstream and, in response, output a modified (e.g., differently encoded) audio bitstream (e.g., by decoding the input stream and re-encoding the decoded stream in a different encoding format). When the transcoder is configured according to an exemplary embodiment of the present invention, the audio bitstream output from the transcoder includes the encoded audio data as well as loudness processing state metadata (and typically other metadata as well), which may be included in the bitstream.

[0070] The decoder of Figure 1 accepts as input an encoded (e.g. compressed) bitstream and (in response) outputs a stream of decoded PCM audio samples. When the decoder is configured according to an exemplary embodiment of the invention, the output of the decoder in exemplary operation is or includes any of the following: a stream of audio samples and a corresponding stream of loudness processing state metadata (and typically other metadata as well) extracted from the input encoded bitstream; or a stream of audio samples and a corresponding stream of control bits determined from loudness processing state metadata (and typically other metadata as well) extracted from the input encoded bitstream; or a stream of audio samples without a corresponding stream of processing state metadata or control bits determined from the processing state metadata. In this last case, the decoder may extract loudness processing state metadata (and / or other metadata) from the input encoded bitstream and perform at least one operation on the extracted metadata (e.g., validity checking) without outputting the extracted metadata or control bits determined therefrom.

[0071] By configuring the post-processing unit of Figure 1 in accordance with an exemplary embodiment of the present invention, the post-processing unit is configured to accept a stream of decoded PCM audio samples and perform post-processing thereon (e.g., volume leveling of the audio content) using the loudness processing state metadata (and typically other metadata as well) received with the samples or control bits (determined by the decoder from the loudness processing state metadata and typically other metadata as well). The post-processing unit is also typically configured to render the post-processed audio content for playback over one or more speakers.

[0072] An exemplary embodiment of the present invention provides an improved audio processing chain in which audio processing units (e.g. encoders, decoders, transcoders and pre-processing and post-processing units) adapt their respective processing according to the concurrent state of media data as indicated by loudness processing state metadata received by each audio processing unit.

[0073] Audio data input to any audio processing unit of the system of Figure 1 (e.g., an encoder or transcoder of Figure 1) may include loudness processing state metadata (and optionally other metadata) in addition to the audio data (e.g., encoded audio data). According to an embodiment of the invention, this metadata may have been included in the input audio by other elements of the system of Figure 1 (or other sources not shown in Figure 1). This processing unit receiving the input audio (with metadata) may perform at least one action (e.g., validation) on or in response to the metadata (e.g., adaptive processing of the input audio) and may typically also be configured to include in its output audio the metadata, a processed version of the metadata, or control bits determined from the metadata.

[0074] Exemplary embodiments of the audio processing unit (or audio processor) of the present invention are configured to perform adaptive processing of the audio data based on a state of the audio data indicated by loudness processing state metadata corresponding to the audio data. In some embodiments, the adaptive processing is (or includes) a loudness processing (if the metadata indicates that no loudness processing or similar processing has already been performed on the audio data), but is not (or does not include) a loudness processing (if the metadata indicates that such loudness processing or similar processing has already been performed on the audio data). In some embodiments, the adaptive processing is or includes a metadata validation (e.g. performed in a metadata validation subunit) to ensure that the audio processing unit performs other adaptive processing of the audio data based on the state of the audio data indicated by the loudness processing state metadata. In some embodiments, the validation determines the reliability of loudness processing state metadata associated with the audio data (e.g. included in the bitstream together with the audio data). For example, if the metadata is validated as reliable, results from a previously performed audio processing of a certain type may be reused and new execution of the same type of audio processing may be avoided. On the other hand, if the metadata is found to be tampered with (or otherwise unreliable), then the media processing of the type that was purportedly performed previously (as indicated by the unreliable metadata) may be repeated by the audio processing unit and / or other processing may be performed on the metadata and / or audio data by the audio processing unit.The audio processing unit may be configured to signal to other downstream audio processing units in the enhanced media processing chain that the loudness processing state metadata (e.g. present in the media bitstream) is valid if the audio processing unit determines that the processing state metadata is valid (e.g. based on a match of the extracted cryptographic value and the reference cryptographic value).

[0075] 2 is a block diagram of an encoder (100) that is an embodiment of an audio processing unit of the present invention. Any of the components or elements of the encoder 100 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA or other integrated circuits) in hardware, software or a combination of hardware and software. The encoder 100 includes a frame buffer 110, a parser 111, a decoder 101, an audio state validator 102, a loudness processing stage 103, an audio stream selection stage 104, an encoder 105, a stuffer / formatter stage 107, a metadata generation stage 106, a dialogue loudness measurement subsystem 108 and a frame buffer 109, connected as shown. The encoder 100 typically also includes other processing elements (not shown).

[0076] The encoder 100 (which is a transcoder) is configured to convert an input audio bitstream (which may be, for example, one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream) into an encoded output audio bitstream (which may be, for example, another of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream), including by performing adaptive and automated loudness processing using loudness processing state metadata included in the input bitstream. For example, the encoder 100 may be configured to convert an input Dolby E bitstream (a format typically used in production and broadcast facilities, but not typically used in consumer devices that receive broadcasted audio programs) into an encoded output audio bitstream in AC-3 or E-AC-3 format (suitable for broadcast to consumer devices).

[0077] 2 also includes an encoded audio delivery subsystem 150 (which stores and / or delivers the encoded bitstream output from encoder 100), and a decoder 152. The encoded audio bitstream output from encoder 100 may be stored by subsystem 150 (e.g., in the form of a DVD or Blu-ray disc), or may be transmitted by subsystem 150 (which may implement a transmission link or network), or may be both stored and transmitted by subsystem 150. Decoder 152 is configured to decode the encoded audio bitstream (produced by encoder 100) received via subsystem 150, including by extracting loudness processing state metadata (LPSM) from each frame of the bitstream (and optionally also program boundary metadata from the bitstream) to generate decoded audio data. Typically, decoder 152 is configured to perform adaptive loudness processing on the decoded audio data using the LPSM (and optionally also the program boundary metadata) and / or forward the decoded audio data and the LPSM to a post-processor configured to perform adaptive loudness processing on the decoded audio data using the LPSM (and optionally also the program boundary metadata). Typically, decoder 152 includes a buffer that stores (e.g., in a non-temporary manner) the encoded audio bitstream received from subsystem 150.

[0078] Various implementations of the encoder 100 and decoder 152 are configured to perform various embodiments of the method of the present invention.

[0079] Frame buffer 110 is a buffer memory coupled to receive the input encoded audio bitstream. In operation, buffer 110 stores (e.g., in a non-transient manner) at least one frame of the encoded audio bitstream, and a sequence of frames of the encoded audio bitstream are presented from buffer 110 to parser 111.

[0080] The parser 111 is coupled to and configured to extract loudness processing metadata (LPSM), and optionally program boundary metadata (and / or other metadata) from each frame of encoded input audio containing such metadata, provide at least the LPSM (and optionally the program boundary metadata and / or other metadata) to the audio state validator 102, the loudness processing stage 103, stage 106 and subsystem 108, extract audio data from the encoded input audio and provide the audio data to the decoder 101. The decoder 101 of the encoder 100 is configured to decode the audio data to generate decoded audio data and provide the decoded audio data to the loudness processing stage 103, the audio stream selection stage 104, subsystem 108 and typically also to the state validator 102.

[0081] The state validator 102 is configured to authenticate and validate the LPSM (and optionally other metadata) presented to it. In some embodiments, the LPSM is (or is included in) a data block that was included in the input bitstream (e.g., in accordance with an embodiment of the present invention). The block may include a cryptographic hash (a Hash-Based Message Authentication Code or "HMAC") for processing the LPSM (and optionally other metadata) and / or the underlying audio data (provided to the validator 102 from the decoder 101). The data block may, in these embodiments, be digitally signed, so that downstream audio processing units may relatively easily authenticate and validate the processing state metadata.

[0082] For example, an HMAC may be used to generate a digest that the protection value(s) included in the bitstream of the present invention may include. For an AC-3 frame, the digest may be generated as follows: 1. After the AC-3 data and the LPSM are encoded, the frame data bytes (concatenated frame data #1 and frame data #2) and the LPSM data bytes are used as input for the hash function HMAC. Any other data that may be present in the ancillary data field is not taken into account for computing this digest. Such other data may be bytes that do not belong to either the AC-3 data or the LSPSM data. The protection bits contained in the LPSM may not be taken into account for computing the HMAC digest. 2. After the digest is calculated, it is written to the bitstream in the field reserved for the protection bits. 3. The final step in the generation of a complete AC-3 frame is the calculation of the CRC check, which is written at the very end of the frame and takes into account all the data belonging to this frame, including the LPSM bits.

[0083] Other cryptographic methods, including but not limited to any one or more non-HMAC cryptographic methods, may be used for validation of the LPSM (e.g., in validator 102) to ensure secure transmission and receipt of the LPSM and / or the underlying audio data. For example, validation (using such cryptographic methods) may be performed at each audio processing unit receiving an embodiment of an audio bitstream of the present invention to determine whether loudness processing state metadata and corresponding audio data included in the bitstream have been subjected to (and / or result from) a particular loudness processing (as indicated by the metadata) and have not been modified since such particular loudness processing was performed.

[0084] The status validator 102 presents control data to the audio stream selection stage 104, the metadata generator 106 and the dialogue loudness measurement subsystem 108 to indicate the result of the validation operation. In response to the control data, stage 104 can select (and communicate to the encoder 105) either: the adaptively processed output of loudness processing stage 103 (e.g. when LPSM indicates that the audio data output from decoder 101 has not been subjected to a particular type of loudness processing and the control bit from validator 102 indicates that LPSM is enabled); or (e.g., when LPSM indicates that the audio data output from decoder 101 has already undergone a particular type of loudness processing to be performed by stage 103 and a control bit from validity checker 102 indicates that LPSM is valid) said audio data output from decoder 101.

[0085] Stage 103 of encoder 100 is configured to perform adaptive loudness processing on the decoded audio data output from decoder 101 based on one or more audio data characteristics indicated by the LPSM extracted by decoder 101. Stage 103 may be an adaptive transform-domain real-time loudness and dynamic range control processor. Stage 103 may receive user input (e.g., user target loudness / dynamic range values ​​or dialnorm values) or other metadata input (e.g., one or more types of third party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, etc.) and / or other information (e.g., from a fingerprinting process) and use such input to process the decoded audio data output from decoder 101. Stage 103 may perform adaptive loudness processing on decoded audio data (output from decoder 101) that indicates a single audio program (indicated by the program boundary metadata extracted by parser 111), and may reset the loudness processing in response to receiving decoded audio data (output from decoder 101) that indicates a different audio program as indicated by the program boundary metadata extracted by parser 111.

[0086] The dialogue loudness measurement subsystem 108 may operate to determine the loudness of segments of decoded audio (from decoder 101) that indicate dialogue (or other speech) using, for example, the LPSM (and / or other metadata) extracted by decoder 101 if the control bit from validity checker 102 indicates that LPSM is disabled. If the control bit from validity checker 102 indicates that LPSM is enabled, operation of the dialogue loudness measurement subsystem 108 may be disabled when the LPSM indicates a previously determined loudness of a dialogue (or other speech) segment of the decoded audio (from decoder 101). Subsystem 108 may perform loudness measurements on decoded audio data that indicates a single audio program (indicated by program boundary metadata extracted by parser 111) and may reset said measurements in response to receiving decoded audio data that indicates a different audio program as indicated by such program boundary metadata.

[0087] There are available tools (e.g., the Dolby LM100) for conveniently and easily measuring the level of dialogue in audio content. Some embodiments of the APU of the present invention (e.g., stage 108 of encoder 100) are implemented to include such tools (or to perform the functionality of such tools) for measuring the average dialogue loudness of the audio content of an audio bitstream (e.g., the decoded AC-3 bitstream presented to stage 108 from decoder 101 of encoder 100).

[0088] If stage 108 is implemented to measure the true average dialogue loudness of the audio data, the measurement may include isolating segments of the audio content that contain primarily speech. The primarily speech audio segments are then processed according to a loudness measurement algorithm. For audio data decoded from an AC-3 bitstream, this algorithm may be the standard K-weighted loudness measure (in accordance with international standard ITU-R BS.1770). Alternatively, other loudness measures (e.g. based on a psychoacoustic model of loudness) may be used.

[0089] Isolating speech segments is not essential for measuring the average dialogue loudness of audio data, but it improves the accuracy of the metric and typically gives more satisfying results from the listener's point of view. Since not all audio content contains dialogue, a loudness metric for the entire audio content may provide a sufficient approximation of the dialogue level of that audio if speech was present.

[0090] The metadata generator 106 generates (and / or passes to) stage 107 metadata to be included by stage 107 in the encoded bitstream output from the encoder 100. The metadata generator 106 may pass to stage 107 the LPSMs (and optionally also the program boundary metadata and / or other metadata) extracted by the encoder 101 and / or parser 111 (e.g., if the control bits from the validity checker 102 indicate that the LPSMs and / or other metadata are valid), or may generate new LPSMs (and optionally also the program boundary metadata and / or other metadata) and present the new metadata to stage 107 (e.g., if the control bits from the validity checker 102 indicate that the LPSMs and / or other metadata extracted by the decoder 101 are invalid). Alternatively, the metadata generator 106 may present to stage 107 a combination of the metadata extracted by the decoder 101 and / or parser 111 and the newly generated metadata. Metadata generator 106 may include the loudness data generated by subsystem 108 and at least one value indicative of the type of loudness processing performed by subsystem 108 in an LPSM that is submitted to stage 107 for inclusion in the encoded bitstream output from encoder 100.

[0091] The metadata generator 106 may generate protection bits (which may consist of or include a hash-based message authentication code or "HMAC") useful for at least one of decrypting, authenticating, or validating the LPSM (and optionally other metadata) to be included in the encoded bitstream and / or the underlying audio data to be included in the encoded bitstream.

[0092] In typical operation, the dialogue loudness measurement subsystem 108 processes the audio data output from the decoder 101 and, in response, generates loudness values ​​(e.g., gated and ungated dialogue loudness values) and dynamic range values. In response to these values, the metadata generator 106 may generate loudness processing state metadata (LPSM) for inclusion (by the stuffer / formatter 107) in the encoded bitstream output from the encoder 100.

[0093] Additionally, optionally or alternatively, subsystems 106 and / or 108 of encoder 100 may perform additional analysis of the audio data to generate metadata indicative of at least one characteristic of the audio data for inclusion in the encoded bitstream output from stage 107.

[0094] Encoder 105 encodes the audio data output from selection stage 104 (e.g. by performing compression on it) and presents the encoded audio to stage 107 for inclusion in an encoded bitstream output from stage 107.

[0095] Stage 107 multiplexes the encoded audio from encoder 105 and the metadata (including LPSM) from generator 106 to generate an encoded bitstream that is output from stage 107. Preferably, the encoded bitstream has a format specified by a preferred embodiment of the present invention.

[0096] Frame buffer 109 is a buffer memory that stores (e.g., in a non-transient manner) at least one frame of the encoded audio bitstream output from stage 107. A sequence of those frames of the encoded audio bitstream are then presented from buffer 109 to a delivery system 150 as output from encoder 100.

[0097] The LPSM generated by metadata generator 106 and included in the encoded bitstream by stage 107 indicates the loudness processing state of the corresponding audio data (e.g. what type(s) of loudness processing have been performed on the audio data) and the loudness of the corresponding audio data (e.g. measured dialogue loudness, gated and / or ungated loudness and / or dynamic range).

[0098] In this document, "gating" level measurements performed on loudness and / or audio data refers to a particular level or loudness threshold such that any calculated value or values ​​above the threshold are included in the final measurement (e.g., ignoring short-term loudness values ​​below -60 dBFS in the final measured value). Gating on an absolute value refers to a fixed level or loudness, while gating on a relative value refers to a value that is dependent on the current "ungated" measurement.

[0099] In some implementations of the encoder 100, the encoded bitstream buffered in the memory 109 (and output to the delivery system 150) is an AC-3 or E-AC-3 bitstream and includes audio data segments (e.g., segments AB0-AB5 of the frame shown in FIG. 4) and metadata segments, where the audio data segments represent audio data and where each of at least some of the metadata segments includes loudness processing state metadata (LPSM). Stage 107 inserts the LPSMs (and optionally also the program boundary metadata) into the bitstream in the following format: Each metadata segment that includes the LPSMs (and optionally also the program boundary metadata) is included in an extra bits segment of the bitstream (e.g., the extra bits segment "W" shown in FIG. 4 or FIG. 7) or in an "addbsi" field of a bitstream information ("BSI") segment of a frame of the bitstream or in an auxiliary data field at the end of a frame of the bitstream (e.g., the AUX segment shown in FIG. 4 or FIG. 7). A frame of the bitstream may contain one or two metadata segments, each containing an LPSM, and if a frame contains two metadata segments, one may be present in the addbsi field of the frame and the other in the AUX field of the frame. In some embodiments, each metadata segment containing an LPSM contains an LPSM payload (or container) segment with the following format: a header (typically including a synchronization word identifying the beginning of the LPSM payload, followed by at least one identification value, such as the LPSM format version, length, period, count and sub-stream association values ​​shown in Table 2 below); After the header, at least one dialogue indication value (e.g., parameter "Dialogue Channels" in Table 2) indicating whether the corresponding audio data indicates dialogue or not (e.g., which channels of the corresponding audio data indicate dialogue); at least one loudness regulation compliance value indicating whether the corresponding audio data complies with an indicated set of loudness regulations (e.g., the parameter "Loudness Regulation Type" in Table 2); At least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data (e.g., one or more of the parameters "Dialogue-Gated Loudness Compensation Flag" and "Loudness Compensation Type" in Table 2); and at least one loudness value indicating at least one loudness (e.g. peak or average loudness) characteristic of the corresponding audio data (e.g. one or more of the parameters "ITU relative gated loudness", "ITU speech gated loudness", "ITU (EBU3341) short-term 3s loudness" and "true peak");

[0100] In some embodiments, each metadata segment containing LPSM and program boundary metadata includes a Core Header (and optionally additional Core Elements) followed by an LPSM payload (or container) segment having the following format: A header, which typically contains at least one identification value (e.g., LPSM format version, length, period, count and sub-stream association values, such as those shown in Table 2 herein); After the header comes the LPSM and program boundary metadata. The program boundary metadata may include a program boundary frame count, a code value (e.g., an "offset_exist" value) indicating whether the frame contains only the program boundary frame count or both the program boundary frame count and an offset value, and (optionally) an offset value.

[0101] In some implementations, each of the metadata segments inserted by stage 107 into the "addbsi" field or ancillary data field of the redundant bits segment or frame of the bitstream has the following format: A Core Header (typically containing a sync word identifying the start of a metadata segment, followed by identifying information values, such as the Core Element Version, Length and Period, Extension Element Count and Sub-Stream Association values ​​shown in Table 1 below); and after the core header, at least one protection value useful for at least one of decryption, authentication, and validation of the loudness processing state metadata and / or the corresponding audio data (e.g., the HMAC digest and audio fingerprint values ​​of Table 1); and Also after the core header, if the metadata segment contains an LPSM, an LPSM payload identification ("ID") and LPSM payload size value that identifies the following metadata as an LPSM payload and indicates the size of the LPSM payload.

[0102] An LPSM payload (or container) segment (preferably having the format specified above) follows the LPSM Payload ID and LPSM Payload Size values.

[0103] In some embodiments, each metadata segment in a frame's ancillary data field (or "addbsi" field) has a three-level structure: A high-level structure that contains a flag indicating whether the ancillary data (or addbsi) field contains metadata, at least one ID value indicating what type(s) of metadata are present, and typically also a value indicating how many bits of metadata (e.g. of each type) are present (if metadata is present). One type of metadata that can be present is LPSM, another type of metadata that can be present is program boundary metadata, and another type of metadata that can be present is media research metadata (e.g. Nielsen Media Research metadata); an intermediate level structure, which contains a core element for each identified type of metadata (e.g., a core header, protection value, and LPSM payload ID and LPSM payload size values ​​as described above for each identified type of metadata); and A low-level structure that includes each payload for a core element (e.g., an LPSM payload if the core element identifies that an LPSM payload is present and / or another type of metadata payload if the core element identifies that a metadata payload is present).

[0104] The data values ​​in such a three-level structure may be nested. For example, a protection value(s) for the LPSM payload and / or another metadata payload identified by the core element may be included after each payload identified by the core element (and thus after the core header of the core element). In one example, the core header may identify the LPSM payload and another metadata payload, a payload ID and payload size value for a first payload (e.g., an LPSM payload) may follow the core header, the first payload itself may follow the ID and size values, a payload ID and payload size value for a second payload may follow the first payload, the second payload itself may follow these ID and size values, and protection bit(s) for both payloads (or for the core element value and both payloads) may follow the last payload.

[0105] In some embodiments, when the decoder 101 receives an audio bitstream generated according to an embodiment of the present invention having a cryptographic hash, the decoder is configured to parse and extract the cryptographic hash from a determined data block from the bitstream, the block including loudness processing state metadata (LPSM) and optionally program boundary metadata. The validator 102 may use the cryptographic hash to validate the received bitstream and / or associated metadata. For example, if the validator 102 finds that the LPSM is valid based on a match between a reference cryptographic hash and the cryptographic hash extracted from the data block, the validator 102 may disable the operation of the processor 103 on the corresponding audio data and may pass the audio data through (without modification) to the selection stage 104. Additionally, optionally or alternatively, other types of cryptographic techniques may be used instead of cryptographic hash-based methods.

[0106] 2 may determine (in response to the LPSM extracted by decoder 101, and optionally also to program boundary metadata) that a post / pre-processing unit has performed some type of loudness processing on the audio data to be encoded (in elements 105, 106 and 107), and may thus generate (in generator 106) loudness processing state metadata that includes certain parameters used in and / or derived from the previously performed loudness processing. In some implementations, encoder 100 may generate (and include in the output encoded bitstream from) processing state metadata that indicates the processing history on the audio content, so long as the encoder recognizes the type of processing that has been performed on the audio content.

[0107] 3 is a block diagram of a decoder (200) and a post-processor (300) coupled thereto, which are an embodiment of an audio processing unit of the present invention. The post-processor (300) is also an embodiment of an audio processing unit of the present invention. Any of the components or elements of the decoder 200 and the post-processor 300 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA or other integrated circuits) in hardware, software or a combination of hardware and software. The decoder 200 includes a frame buffer 201, a parser 205, an audio decoder 202, an audio state validity check stage (validity checker) 203 and a control bit generation stage 204, connected as shown. Typically, the decoder 200 also includes other processing elements (not shown).

[0108] The frame buffer 201 (buffer memory) stores (e.g., in a non-transient manner) at least one frame of the encoded audio bitstream received by the decoder 200. A sequence of frames of the encoded audio bitstream are presented from the buffer 201 to a parser 205.

[0109] The parser 205 is coupled and configured to extract loudness processing metadata (LPSM), and optionally also program boundary metadata and other metadata, from each frame of the encoded input audio, provide at least the LPSM (and also the program boundary metadata if program boundary metadata is extracted) to the audio state validator 203 and stages 204, provide the LPSM (and optionally also the program boundary metadata) as output (e.g. to the post-processor 300), extract audio data from the encoded input audio, and provide the extracted audio data to the decoder 202.

[0110] The encoded audio bitstream input to the decoder 200 may be one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream.

[0111] The system of Fig. 3 also includes a post-processor 300. The post-processor 300 has a frame buffer 301 and other processing elements (not shown) including at least one processing element coupled to the buffer 301. The frame buffer 301 stores (e.g. in a non-temporary manner) at least one frame of the decoded audio bitstream received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are coupled and configured to receive a sequence of frames of the decoded audio bitstream output from the buffer 301 and adaptively process using metadata (including LPSM) output from the decoder 202 and / or control bits output from stage 204 of the decoder 200. Typically, the post-processor 300 is configured to perform adaptive loudness processing on the decoded audio data using the LPSM values ​​and optionally also the program boundary metadata (e.g. based on a loudness processing state and / or one or more audio data characteristics indicated by the LPSM for the audio data indicating a single audio program).

[0112] Various implementations of the decoder 200 and post-processor 300 are configured to perform various embodiments of the method of the present invention.

[0113] The audio decoder 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 to generate decoded audio data and to present the decoded audio data as output (e.g. to a post-processor 300).

[0114] The state validator 203 is configured to authenticate and validate the LPSM (and optionally other metadata) presented to it. In some embodiments, the LPSM is (or is included in) a data block that is included in the input bitstream (e.g., in accordance with an embodiment of the present invention). The block may include a cryptographic hash (Hash-Based Message Authentication Code or "HMAC") for processing the LPSM (and optionally other metadata) and / or the underlying audio data (provided to the validator 203 from the parser 205 and / or the decoder 202). The data block may, in these embodiments, be digitally signed, so that downstream audio processing units may relatively easily authenticate and validate the processing state metadata.

[0115] Other cryptographic methods, including but not limited to any one or more non-HMAC cryptographic methods, may be used for validation of the LPSM (e.g., in validator 203) to ensure secure transmission and reception of the LPSM and / or underlying audio data. For example, validation (using such cryptographic methods) may be performed at each audio processing unit receiving an embodiment of an audio bitstream of the present invention to determine whether loudness processing state metadata and corresponding audio data included in the bitstream have been subjected to (and / or result from) a particular loudness processing (as indicated by the metadata) and have not been modified since such particular loudness processing was performed.

[0116] State validator 203 provides control data to control bit generator 204 and / or provides the control data as output (e.g. to post-processor 300) to indicate the result of the validation operation. In response to the control data (and optionally other metadata extracted from the input bitstream), stage 204 may generate (and provide to post-processor 300) any of the following: a control bit indicating that the decoded audio data output from the decoder 202 has undergone a particular type of loudness processing (e.g., when LPSM indicates that the audio data output from the decoder 202 has undergone the particular type of loudness processing and the control bit from the validity checker 203 indicates that LPSM is enabled); or A control bit indicating that the decoded audio data output from decoder 203 should be subjected to a particular type of loudness processing (for example, when LPSM indicates that the audio data output from decoder 202 has not been subjected to a particular type of loudness processing, or when LPSM indicates that the audio data output from decoder 202 has been subjected to a particular type of loudness processing but the control bit from validity checker 203 indicates that LPSM is not valid).

[0117] Alternatively, the decoder 200 presents the metadata extracted by the decoder 202 from the input bitstream and the LPSM (and optionally also the program boundary metadata) extracted by the parser 205 from the input bitstream to a post-processor 300 which uses the LPSM (and optionally also the program boundary metadata) to perform loudness processing on the decoded audio data, performs a validity check of the LPSM, and then, if the validity check indicates that the LPSM is valid, performs loudness processing on the decoded audio data using the LPSM (and optionally also the program boundary metadata).

[0118] In some embodiments, when the decoder 200 receives an audio bitstream generated according to an embodiment of the present invention having a cryptographic hash, the decoder is configured to parse and extract the cryptographic hash from a determined data block from the bitstream, the block including loudness processing state metadata (LPSM). The validator 203 may use the cryptographic hash to validate the received bitstream and / or associated metadata. For example, if the validator 203 finds the LPSM to be valid based on a match between a reference cryptographic hash and the cryptographic hash extracted from the data block, the validator 203 may signal a downstream audio processing unit (e.g., a post-processor 300 that is or may include a volume leveling unit) to pass through the audio data of the bitstream (without modification). Additionally, optionally or alternatively, other types of cryptographic techniques may be used instead of a cryptographic hash-based method.

[0119] In some implementations of the decoder 200, the received encoded bitstream (and buffered in the memory 201) is an AC-3 or E-AC-3 bitstream and includes audio data segments (e.g., the AB0-AB5 segments of the frame shown in FIG. 4) and metadata segments, where the audio data segments represent audio data and where each of at least some of the metadata segments also includes loudness processing state metadata (LPSM) and optionally program boundary metadata. The decoder stage 202 (and / or the parser 205) is configured to extract from the bitstream the LPSM (and optionally also the program boundary metadata) having the following format: Each of the metadata segments including the LPSM (and optionally also the program boundary metadata) is included in the extra bits segment of the frame of the bitstream or in the "addbsi" field of the bitstream information ("BSI") segment of the frame of the bitstream, or in an auxiliary data field at the end of the frame of the bitstream (e.g., the AUX segment shown in FIG. 4). A frame of the bitstream may contain one or two metadata segments each containing an LPSM, and if a frame contains two metadata segments, one may be present in the addbsi field of the frame and the other may be present in the AUX field of the frame. In some embodiments, each metadata segment containing an LPSM contains an LPSM payload (or container) segment having the following format:

[0120] a header (which typically includes a synchronization word identifying the beginning of the LPSM payload, followed by identification values, such as the LPSM format version, length, period, count and sub-stream association values ​​shown in Table 2 below); After the header, at least one dialogue indication value (e.g., parameter "Dialogue Channels" in Table 2) that indicates whether the corresponding audio data indicates dialogue or not (e.g., which channels of the corresponding audio data indicate dialogue); at least one loudness regulation compliance value indicating whether the corresponding audio data complies with an indicated set of loudness regulations (e.g. the parameter "Loudness regulation type" in Table 2); At least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data (e.g., one or more of the parameters "Dialogue-Gated Loudness Compensation Flag" and "Loudness Compensation Type" in Table 2); and At least one loudness value indicating at least one loudness (e.g. peak or average loudness) characteristic of the corresponding audio data (e.g. one or more of the parameters "ITU relative gated loudness", "ITU speech gated loudness", "ITU (EBU3341) short-term 3s loudness" and "true peak" in Table 2).

[0121] In some embodiments, each metadata segment containing LPSM and program boundary metadata includes a Core Header (and optionally additional Core Elements) followed by an LPSM payload (or container) segment having the following format: A header, which typically contains at least one identification value (e.g., LPSM format version, length, period, count and sub-stream association values, as shown in Table 2 below); After the header comes the LPSM and program boundary metadata. The program boundary metadata may include a program boundary frame count, a code value (e.g., an "offset_exist" value) indicating whether the frame contains only the program boundary frame count or both the program boundary frame count and an offset value, and (optionally) an offset value.

[0122] In some implementations, the parser 205 (and / or the decoder stage 202) is configured to extract from the extra bits segment or the "addbsi" field or the ancillary data field of a frame of the bitstream, each metadata segment having the following format: A Core Header (typically including a sync word identifying the start of a metadata segment, followed by at least one identifying information value, such as the Core Element Version, Length and Period, Extension Element Count and Sub-Stream Association values ​​shown in Table 1 below); and after the core header, at least one protection value useful for at least one of decryption, authentication, and validation of the loudness processing state metadata and / or the corresponding audio data (e.g., the HMAC digest and audio fingerprint values ​​of Table 1); and Also after the core header, if the metadata segment contains an LPSM, an LPSM payload identification ("ID") and LPSM payload size value that identifies the following metadata as an LPSM payload and indicates the size of the LPSM payload.

[0123] An LPSM payload (or container) segment (preferably having the format specified above) follows the LPSM Payload ID and LPSM Payload Size values.

[0124] More generally, encoded audio bitstreams produced by preferred embodiments of the present invention have a structure that provides a mechanism for labeling metadata elements and sub-elements as core (mandatory) or extension (optional elements). This allows the data rate of the bitstream (including metadata) to be scaled across a multitude of applications. Core (mandatory) elements of the preferred bitstream syntax should also be able to signal the presence (in-band) and / or remote (out of band) locations of extension (optional) elements associated with the audio content.

[0125] A core element or elements are required to be present in all frames of the bitstream. Some sub-elements of a core element are optional and may be present in any combination. Extension elements are not required to be present in all frames (to limit bitrate overhead). Thus, an extension element may be present in some frames and absent in other frames. Some sub-elements of an extension element are optional and may be present in any combination, while some sub-elements of an extension element may be mandatory (i.e., mandatory if the extension element is present in a frame of the bitstream).

[0126] In one class of embodiments, an encoded audio bitstream is generated (e.g., by an audio processing unit embodying the present invention) that includes a sequence of audio data segments and metadata segments. The audio data segments represent audio data, and each of at least some of the metadata segments also includes loudness processing state metadata (LPSM) and optionally program boundary metadata, and the audio data segments are time division multiplexed with the metadata segments. In a preferred embodiment of this class, each of the metadata segments has a preferred format described herein.

[0127] In one preferred format, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each metadata segment containing an LPSM is included as additional bitstream information in an "addbsi" field (shown in FIG. 6) of a bitstream information ("BSI") segment of a frame of the bitstream, or in an auxiliary data field of a frame of the bitstream, or in an extra bits segment of a frame of the bitstream (e.g., by stage 107 of a preferred implementation of encoder 100).

[0128] In the preferred format described above, each frame contains a core element in the addbsi field (or redundant bits segment) of the frame, which has the format shown in Table 1 below.

[0129] [Table 1] In the preferred format, each addbsi (or ancillary data) field or redundancy bit segment that contains an LPSM includes a core header (and optionally additional core elements) and, after the core header (or after the core header and other core elements), the following LPSM values ​​(parameters): Payload ID (identifies the metadata as an LPSM), which follows the Core Element Value (e.g., as specified in Table 1); Payload Size (indicates the size of the LPSM payload). This follows the Payload ID; LPSM data (following the Payload ID and Payload Size values), which has the format shown in the following table (Table 2).

[0130] [Table 2-1] [Table 2-2] In another preferred format of an encoded bitstream generated in accordance with the present invention, the bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each of the metadata segments containing the LPSM (and optionally also the program boundary metadata) is included (e.g., by stage 107 in a preferred implementation of encoder 100) in any of: an extra bits segment of a frame of the bitstream; or an "addbsi" field of a Bitstream Information ("BSI") segment of a frame of the bitstream (as shown in FIG. 6); or an auxiliary data field at the end of a frame of the bitstream (e.g., the AUX field shown in FIG. 4). A frame may contain one or two metadata segments, each containing an LPSM, and if a frame contains two metadata segments, one may be present in the addbsi field of the frame and the other may be present in the AUX field of the frame. Each metadata segment containing an LPSM has the format specified above with reference to Tables 1 and 2 above (i.e., it contains the core elements specified in Table 1, followed by a payload ID (which identifies the metadata as an LPSM) and a payload size value as specified above, followed by the payload (the LPSM data having the format shown in Table 2).

[0131] In another preferred format, the encoded bitstream is a Dolby E bitstream, and each of the metadata segments containing LPSM (and optionally also program boundary metadata) is the first N sample locations of a Dolby E guard band interval. A Dolby E bitstream containing such a metadata segment containing LPSM preferably includes a value indicating the LPSM payload length signaled in a Pd word of the SMPTE 337M preamble (the SMPTE 337M Pa word repetition rate preferably remains the same as the associated video frame rate).

[0132] In a preferred format in which the encoded bitstream is an E-AC-3 bitstream, each of the metadata segments containing LPSMs (and optionally also program boundary metadata) is included (e.g., by stage 107 of a preferred implementation of encoder 100) as additional bitstream information in an extra bits segment or in an "addbsi" field of a bitstream information ("BSI") segment of a frame of the bitstream. Further aspects of encoding an E-AC-3 bitstream with LPSMs in this preferred format are now described.

[0133] 1. During the generation of an E-AC-3 bitstream, while the E-AC-3 encoder (inserting LPSM values ​​into the bitstream) is "active", for every frame generated (sync frame), the bitstream should contain a metadata block (containing LPSM) carried in the addbsi field (or extra bits segment) of the frame. The bits required to carry the metadata block should not increase the encoder bitrate (frame length).

[0134] 2. All metadata blocks (including LPSM) should contain the following information: loudness_correction_type_flag: where "1" indicates that the loudness of the corresponding audio data has been corrected upstream of the encoder, and "0" indicates that the loudness has been corrected by a loudness corrector integrated into the encoder (e.g., loudness processor 103 of encoder 100 of FIG. 2); speech_channel: indicates which source channel(s) contain speech (during the previous 0.5 seconds). If no speech is detected, this is indicated; speech_loudness: indicates the integrated speech loudness (over the previous 0.5 seconds) of each corresponding audio channel that contains speech; ITU_loudness: indicates the combined ITU BS.1770-3 loudness of each corresponding audio channel; Gain: The loudness complex gain(s) to invert at the decoder (to demonstrate reversibility).

[0135] 3. (Insert LPSM values ​​into the bitstream) While an E-AC-3 encoder is “active” and is receiving an AC-3 frame with the “trusted” flag, the loudness controller in the encoder (e.g. loudness processor 103 of encoder 100 in FIG. 2) should be bypassed. The “trusted” source dialnorm and DRC values ​​should be passed (e.g. by generator 106 of encoder 100) to the E-AC-3 encoder component (e.g. stage 107 of encoder 100). LPSM block generation continues and loudness_correction_type_flag is set to “1”. The loudness controller bypass sequence should be synchronized to the beginning of a decoded AC-3 frame where the “trusted” flag appears. The loudness controller bypass sequence should be implemented as follows: The leveler_amount control is decremented from value 9 to value 0 over 10 audio block periods (i.e., 53.3 msec) and the leveler_back_end_meter control is put into bypass mode (this should give a seamless transition). The term "trusted" bypass of the leveler implies that the dialnorm value of the source bitstream is also reused at the encoder's output (e.g., if a "trusted" source bitstream has a dialnorm value of -30, the encoder's output should use -30 for the outgoing dialnorm value). While an E-AC-3 encoder (which inserts LPSM values ​​into the bitstream) is "active" and is receiving AC-3 frames without the "trusted" flag, the loudness controller built into that encoder (e.g., loudness processor 103 of encoder 100 in FIG. 2) should be active. LPSM block generation continues and loudness_correction_type_flag is set to '0'. The loudness controller activation sequence should be synchronized to the beginning of the decoded AC-3 frame where the 'confidence' flag disappears.The loudness controller activation sequence should be implemented as follows: the leveler_amount control is incremented from value 0 to value 9 over one audio block period (i.e. 5.3 msec), and the leveler_back_end_meter control is put into "active" mode (this action should give a seamless transition and include an integral reset of the back_end_meter).

[0136] 5. During encoding, the Graphical User Interface (GUI) should show the following parameters to the user: "Input Audio Program [Trusted / Not Trusted]" - the state of this parameter is based on the presence of the "Trusted" flag in the input signal; and "Real-time Loudness Correction: [Enabled / Disabled]" - the state of this parameter is based on whether the loudness controller built into the encoder is active or not.

[0137] When decoding an AC-3 or E-AC-3 bitstream with LPSMs included in the extra bits segment or the "addbsi" field of the Bitstream Information ("BSI") segment of each frame of the bitstream (in the preferred format above), the decoder SHOULD parse the LPSM block data (in the extra bits segment or addbsi field) and pass all of the extracted LPSM values ​​to the Graphical User Interface (GUI). The set of extracted LPSM values ​​is refreshed every frame.

[0138] In another preferred format of an encoded bitstream generated in accordance with the present invention, the encoded bitstream is an AC-3 or E-AC-3 bitstream, and each of the metadata segments that contains an LPSM is included (e.g., by stage 107 of a preferred implementation of encoder 100) in an Extra Bits segment, or in an Aux segment, or as additional bitstream information in an "addbsi" field (shown in FIG. 6) of a Bitstream Information ("BSI") segment of a frame of the bitstream. In this format (which is a variation on the formats described above with reference to Tables 1 and 2), each of the addbsi (or Aux or Extra Bits) fields that contains an LPSM contains the following LPSM value:

[0139] The core elements are as specified in Table 1, followed by a payload ID (which identifies the metadata as an LPSM) and a payload size value, followed by the payload (the LPSM data). The LPSM data has the following format (similar to the required elements shown in Table 2 above):

[0140] LPSM Payload Version: A 2-bit field indicating the version of the LPSM payload.

[0141] dialchan: a 3-bit field indicating whether the left, right and / or center channel of the corresponding audio data contains spoken dialogue. The bit assignment of the dialchan field may be as follows: bit 0, indicating the presence of dialogue in the left channel, is stored in the most significant bit of the dialchan field, and bit 2, indicating the presence of dialogue in the center channel, is stored in the least significant bit of the dialchan field. Each bit of the dialchan field is set to "1" if the corresponding channel contains dialogue spoken during the preceding 0.5 seconds of the program.

[0142] loudregtyp: A 4-bit field indicating which regulatory standard the program loudness complies with. Setting the "loudregtyp" field to "000" indicates that the LPSM does not indicate loudness regulatory compliance. For example, one value of this field (e.g., 0000) may indicate that compliance with a loudness regulatory standard is not indicated, another value of this field (e.g., 0001) may indicate that the audio data of the program complies with the ATSC A / 85 standard, and another value of this field (e.g., 0010) may indicate that the audio data of the program complies with the EBU R128 standard. In this example, if this field is set to any value other than "0000", the loudcorrdialgat and loudcorrtyp fields should follow the payload.

[0143] loudcorrdialgat: A 1-bit field indicating whether dialogue-gated loudness correction has been applied. If the program's loudness has been corrected using dialogue gating, the value of the loudcorrdialgat field is set to "1". Otherwise, its value is set to "0".

[0144] loudcorrtyp: A 1-bit field indicating the type of loudness correction applied to the program. If the program's loudness has been corrected with an infinite look-ahead (file-based) loudness correction process, the loudcorrtyp field is set to "0". If the program's loudness has been corrected using a combination of real-time loudness measurement and dynamic range control, the value of this field is set to "1".

[0145] loudrelgate: A 1-bit field indicating whether relative gated loudness data (ITU) is present. If the loudrelgate field is set to "1", it should be followed in the payload by the 7-bit ituloudrelgat field.

[0146] loudrelgat: A 7-bit field indicating the relative gated program loudness (ITU). This field indicates the integrated loudness of the audio program, measured according to ITU-R BS.1770-3, without any gain adjustments due to dialnorm and dynamic range compression applied. Values ​​from 0 to 127 are interpreted as -58LKFS to +5.5LKFS in 0.5LKFS steps.

[0147] loudspchgate: A 1-bit field indicating whether speech-gated loudness data (ITU) is present. If the loudspchgate field is set to "1", it should be followed in the payload by the 7-bit loudspchgat field.

[0148] loudspchgat: A 7-bit field indicating the speech-gated program loudness. This field indicates the integrated loudness of the entire corresponding audio program, measured according to formula (2) of ITU-R BS.1770-3, without the application of any gain adjustments due to dialnorm and dynamic range compression. Values ​​from 0 to 127 are interpreted as -58LKFS to +5.5LKFS in steps of 0.5LKFS.

[0149] loudstrm3se: A 1-bit field indicating whether short-term (3 second) loudness data is present. If this field is set to "1", it should be followed by a 7-bit loudstrm3s field in the payload.

[0150] loudstrm3s: A 7-bit field indicating the ungated loudness of the preceding 3 seconds of the corresponding audio program, measured according to ITU-R BS.1771-1, without any gain adjustments due to dialnorm and dynamic range compression applied. Values ​​0 to 256 are interpreted as -116LKFS to +11.5LKFS, in steps of 0.5LKFS.

[0151] truepke: A 1-bit field that indicates whether true peak loudness data is present. If the truepke field is set to "1", it should be followed by an 8-bit truepk field in the payload.

[0152] truepk: An 8-bit field indicating the true peak sample value of the program, measured in accordance with Annex 2 of ITU-R BS.1770-3, without any gain adjustments due to dialnorm and dynamic range compression being applied. Values ​​0 to 256 are interpreted as -116LKFS to +5.5LKFS, in steps of 0.5LKFS.

[0153] In some embodiments, a core element of a metadata segment in an extra bits segment or ancillary data (or "addbsi") field of a frame of an AC-3 or E-AC-3 bitstream includes a core header (typically an identification value, e.g., a core element version), followed by: a value indicating whether fingerprint data (or other protection value) is included for the metadata of the metadata segment, a value indicating whether external data (related to the audio data corresponding to the metadata of the metadata segment) is present, a payload ID and payload size value for each type of metadata identified by the core element (e.g., LPSM and / or non-LPSM types of metadata), and a protection value for at least one type of metadata identified by the core element. The metadata payload(s) of the metadata segment follow the core header and are (possibly) nested within the value of the core element.

[0154] Exemplary embodiments of the present invention include program boundary metadata within an encoded audio bitstream in an efficient manner that allows accurate and robust determination of at least one boundary between successive audio programs represented by the bitstreams. Exemplary embodiments allow accurate and robust determination of program boundaries in the sense that they allow accurate program boundary determination even when bitstreams representing different programs are spliced ​​together (to produce a bitstream of the present invention) in a manner that truncates one or both of the spliced ​​bitstreams (thereby discarding program boundary metadata that was included in at least one of the bitstreams prior to the splicing).

[0155] In a typical embodiment, the program boundary metadata for a frame of a bitstream of the present invention is a program boundary flag indicating a frame count. Typically, this flag indicates the number of frames between the current frame (the frame containing the flag) and a program boundary (the beginning or end of the current audio program). In some preferred embodiments, program boundary flags are inserted in a symmetric and efficient manner at the beginning and end of each bitstream segment representing a single program (i.e., in frames occurring within some predetermined number of frames after the beginning of the segment and in frames occurring within some predetermined number of frames before the end of the segment). Thus, when two such bitstreams are concatenated (thereby representing a sequence of two programs), program boundary metadata can be present on both sides of the boundary between the two programs (e.g., symmetrically).

[0156] Although maximum robustness can be achieved by inserting program boundary flags into every frame of the bitstream that represents a program, this is typically impractical due to the associated increased data rate. In a typical embodiment, program boundary flags are inserted into only a subset of frames of the encoded audio bitstream (which may represent an audio program or a sequence of audio programs), and the boundary flag insertion rate is a non-increasing function of the increasing distance of each frame of the bitstream (in which the flag is inserted) from its nearest program boundary. Here, the "boundary flag insertion rate" refers to the average ratio of the number of frames (representing a program) that contain a program boundary flag to the number of frames (representing the program) that do not contain a program boundary flag, where the average is a running average over a number (e.g., a relatively small number) of consecutive frames of the encoded audio bitstream.

[0157] Increasing the boundary flag insertion rate (at positions in the bitstream closer to the program boundary) increases the data rate required for delivery of the bitstream. To compensate for this, the size (number of bits) of each inserted flag is preferably decreased as the boundary flag insertion rate increases (so that the size of the program boundary flag in the Nth frame of the bitstream is a non-increasing function of the distance (number of frames) between the Nth frame and the nearest program boundary, where N is an integer). In one class of embodiments, the boundary flag insertion rate is a logarithmically decreasing function of increasing distance (of each flag insertion position) from the nearest program boundary, and for each flag-containing frame that contains one of the flags, the size of the flag in the flag-containing frame is equal to or greater than the size of each flag in a frame located closer to the nearest program boundary than the flag-containing frame. Typically, the size of each flag is determined by an increasing function of the number of frames from the flag insertion position to the nearest program boundary.

[0158] For example, consider the embodiments of Figures 8 and 9, where each column identified by a frame number (top row) represents a frame of the encoded audio bitstream. The bitstream represents an audio program with a first program boundary (indicating the beginning of that program) appearing immediately to the left of the column identified by frame number "17" on the left side of Figure 9, and a second program boundary (indicating the end of that program) appearing immediately to the right of the column identified by frame number "1" on the right side of Figure 8. The program boundary flags included in the frames shown in Figure 8 count down the number of frames between the current frame and the second program boundary. The program boundary flags included in the frames shown in Figure 9 count up the number of frames between the current frame and the first program boundary.

[0159] In the embodiments of Figures 8 and 9, the program boundary flag is set to 2 of the first X frames of the encoded bitstream after the start of the audio program indicated by the bitstream. N th frame, and closest to the end of the program indicated by the bitstream (the last X frames of the bitstream). N 8 and 9), program boundary flags are inserted in the second frame (N=1) of the bitstream (the flag-containing frame closest to the start of the program), the fourth frame (N=2), the eighth frame (N=3), etc., as well as in the eighth frame from the end of the bitstream, the fourth frame from the end of the bitstream, and the second frame from the end of the bitstream (the flag-containing frame closest to the end of the program). In this example, the second frame from the start (or end) of the program is inserted into the bitstream, and the fourth frame from the end of the program is inserted into the bitstream. N The program boundary flag in the th frame is calculated as log2(2 N+2 ) binary bits. Thus, the program boundary flag in the second to last frame (N=1) of a program is log2(2 N+2 )=log2(2 3 ) = 3 binary bits, and the flag in the fourth frame from the beginning (or end) of the program (N = 2) is log2(2 N+2 )=log2(2 4 ) = 4 binary bits, and so on.

[0160] In the examples of Figures 8 and 9, the format of each program boundary flag is as follows: each program boundary flag consists of a leading "1" bit, a sequence of "0" bits after the leading "1" bit (either zero "0" bits or one or more consecutive "0" bits), and a two-bit trailing code. For flags in the last X frames of the bitstream (the frames closest to the end of the program), the trailing code is "11", as shown in Figure 8. For flags in the first X frames of the bitstream (the frames closest to the start of the program), the trailing code is "10", as shown in Figure 9. Thus, to read (decode) each flag, the number of zeros between the leading "1" bit and the trailing code is counted. If the trailing code is identified as "11", the flag is determined to have (2) zeros between the current frame (the frame containing the flag) and the end of the program. Z+1 -1) frames, where Z is the number of zeros between the leading "1" bit of this flag and the trailing code. This decoder can be efficiently implemented to ignore the first and last bits of each such flag, determine the inverse of the sequence of the flag's other (middle) bits (e.g., if the sequence of middle bits is "0001" and the "1" bit is the last bit of the sequence, then the inverted sequence of middle bits will be "1000", with the "1" bit being the first bit of the inverted sequence), and identify the binary value of the inverted sequence of middle bits as the index of the current frame (the frame containing the flag) relative to the end of the program. For example, if the inverted sequence of middle bits is "1000", then this inverted sequence will have the binary value 2 4 =16, identifying this frame as the 16th frame before the end of the program (as shown in FIG. 8 in the column describing frame "0").

[0161] If the trailing code is identified as "10", the flag is set (2 frames) between the start of the program and the current frame (the frame that contains the flag).Z+1 0001", where Z is the number of zeros between the leading "1" bit and the trailing code of this flag. This decoder can be efficiently implemented to ignore the first and last bits of each such flag, determine the inverse of the sequence of the flag's middle bits (e.g., if the sequence of middle bits is "0001" and the "1" bit is the last bit in the sequence, then the inverted sequence of middle bits will be "1000", with the "1" bit being the first bit in the inverted sequence), and identify the binary value of the inverted sequence of middle bits as the index of the current frame (the frame in which the flag is contained) relative to the beginning of the program. For example, if the inverted sequence of middle bits is "1000", then this inverted sequence will have the binary value 2 4 =16, identifying this frame as the 16th frame after the beginning of the program (as shown in FIG. 9, in the column describing frame "32").

[0162] In the examples of Figures 8 and 9, the program boundary flag is set to 2 of the first X frames of the encoded bitstream after the start of the audio program indicated by the bitstream. N for each of the 2 th frames and closest to the end of the audio program indicated by the bitstream (the last X frames of the bitstream). N th frame, a program has Y frames, where X is an integer less than or equal to Y / 2, and N is a positive integer ranging from 1 to log2X. The inclusion of the program boundary flag only adds an average bitrate of 1.875 bits / frame to the bitrate required to transmit the bitstream without the flag.

[0163] In a typical implementation of the embodiment of Figures 8 and 9 where the bitstream is an AC-3 encoded audio bitstream, each frame contains audio content and metadata for 1536 samples of digital audio. For a sampling rate of 48 kHz, this represents a rate of 32 milliseconds of digital audio or 31.25 frames of audio per second. Thus, in such an embodiment, a program boundary flag in a frame that is spaced a certain number of frames ("X" frames) from a program boundary indicates that the boundary occurs 32X milliseconds after the end of the flag-containing frame (or 32X milliseconds before the start of the flag-containing frame).

[0164] In a typical implementation of the embodiment of Figures 8 and 9, where the bitstream is an E-AC-3 encoded audio bitstream, each frame of the bitstream contains audio content and metadata for 256, 512, 768 or 1536 samples of digital audio, respectively, depending on whether the frame contains one, two, three or six blocks of audio data. For a sampling rate of 48 kHz, this represents 5.333, 10.667, 16 or 32 milliseconds of digital audio, respectively, or a rate of 189.9, 93.75, 62.5 or 31.25 frames per second of audio, respectively. Thus, in such an embodiment, a program boundary flag in a frame that is spaced a certain number of frames ("X" frames) from a program boundary (where each frame represents 32 milliseconds of digital audio) indicates that the boundary occurs 32X milliseconds after the end of the flag-containing frame (or 32X milliseconds before the start of the flag-containing frame).

[0165] In some embodiments where program boundaries can occur within frames of the audio bitstream (i.e., without aligning with the beginning or end of a frame), the program boundary metadata included with the frames of the bitstream includes a program boundary frame count (i.e., metadata indicating the number of complete frames between the beginning or end of the frame count-containing frame and the program boundary) and an offset value indicating an offset (typically in number of samples) between the beginning or end of the program boundary-containing frame and the actual location of the program boundary within the program boundary-containing frame.

[0166] An encoded audio bitstream may represent a sequence of programs (soundtracks) in a corresponding sequence of video programs, and such audio program boundaries tend to occur at the edges of video frames rather than at the edges of audio frames. Also, some audio codecs (e.g., the E-AC-3 codecs) use audio frame sizes that are not aligned with the video frames. Also, in some cases, an originally encoded audio bitstream is transcoded to produce a transcoded bitstream, and the originally encoded bitstream has a different frame size than the transcoded bitstream. Thus, program boundaries (determined by the originally encoded bitstream) are not guaranteed to appear at frame boundaries in the transcoded bitstream. For example, if the originally encoded bitstream (e.g., bitstream "IEB" in FIG. 10) has a frame size of 1536 samples per frame, and the transcoded bitstream (e.g., bitstream "TB" in FIG. 10) has a frame size of 1024 samples per frame, the transcoding process may result in the actual program boundaries not appearing on frame boundaries in the transcoded bitstream, but somewhere within its frames (e.g., 512 samples into the frame of the transcoded bitstream, as shown in FIG. 10), due to the different frame sizes of the different codecs. An embodiment of the present invention in which the program boundary metadata included with frames of the encoded audio bitstream includes an offset value in addition to the program boundary frame count is useful in the three cases described in this paragraph (and others).

[0167] The embodiment described above with reference to Figures 8 and 9 does not include an offset value (e.g., an offset field) in any of the frames of the encoded bitstream. In variations on this embodiment, an offset value is included in each frame of the encoded audio bitstream that includes a program boundary flag (e.g., in frames corresponding to frames numbered 0, 8, 12, and 14 in Figure 8 and frames numbered 18, 20, 24, and 32 in Figure 9).

[0168] In one class of embodiments, a data structure (in each frame of an encoded bitstream containing the program boundary metadata of the present invention) includes a code value that indicates whether the frame contains only a program boundary frame count or both a program boundary frame count and an offset value. For example, the code value may be the value of a single bit field (referred to herein as an "offset_exist" field). A value of offset_exist=0 may indicate that the frame does not contain an offset value, and a value of offset_exist=1 may indicate that the frame contains a program boundary frame count and an offset value.

[0169] In some embodiments, at least one frame of an AC-3 or E-AC-3 encoded audio bitstream includes a metadata segment that includes LPSM and program boundary metadata (and optionally other metadata) for an audio program determined by the bitstream. Each such metadata segment (which may be included in an addbsi field or an ancillary data field or an extra bits segment of the bitstream) includes a Core Header (and optionally additional Core Elements) and, after the Core Header (or after the Core Header and other Core Elements), an LPSM Payload (or container) segment having the following format:

[0170] a header (which typically contains at least one identifying value, e.g., LPSM format version, length, period, count and sub-stream association values); After the header is program boundary metadata (which may include a program boundary frame count, a code value (e.g., an "offset_exist" value) indicating whether the frame contains only a program boundary frame count or both a program boundary frame count and an offset value, and possibly an offset value) and an LPSM. The LPSM may include: at least one dialogue indication value indicating whether the corresponding audio data indicates dialogue or not (e.g., which channels of the corresponding audio data indicate dialogue), the dialogue indication value(s) may indicate whether dialogue is present in any combination or all of the channels of the corresponding audio data; at least one loudness regulations compliance value indicating whether the corresponding audio data complies with an indicated set of loudness regulations; at least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data; and At least one loudness value indicating at least one loudness characteristic of the corresponding audio data (e.g., peak or average loudness).

[0171] In some embodiments, the LPSM payload segment includes a coded value (e.g., an "offset_exist" value) that indicates whether the frame includes only a program boundary frame count or both a program boundary frame count and an offset value. For example, in one such embodiment, when such a coded value (e.g., offset_exist=1) indicates that the frame includes a program boundary frame count and an offset value, the LPSM payload segment may include an offset value that is an 11-bit unsigned integer (i.e., having a value between 0 and 2048) and indicates the number of additional audio samples between the signaled frame boundary (the boundary of the frame that includes the program boundary) and the actual program boundary. If the program boundary frame count indicates the number of frames (at the current frame rate) until the program boundary-containing frame, then the precise location (in samples) of the program boundary (relative to the start or end of the frame that includes the LPSM payload segment) may be: S=(frame_counter*frame size)+offset where S is the number of samples to a program boundary (from the beginning or end of the frame containing the LPSM payload segment), "frame_counter" is the frame count indicated by the program boundary frame count, "frame size" is the number of samples per frame, and "offset" is the number of samples indicated by the offset value.

[0172] Some embodiments in which the program boundary flag insertion rate increases near actual program boundaries implement a rule that a frame will never contain an offset value if it is a certain number ("Y") frames or less from the frame containing the program boundary. Typically, Y=32. For E-AC-3 encoders that implement this rule (as Y=32), the encoder will never insert an offset value in the last second of an audio program. In this case, the receiving device is responsible for maintaining a timer and thereby performing its own offset calculations (in response to program boundary metadata containing an offset value for frames of the encoded bitstream that are more than Y frames away from the program boundary-containing frame).

[0173] For programs where the audio program is known to be "frame-aligned" to the video frames of the corresponding video program (e.g., a typical contributing feed with Dolby E encoded audio), including an offset value in the encoded bitstream that represents the audio program is redundant, and thus an offset value is typically not included in such encoded bitstreams.

[0174] Referring to FIG. 11, consider next how encoded audio bitstreams are spliced ​​together to generate an audio bitstream embodiment of the present invention.

[0175] The top bitstream of FIG. 11 (labeled "Scenario 1") shows an entire first audio program (P1) including program boundary metadata (program boundary flags F) followed by an entire second audio program (P2) also including program boundary metadata (program boundary flags F). The program boundary flags at the end of the first program (some of which are shown in FIG. 11) are the same as or similar to those described with reference to FIG. 8 and determine the location of the boundary between the two programs (i.e., the boundary at the beginning of the second program). The program boundary flags at the beginning of the second program (some of which are shown in FIG. 11) are the same as or similar to those described with reference to FIG. 9 and also determine the location of the boundary. In a typical embodiment, the encoder or decoder implements a timer (calibrated by a flag in the first program) that counts down to a program boundary, and the same timer (calibrated by a flag in the second program) counts up from the same program boundary. As shown by the boundary timer graph for Scenario 1 in Figure 11, the countdown of such a timer (calibrated by a flag in the first program) reaches 0 at the boundary, and the countup of the timer (calibrated by a flag in the second program) refers to the same position on the boundary.

[0176] The second bitstream from the top in FIG. 11 (labeled "Scenario 2") shows an entire first audio program (P1) including program boundary metadata (program boundary flag F), followed by an entire second audio program (P2) without program boundary metadata. The program boundary flags at the end of the first program (some of which are shown in FIG. 11) are the same or similar to those described with reference to FIG. 8 and, as in Scenario 1, determine the location of the boundary between the two programs (i.e., the boundary at the beginning of the second program). In a typical embodiment, the encoder or decoder implements a timer (calibrated by a flag in the first program) that counts down to the program boundary, and the same timer continues to count up from the program boundary (without further calibration) (as shown by the boundary timer graph in Scenario 2 in FIG. 11).

[0177] The third bitstream from the top in FIG. 11 (labeled “Scenario 3”) shows a truncated first audio program (P1) containing program boundary metadata (program boundary flag F) spliced ​​with an entire second audio program (P2) also containing program boundary metadata (program boundary flag F). The splicing removes the last “N” frames of the first program. Program boundary flags (some of which are shown in FIG. 11) at the beginning of the second program are the same or similar to those described with reference to FIG. 9 and determine the location of the boundary (splice) between the truncated first program and the complete second program. In a typical embodiment, the encoder or decoder implements a timer (calibrated by a flag in the first program) that counts down to the end of the untruncated first program, and the same timer (calibrated by a flag in the second program) counts up from the beginning of the second program. The beginning of the second program is the program boundary in Scenario 3. As shown by the boundary timer graph for scenario 3 in Figure 11, the countdown of such a timer (calibrated by the program boundary metadata in the first program) is reset (in response to the program boundary metadata in the second program) before it reaches zero (in response to the program boundary metadata in the first program). Thus, abortion of the first program (by splicing) prevents the timer from identifying the program boundary between the aborted first program and the start of the second program in response to (i.e., under calibration by) only the program boundary metadata in the first program, but the program metadata in the second program resets the timer, such that the reset timer correctly indicates the location of the program boundary between the aborted first program and the start of the second program (as the location corresponding to the "0" count of the reset timer).

[0178] The fourth bitstream (labeled "Scenario 4") illustrates a truncated first audio program (P1) including program boundary metadata (program boundary flags F) and a truncated second audio program (P2) including program boundary metadata (program boundary flags F) spliced ​​with a portion (non-truncated portion) of the first audio program. The program boundary flags at the beginning of the entire second program (before truncation), some of which are shown in FIG. 11, are the same as or similar to those described with reference to FIG. 9, and the program boundary flags at the end of the entire first program (before truncation), some of which are shown in FIG. 11, are the same as or similar to those described with reference to FIG. 8. The splicing removes the last "N" frames of the first program (and thus the portion of the program boundary flags included before the splice) and the first "M" frames of the second program (and thus the portion of the program boundary flags included before the splice). In a typical embodiment, an encoder or decoder implements a timer (calibrated by a flag in the aborted first program) that counts down toward the end of the first unaborted program, and the same timer (calibrated by a flag in the aborted second program) counts up from the beginning of the second unaborted program. As illustrated by the boundary timer graph in Scenario 4 of FIG. 11, such a timer's countdown (calibrated by program boundary metadata in the first program) is reset (in response to program boundary metadata in the second program) before it reaches zero (in response to program boundary metadata in the first program). Abortion of the first program (by splicing) prevents the timer from identifying a program boundary between the aborted first program and the beginning of the aborted second program in response to (i.e., under calibration by) only the program boundary metadata in the first program.However, the reset timer will not correctly indicate the location of the program boundary between the end of the first aborted program and the beginning of the second aborted program. Thus, the abortion of both bitstreams being spliced ​​together may prevent accurate determination of the boundary between them.

[0179] Embodiments of the invention may be implemented in hardware, firmware or software or a combination of both (e.g., as a programmable logic array). Unless otherwise specified, the algorithms or processes included as part of the invention are not inherently related to any particular computer or other apparatus. In particular, various general purpose machines may be used with programs written in accordance with the teachings of the present application, or it may prove more convenient to construct a more specialized apparatus (e.g., an integrated circuit) to perform the required method steps. Thus, the invention may be implemented in one or more computer programs executing on one or more programmable computer systems (e.g., an implementation of any of the elements of FIG. 1 or the encoder 100 (or elements thereof) of FIG. 2 or the decoder 200 (or elements thereof) of FIG. 3 or the post-processor (or elements thereof) of FIG. 3). Each computer system has at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices in a known manner.

[0180] Each such program may be implemented in any desired computer language (including machine, assembly, or high level procedural, logical or object-oriented programming languages) to communicate with a computer system, and in any case, the language may be a compiled or interpreted language.

[0181] For example, when implemented by a sequence of computer software instructions, various functions and steps of embodiments of the present invention may be implemented by a multi-threaded sequence of software instructions executed on suitable digital signal processing hardware, in which case various units, steps and functions of the embodiments may correspond to portions of the software instructions.

[0182] Each such computer program is preferably stored or downloaded onto a general purpose or special purpose programmable computer readable storage medium or device (e.g., semiconductor memory or media or magnetic or optical media) that, when read by a computer system, configures or operates a computer to perform the procedures described herein. The system of the present invention may be implemented as a computer readable storage medium configured with (i.e., having stored thereon) a computer program, the storage medium so configured causing the computer system to operate in a specific predefined manner to perform the functions described herein.

[0183] Although several embodiments of the present invention have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present invention. Numerous modifications and variations of the present invention are possible in light of the above teachings. It will be understood that, within the scope of the appended claims, the present invention may be practiced otherwise than as specifically described herein.

[0184] Several aspects will be described. [Aspect 1] 13. An audio processing unit comprising: a buffer memory for storing at least one frame of an encoded audio bitstream, the encoded audio bitstream including audio data and a metadata container, the metadata container including a header, one or more metadata payloads and protection data; an audio decoder coupled to the buffer memory for decoding the audio data; a parser coupled to or integrated with the audio decoder for parsing the encoded audio bitstream; the header includes a synchronization word identifying the beginning of the metadata container, the one or more metadata payloads describing an audio program associated with the audio data, the protection data following the one or more metadata payloads, the protection data being usable to verify the integrity of the metadata container and the one or more payloads within the metadata container. Audio processing unit. [Aspect 2] 2. The audio processing unit of embodiment 1, wherein the metadata container is stored in an AC-3 or E-AC-3 reserved data space selected from the group consisting of a skip field, an auxiliary data field, an addbsi field, and combinations thereof. [Aspect 3] 3. The audio processing unit of aspect 1 or 2, wherein the one or more metadata payloads include metadata indicating at least one boundary between successive audio programs. Aspect 4 3. The audio processing unit of aspect 1 or 2, wherein the one or more metadata payloads include a program loudness payload containing data indicative of a measured loudness of an audio program. Aspect 5 5. The audio processing unit of embodiment 4, wherein the program loudness payload includes a field indicating whether the audio channel contains spoken dialogue. Aspect 6 5. The audio processing unit of embodiment 4, wherein the program loudness payload includes a field indicating a loudness measurement method used to generate the loudness data included in the program loudness payload. Aspect 7 5. The audio processing unit of embodiment 4, wherein the program loudness payload includes a field indicating whether the loudness of the audio program has been corrected using dialogue gating. Aspect 8 5. The audio processing unit of embodiment 4, wherein the program loudness payload includes a field indicating whether the loudness of the audio program has been corrected using an infinite look-ahead or a file-based loudness correction process. Aspect 9 5. The audio processing unit of embodiment 4, wherein the program loudness payload includes a field indicating an integrated loudness of the audio program without any gain adjustments that may result in dynamic range compression. Aspect 10 5. The audio processing unit of embodiment 4, wherein the program loudness payload includes a field indicating the integrated loudness of the audio program without any gain adjustments that can be reduced to dialog normalization. Aspect 11 5. The audio processing unit of embodiment 4, configured to perform adaptive loudness processing using the program loudness payload. Aspect 12 12. The audio processing unit of any one of aspects 1-11, wherein the encoded audio bitstream is an AC-3 bitstream or an E-AC-3 bitstream. Aspect 13 12. The audio processing unit of any one of aspects 4 to 11, configured to extract the program loudness payload from the encoded audio bitstream and authenticate or validate the program loudness payload. Aspect 14 14. The audio processing unit of any one of aspects 1 to 13, wherein each of the one or more metadata payloads includes a unique payload identifier, the unique payload identifier being located at the beginning of each metadata payload. Aspect 15 14. The audio processing unit of any one of aspects 1-13, wherein the synchronization word is a 16-bit synchronization word having a value of 0x5838. Aspect 16 1. A method for decoding an encoded audio bitstream, comprising: receiving an encoded audio bitstream, the audio bitstream being segmented into one or more frames; extracting audio data and a metadata container from the encoded audio bitstream, the metadata container including a header followed by one or more metadata payloads followed by protection data; verifying the integrity of the container and the one or more metadata payloads through use of the protected data; the one or more metadata payloads include a program loudness payload containing data indicative of a measured loudness of an audio program associated with the audio data; method. Aspect 17 17. The method of embodiment 16, wherein the encoded audio bitstream is an AC-3 bitstream or an E-AC-3 bitstream. Aspect 18 17. The method of embodiment 16, further comprising: performing adaptive loudness processing on audio data extracted from the encoded audio bitstream using the program loudness payload. Aspect 19 17. The method of embodiment 16, wherein the container is located in and extracted from a reserved data space of AC-3 or E-AC-3 selected from the group consisting of a skip field, an auxiliary data field, an addbsi field, and combinations thereof. Aspect 20 17. The method of embodiment 16, wherein the program loudness payload includes a field indicating whether an audio channel contains spoken dialogue. Aspect 21 17. The method of claim 16, wherein the program loudness payload includes a field indicating a loudness measurement method used to generate the loudness data included in the program loudness payload. Aspect 22 17. The method of embodiment 16, wherein the program loudness payload includes a field indicating whether the loudness of the audio program has been corrected using dialogue gating. Aspect 23 17. The method of embodiment 16, wherein the program loudness payload includes a field indicating whether the loudness of the audio program has been corrected using an infinite look-ahead or a file-based loudness correction process. Aspect 24 17. The method of embodiment 16, wherein the program loudness payload includes a field indicating an integrated loudness of the audio program without any gain adjustments due to dynamic range compression. Aspect 25 17. The method of embodiment 16, wherein the program loudness payload includes a field indicating the integrated loudness of the audio program without any gain adjustments that can be reduced to dialog normalization. Aspect 26 17. The method of embodiment 16, wherein the metadata container includes metadata indicating at least one boundary between successive audio programs. Aspect 27 17. The method of embodiment 16, wherein the metadata container is stored in one or more skip fields or redundant bit segments of a frame.

Claims

1. 13. An audio processing unit comprising: a buffer memory configured to store an encoded audio bitstream, the encoded audio bitstream including audio data and metadata, the metadata including a payload of loudness metadata; a parser coupled to the buffer memory and configured to extract the audio data and the loudness metadata payload from the encoded audio bitstream; a decoder coupled to the parser and configured to decode the audio data to generate decoded audio data; a subsystem coupled to the parser and the decoder and configured to receive a target loudness value and perform post-processing on the decoded audio data in response to the loudness metadata and the target loudness value; the loudness metadata includes metadata indicating the presence of an audio program loudness in a loudness metadata payload, and when the audio program loudness is present in the loudness metadata payload, the loudness metadata further includes an indication of a measurement method used to determine the audio program loudness. Audio processing unit.

2. 2. The audio processing unit of claim 1, wherein the measurement method is defined in ITU-R BS.1770.

3. 13. A method of audio processing performed by an audio processing unit, comprising: receiving, by the audio processing unit, an encoded audio bitstream, the encoded audio bitstream including audio data and metadata, the metadata including a payload of loudness metadata; extracting, by a parser of the audio processing unit, a payload of the audio data and the loudness metadata from the encoded audio bitstream; decoding, by a decoder of the audio processing unit, the audio data to generate decoded audio data; receiving, by a subsystem of the audio processing unit, a target loudness value; performing, by the subsystems of the audio processing unit, post-processing on the decoded audio data in response to the loudness metadata and the target loudness value; the loudness metadata includes metadata indicating the presence of an audio program loudness in a loudness metadata payload, and when the audio program loudness is present in the loudness metadata payload, the loudness metadata further includes an indication of a measurement method used to determine the audio program loudness. Audio processing methods.

4. 4. The audio processing method according to claim 3, wherein the measurement method is defined in ITU-R BS.1770.

5. A non-transitory medium storing a computer program, the computer program causing a computer to: receiving an encoded audio bitstream, the encoded audio bitstream including audio data and metadata, the metadata including a payload of loudness metadata; extracting the audio data and the loudness metadata payload from the encoded audio bitstream; decoding the audio data to generate decoded audio data; receiving a target loudness value; and performing post-processing on the decoded audio data in response to the loudness metadata and the target loudness value; the loudness metadata includes metadata indicating the presence of an audio program loudness in a loudness metadata payload, and when the audio program loudness is present in the loudness metadata payload, the loudness metadata further includes an indication of a measurement method used to determine the audio program loudness. Non-temporary medium.

6. The non-transitory medium of claim 5 , wherein the measurement method is defined in ITU-R BS.1770.

Citation Information

Patent Citations

  • Audio metadata confirmation

    JP2008536193A

  • System for combining loudness measurements in a single playback mode

    WO2011110525A1

  • Encoder / decoder for multidimensional sound fields

    US5583962A

  • Encoder / decoder for multidimensional sound fields

    US5632005A

  • Method and apparatus for adjusting dynamic range and gain in an encoder / decoder for multidimensional sound fields

    US5633981A