Audio processing unit and method for audio processing

TWI939338BActive Publication Date: 2026-09-11DOLBY LABORATORIES LICENSING CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
TW115108363
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-06-19
Filing Date
2014-05-29
Publication Date
2026-09-11
Estimated Expiration
2034-05-28

AI Technical Summary

Technical Problem

Existing audio processing systems fail to optimally handle audio data when multiple units are distributed across networks, leading to unnecessary processing and degradation of audio content due to lack of awareness of previous processing history, especially in formats like Dolby Digital (AC-3) and Dolby Digital+ (E-AC-3).

Method used

Incorporating substream structure metadata (SSM) and program information metadata (PIM) within audio bitstreams, allowing decoders to extract and process metadata alongside audio data, enabling adaptive processing and error correction across cascaded audio processing units.

Benefits of technology

Ensures optimal audio processing by adapting to previous processing history, preventing unnecessary operations and maintaining audio quality through metadata-driven adaptive processing and error detection/correction, particularly in Dolby Digital and Dolby Digital+ formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001910996_001
    Figure TWG2TB001910996_001
  • Figure TWG2TB001910996_002
    Figure TWG2TB001910996_002
  • Figure TWG2TB001910996_003
    Figure TWG2TB001910996_003
Patent Text Reader

Abstract

An audio processing unit includes a buffer memory storing a portion of an encoded audio bitstream, wherein the encoded audio bitstream is segmented into frames, and at least one frame contains program information metadata in a metadata segment of the at least one frame, and audio data in another segment of the at least one frame; and a processing subsystem coupled to the buffer memory, wherein the processing subsystem is configured to decode the encoded audio bitstream, wherein the metadata segment contains at least one metadata payload, the metadata payload including a header, and at least some of the program information metadata following the header.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to audio signal processing, and more specifically, to the encoding and decoding of audio data bitstreams, using metadata to represent the substream structure and / or program information relating to the audio content represented by the bitstream. Some embodiments of this invention generate or decode audio data in any format referred to as Dolby Digital (AC-3), Dolby Digital+ (Enhanced AC-3 or E-AC-3), or Dolby E. Prior Technology

[0002] Dolby, Dolby Digital, Dolby Digital+, and Dolby E are trademarks of Dolby Laboratories' licensees. Dolby Laboratories provides proprietary implementations of AC-3 and E-AC-3, respectively, called Dolby Digital and Dolby Digital+.

[0003] Audio processing units typically operate blindly, unaware of the audio processing history that occurred before the data was received. This can also work within a processing framework where a single entity performs all audio processing and encoding for various target media playback devices, while the target media playback devices perform all decoding and playback of the encoded audio. However, this blind processing does not work well (or not at all) when multiple audio processing units are distributed across different networks or cascaded (i.e., linked) and expected to perform their individual types of audio processing optimally. For example, some audio may be encoded for high-performance media systems and may need to be converted to a scaled-down format suitable for mobile devices along the media processing chain. Therefore, an audio processing unit may not necessarily perform a certain type of processing on the audio already processed. For example, a volume leveling unit may process an input audio clip regardless of whether the same or similar volume level has been previously applied to that input audio clip. As a result, the volume leveling unit may perform leveling even when it is not necessary. This unnecessary processing may also result in the degradation and / or removal of certain characteristics in the content of the performance audio data. Summary of the Invention

[0004] In one set of embodiments, the present invention is an audio processing unit capable of decoding an encoded bitstream, the encoded bitstream comprising substream structure metadata and / or program information metadata (and optionally other metadata, such as loudness processing status metadata) in at least one segment of at least one frame of the bitstream and audio data in at least one other segment of the frame. Here, substream structure metadata (or SSM) represents metadata of an encoded bitstream (or a group of encoded bitstreams), representing the substream structure of the audio content of the encoded bitstream, and "program information metadata (or PIM)" represents metadata of the encoded audio bitstream, representing at least one audio program (e.g., two or more audio programs), wherein the program information metadata represents at least one characteristic or feature of the audio content of at least one program (e.g., metadata representing the type or parameters of processing performed on the audio data of the program, or metadata representing which channel's program is the operating channel).

[0005] In typical cases (e.g., where the encoded bitstream is AC-3 or E-AC-3), Program Information Metadata (PIM) represents program information that cannot be actually carried in other parts of the bitstream. For example, PIM can represent the processing applied to the PCM audio before encoding (e.g., AC-3 or E-AC-3 encoding) and the compression profile used to establish Dynamic Range Compression (DRC) data in the bitstream, where the frequency bands of the audio program have already been encoded using a specific audio coding technique.

[0006] In other group embodiments, a method includes multiplexing encoded audio data in SSM and / or PIM within each frame (or at least a portion of each frame) of a bitstream. In typical decoding, a decoder extracts the SSM and / or PIM (including parsing and demultiplexing the SSM and / or PIM and the audio data) from the bitstream and processes the audio data to produce a decoded audio data stream (and in some cases, also performs adaptation processing on the audio data). In some embodiments, the decoded audio data and the SSM and / or PIM are transmitted by the decoder to a post-processor configured to perform adaptation processing on the decoded audio data using the SSM and / or PIM.

[0007] In a group of embodiments, the encoding method of the present invention produces an encoded audio bitstream (e.g., an AC-3 or E-AC-3 bitstream) comprising audio data segments (e.g., segments AB0-AB5 shown in the frame of FIG. 4 or all or part of segments AB0-AB5 shown in the frame of FIG. 7), containing encoded audio data, and metadata segments time-multiplexed by audio data segments (including SSM and / or PIM, or optionally also including other metadata). In some embodiments, each metadata segment (sometimes referred to herein as a “box”) has a format comprising a metadata segment header (and optionally also including other mandatory or “core” elements), and one or more metadata payloads following the metadata segment header. If present, the SIM is included in one metadata payload (identified by the payload header and typically having a first type of format). If present, the PIM is included in another metadata payload (identified by the payload header and typically having a second type of format). Similarly, (if any) other types of metadata are included in a further metadata payload (identified by the payload header and typically having a format specific to that type of metadata). This illustrative format allows (e.g., a post-processor after decoding, or a processor configured to identify the metadata, without performing the entire decoding of the encoded bitstream) convenient access to the SSM, PIM, and other metadata, and convenient access to other metadata at times other than decoding, and allows convenient and effective (e.g., secondary stream identification) error detection and correction during bitstream decoding. For example, without taking the illustrative format of the SSM, the decoder may incorrectly identify the correct number of secondary streams for a program. One metadata payload in a metadata segment may contain the SSM, another metadata payload in a metadata segment may contain the PIM, and optionally at least one other metadata payload in a metadata segment may contain other metadata (e.g., loudness processing status metadata or "LPSM"). Simple Explanation of the Diagram

[0008] Figure 1 is a block diagram of an embodiment of a system configured to perform an embodiment of the method of the present invention.

[0009] Figure 2 is a block diagram of the encoder of an embodiment of the audio processing unit of the present invention.

[0010] Figure 3 is a block diagram of a decoder of an embodiment of the audio processing unit of the present invention, and a post-processor of another embodiment of the audio processing unit of the present invention coupled thereto.

[0011] Figure 4 is a schematic diagram of the AC-3 frame, which includes the segmented sections.

[0012] Figure 5 is a schematic diagram of the synchronization information (SI) segment of the AC-3 frame, which includes the segmented segments.

[0013] Figure 6 is a schematic diagram of the bit stream information (BSI) segment of the AC-3 frame, which includes the segmented segments.

[0014] Figure 7 is a schematic diagram of the E-AC-3 frame, which includes the segmented sections.

[0015] Figure 8 is a block diagram of the metadata segment of the encoded bitstream generated according to an embodiment of the present invention. It includes a metadata segment header, which includes a box synchronization character (identified as "box synchronization" in Figure 8) and a version and key ID value, followed by a plurality of metadata payloads and protection bits. Implementation Identification and Naming Methods

[0016] Throughout the specification, including the claims, the notation of performing an operation "on" a signal or data (e.g., filtering, scaling, conversion, or applying gain to a signal or data) is used in a broad sense to indicate performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., a preprocessed version of the signal that has been initially filtered or on which an operation is performed).

[0017] Throughout the specification, including the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem implementing a decoder can also be called a decoder system, and a system containing such a primary system (e.g., a system that generates an X output signal in response to multiple inputs, wherein the secondary system generates an M input and other XM inputs are received by an external source) can also be called a decoder system.

[0018] Throughout this specification, including the claims, the term "processor" is used broadly to refer to a system or device that can be (e.g., in software or firmware) programmed or configured to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or configured to perform pipelined processing on audio or other sound data, general-purpose processors or computers, and microprocessor chipsets or chipsets.

[0019] Throughout this specification, including the claims, the terms "audio processor" and "audio processing unit" are used interchangeably, broadly referring to a system configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, preprocessing systems, post-processing systems, and bitstream processing systems (sometimes called bitstream processing tools).

[0020] Throughout the specification, including the scope of the patent application, the notation of "metadata" (for encoded audio bitstreams) indicates separate and distinct data from the corresponding audio data in the bitstream.

[0021] In this case, which includes the scope of the patent application, the notation "substream structure metadata (SSM)" refers to the metadata of an encoded audio bitstream (or a group of encoded audio bitstreams), and to the substream structure of the audio content of the encoded bitstream.

[0022] In this case, which includes the scope of the patent application, the notation “Program Information Metadata” (or “PIM”) refers to the metadata of the encoded audio bitstream of at least one audio program (e.g., two or more audio programs), wherein the metadata represents at least one characteristic or feature of the audio content of at least one program (e.g., the metadata represents the type or parameters of processing performed on the audio data of the program, or the metadata representing which channels of the program are the operating channels).

[0023] In this case, which includes the scope of the patent application, the notation "processor state metadata" (e.g., denoted as "loudness processing state metadata") represents metadata relating to audio data (encoded audio bitstream), indicating the processing state of the relative (related) audio data (e.g., what type of processing has been performed on the audio data), and typically represents at least one characteristic or feature of the audio data. The correlation between the processing state metadata and the audio data is time-synchronized. Therefore, current (latest received or updated) processing state metadata represents the corresponding audio data and simultaneously includes the result of the representation type of audio data processing. In some examples, the processing state metadata may include processing history and / or some or all of the parameters used for and / or derived from the represented type of processing. In addition, the processing state metadata may include at least one characteristic or feature of the corresponding audio data that has been calculated or captured by the audio data. The processing state metadata may also include other metadata that is unrelated or not derived from the processing of the corresponding audio data. For example, third-party data, tracking information, identification codes, proprietary or standard information, user annotations, user preference data, etc., can be added by a specific audio processing unit and transmitted to other audio processing units.

[0024] In this case, which includes the scope of the patent application, the notation "loudness processing state metadata" (or "LPSM") refers to processing state metadata that indicates the loudness processing state of the corresponding audio data (e.g., what type of loudness processing has been performed on the audio data) and typically corresponds to at least one characteristic or feature of the audio data (e.g., loudness). Loudness processing state metadata may include data (e.g., other metadata) that are not loudness processing state metadata (i.e., when considered alone).

[0025] In this case, which includes the scope of the patent application, the term "channel" (or "audio channel") refers to a single-tone audio signal.

[0026] In this case, which includes the scope of the patent application, the representation “audio program” refers to a group of one or more audio channels and the selected location also includes related metadata (e.g., metadata describing the desired spatial audio representation, and / or PIM, and / or SSM, and / or LPSM, and / or program boundary metadata).

[0027] In this case, which includes the scope of the patent application, the notation "program boundary metadata" represents metadata of an encoded audio bitstream, wherein the encoded audio bitstream represents at least one audio program (e.g., two or more audio programs), and the program boundary metadata represents the position of the bitstream at at least one boundary (start and / or end) of at least one of the audio programs. For example, the program boundary metadata (representing the encoded audio bitstream of an audio program) may include metadata indicating the start of the program (e.g., the start of the "N"th frame of the bitstream, or the "M"th sample position of the "N"th frame of the bitstream), and other metadata indicating the end of the program (e.g., the start of the "J"th frame of the bitstream, or the "K"th sample position of the "J"th frame of the bitstream).

[0028] In this case, which includes the scope of the patent application, the terms “coupled” or “coupled” are used to indicate a direct or indirect connection. Thus, if the first device is coupled to the second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0029] A typical audio stream contains audio content (e.g., audio content from one or more channels) and metadata representing at least one characteristic of that audio content. For example, in an AC-3 bitstream, there are several audio metadata parameters that are specifically intended to alter the sound of the program input to the listening environment. One such metadata parameter is the DIALNORM parameter, which is intended to represent the average bit level of the dialogue in the audio program and is used to determine the bit level of the audio playback signal.

[0030] When playing a bitstream containing sequentially coded audio segments (each with different DIALNORM parameters), the AC-3 decoder uses the DIALNORM parameters of each segment to perform a type of loudness processing, where it modifies the playback level or loudness so that the listening loudness of the dialogue in that sequence of segments is at a consistent level. The individual coded audio segments (items) in a sequentially coded audio project will (typically) have different DIALNORM parameters, and the decoder will scale the level of each item so that the playback level or loudness of the dialogue in each item is the same or very similar, but this may require applying different amounts of gain to different items during playback.

[0031] Although DIALNORM is typically set by the user and not automatically generated, there are still default DIALNORM values ​​if no user-defined value exists. For example, a content creator can perform loudness measurement using a device outside the AC-3 encoder and then send the results (representing the loudness of spoken dialogue in an audio program) to the encoder to set the DIALNORM value. Therefore, the content creator has confidence in correctly setting the DIALNORM parameters.

[0032] There are several different reasons why the DIALNORM parameter in an AC-3 bitstream might be incorrect. First, if the DIALNORM value is not set by the content creator, each AC-3 encoder has a default DIALNORM value used during bitstream generation. This default value can differ significantly from the actual dialogue loudness level of the audio. Second, even if the content creator measures loudness and sets the DIALNORM value, a loudness measurement algorithm or table that does not conform to the recommended AC-3 loudness measurement method may have been used, resulting in an incorrect DIALNORM value. Third, even if the AC-3 bitstream has been built with the measured DIALNORM value and correctly set by the content creator, it may change to an incorrect value during bitstream transmission and / or storage. For example, it is not uncommon for television broadcasting applications to use incorrect DIALNORM metadata information to decode, modify, and re-encode AC-3 bitstreams. Therefore, the DIALNORM value contained in an AC-3 bitstream can be incorrect or inaccurate, which can negatively impact the quality of the listening experience.

[0033] Furthermore, the DIALNORM parameter does not indicate the loudness processing status of the corresponding audio data (e.g., what type of loudness processing has been performed on the audio data). The loudness processing status metadata (in the format provided in some embodiments of the present invention) is used to facilitate the efficient and adaptive loudness processing of the audio bitstream and / or to verify the validity of the loudness processing status and the loudness of the audio content.

[0034] While the present invention is not limited to the use of AC-3 bitstreams, E-AC-3 bitstreams, or Dolby E bitstreams, for convenience, embodiments of generating, decoding, or processing such bitstreams will be described.

[0035] The AC-3 encoded bitstream contains channels one through six, which consist of metadata and audio content. The audio content is compressed audio data using perceptual audio coding. The metadata includes several audio metadata parameters that are intended to modify the sound of the program delivered to the listening environment.

[0036] Each frame of an AC-3 encoded audio bitstream contains the audio content and metadata used for 1536-sampled digital audio. For a sampling rate of 48kHz, this represents 32 milliseconds of digital audio or 31.25 frames per second.

[0037] Depending on whether the frame contains one, two, three, or six blocks of audio data, each frame of the E-AC-3 encoded audio bitstream contains audio content and metadata for 256, 512, 768, or 1536 sampled digital audio. For a 48kHz sampling rate, this represents 5.333, 10.667, 16, or 32 milliseconds of digital audio, or 189.9, 93.75, 62.5, or 31.25 frame rates per second, respectively.

[0038] As shown in Figure 4, each AC-3 frame is divided into regions (segments), including: a synchronization information (SI) region, which includes (as shown in Figure 5) the synchronization character (SW) and the first of the two error correction characters (CRC1); a bit stream information (BSI) region, which contains most of the metadata; six audio blocks (AB0-AB5), which contain data-compressed audio content (and also metadata), and its discard bit segment (W) (also known as the "skip bar"), which contains the unused bits remaining after the audio content is compressed; an auxiliary (AUX) information segment that may contain more metadata; and the second of the two error correction characters (CRC2).

[0039] As shown in Figure 7, each E-AC-3 frame is divided into multiple regions (segments), including: a synchronization information (SI) region including (as shown in Figure 5) synchronization characters (SW); a bit stream information (BSI) region including most of the metadata; one to six audio blocks (AB0 to AB5) containing compressed audio content (and possibly metadata); a waste bit segment (W) including the remaining unused bits after the audio content is compressed (also known as a "skip bar") (although only one waste bit segment is displayed, different waste bit or skip bar segments may typically follow each audio block); an auxiliary (AUX) information segment that may include more metadata; and an error correction character (CRC).

[0040] In AC-3 (or E-AC-3) bitstreams, there are several audio metadata parameters that are specifically designed to alter the sound of the program delivered to the listening environment. One of these metadata parameters is the DIALNORM parameter, which is included in the BSI segment.

[0041] As shown in Figure 6, the BSI segment of the AC-3 frame includes a five-bit parameter (“DIALNORM”) representing the DIALNORM value used for that program. If the audio encoding mode (acmod) of the AC-3 frame is “0”, it includes a five-bit parameter (DIALNORM2) representing the DIALNORM value used for a second audio program carried in the same AC-3 frame, indicating that a “dual-single” or “1+1” channel configuration is being used.

[0042] The BSI segment also includes a flag (“addbsie”) indicating the presence (or absence) of additional bitstream information following the “addbsie” bit; and a parameter (addbsil) indicating the length of any additional bitstream information following the “addbsil” value, and a maximum of 64 bits of additional bitstream information (addbsi) following the “addbsil” value.

[0043] The BSI section includes other metadata values ​​not explicitly shown in Figure 6.

[0044] According to a set of embodiments, an encoded audio bitstream represents the audio content of multiple substreams. In some cases, the substreams represent the audio content of a multi-channel program, and each substream represents one or more program channels. In other cases, the multiple streams of the encoded audio bitstream represent the audio content of several audio programs, typically a "main" audio program (which may be a multi-channel program) and at least one other audio program (e.g., a commentary program on the main audio program).

[0045] An encoded audio bitstream representing at least one audio program necessarily includes audio content from at least one "independent" substream. An independent substream represents at least one channel of the audio program (for example, an independent substream could represent a traditional 5.1 channel audio program with five full-range channels). This audio program is referred to here as the "main" program.

[0046] In some group embodiments, the encoded audio bitstream represents two or more audio programs (a "main" program and at least one other audio program). In such cases, the bitstream comprises two or more independent substreams: a first independent substream representing at least one channel of the main program; and at least one other independent substream representing at least one channel of another audio program (a program different from the main program). Each independent bitstream can be decoded independently, and a decoder can operate to decode only (not all) subgroups of the independent substreams of the encoded bitstream.

[0047] In a typical example representing two independent substreams of encoded audio bitstreams, one independent substream represents the standard format speaker channel of a multi-channel main program (e.g., the left, right, center, left surround, and right surround full-range speaker channels of a 5.1 channel main program), and the other independent substreams represent annotation monotone audio messages on the main program (e.g., director's notes on a film, where the main program is the film's audio channel). In another example representing multiple independent substreams of encoded audio bitstreams, one independent substream represents the standard format speaker channel of a multi-channel main program (e.g., a 5.1 channel main program), which contains dialogue in the first language (e.g., one of the main program's speaker channels could represent this dialogue), and the various other independent substreams represent monotone translations of this dialogue (in different languages).

[0048] Alternatively, it may indicate that the encoded audio bitstream of the main program (and at least one other audio program in the selection area) contains at least one "dependent" substream containing audio content. Each dependent substream is associated with an independent substream of the bitstream and represents at least one additional channel of the program (e.g., the main program) whose content is represented by the associated independent substream (i.e., the dependent substream represents at least one channel in the program that is not represented by the associated independent substream, and the associated independent substream represents at least one channel of the program).

[0049] In an example of a coded bitstream that includes a separate secondary stream (representing at least one channel of the main program), the bitstream also contains (related to the separate bitstream) a corresponding secondary stream that represents one or more additional speaker channels of the main program. These additional speaker channels are additional to the main program channel represented by the separate secondary stream. For example, if the separate secondary stream represents the standard format full-range speaker channels (left, right, center, left surround, right surround) of the 7.1 channel main program, then the corresponding secondary stream could represent those other two full-range speaker channels of the main program.

[0050] According to the E-AC-3 standard, an E-AC-3 bitstream must represent at least one independent substream (e.g., a single AC-3 bitstream) and can represent up to eight independent substreams. Each independent substream of an E-AC-3 bitstream can be associated with up to eight phase substreams.

[0051] E-AC-3 bitstreams include metadata representing the substream structure of the bitstream. For example, the "chanmap" field in the Bitstream Information (BSI) area of ​​an E-AC-3 bitstream determines the channel map representing the program channels of the corresponding substreams of that bitstream. However, the metadata representing the substream structure is traditionally included in the E-AC-3 bitstream in a format that makes it convenient only for access and use by the E-AC-3 decoder (during decoding of the encoded E-AC-3 bitstream); and not accessed or used after decoding (e.g., by a post-processor) or before decoding (e.g., by a processor configured to identify the metadata). At the same time, there is a risk that the decoder may use conventionally included metadata to incorrectly identify the secondary stream of the conventional E-AC-3 encoded bitstream, and this was unknown until the present invention. It is possible to include secondary stream structure metadata in the encoded bitstream (e.g., encoded E-AC-3 bitstream) in a format that allows for convenient and effective detection and correction of errors in secondary stream identification during the decoding of the bitstream.

[0052] E-AC-3 bitstreams can also contain metadata about the audio content of an audio program. For example, an E-AC-3 bitstream representing an audio program may contain metadata indicating the minimum and maximum frequencies of the spectrum extension processing (and channel coupling coding) used to encode the program content. However, this metadata is typically included in the E-AC-3 bitstream in a format that is only convenient for the E-AC-3 decoder to access and use (during decoding of the encoded E-AC-3 bitstream); it is inconvenient to access and use after decoding (e.g., by a later processor) or before decoding (e.g., by a processor configured to identify the metadata). Furthermore, this metadata is not included in the E-AC-3 bitstream in a format that allows for convenient and effective error detection and correction for the identification of this metadata during decoding.

[0053] According to a typical embodiment of the present invention, the PIM and / or SSM (and optionally other metadata, such as loudness processing status metadata or "LPSM") are embedded in one or more reserved columns (or slots) of the metadata segment of the audio bitstream, which also includes audio data in other segments (audio data segments). Typically, at least one segment of each frame of the bitstream contains the PIM or SSM, and at least another segment of the frame contains the corresponding audio data (i.e., audio data whose secondary stream structure is represented by the SSM and / or at least one feature or characteristic represented by the PIM).

[0054] In a set of embodiments, each metadata segment is a data structure (sometimes referred to herein as a box), which may contain one or more metadata payloads. Each payload includes a header with a specific payload identifier (and payload configuration data) to provide an explicit indication of the type of metadata appearing in the payload. The payload order within the box is not defined, allowing payloads to be stored in any order, and the parser must be able to parse the entire box to retrieve relevant payloads and ignore irrelevant or unsupported payloads. Figure 8 (described below) illustrates the structure of such a box and the payloads within it.

[0055] When two or more audio processing units need to operate cascaded over the entire processing chain (or content lifecycle), transmitting metadata (e.g., SSM and / or PIM and / or LPSM) in the audio data processing chain is particularly useful. Without metadata in the audio bitstream, serious media processing problems such as quality, level, and spatial degradation can occur, for example, when two or more audio codecs are used in the chain and when the single-ended volume level is applied more than once (or at the performance point of the audio content in the bitstream) during the bitstream path to the media consumption device.

[0056] According to some embodiments of the present invention, the loudness processing status metadata (LPSM) embedded in the audio bitstream can be identified and verified, for example, to enable loudness management agencies to verify whether the loudness of a particular program is within a specified range and whether the relevant audio data itself has been modified (to ensure compliance with applicable regulations). The loudness value contained within a data block having the loudness processing status metadata can be read out to verify this, rather than recalculating the loudness. In response to the LPSM, (as the LPSM indicates) management agencies can determine whether the relevant audio content complies with loudness regulations and / or management requirements (e.g., regulations under the Commercial Advertising Loudness Mitigation Act, also known as the "CALM" Act) without having to calculate the loudness of the audio content.

[0057] Figure 1 is a block diagram illustrating an audio processing chain (audio data processing system), wherein one or more components of the system can be configured according to embodiments of the present invention. The system includes the following components, coupled together as shown: a preprocessing unit, an encoder, a signal analysis and metadata correction unit, a transcoder, a decoder, and a post-processing unit. In variations of the system shown, one or more components are omitted or other audio data processing units are also included.

[0058] In some implementations, the preprocessing unit of FIG1 is configured to accept PCM (time domain) samples containing audio content as input and output processed PCM samples. The encoder can be configured to accept PCM samples as input and output an audio bitstream representing the encoded (e.g., compressed) audio content. The data representing the bitstream of the audio content is sometimes referred to herein as "audio data." If the encoder is configured according to a typical embodiment of the invention, the audio bitstream output by the autoencoder includes PIM and / or SSM (and preferably also loudness processing state metadata and / or other metadata) and audio data.

[0059] The signal analysis and metadata correction unit in Figure 1 can accept one or more coded audio bitstreams as input and determine (e.g., verify) whether the metadata (e.g., processing status metadata) in each coded audio bitstream is correct by performing signal analysis (e.g., using program boundary metadata in the coded audio bitstream). If the signal analysis and metadata correction unit finds the included metadata to be invalid, it typically replaces the incorrect value with the correct value obtained from the signal analysis. Therefore, each coded audio bitstream output from the signal analysis and metadata correction unit contains corrected (or uncorrected) processing status metadata and coded audio data.

[0060] The transcoder of Figure 1 can accept an encoded audio bitstream as input and respond (e.g., by decoding the input stream and then re-encoding the decoded stream in a different encoding format) to output a modified (e.g., differently encoded) audio bitstream. If the transcoder is configured according to a typical embodiment of the present invention, the audio bitstream output from the transcoder includes SSM and / or PIM (and typically also other metadata) and encoded audio data. The metadata may also be included in the input bitstream.

[0061] The decoder of Figure 1 can accept an encoded (e.g., compressed) audio bitstream as input and output a stream of decoded PCM audio samples. If the decoder is configured according to a typical embodiment of the invention, the decoder's output in typical operation is any of the following or includes any of the following:

[0062] Audio sample stream, and at least one corresponding stream's SSM and / or PIM (and typically other metadata) captured from the input coded bit stream; or

[0063] The audio sample stream, and the corresponding stream of control bits determined by the SSM and / or PIM (and typically other metadata, such as LPSM) captured from the input coded bit stream; or

[0064] An audio sample stream that does not have a corresponding stream of metadata or control bits determined by the metadata. In the latter, the decoder can extract metadata from the input encoded bit stream and perform at least one operation (e.g., verification) on the extracted metadata, even if it does not output the extracted metadata or control bits determined therein.

[0065] By means of the post-processing unit according to the configuration diagram 1 of a typical embodiment of the present invention, the post-processing unit is configured to receive a decoded PCM audio sample stream and perform post-processing (e.g., volume level of the audio content) on it using the SSM and / or PIM (and typically other metadata, such as LPSM) received along with the sample, or, control bits of the metadata received along with the sample determined by the decoder. The post-processing unit is also typically configured to play the post-processed audio content through one or more speakers.

[0066] Typical embodiments of the present invention provide an enhanced audio processing chain, wherein audio processing units (e.g., encoders, decoders, transcoders, and pre- and post-processing units) adapt their individual processing to the audio data based on the simultaneous state of the media data represented by the metadata individually received by the audio processing units.

[0067] Audio data input to any audio processing unit of the system of FIG1 (e.g., the encoder or transcoder of FIG1) may include SSM and / or PIM (and optionally other metadata) and audio data (e.g., encoded audio data). According to embodiments of the present invention, this metadata may be included in the input audio by another unit of the system of FIG1 (or another source not shown in FIG1). The processing unit receiving the input audio (and metadata) may be configured to perform at least one operation (e.g., verification) on the metadata or respond to the metadata (e.g., adaptive processing of the input audio), and typically includes the metadata, a processed version of the metadata, or control bits determined by the metadata in its output audio.

[0068] A typical embodiment of the audio processing unit (or audio processor) of the present invention is configured to perform adaptive processing of the audio data based on the state of the audio data represented by metadata relating to the audio data. In some embodiments, the adaptive processing is (or includes) loudness processing (if the metadata indicates that loudness processing or similar processing has not been performed on the audio data), but is not (and does not include) loudness processing (if the metadata indicates that this loudness processing or similar processing has been performed on the audio data). In some embodiments, the adaptive processing is or includes metadata verification (e.g., performed in a metadata verification subunit) to ensure that the audio processing unit performs other adaptive processing on the audio data based on the state of the audio data represented by the metadata. In some embodiments, verification determines the reliability of metadata relating to the audio data (e.g., contained in a bitstream). For example, if the metadata is verified to be reliable, then Results from previously executed audio processing types can be reused, and the re-execution of the same type of audio processing can be avoided. On the other hand, if metadata is deemed to have been tampered with (or unreliable), the media processing type claimed to have been previously executed (represented by unreliable metadata) can be repeated by the audio processing unit, and / or additional processing can be performed on the metadata and / or audio data by the audio processing unit. Audio processing units can also be configured to send a message to other audio processing units downstream in the enhanced media processing chain, informing them (e.g., appearing in the media bitstream) that the metadata is valid if the unit determines the metadata is valid (e.g., based on a match between the captured cipher value and a reference cipher value).

[0069] Figure 2 is a block diagram of the encoder (100) of an embodiment of the audio processing unit of the present invention. Any element or unit of the encoder 100 may be implemented as one or more programs and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits), becoming hardware, software, or a combination of hardware and software. The encoder 100 includes a frame buffer 110, a parser 111, a decoder 101, an audio state verifier 102, a loudness processing level 103, an audio stream selection level 104, an encoder 105, a filler / formatting level 107, a metadata generator 106, a dialogue loudness measurement system 108, and a frame buffer 109, and is connected as shown. Typically, the encoder 100 also includes other processing elements (not shown).

[0070] The encoder 100 (as a transcoder) is configured to convert an input audio bitstream (e.g., one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream) into an encoded output audio bitstream (e.g., another of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream), which includes performing adaptive and automatic loudness processing by using loudness processing state metadata included in the input bitstream. For example, the encoder 100 can be configured to convert an input Dolby E bitstream (a format typically used in production and broadcasting facilities, rather than for consumer devices receiving audio programs already broadcast to it) into an encoded output audio bitstream in AC-3 or E-AC-3 format (suitable for broadcast to consumer devices).

[0071] The system in Figure 2 also includes an encoded audio transmission subsystem 150 (which stores and / or transmits the encoded bitstream output by the autoencoder 100) and a decoder 152. The encoded audio bitstream output by the autoencoder 100 may be stored in the subsystem 150 (e.g., in DVD or Blu-ray disc format), or transmitted by the subsystem 150 (which may implement a transmission link or network), or may be stored and transmitted by the subsystem 150. The decoder 152 is configured to decode the encoded audio bitstream (generated by the encoder 100) received via the subsystem 150, which includes: extracting metadata (PIM and / or SSM, and optionally loudness processing status metadata and / or other metadata) from each frame of the bitstream (and optionally extracting program boundary metadata from the bitstream); and generating encoded audio data. Typically, decoder 152 is configured to perform adaptation processing on decoded audio data using PIM and / or SSM, and / or LPSM (and optional program boundary metadata), and / or transmit decoded audio data and metadata to a postprocessor configured to perform adaptation processing on the decoded audio data using metadata. Typically, decoder 152 includes a buffer that stores (in a non-transitory manner) the encoded audio bitstream received from subsystem 150.

[0072] Various implementations of encoder 100 and decoder 152 are configured to perform different embodiments of the method of the present invention.

[0073] Frame buffer 110 is a buffer memory coupled to receive encoded input audio bitstream. In operation, buffer 110 stores (e.g., in a non-transient manner) at least one frame of the encoded audio bitstream, and a sequence of frames of the encoded audio bitstream is prompted to parser 111 by buffer 110.

[0074] The parser 111 is coupled and configured to extract PIM and / or SSM, loudness processing status metadata (LPSM), and selected program boundary metadata (and / or other metadata) from each frame of the encoded input audio containing this metadata, to prompt at least the LPSM (and selected program boundary metadata and / or other metadata) to the audio status verifier 102, loudness processing level 103, metadata generator 106, and subsystem 108 to extract audio data from the encoded input audio and prompt the decoder 101 with the audio data. The decoder 101 of the encoder 100 is configured to decode the audio data to generate decoded audio data and prompt the decoded audio data to the loudness processing level 103, audio stream selection level 104, subsystem 108, and typically status verifier 102.

[0075] The status verifier 102 is configured to authenticate and verify the LPSM (and other metadata selected) indicated to it. In some embodiments, the LPSM is (or is contained in) a data block already included in the input bitstream (e.g., according to embodiments of the present invention). This block may contain a cryptographic hash (hash main information authentication code or "HMAC") for processing the LPSM (and other metadata selected) and / or (provided to the verifier 102 by the decoder 101) embedded audio data. In these embodiments, the data block may be digitally signed, allowing downstream audio processing units to readily authenticate and verify the processed status metadata.

[0076] For example, HMAC is used to generate a digest, and the protection value contained in the bit stream of this invention may include this digest. This digest can be generated for AC-3 frames as follows:

[0077] 1. After the AC-3 data and LPSM are encoded, the frame data bytes (sequential frame_data #1 and frame_data #2) and LPSM data bytes are used as inputs to the HMAC hash function. Other data that may appear in the auxdata field is not considered in calculating the digest. This other data can be bytes that are not AC-3 data or LPSM data. Guard bits included in the LPSM may be ignored in calculating the HMAC digest.

[0078] 2. After the summary is calculated, it is written into the column of the bit stream reserved for the protection bits.

[0079] 3. The final step in generating a complete AC-3 frame is to calculate the CRC check. This is taken into account all data written to the end of the frame and belonging to this frame, including the LPSM bits.

[0080] Other cryptographic methods, including but not limited to one or more non-HMAC cryptographic methods, may be used to verify LPSM and / or other metadata (e.g., in verifier 102) to ensure the secure transmission and reception of metadata and / or embedded audio data. For example, verification (using this cryptographic method) may be performed in various audio processing units that receive embodiments of the audio bitstream of the present invention to determine whether the metadata and associated audio data contained in the bitstream have been (as indicated by the metadata) specifically processed (and / or have a result), and have not been modified after performing this specific processing.

[0081] The state verifier 102 provides control data to the audio stream selection stage 104, the metadata generator 106, and the audio response measurement system 108 to indicate the result of the verification operation. In response to the control data, stage 104 can select (and transmit to encoder 105):

[0082] The adaptive processing output of loudness processing level 103 (e.g., when LPSM indicates that the audio data output from decoder 101 has not undergone a specific type of loudness processing, and control bits from verifier 102 indicate that LPSM is valid); or

[0083] The audio data output from decoder 101 (e.g., when LPSM indicates that the audio data output from decoder 101 has been subjected to a specific type of loudness processing, which will be performed by loudness processing level 103, and control bits from verifier 102 indicate that LPSM is valid).

[0084] The loudness processing stage 103 of encoder 100 is configured to perform adaptive loudness processing on the decoded audio data output from decoder 101, based on one or more audio data characteristics represented by the LPSM captured by decoder 101. Loudness processing stage 103 may be an adaptive transposition real-time loudness and dynamic range control processor. Loudness processing stage 103 may receive user input (e.g., user-targeted loudness / dynamic range value or dialnorm value), or other metadata input (e.g., one or more types of third-party data, tracking information, identification codes, proprietary or standard information, user annotation data, user preference data, etc.) and / or other input (e.g., from fingerprint processing), and use this input to process the decoded audio data output from decoder 101. Loudness processing level 103 can perform adaptive loudness processing on decoded audio data (output from decoder 101) representing a single audio program (as represented by program boundary data captured by parser 111); and can reset loudness processing in response to receiving decoded audio data (output from decoder 101) representing different audio programs (as represented by program boundary data captured by parser 111).

[0085] When control bits from verifier 102 indicate that the LPSM is invalid, the dialogue loudness measurement system 108 can, for example, use the LPSM (and / or other metadata) captured by decoder 101 to determine the loudness of a segment of decoded audio representing dialogue (or other speech) (from the decoder). When control bits from verifier 102 indicate that the LPSM is valid, the dialogue loudness measurement system 108 can operate when the LPSM indicates that the previously determined dialogue (or other speech) segment of decoded audio (from decoder 101) is disabled. The system 108 can perform loudness measurement on decoded audio data representing a single audio program (as represented by program boundary metadata captured by parser 111) and can reset the measurement in response to receiving decoded audio data representing a different audio program represented by the program boundary metadata.

[0086] There are existing useful tools for conveniently and easily measuring the level of dialogue in audio content (e.g., the Dolby LM100 loudness meter). Some embodiments of the present invention's APU (e.g., stage 108 of encoder 100) are implemented to include this tool (or perform the function of this tool) to measure audio bitstreams (e.g., decoded AC-3 bitstreams prompted by decoder 101 of encoder 100 to stage 108).

[0087] If level 108 is implemented to measure the true average dialogue loudness of the audio data, the measurement method may include the step of isolating segments of audio content that primarily contain speech. These speech-centric audio segments are then processed according to a loudness measurement algorithm. For audio data decoded from an AC-3 bitstream, this algorithm can be a standard K-weighted loudness measurement (e.g., according to international standard ITU-R BS.1770). Alternatively, other loudness measurement methods may be used (e.g., based on a psychoacoustic model of loudness).

[0088] Isolation of speech segments is not necessary for measuring the average dialogue loudness of audio data. However, this improves the accuracy of the measurement and typically provides more satisfactory results for the listener. Because not all audio content contains dialogue (speech), loudness measurement of the entire audio content can provide a sufficiently approximation of the dialogue level where speech has already occurred.

[0089] Metadata generator 106 generates (and / or transmits through stage 107) metadata contained in the encoded bitstream as output by encoder 100. Metadata generator 106 may transmit LPSM (and optionally LIM and / or PIM and / or program boundary metadata and / or other metadata) captured by decoder 101 and / or parser 111 to stage 107 (e.g., when control bits from verifier 102 indicate that LPSM and / or other metadata are valid), or generate new LIM and / or PIM and / or LPSM and / or program boundary metadata and / or other metadata to prompt stage 107 with the new metadata (e.g., when control bits from verifier 102 indicate that metadata captured by decoder 101 is invalid), or prompt stage 107 with a combination of metadata captured by decoder 101 and / or parser 111 and newly generated metadata. The metadata generator 106 may contain loudness data generated for the subsystem 108, the at least one value indicating the type of loudness processing performed by the subsystem 108, the LPSM prompted by the level 107 for inclusion in the encoded bit stream output by the encoder 100.

[0090] Metadata generator 106 can generate protection bits (which may contain a hash master information authentication cipher or "HMAC" or something similar) for decrypting, authenticating, or verifying LPSM (and optionally other metadata) and / or embedded audio data contained in the coded bitstream. Metadata generator 106 can provide such protection bits to level 107 for inclusion in the coded bitstream.

[0091] In typical operation, the dialogue loudness measurement system 108 processes the audio data output from the decoder 101 in response to generate loudness values ​​(such as gated or ungated dialogue loudness values) and dynamic range values. In response to these values, the metadata generator 106 can generate loudness processing state metadata (LPSM) to be included in the encoded bitstream output by the encoder 100 (for filler / formatting level 107).

[0092] Alternatively, the subsystems 106 and / or 108 of encoder 100 may perform additional analysis on the audio data to generate metadata representing at least one feature of the audio data contained in the coded bitstream output by stage 107.

[0093] Encoder 105 encodes (e.g., by performing compression) the audio data output from selector 104 and prompts 107 with encoded audio data to be included in the encoded bitstream output by 107.

[0094] Stage 107 multiplexes encoded audio from encoder 105 and metadata (including PIM and / or SSM) from metadata generator 106 to generate an encoded bitstream output by stage 107, preferably such that the encoded bitstream has a format as specified in the preferred embodiment of the present invention.

[0095] The frame buffer 109 is a buffer memory that (e.g., in a non-transient manner) stores at least one frame of the encoded bitstream output from stage 107, and a sequence of frames of the encoded audio bitstream, which are then prompted by the buffer 109 as output from encoder 100 to be delivered to system 150.

[0096] The LPSM system generated by the metadata generator 106 and included in the encoded bitstream of stage 107 typically represents the loudness processing status of the corresponding audio data (e.g., the type of loudness processing already performed on the audio data) and the loudness of the related audio data (e.g., measuring dialogue loudness, gated and / or ungated loudness, and / or dynamic range).

[0097] Here, "gluing on" loudness and / or level measurements on audio data represents a specific level or loudness threshold, and calculated values ​​exceeding this threshold are included in the final measurement (e.g., short-term loudness values ​​below -60 dBFS are ignored in the final measurement). Gluing on absolute values ​​represents a fixed level or loudness, while gluing on relative values ​​represents a value dependent on the current "unglued" measurement.

[0098] In some embodiments of encoder 100, the encoded bitstream buffered in memory 109 (and output to transport system 150) is an AC-3 bitstream or an E-AC-3 bitstream, and includes audio data segments (e.g., segments AB0-AB5 shown in the frame of FIG4) and metadata segments, wherein the audio data segments represent audio data, and at least a portion of each metadata segment contains PIM and / or SSM (and optionally other metadata). Stage 107 inserts the metadata segments (containing metadata) into the bitstream in the following format: Each metadata segment containing PIM and / or SSM is included in the discarded bitstream segment of the bitstream (e.g., the discarded bitstream segment "W" shown in FIG4 or FIG7) or in the "addbsi" field of the Bitstream Stream Information (BSI) segment of the frame of the bitstream, or in the auxdata field at the end of the frame of the bitstream (e.g., the AUX segment shown in FIG4 or FIG7). A frame of a bitstream can contain one or two metadata segments, each containing metadata. If the frame contains two metadata segments, one can appear in the addbsi field of the frame, and the other can appear in the AUX field of the frame.

[0099] In some embodiments, each metadata segment (sometimes referred to as a "box") inserted for level 107 has a format that includes a metadata segment header (and optionally other mandatory or "core" elements), and one or more metadata payloads following the metadata segment header. If present, a SIM is included in one of the metadata payloads (specified in the payload header and typically having a first type format). If present, a PIM is included in another metadata payload (specified in the payload header and typically having a second type format). Similarly, each type of metadata (if present) is included in another metadata payload (specified in the payload header and typically having a format specific to that metadata type). The exemplified format allows convenient access to SSM, PIM, and other metadata at times other than decoding (e.g., in a post-decoding post-processor, or by configuring a processor to identify metadata without performing full decoding of the entire encoded bitstream), and allows convenient and effective error detection and correction (e.g., substream identification) during bitstream decoding. For example, when the SSM is not accessed in the exemplary format, the decoder may incorrectly identify the correct number of substreams for a program. One metadata payload in a metadata segment may contain the SSM, another metadata payload in a metadata segment may contain the PIM, and optionally, at least one other metadata payload in a metadata segment may contain other metadata (e.g., loudness processing status metadata or "LPSM").

[0100] In some embodiments, the substream structure metadata (SSM) payload contained in the frame of the coded bitstream (e.g., representing an E-AC-3 bitstream of at least one audio program) in level 107 includes an SSM in the following format:

[0101] The payload header typically includes at least one identification value (e.g., a 2-bit value indicating the SSM format version, and selected length, period, count, and sub-stream related values); and Following the letter header: Independent substream data, represented by the number of independent substreams of the program represented by the bitstream; and The sequential stream data indicates whether each independent substream of the program has at least one related sequential stream (i.e., whether at least one sequential stream is related to each independent substream), and if so, the number of sequential streams is related to each independent substream of the program.

[0102] It is conceivable that an independent substream of a coded bitstream can represent a set of speaker channels for an audio program (e.g., speaker channels for a 5.1 speaker channel audio program), and one or more sequential streams (as indicated by the sequential stream metadata) can represent the target channel of the program. However, typically, an independent substream of a coded bitstream represents a set of speaker channels for a program, and each sequential stream of the independent substream (as indicated by the sequential stream metadata) represents at least one additional speaker channel for the program.

[0103] In some embodiments (for level 107), the Program Information Metadata (PIM) payload contained in a frame of a coded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) has the following format: The payload header typically includes at least one identification value (e.g., a value indicating the PIM format version, and also values ​​related to length, period, count, and sub-stream); and Following this header, the PIM format is as follows:

[0104] The actuation channel metadata indicates the various silent and non-silent channels of the audio program (i.e., which channels of the program contain audio information, and (if any) which contain only silence (typically during the frame)). In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the actuation channel metadata in the frame of the bitstream can be used in conjunction with additional metadata of the bitstream (e.g., the audio encoding mode (acmod) field of the frame, if any, in the chanmap field of that frame or related sequential stream frames) to determine which channels of the program contain audio information and which contain silence. The "acmod" field of an AC-3 or E-AC-3 frame indicates the number of channels in the full range of the audio program represented by the audio content of that frame (e.g., the program is a 1.0 channel mono program, a 2.0 channel stereo program, or a program containing the full range of L, R, C, Ls, Rs channels), or the frame represents two independent 1.0 channel mono programs. The "chanmap" of an E-AC-3 bitstream refers to the channel map of the sequential streams indicated by that bitstream. Actuating channel metadata can be used for upmixing (in the post-processor) downstream of the decoder, for example, adding audio to the decoder output onto a channel containing silence.

[0105] The downmixing status metadata indicates whether the program (before or during encoding) is downmixed, and if so, the type of downmixing applied. The downmixing status metadata may include upmixing for downstream implementation of the decoder (in the post-processor), for example, using parameters that best match the applied downmixing type to upmix the program's audio content. In embodiments where the encoded bitstream is AC-3 or E-AC-3 bitstream, the downstream processing status metadata can be used in conjunction with the frame's audio encoding mode (acmod) field to determine the type of downmixing (if any) applied to the channel of the program.

[0106] The upmixing status metadata indicates whether the program was upmixed (e.g., from a smaller number of channels) before or during encoding, and if so, the type of upmixing applied. The upmixing status metadata may include downmixing (in the post-processor) for implementing downstream decoding, such as downmixing the program's audio content to match the type of upmixing applied to the program (e.g., Dolby Pro Logic, or Dolby Pro Logic II Cinema Mode, or Dolby Pro Logic II Music Mode, or Dolby Pro Upmixer). In embodiments where the encoded bitstream is an E-AC-3 bitstream, the upmixing status metadata may be used in conjunction with other metadata (e.g., the value of the "strmtyp" field in the frame) to determine, if any, the type of upmixing applied to the program channel. The value of the “strmtyp” field (the BSI segment of a frame in an E-AC-3 bitstream) indicates whether the audio content of the frame belongs to an independent stream (which determines the program) or an independent substream (containing or relating to a program with multiple substreams), and therefore can be decoded independently of any other substream represented by the E-AC-3 bitstream; or, the audio content of the frame belongs to a related substream (containing or relating to a program with multiple substreams), and therefore must be decoded in conjunction with its associated independent substream; and

[0107] The preprocessing status metadata indicates whether preprocessing has been performed on the audio content of this frame (before encoding the audio content to generate the encoded bitstream), and if so, the type of preprocessing performed.

[0108] In some implementations, preprocessed state metadata is represented as:

[0109] Whether to apply surround attenuation (e.g., whether the surround channels of an audio program are attenuated by 3dB before encoding).

[0110] Whether to apply a 90-degree phase shift (e.g., on the surround channels Ls and Rs of the audio program before encoding).

[0111] Whether a low-pass filter is applied to the LFE channel of the audio program before encoding,

[0112] Was the LFE channel's bitrate monitored during production? If so, what was the monitoring bitrate of the LFE channel relative to the bitrate of the program's full-range audio channels?

[0113] Whether dynamic range compression should be performed (e.g., in the decoder) on each block of the decoded audio content of the program, and if so, the type (and / or parameters) of dynamic range compression to be performed (e.g., this type of preprocessing state metadata could indicate which of the following compression distribution types is assumed by the encoder to produce dynamic range compression control values ​​contained in the encoded bitstream: cinematic standard, cinematic light, music standard, music light, or speech. Alternatively, this type of preprocessing state metadata could indicate that re-dynamic range compression (“compr” compression) should be performed on each frame of the decoded audio content of the program in a manner determined by the dynamic range compression control values ​​contained in the encoded bitstream).

[0114] Whether spectrum spreading and / or channel coupling coding are used to encode a specific frequency range of the program content, and if so, the minimum and maximum frequencies of the frequency components of the content to be processed by spectrum spreading coding, and the minimum and maximum frequencies of the frequency components to be processed by channel coupling coding. This type of preprocessing state metadata can be used for equalization (in the post-processor) of the downstream of the decoder. Frequency coupling and spectrum spreading information are used to optimize the quality during transcoding operations and applications. For example, the encoder can optimize its behavior (including employing preprocessing steps, such as headphone virtualization, upmixing, etc.) based on the state of parameters, such as spectrum spreading and channel coupling information. Furthermore, the encoder can dynamically adapt its coupling and spectrum spreading parameters to match and / or optimize values ​​based on the state of the incoming (and identified) metadata.

[0115] Whether the dialogue enhancement adjustment range data is included in the encoded bitstream, and if so, the adjustment range available during the execution of the dialogue enhancement process (e.g., downstream of the decoder's post-processor) to adjust the bit level of the dialogue content relative to the bit level of the non-dialogue content in the audio program.

[0116] In some implementations, additional preprocessed state metadata (e.g., metadata representing headphone-related parameters) is included in the PIM payload of the encoded bit stream output by encoder 100 (level 107).

[0117] In some embodiments, the LPSM payload contained in the frame of the coded bitstream (e.g., representing an E-AC-3 bitstream of at least one audio program) in level 107 contains LPSM in the following format:

[0118] (Typically includes a header that indicates the start of the LPSM payload, followed by at least one identifying value, such as the LPSM format version, length, period, count, and the secondary stream-related values ​​shown in Table 2 below; and)

[0119] After the letterhead, At least one dialogue indication value (e.g., the parameter "Dialogue Channel" in Table 2) indicates whether the relevant audio data indicates a dialogue or not (e.g., which channels of the relevant audio data represent dialogue);

[0120] At least one loudness regulation compliance value (e.g., the parameter "Loudness Regulation Type" in Table 2) indicates whether the corresponding audio data complies with the loudness regulation of the specified group;

[0121] At least one loudness processing value (e.g., one or more of the parameters "Dialogue Gate Loudness Correction Flag" and "Loudness Correction Type" in Table 2) indicates the type of loudness processing that has been performed on the corresponding audio data; and

[0122] At least one loudness value (e.g., one or more of the parameters “ITU relative gated loudness”, “ITU voice gated loudness”, “ITU (EBU3341) short-term 3s loudness”, and “true peak” in Table 2) represents at least one loudness (e.g., peak or average loudness) characteristic of the relevant audio data.

[0123] In some embodiments, each metadata segment containing PIM and / or SSM (and optionally other metadata) includes a metadata segment header (and optionally other additional core elements), and after the metadata segment header (or metadata segment signal and other core elements), at least one metadata payload segment has the following format:

[0124] The load signal typically includes at least one identification value (e.g., SSM or PIM format version, length, period, count, and secondary current related value), and

[0125] Following the payload header is SSM or PIM (or another type of metadata).

[0126] In some implementations, the individual metadata segments (sometimes referred to here as "metadata boxes" or "boxes") of the discard bit / jump column segment (or "addbsi" column or "auxdata" column) of the frame of the bit stream inserted for level 107 have the following format:

[0127] The metadata segment header (typically includes a syncword indicating the start of the metadata segment, followed by identification values ​​such as version, length, period, expansion element count, and secondary stream-related values ​​as indicated in Table 1 below); and

[0128] Following the header of the metadata segment, at least one protection value (e.g., the HMACC digest and audio fingerprint value in Table 1) is included, which is at least one of the metadata used to decrypt, authenticate, or verify at least one piece of metadata in the metadata segment or the corresponding audio data; and

[0129] Additionally, following the metadata segment header, there is a metadata payload identification (ID) and payload configuration value, which indicate the metadata type in each of the following metadata payloads and indicate at least one aspect of the configuration of each payload (e.g., size).

[0130] Each metadata payload is accompanied by its corresponding payload ID and payload configuration value.

[0131] In some embodiments, the metadata segments in the discarded bit segment (or auxdata column or "addbsi" column) of the frame have a three-layer structure:

[0132] The high-level structure (e.g., metadata segment header) includes a flag indicating whether a bit (or auxdata or addbsi) is discarded. The field contains metadata, with at least one ID value indicating the type of metadata appearing, and typically, a value indicating how many (e.g., each type of) metadata bits appear (if any). One type of metadata that can appear is PIM, another type is SSM, and yet another type is LPSM, and / or program boundary metadata, and / or media research metadata.

[0133] The middle layer contains data about various specified types of metadata (e.g., metadata payload headers, protection values, payload IDs, and payload configuration values ​​used for each specified type of metadata); and

[0134] The low-level structure contains metadata payloads for each specified type of metadata (e.g., a sequential PIM value if PIM is specified as present, and / or another type of metadata value (e.g., SSM or LPSM if this type of metadata is specified as present).

[0135] Data values ​​in this three-layer structure can be nested. For example, protection values ​​identified for each payload (e.g., each PIM, or SSM, or other metadata payload) by the high and middle layers can be included after the payload (and thus after the metadata payload header of the payload), or protection values ​​for all metadata payloads identified for the high and middle layers can be included after the final metadata payload in the metadata segment (and thus after the metadata payload header of all payloads in the metadata segment).

[0136] In one embodiment (described with reference to the metadata segment or “box” in FIG8), a metadata segment header identifies four metadata payloads. As shown in FIG8, the metadata segment header includes a box synchronization character (identified as “box synchronization”) and a version and key ID value. The metadata segment header is followed by the four metadata payloads and protection bits. The payload ID and payload configuration (e.g., payload size) values ​​for the first payload (e.g., PIM payload) follow the metadata segment header, and the first payload itself follows the ID and configuration values; the payload ID and payload configuration (e.g., payload size) values ​​for the second payload (e.g., SSM payload) follow the first payload; the second payload itself follows these IDs and configuration values; the payload ID and payload configuration (e.g., payload size) values ​​for the third payload (e.g., LPSM payload) follow the second payload; and the third payload itself follows these IDs and configuration values; the payload ID and payload configuration (e.g., payload size) values ​​for the fourth payload follow the third payload; the fourth payload itself follows these IDs and configuration values; and the protection values ​​(identified as "protection data" in Figure 8) for all and some payloads (for high and mid-level structures and all or some payloads) follow the last payload.

[0137] In some embodiments, if decoder 101 receives an audio bitstream with a cryptographic hash generated according to embodiments of the present invention, the decoder is configured to parse and retrieve the cryptographic hash using data blocks determined by the bitstream, wherein the blocks contain metadata. Verifier 102 can use the cryptographic hash to verify the received bitstream and / or related metadata. For example, if verifier 102 considers the metadata valid based on a match between a reference cryptographic hash and a self-retrieved data block cryptographic hash, it disables the loudness processing level 103's operation on the related audio data and allows selection level 104 to pass (unmodified) the audio data. Alternatively, other types of cryptographic techniques may be used instead of the cryptographic hash method.

[0138] The encoder 100 of Figure 2 can (in response to the LPSM and optionally the program boundary metadata captured by the decoder 101) determine that the post / preprocessing unit has performed a type of loudness processing on the encoded audio data (in elements 105, 106, and 107) and can therefore (in the metadata generator 106) establish loudness processing state metadata containing specific parameters for previously performed loudness processing and / or derived therefrom. In some embodiments, the encoder 100 (and the encoded bitstream output contained therein) can establish metadata representing the processing history of the audio content, provided that the encoder is aware of the type of processing performed on the audio content.

[0139] Figure 3 is a block diagram of a decoder (200), which is an embodiment of the audio processing unit of the present invention, and a post-processor (300) coupled thereto. The post-processor (300) is also an embodiment of the audio processing unit of the present invention. Any element or component of the decoder 200 and the post-processor 300 may be implemented as one or more programs and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits), as hardware, software, or a combination of hardware and software. The decoder 200 includes a frame buffer 201, a parser 205, an audio decoder 202, an audio state verifier (verification stage) 203, and a control bit generator (generation stage) 204, and is connected as shown. Typically, the decoder 200 includes other processing elements (not shown).

[0140] Frame buffer 201 (buffer memory) stores (e.g., in a non-transient manner) at least one frame of the encoded audio bitstream received by decoder 200. A sequence of frames of the encoded audio bitstream is prompted to parser 205 by buffer 201.

[0141] The parser 205 is coupled and configured to capture PIM and / or SSM (and other metadata, such as LPSM, by choice) from each frame of the encoded input audio, to prompt at least a portion of the metadata (e.g., LPSM and program boundary metadata (if any of which is captured), and / or PIM and / or SSM) to the audio state verifier 203 and the control bit generator 204, to prompt the captured metadata as output (e.g., to the post-processor 300), to capture audio data from the self-encoded input audio, and to prompt the captured audio data to the decoder 202.

[0142] The encoded audio bitstream input to decoder 200 can be one of AC-3 bitstream, E-AC-3 bitstream, or Dolby E bitstream.

[0143] The system in Figure 3 also includes a post-processor 300. The post-processor 300 includes a frame buffer 301 and another processing element (not shown), which includes at least one processing element coupled to the buffer 301. The frame buffer 301 stores (e.g., in a non-transient manner) at least one frame in the decoded audio bitstream received by the post-processor 300 from the decoder 200. The processing element of the post-processor 300 is coupled and configured to receive and adaptively use metadata output from the decoder 200 and / or control bits output from the control bit generator 204 of the decoder 200 to process a sequence of frames of the encoded audio bitstream output from the buffer 301. Typically, the post-processor 300 is configured to perform adaptation processing on the decoded audio data using metadata from the decoder 200 (e.g., adapting the decoded audio data to loudness using LPSM values ​​and selected program boundary metadata, wherein the adaptation processing can be based on loudness processing status and / or one or more audio data characteristics, for audio data representing a single audio program as indicated by LPSM).

[0144] Various implementations of the decoder 200 and the post-processor 300 are configured to perform various different embodiments of the method of the present invention.

[0145] The audio decoder 202 of the decoder 200 is configured to decode the audio data captured by the parser 205 to generate decoded audio data and display the decoded audio data as output (e.g., to the post-processor 300).

[0146] Audio status verifier 203 is configured to authenticate and verify metadata prompted to it. In some embodiments, the metadata is (or is contained in) a data block that has been (e.g., according to embodiments of the invention) included in the input bitstream. This block may contain a cryptographic hash (hash master information authentication code or "HMAC") for processing metadata and / or embedded audio data (provided to audio status verifier 203 by parser 205 and / or decoder 202). In these embodiments, the data block may be digitally signed, making it relatively easy for downstream audio processing to authenticate and verify the processing status metadata.

[0147] Other cryptographic methods include, but are not limited to, any one or more non-HMAC cryptographic methods that can be used to verify metadata (e.g., in audio state verifier 203) to ensure secure transmission and reception of metadata and / or embedded audio data. For example, verification (using this cryptographic method) can be performed at various audio processing units that receive embodiments of the audio bitstream of the present invention to determine whether loudness processing state metadata and associated audio data contained in the bitstream have been subjected to (and / or resulted in) a specific loudness processing (as indicated by the metadata), and whether this specific loudness processing has not been corrected after its execution.

[0148] The audio status verifier 203 prompts control data, using the control bit generator 204 and / or the prompt control data as output (e.g., to the post-processor 300) to indicate the result of the verification operation. In response to the control data (and optionally other metadata retrieved from the input bit stream), the control bit generator 204 can generate (and prompt the post-processor 300) the following:

[0149] Control bits indicate that the decoded audio data output from decoder 202 has undergone a specific type of loudness processing (when LPSM indicates that the audio data output from decoder 202 has undergone a specific type of loudness processing, the control bits from audio state verifier 203 indicate that LPSM is valid); or

[0150] The control bits of the decoded audio data output by the decoder 202 should be subject to a specific type of loudness processing (e.g., when LPSM indicates that the audio data output by the decoder 202 is not subject to that specific type of loudness processing, or when LPSM indicates that the audio data output by the decoder 202 has been subject to a specific type of loudness processing, but the control bits from the audio status verifier 203 indicate that LPSM is not valid).

[0151] Alternatively, decoder 200 may provide the metadata extracted from the input bitstream by decoder 202 and the metadata extracted from the input bitstream by parser 205 to postprocessor 300, and postprocessor 300 may use the metadata to perform adaptation processing on the decoded audio data, or perform metadata verification and if the verification indicates that the metadata is valid, perform adaptation processing on the decoded audio data using the metadata.

[0152] In some embodiments, if decoder 200 receives an audio bitstream generated according to an embodiment of the present invention, with a cryptographic hash, the decoder is configured to parse and retrieve a cryptographic hash from a data block determined from the bitstream, the data block containing loudness processing status metadata (LPSM). Audio status verifier 203 can use the cryptographic hash to verify the received bitstream and / or related metadata. For example, if audio status verifier 203 finds the LPSM to be valid based on a match between a reference cryptographic hash and a cryptographic hash retrieved from a data block, it can signal to a downstream audio processing unit (e.g., post-processor 300, which may or may include a volume level unit) to pass the (unaltered) audio data of the bitstream. Alternatively, other types of cryptographic techniques may be used instead of the cryptographic hash method.

[0153] In some implementations of decoder 200, the received (and buffered in memory 201) encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and includes audio data segments (e.g., AB0-AB5 segments of the frame shown in Figure 4) and metadata segments, wherein the audio data segments represent audio data, and each of at least some metadata segments contains PIM or SSM (or other metadata). Decoder stage 202 (and / or parser 205) is configured to extract metadata from the bitstream. Each metadata segment containing PIM and / or SSM (and optionally other metadata) is included in the discarded bit segment of the frame of the bitstream, or in the "addbsi" field of the Bitstream Information (BSI) segment of the frame of the bitstream, or in the auxdata field at the end of the frame of the bitstream (e.g., the AUX segment shown in Figure 4). A frame of a bitstream can contain one or two data segments, each containing metadata. If the frame contains two data segments, one can appear in the addbsi field of the frame, and the other can appear in the AUX field of the frame.

[0154] In some embodiments, each metadata segment (sometimes referred to herein as a "box") of the bitstream buffered in buffer 201 has a format that includes a metadata segment header (and optionally other mandatory or "core" elements), and one or more metadata payloads, following the payload segment header. If present, a SIM is contained in a metadata payload (identified by the payload header, typically having a first type of format). If present, a PIM is contained in another metadata payload (identified by the payload header and typically having a second type of format). Similarly, each other type of metadata (if present) is contained in another metadata payload (identified by the payload header and typically having a format of a particular metadata type). The example format allows for convenient access to SSM, PIM, and other metadata at times other than decoding (e.g., in the post-processor 300 after decoding, or by a processor configured to identify metadata, without having to perform full decoding on the encoded bitstream), and allows for convenient and effective error detection and correction (e.g., secondary stream identification) during the decoding of the bitstream. For example, without access to the example format's SSM, decoder 200 may incorrectly identify the correct number of secondary streams for a program. One metadata payload in a metadata segment may contain an SSM, another metadata payload in a metadata segment may contain a PIM, or at least one other metadata payload in a metadata segment may contain other metadata (e.g., loudness processing status metadata or "LPSM").

[0155] In some embodiments, the secondary stream structure metadata (SSM) payload buffered in buffer 201 within a frame containing a coded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) includes SSMs in the following formats:

[0156] The payload header typically includes at least one identification value (e.g., a 2-bit value indicating the SSM format version, and selected length, period, count, and sub-stream related values); and After the letterhead:

[0157] Independent substream data is represented by the number of independent substreams of the program represented by that bitstream; and

[0158] The relative stream metadata indicates whether each independent substream of a program has at least one associated relative stream; if so, the number of relative streams is related to each independent substream of the program.

[0159] In some embodiments, a Program Information Metadata (PIM) payload buffered in buffer 201 and contained within a frame of a coded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program) has the following format:

[0160] The payload header typically includes at least one identification value (e.g., a value indicating the PIM format version, and optionally also values ​​related to length, period, count, and sub-stream); and Following the header, PIM takes the following format:

[0161] The active channel metadata for each silent and non-silent channel of the audio program (i.e., which channels of the program contain audio information, and if so, which are only muted (typically only during the frame)). In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the active channel metadata in the frame of the bitstream can be used in conjunction with additional metadata of the bitstream (e.g., the audio encoding mode (“acmod”) field of the frame, and, if so, the chanmap field in the frame or related sequential stream frames) to determine which channels of the program contain audio information and which are muted;

[0162] The downmixing processing level metadata indicates whether the program is downmixed (before or during encoding), and if so, the type of downmixing applied. The downmixing processing status metadata may include downstream upmixing for implementing the decoder (e.g., in post-processor 300), for example, using parameters that nearly match the applied downmixing type to upmix the program's audio content. In embodiments where the encoded bitstream is AC-3 or E-AC-3 bitstream, the downstream processing status metadata can be used in conjunction with the frame's audio encoding mode ("acmod") field to determine, if any, the type of downmixing applied to the program's channels;

[0163] The upmixing status metadata indicates whether a program (before or during encoding) is upmixed (e.g., by a smaller number of channels), and if so, the type of upmixing applied. The upmixing status metadata can be used to perform downmixing downstream of the decoder (in the post-processor), for example, downmixing the audio content of the program to conform to the type of upmixing applied to that program (e.g., Dolby Pro Logic, or Dolby Pro Logic II Cinema Mode, or Dolby Pro Logic II Music Mode, or Dolby Pro Upmixer). In embodiments where the encoded bitstream is an E-AC-3 bitstream, the upmixing status metadata can be used in conjunction with other metadata (e.g., the value of the “strmtyp” field in the frame) to determine, if any, the type of upmixing applied to the channels of the program. (In the BSI section of an E-AC-3 bitstream frame) the value of the "strmtyp" field indicates whether the audio content of the frame belongs to an independent stream (which determines a program) or an independent substream (containing multiple substreams or programs related to multiple substreams), and therefore can be independently decoded to any other substream represented by the E-AC-3 bitstream; or whether the audio content of the frame belongs to a sequential stream (or contains programs related to multiple substreams), and therefore must be decoded in conjunction with the associated independent substream; and

[0164] Preprocessing status metadata indicates whether preprocessing is performed on the audio content of this frame (generating a bitstream before encoding the audio content), and if so, the type of preprocessing performed.

[0165] In some embodiments, the preprocessed state metadata is represented as:

[0166] Whether surround attenuation is applied (e.g., whether the surround channels of an audio program are attenuated by 3dB before encoding).

[0167] Whether to apply a 90-degree phase shift (e.g., before encoding, around channels Ls and Rs).

[0168] Before encoding, is a low-pass filter applied to the LFE channel of the audio program?

[0169] During production, is the LFE channel level of the program monitored? If so, what is the monitoring level of the LFE channel relative to the level of the program's full-range audio channels?

[0170] Whether dynamic range compression should be performed (e.g., in the decoder) on each frame of the decoded audio content of the program, and if so, the type (and / or parameters) of dynamic range compression to be performed (e.g., this type of preprocessing state metadata could indicate which of the following compression distribution types is prompted by the encoder to produce dynamic range compression control values ​​contained in the encoded bitstream: cinematic standard; cinematic light; music standard; music light; or speech). Alternatively, this type of preprocessing state metadata could indicate that re-dynamic range compression (“compr” compression) should be performed on each frame of the decoded audio content of the program, in a manner determined by the dynamic range compression control values ​​contained in the encoded bitstream.

[0171] Whether spectrum spreading and / or channel coupling coding is used to encode a specific frequency range of program content, and if so, the minimum and maximum frequencies of the frequency components of the content executed by spectrum spreading coding, and the minimum and maximum frequencies of the frequency components of the content executed by channel coupling coding. This type of preprocessing state metadata information can be used to perform equalization downstream of the decoder (in the post-processor). Channel coupling and spectrum spreading information are also used for quality optimization during transcoding operations and applications. For example, the encoder can optimize its behavior (including adapting to preprocessing steps such as headphone virtualization, upmixing, etc.) based on the state of parameters such as spectrum spreading and channel coupling information. Furthermore, the encoder can dynamically adapt its coupling and spectrum spreading parameters to match and / or optimize values ​​based on the state of the incoming (and identified) metadata.

[0172] Whether the dialogue enhancement adjustment range data is included in the encoded bitstream, if so, then the range adjustment available during the execution of dialogue enhancement processing (e.g., downstream of the decoder's post-processor) is used to adjust the dialogue content level relative to the level of non-dialogue content in the audio program.

[0173] In some embodiments, the LPSM payload buffered in buffer 201, comprising a frame containing an encoded bitstream (e.g., an E-AC-3 bitstream representing at least one audio program), includes LPSM in the following formats:

[0174] The header (typically containing a syncword identifying the start of the LPSM payload, followed by at least one identification value, such as LPSM format version, length, period, count, and substream-related values ​​as shown in Table 2 below); and Following the letterhead,

[0175] At least one dialogue indicator value (e.g., the parameter "Dialogue Channel" in Table 2) indicates whether the corresponding audio data represents a dialogue or does not contain a dialogue (e.g., which channels' corresponding audio data represent dialogue).

[0176] At least one loudness regulation compliance value (e.g., the parameter "Loudness Regulation Type" in Table 2) indicates whether the corresponding audio data complies with the loudness regulation of the instruction group;

[0177] At least one loudness processing value (e.g., one or more parameters in Table 2, "Dialogue Gate Loudness Correction Flag", "Loudness Correction Type") indicates at least one type of loudness processing that has been applied to the corresponding audio data; and

[0178] At least one loudness value (e.g., one or more parameters in Table 2, such as “ITU relative gated loudness”, “ITU voice gated loudness”, “ITU (EBU 3341) short-term 3s loudness”, and “true peak”) represents at least one loudness (e.g., peak or average loudness) characteristic of the corresponding audio data.

[0179] In some embodiments, the parser 205 (and / or decoder stage 202) is configured to extract various metadata segments having the following formats from the discarded bit segments of the frame of the bit stream, or the "addbsi" column, or the auxdata column:

[0180] Metadata segment header (typically includes a syncword identifying the start of the metadata segment, followed by at least one identification value, such as version, length, and period, expander count, and sub-stream related value); and

[0181] Following the header of the metadata segment, at least one protection value (e.g., the HMAC digest and audio fingerprint value in Table 1) is provided for decrypting, authenticating, or verifying at least one of the metadata of the metadata segment or related audio data; and

[0182] Simultaneously, following the metadata segment header, there is a metadata payload identification (ID) and payload configuration value, which identifies the type of each subsequent metadata payload and at least one similar configuration (e.g., size).

[0183] Each metadata payload section (preferably in the format described above) is followed by the corresponding metadata payload ID and payload configuration value.

[0184] Typically, the encoded audio bitstream generated in a preferred embodiment of the present invention has a structure that provides a mechanism to identify metadata elements and sub-elements as core (mandatory) or extension (optional) elements or sub-elements. This allows the data rate of the bitstream (containing its metadata) to be scaled to various applications. The core (mandatory) of the preferred bitstream syntax should also be able to signal the appearance (in-band) and / or a distant (out-of-band) location of the extension (optional) elements related to the audio content.

[0185] The core element needs to appear in every frame of the bit stream. Some sub-elements of the core element are optional and can appear in any combination. Extension elements do not need to appear in every frame (to limit bit rate load). Therefore, extension elements can appear in some frames but not others. Some sub-elements of the extension element are optional and can appear in any combination, while some sub-elements of the extension element can be mandatory (i.e., if the extension element appears in a frame of the bit stream).

[0186] In a group of embodiments (e.g., an audio processing unit implementing the present invention), an encoded audio bitstream comprising a sequence of audio data segments and metadata segments is generated. The audio data segments represent audio data, and each metadata segment, at least a portion thereof, comprises PIM and / or SSM (and optionally at least another type of metadata), and the audio data segments and metadata segments are time-division multiplexed. In a preferred embodiment of this group, each metadata segment has a preferred format as described herein.

[0187] In a preferred format, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and includes additional bitstream information in the “addbsi” field (as shown in Figure 6) of the Bitstream Stream Information (BSI) field of the frame of the bitstream, or the auxdata field of the frame of the bitstream, or the discarded bit field of the frame of the bitstream, in the various metadata segments containing SSM and / or PIM (e.g., level 107 of a preferred embodiment of encoder 100).

[0188] In a preferred format, each frame contains a metadata segment (sometimes referred to herein as a metadata box, or box) in the frame's discarded bit segment (or addbsi column). The metadata segment has mandatory elements (collectively referred to as "core elements"), as shown in Table 1 below (and may contain optional elements as shown in Table 1). At least a portion of the required elements shown in Table 1 are contained within the metadata segment information of the metadata segment, but some may be contained elsewhere within the metadata segment:

[0189] In a preferred format, each metadata segment (in the discarded bit segment or addbsi or auxdata column of the frame of the encoded bitstream), which contains an SSM, PIM, or LPSM, contains a metadata segment header (and optionally other core elements), and after the metadata segment header (or the metadata segment header and other core elements), one or more metadata payloads. Each metadata payload contains a metadata payload header indicating the specific type of metadata (e.g., SSM, PIM, or LPSM) included in the payload, followed by that specific type of metadata. Typically, the metadata payload header contains the following values ​​(parameters): Payload ID (identifying the metadata type, such as SSM, PIM, or LPSM), follows the metadata segment header (which may contain values ​​specified in Table 1); The payload configuration value following the payload ID (typically representing the payload size); And selected location, additional payload configuration values ​​(e.g., a compensation value indicating the number of audio samples from the start of the frame to the first audio sample to which the payload belongs, and a payload priority value, e.g., indicating a payload that can be abandoned).

[0190] Typically, the metadata carried has one of the following formats: The metadata of the payload is SSM, which includes independent substream metadata, representing the number of independent substreams of the program represented by the bitstream; and related substream metadata, representing whether each independent substream of the program has at least one associated related substream, and if so, the number of related related substreams of each independent substream of the program; The payload metadata is PIM, which includes operating channel metadata, indicating which channels of the audio program contain audio information, and (if any) only silence (typically used for the duration of the frame); downmixing status metadata, indicating whether the program is downmixed (before or during encoding); if so, the type of downmixing applied; upmixing status metadata, indicating whether the program is upmixed (e.g., by a minimum number of channels) before or during encoding; if so, the type of upmixing applied; and preprocessing metadata, indicating whether preprocessing is performed on the audio content of the frame (before encoding the audio content to produce the encoded bitstream); if so, the type of preprocessing performed; or The metadata of the payload is LPSM, with the format indicated in the following table (Table 2):

[0191] In another preferred format of the encoded bitstream generated according to the present invention, the bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each metadata segment containing PIM and / or SSM (and selecting at least another type of metadata) is (e.g., level 107 of a preferred embodiment of encoder 100) contained in any of the following: a discarded bit segment of a frame of the bitstream; or the “addbsi” field of the Bitstream Information (BSI) segment of the frame of the bitstream (as shown in FIG. 6); or the auxdata field at the end of the frame of the bitstream (e.g., the AUX segment shown in FIG. 4). A frame may contain one or two metadata segments, each segment containing PIM and / or SSM, and (in some embodiments), if the frame contains two metadata segments, one may appear in the addbsi field of the frame and the other may appear in the AUX field of the frame. Each metadata segment preferably has the format specified in Table 1 above (i.e., it contains the core element specified in Table 1, followed by a payload ID (identifying the metadata type in each payload of the metadata segment) and a payload configuration value, and each metadata payload). Each metadata segment containing LPSM preferably has the format specified in Tables 1 and 2 above (i.e., it contains the core element specified in Table 1, followed by a payload ID (indicating the metadata is LPSM) and a payload configuration value, followed by a payload (LPSM data, having the format indicated in Table 2)).

[0192] In another preferred format, the encoded bitstream is a Dolby E bitstream, and each metadata segment containing PIM and / or SSM (and other metadata as selected) is the first N sampling positions of the Dolby E guard band spacing. The Dolby E bitstream containing this metadata segment (including LPSM) preferably contains a value indicating the LPSM payload length, which is transmitted in the Pd character of the SMPTE 337M preamble (the SMPTE 337M Pa character repeatability is preferably maintained at the same as the relevant video frame rate).

[0193] In the preferred format of encoding a bitstream as an E-AC-3 bitstream, each metadata segment containing PIM and / or SSM (and optionally LPSM and / or other metadata) is (e.g., level 107 of a preferred embodiment of encoder 100) included as additional bitstream information in the "addbsi" column of the Bitstream Information (BSI) segment of the bitstream frame in a discarded bit segment. Next, additional aspects of encoding an E-AC-3 bitstream are described, including LPSM with the following preferred format:

[0194] 1. When an E-AC-3 bitstream is generated, and the E-AC-3 encoder (which inserts LPSM values ​​into the bitstream) is "operating," for each generated syncframe, the bitstream should contain a metadata block (including LPSM) carried in the addbsi field (or discarded bit section) of that syncframe. These bits that need to carry the metadata block should not increase the encoder bitrate (syncframe length);

[0195] 2. Each metadata block (including LPSM) should contain the following information: Loudness_Correction_Type_Flag: where '1' indicates that the loudness of the corresponding audio data is corrected upstream of the encoder, and '0' indicates that the loudness is corrected by the loudness corrector built into the encoder (e.g., loudness processing level 103 of encoder 100 in Figure 2). Voice_Channel: Indicates which source channels contain voice (exceeding the previous 0.5 seconds). If no voice is detected, this should be as follows: Voice Loudness: Represents the combined voice loudness of each corresponding audio channel, including voice (0.5 seconds beyond the previous duration). ITU Loudness: Represents the integrated ITU BS.1770-3 loudness of each corresponding audio channel; and Gain: In the decoder, the inverse loudness composite gain (exhibiting reversibility);

[0196] 3. Although the E-AC-3 encoder (which inserts LPSM values ​​into the bitstream) is "activated" and receiving AC-3 frames with the "trust" flag, the loudness controller in the encoder (e.g., loudness processing stage 103 of encoder 100 in Figure 2) should be bypassed. The "trust" source dialnorm and DRC value should be transmitted (by the metadata generator 106 of encoder 100) to the E-AC-3 encoder element (e.g., stage 107 of encoder 100). The LPSM block generation duration and loudness_correction_type_ flag are set to '1'. The loudness controller bypass sequence must be synchronized with the start of the decoded AC-3 frame where the "trust" flag appears. The loudness controller bypass sequence should be implemented as follows: during the 10 audio block period (i.e., 53.5 milliseconds), the level_quantity control is decremented from a value of 9 to a value of 0, and the level_post_end_table control is placed in bypass mode (this operation should result in a seamless transition). The term "trust" bypass for the leveler implies that the dialnor value of the source bitstream is also reused at the encoder output. (For example, if the "trust" source bitstream has a dialnor value of -30, then the encoder output should use -30 as the outward dialnor value).

[0197] 4. Although the E-AC-3 encoder (which inserts LPSM values ​​into the bitstream) is "Action" and receiving AC-3 frames without the "Trust" flag, the loudness controller built into the encoder (e.g., loudness processing stage 103 of encoder 100 in Figure 2) should be actuated. The LPSM block generation duration and the loudness_correction_type_ flag are set to '0'. The loudness controller startup sequence should be synchronized to the start of the decoded AC-3 frame where the "Trust" flag disappears. The loudness controller startup sequence should be implemented as follows: during the first audio block (i.e., 5.3 milliseconds), the level_quantity control is incremented from 0 to 9, and the level_back_end_table control is placed in "Action" mode (this operation should result in a seamless transition and include a back_end_table integration reset); and

[0198] 5. During encoding, the graphical user interface (GUI) should display the following parameters to the user: "Input audio program: [Trust / Untrusted]" - the status of this parameter depends on the presence of the "Trust" flag in the input signal; and "Real-time loudness correction: [Enable / Deactivate]" - the status of this parameter depends on whether the loudness controller built into the encoder is activated.

[0199] When decoding an AC-3 or E-AC-3 bitstream containing LPSM (preferably) in the discarded bits or skipped segments of each frame of the bitstream, or in the "addbsi" field of the Bitstream Information (BSI) segment, the decoder should parse the LPSM block data (in the discarded bit segments or addbsi field) and transmit all captured LPSM values ​​to the graphical user interface (GUI). This set of captured LPSM values ​​is regenerated for each frame.

[0200] In another preferred format of the encoded bitstream generated according to the present invention, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each metadata segment containing PIM and / or SSM (and optionally also LPSM and / or other metadata) (e.g., stage 107 of a preferred embodiment of encoder 100) contains additional bitstream information in the "addbsi" column of the Bitstream Information (BSI) segment (as shown in FIG. 6) in the discarded bit segment, or in the AUX segment, or as a frame of the bitstream. In this format (which is a variation of the format described in reference Tables 1 and 2 above), each addbsi (or AUX or discarded bit) column containing LPSM contains the following LPSM values:

[0201] The core components specified in Table 1 are followed by a payload ID (indicating the metadata is LPSM) and a payload configuration value, followed by payload (LPSM data) in the following format (similar to the format of forced components in Table 2 above):

[0202] LPSM payload version: A 2-bit field that specifies the version of the LPSM payload;

[0203] `dialchan`: A 3-bit field indicating whether the left, right, and / or center channels of the corresponding audio data contain voice dialogue. The bit configuration of the `dialchan` field can be as follows: bit 0, indicating dialogue in the left channel, is stored in the most efficient bit of the `dialchan` field; and bit 2, indicating dialogue in the center channel, is stored in the least efficient bit of the `dialchan` field. During the first 0.5 seconds of the program, if the corresponding channel contains dialogue, all bits in the `dialchan` field are set to '1'.

[0204] loudregtyp: A four-bit field indicating which loudness regulation standard the program's loudness complies with. Setting the "loudregtyp" field to "000" indicates that the LPSM does not specify loudness regulation compliance. For example, one value for this field (e.g., 0000) could indicate compliance with an unspecified loudness regulation standard, another value (e.g., 0001) could indicate that the program's audio data complies with the ATSC A / 85 standard, and yet another value (e.g., 0010) could indicate that the program's audio data complies with the EBU R128 standard. In this example, if this field is set to any value other than '0000', then the loudcorrdialgat and loudcorrtyp fields should follow in the payload;

[0205] loudcorrdialgat: This is a bit field indicating whether dialogue gated loudness correction has been applied. If dialogue gated loudness correction has been applied to the program, the value of the loudcorrdialgat field is set to '1'; otherwise, it is set to '0'.

[0206] loudcorrtyp: A single-bit field indicating the type of loudness correction applied to this program. If the program's loudness has already been corrected using an effective look-ahead (file-based) loudness correction procedure, the value of the loudcorrtyp field is set to '0'. If the program's loudness has been corrected using a combination of real-time loudness measurement and dynamic range control, the value of this field is set to '1'.

[0207] loudrelgate: A one-bit field indicating whether relevant loudness data (ITU) is present. If the loudrelgate field is set to '1', then the 7-bit loudrelgate field should follow in the payload;

[0208] loudrelgat: This is a 7-bit field representing the associated gated program loudness (ITU). This field indicates the integrated loudness of the audio program measured according to ITU-R BS.1770-3, without any gain adjustment, due to the application of dial norm and dynamic range compression (DRC). Values ​​from 0 to 127 are interpreted in 0.5 LKFS steps from -58 LKFS to +5.5 LKFS.

[0209] loudspchgate: A one-bit field indicating whether loudspeaker loudness data (ITU) is present. If the loudspchgate field is set to '1', the 7-bit loudspchgate field should follow this value.

[0210] loudspchgat: This is a 7-bit field representing the loudness of the voice-gated program. This field represents the integrated loudness of the entire relevant audio program measured according to Formula (2) of ITU-R BS.1770-3, without any gain adjustment, due to the use of dial norm and dynamic range compression. Values ​​from 0 to 127 are interpreted in 0.5LKFS steps from -58 to +5.5LKFS.

[0211] loudstrm3se: Indicates whether short-term (3-second) loudness data is present in a single-bit column. If this column is set to '1', a 7-bit loudstrm3s column will follow in the payload;

[0212] loudstrm3s: Represents the unblocked loudness of the first 3 seconds of a corresponding audio program measured according to ITU-R BS.1771-1, without any gain adjustment, due to the application of dial norm and dynamic range compression. Values ​​from 0 to 256 are interpreted in 0.5LKFS steps from -116LKFS to +11.5LKFS.

[0213] `truepke`: A one-bit field indicating whether true peak loudness data exists. If the `truepke` field is set to '1', then an 8-bit `truepke` field should follow in the payload; and

[0214] `truepk`: Represents the 8-bit column of the true peak sample value of the program measured according to Annex 2 of ITU-R BS.1770-3, without any gain adjustment, due to the application of dial norm and dynamic range compression. Values ​​from 0 to 256 are interpreted in 0.5LKFS steps from -116LKFS to +11.5LKFS.

[0215] In some embodiments, the core element of a metadata segment in a discarded bit segment or in the auxdata (or "addbsi") field of an AC-3 bit stream or E-AC-3 bit stream frame includes a metadata segment header (typically containing an identification value, such as version), and following the metadata segment header: an indication of whether a fingerprint value (or other protection value) is included in the metadata of that metadata segment; an indication of whether external data (related to audio data corresponding to the metadata of the metadata segment) is present; payload IDs and payload configuration values ​​(e.g., PIM and / or SSM and / or LPSM and / or a type of element) for each type of metadata identified by the core element; and protection values ​​(or other core elements of the metadata segment) for at least one type of metadata identified by the metadata segment header. The metadata payloads of the metadata segment follow the metadata segment header and (in some cases) nest within the core element of the metadata segment.

[0216] Embodiments of the present invention may be implemented as hardware, firmware, or software, or a combination of both (e.g., as a programmable logic array). Unless otherwise specified, algorithms or programs incorporated as part of the present invention are not substantially related to any particular computer or other device. More specifically, various general-purpose machines can use the teachings herein with the written programs, which can more conveniently construct more specific devices (e.g., integrated circuits) to perform the desired method steps. Therefore, the present invention can be implemented in one or more computer programs executing in one or more programmable computer systems (e.g., embodiments of any element of FIG. 1, encoder 100 (or elements thereof) of FIG. 2, decoder 200 (or elements thereof) of FIG. 3, or post-processor 300 (or elements thereof) of FIG. 3, each system comprising at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.

[0217] Each of these programs can be implemented in any desired computer language (including machine, composition, or high-level programming, logic, or object-oriented programming languages) to communicate with a computer system. In any case, the language can be a compiled or interpreted language.

[0218] For example, when computer software instructions are executed in sequence, the various functions and steps of the embodiments of the present invention can be implemented by executing a multi-threaded software instruction sequence on appropriate digital signal processing hardware, wherein the various devices, steps and functions of each embodiment can correspond to parts of the software instructions.

[0219] Each of these computer programs is preferably stored or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., solid-state memory or media, or magnetic or optical media) for configuring or operating the computer to execute the program when the storage medium or device is read by the computer system. The invention can also be implemented as a computer-readable medium configured (i.e., storing) a computer program, wherein the storage medium is configured to cause the computer system to operate in a specific predetermined manner to perform the functions described herein.

[0220] Several embodiments of the present invention have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the invention. Various modifications and variations of the invention are still possible under the above teachings. It is understood that, within the scope of the appended claims, the invention can be practiced in ways other than those specifically described herein.

[0221] 100: Encoder 101: Decoder 102: Audio Status Verifier 103: Loudness Processing Level 104: Audio Stream Selection Level 105: Encoder 106: Metadata Generator 107: Filler / Formatting Level 108: Dialogue Resound Measurement System 109: Frame Buffer 110: Frame Buffer 111: Analyzer 150: Conveying System 152: Decoder 200: Decoder 201: Frame Buffer 202: Audio Decoder 203: Audio Status Verifier 204: Control bit generator 205: Analyzer 300: Post-processor 301: Frame Buffer

Claims

1. An audio processing unit comprising: one or more processors; and a memory coupled to the one or more processors and configured to store a plurality of instructions, which, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising: receiving an encoded audio bitstream comprising an audio program, the encoded audio bitstream comprising encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata comprises dynamic range control (DRC) metadata, loudness metadata, and metadata indicating the number of channels in the set of audio channels, wherein the DRC metadata comprises DRC values ​​and DRC distribution metadata indicating a DRC distribution for generating the DRC values, and wherein the loudness metadata comprises metadata indicating the dialogue loudness of the audio program, wherein the dialogue loudness of the audio program is measured according to ITU-R BS.1770; and decoding the encoded audio data to obtain decoded audio data of the set of audio channels. Obtain the DRC value and the metadata indicating the dialogue loudness of the audio program from the metadata of the encoded audio bitstream; and modify the decoded audio data in the group of audio channels in response to the DRC value and the metadata indicating the dialogue loudness of the audio program, wherein modifying the decoded audio data includes loudness control of the decoded audio data using the dialogue loudness of the audio program.

2. The audio processing unit of claim 1, wherein the encoded audio bit stream includes a metadata box, and wherein the metadata box includes a header and one or more metadata payloads following the header, the one or more metadata payloads including the DRC metadata.

3. A method performed via an audio processing unit, comprising: receiving an encoded audio bitstream containing an audio program, the encoded audio bitstream comprising encoded audio data of a group of one or more audio channels and metadata associated with the group of audio channels, wherein the metadata comprises dynamic range control (DRC) metadata, loudness metadata, and metadata indicating the number of channels in the group of audio channels, wherein the DRC metadata comprises DRC values ​​and DRC distribution metadata indicating a DRC distribution used to generate the DRC values, and wherein the loudness metadata comprises metadata indicating the dialogue loudness of the audio program, wherein the dialogue loudness of the audio program is measured according to ITU-R BS.1770; and decoding the encoded audio data to obtain decoded audio data of the group of audio channels; Obtain the DRC value and the metadata indicating the dialogue loudness of the audio program from the metadata of the encoded audio bitstream; and modify the decoded audio data in the group of audio channels in response to the DRC value and the metadata indicating the dialogue loudness of the audio program, wherein modifying the decoded audio data includes loudness control of the decoded audio data using the dialogue loudness of the audio program.

4. A non-transitory computer-readable storage medium having a plurality of instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform operations comprising: receiving an encoded audio bitstream containing an audio program, the encoded audio bitstream containing encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata includes dynamic range control (DRC) metadata, loudness metadata, and metadata indicating the number of channels in the set of audio channels, wherein the DRC metadata includes DRC values ​​and DRC distribution metadata indicating a DRC distribution used to generate the DRC values, and wherein the loudness metadata includes metadata indicating the dialogue loudness of the audio program, wherein the dialogue loudness of the audio program is measured according to ITU-R BS.1770; and decoding the encoded audio data to obtain decoded audio data of the set of audio channels. Obtain the DRC value and the metadata indicating the dialogue loudness of the audio program from the metadata of the encoded audio bitstream; and modify the decoded audio data in the group of audio channels in response to the DRC value and the metadata indicating the dialogue loudness of the audio program, wherein modifying the decoded audio data includes loudness control of the decoded audio data using the dialogue loudness of the audio program.

Citation Information

Patent Citations

  • Recording medium, reproducing device, recording method, and reproducing method

    US20090097821A1