Audio processing unit, method executed by the audio processing unit, and storage medium
The audio processing unit addresses the issue of inefficient reprocessing by using metadata to adaptively manage volume levels, ensuring consistent quality across devices.
Patent Information
- Application Number
- CN201910831687.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2013-06-19
- Filing Date
- 2013-07-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2034-02-10
AI Technical Summary
Existing audio processing units may unnecessarily perform the already performed processing during blind processing, resulting in the degradation and elimination of the characteristics of the audio data, especially inability to operate effectively when spanning multiple audio processing units in a diverse network.
By receiving an encoded audio bitstream including an audio program, decode the audio data and extracting the loudness processing status metadata, the adaptive loudness processing is performed based on the metadata, ensuring that the adaptive processing of the audio data conforms to the actual state of the audio program.
Adaptive processing of audio data is realized, unnecessary duplicate processing is avoided, audio quality and consistency is improved, and audio data is effectively rendered in different devices and networks.
Smart Images

Figure CN110600043B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of July 31, 2013, the application number of "201310329128.8", and the invention name of "Audio Encoder and Decoder Using Program Information or Substream Structure Metadata". Technical Field
[0002] The present invention relates to audio signal processing, and more particularly, to the encoding and decoding of audio data bitstreams having metadata indicating substream structure and / or program information related to the audio content indicated by the bitstream. Some embodiments of the present invention generate or decode audio data in one of the formats known as Dolby Digital (AC-3), Dolby Digital+ (Enhanced AC-3 or E-AC-3), or Dolby E. Background Art
[0003] Dolby, Dolby Digital, Dolby Digital+, and Dolby E are trademarks of Dolby Laboratories Licensing Corporation. Dolby Laboratories provides proprietary implementations of AC-3 and E-AC-3 known as Dolby Digital and Dolby Digital+ respectively.
[0004] Audio data processing units typically operate in a blind fashion and do not concern themselves with the processing history of the audio data that occurred prior to the data being received. This can work in a processing framework where a single entity performs all of the audio data processing and encoding for various target media rendering devices and the target media rendering devices perform all of the decoding and rendering of the encoded audio data. However, this blind processing does not work well (or not at all) in situations where multiple audio processing units are scattered or chained across a diverse network and are expected to perform their respective types of audio processing optimally. For example, some audio data may be encoded for a high-performance media system and may need to be converted into a simplified form suitable for a mobile device along the media processing chain. Thus, an audio processing unit may unnecessarily perform types of processing on the audio data that have already been performed. For example, a volume leveling unit may perform processing on an input audio segment regardless of whether the same or similar volume leveling has been performed on the input audio segment previously. Thus, the volume leveling unit may perform leveling even when it is not necessary. This unnecessary processing may also result in the degradation and / or elimination of specific features when rendering the content of the audio data. Summary of the Invention
[0005] The present invention discloses an audio processing unit, comprising: one or more processors; a memory coupled to the one or more processors and configured to store instructions which, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata includes loudness processing status metadata, and wherein the loudness processing status metadata includes metadata indicating the loudness of the audio program; decoding the encoded audio data to obtain decoded audio data of the set of audio channels; obtaining the loudness processing status metadata from the metadata of the encoded audio bitstream; and performing adaptive loudness processing on the decoded audio data of the set of audio channels based on the loudness processing status metadata.
[0006] The present invention also discloses a method executed by an audio processing unit, comprising: receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata includes loudness processing status metadata, and wherein the loudness processing status metadata includes metadata indicating the loudness of the audio program; decoding the encoded audio data to obtain decoded audio data of the set of audio channels; obtaining the loudness processing status metadata from the metadata of the encoded audio bitstream; and performing adaptive loudness processing on the decoded audio data of the set of audio channels based on the loudness processing status metadata.
[0007] The present invention further discloses a non-transitory computer-readable storage medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata includes loudness processing status metadata, and wherein the loudness processing status metadata includes metadata indicating the loudness of the audio program; decoding the encoded audio data to obtain decoded audio data of the set of audio channels; obtaining the loudness processing status metadata from the metadata of the encoded audio bitstream; and performing adaptive loudness processing on the decoded audio data of the set of audio channels based on the loudness processing status metadata.
[0008] In one class of embodiments, the present invention is an audio processing unit capable of decoding an encoded bitstream that includes substream structure metadata and / or program information metadata (optionally also including other metadata, such as loudness processing status metadata) in at least one segment of at least one frame of the bitstream and audio data in at least one other segment of the frame. Herein, substream structure metadata (or "SSM") represents metadata of an encoded bitstream (or a set of encoded bitstreams) that indicates the substream structure of the audio content of the encoded bitstream, and "program information metadata" (or "PIM") represents metadata of an encoded audio bitstream that indicates at least one audio program (e.g., two or more audio programs), where the program information metadata indicates at least one attribute or characteristic of the audio content of at least one of the programs (e.g., metadata indicating the type or parameters of processing performed on the audio data of the program, or metadata indicating which channels of the program are active channels).
[0009] In a typical case (e.g., where the encoded bitstream is an AC-3 or E-AC-3 bitstream), the program information metadata (PIM) indicates program information that cannot actually be carried in other parts of the bitstream. For example, the PIM can indicate the processing applied to PCM audio before encoding (e.g., AC-3 or E-AC-3 encoding), which frequency bands of the audio program have been encoded using a specific audio coding technique, and the compression profile used to create dynamic range compression (DRC) data in the bitstream.
[0010] In another class of embodiments, the method includes the step of multiplexing encoded audio data with SSM and / or PIM in each frame (or each frame in at least some frames) of the bitstream. In typical decoding, the decoder extracts SSM and / or PIM from the bitstream (including by analyzing and demultiplexing the SSM and / or PIM and the audio data), and processes the audio data to generate a stream of decoded audio data (and in some cases also performs adaptive processing of the audio data). In some embodiments, the decoded audio data and the SSM and / or PIM are forwarded from the decoder to a post-processor that is configured to perform adaptive processing on the decoded audio data using the SSM and / or PIM.
[0011] In one class of embodiments, the encoding method of the present invention generates an audio data segment (e.g., Figure 4 segments AB0 to AB5 of the frame shown or Figure 7an encoded audio bitstream (e.g., an AC-3 or E-AC-3 bitstream) of all or some of segments AB0 to AB5 of the illustrated frame, the audio data segment including encoded audio data and a metadata segment time-division multiplexed with the audio data segment (including SSM and / or PIM, optionally also including other metadata). In some embodiments, each metadata segment (sometimes referred to herein as a "container") has a metadata segment header (optionally also including other mandatory or "core" elements), and one or more metadata payloads following the metadata segment header. If present, the SIM is included in one of the metadata payloads (identified by a payload header and typically having a first type of format). If present, the PIM is included in another of the metadata payloads (identified by a payload header and typically having a second type of format). Similarly, each other type of metadata (if present) is included in another of the metadata payloads (identified by a payload header and typically having a format specific to the type of metadata). The exemplary format allows for convenient access to SSM, PIM, or other metadata at times other than during decoding of the bitstream (e.g., by a post-processor after decoding, or by a processor configured to identify metadata without performing a full decoding of the encoded bitstream), and allows for convenient and efficient error detection and correction during decoding of the bitstream (e.g., for substream identification). For example, without accessing the SSM in the exemplary format, a decoder may incorrectly identify the correct number of substreams associated with a program. One metadata payload in the metadata segment may include SSM, another metadata payload in the metadata segment may include PIM, and optionally, at least one other metadata payload in the metadata segment may include other metadata (e.g., loudness processing status metadata or "LPSM"). BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram of an embodiment of a system that can be configured to perform an embodiment of the method of the present invention.
[0013] Figure 2 is a block diagram of an encoder that is an embodiment of an audio processing unit of the present invention.
[0014] Figure 3 is a block diagram of a decoder that is an embodiment of an audio processing unit of the present invention and a post-processor coupled to the decoder that is another embodiment of an audio processing unit of the present invention.
[0015] Figure 4 is a diagram of an AC-3 frame including segments into which it is divided.
[0016] Figure 5A diagram of the synchronization information (SI) segment of an AC-3 frame including segments into which it is divided.
[0017] Figure 6 A diagram of the bitstream information (BSI) segment of an AC-3 frame including segments into which it is divided.
[0018] Figure 7 A diagram of an E-AC-3 frame including segments into which it is divided.
[0019] Figure 8 A diagram of the metadata segment of an encoded bitstream including a metadata segment header generated according to an embodiment of the present invention, the metadata segment header including a container sync word (identified as "container sync" in Figure 8 ), as well as a version and key ID value, followed by a plurality of metadata payloads and a protection bit.
[0020] Symbols and terms
[0021] Throughout this disclosure including the claims, the expression of performing an operation on a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used in a broad sense to mean directly performing an operation on a signal or data, or on a processed version of the signal or data (e.g., a version of a signal that has undergone preliminary filtering or preprocessing before an operation is performed on the signal).
[0022] Throughout this disclosure including the claims, the expression of "system" is used in a broad sense to mean a device, system, or subsystem. For example, a subsystem that implements a decoder can be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to a plurality of inputs, where the subsystem generates M inputs and the other X - M inputs are received from an external source) can also be referred to as a decoder system.
[0023] Throughout this disclosure including the claims, the term "processor" is used in a broad sense to mean a system or device that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (e.g., audio data, video data, or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chip sets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio data or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chip sets.
[0024] Throughout the present disclosure, including the claims, the expressions "audio processor" and "audio processing unit" are used interchangeably to broadly denote a system configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, preprocessing systems, postprocessing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools).
[0025] Throughout the present disclosure, including the claims, the expression "(of an encoded audio bitstream) 'metadata'" refers to data that is separate and distinct from the corresponding audio data of the bitstream.
[0026] Throughout the present disclosure, including the claims, the expression "substream structure metadata" (or "SSM") denotes metadata of an encoded audio bitstream (or set of encoded audio bitstreams) that indicates the substream structure of the audio content of the encoded bitstream.
[0027] Throughout the present disclosure, including the claims, the expression "program information metadata" (or "PIM") denotes metadata of an encoded audio bitstream that indicates at least one audio program (e.g., two or more audio programs), where the metadata indicates at least one attribute or characteristic of the audio content of at least one of the programs (e.g., metadata indicating the type or parameters of processing performed on the audio data of the program, or metadata indicating which channels of the program are active channels).
[0028] Throughout this disclosure, including the claims, the expression "processing status metadata" (e.g., as in the expression "loudness processing status metadata") refers to metadata associated with the audio data of a bitstream (encoding an audio bitstream), indicating the processing status of the corresponding (associated) audio data (e.g., what type of processing has been performed on the audio data), and generally also indicating at least one feature or characteristic of the audio data. The association of the processing status metadata with the audio data is time-synchronized. Thus, the current (most recently received or updated) processing status metadata indicates that the corresponding audio data simultaneously includes the result of the indicated type of audio data processing. In some cases, the processing status metadata may include some or all of the processing history and / or parameters used in and / or obtained from the indicated type of processing. Additionally, the processing status metadata may include at least one feature or characteristic that has been calculated or extracted from the corresponding audio data. The processing status metadata may also include other metadata that is not related to or obtained from any processing of the corresponding audio data. For example, third-party data, tracking information, identifiers, ownership or standard information, user annotation data, user preference data, etc. may be added by a particular audio processing unit for transfer to other audio processing units.
[0029] Throughout this disclosure, including the claims, the expression "loudness processing status metadata" (or "LPSM") represents processing status metadata that indicates the loudness processing status of the corresponding audio data (e.g., what type of loudness processing has been performed on the audio data), and generally also indicates at least one feature or characteristic of the corresponding audio data (e.g., loudness). The loudness processing status metadata may include data that is not (i.e., when considered alone) loudness processing status metadata (e.g., other metadata).
[0030] Throughout this disclosure, including the claims, the expression "channel" (or "audio channel") represents a single-channel audio signal.
[0031] Throughout this disclosure, including the claims, the expression "audio program" represents a collection of one or more audio channels and optionally also represents associated metadata (e.g., metadata describing a desired spatial audio representation, and / or PIM, and / or SSM, and / or LPSM, and / or program boundary metadata).
[0032] Throughout this disclosure, including the claims, the expression "program boundary metadata" refers to metadata of an encoded audio bitstream, where the encoded audio bitstream indicates at least one audio program (e.g., two or more programs), and the program boundary metadata indicates the position in the bitstream of at least one boundary (start and / or end) of at least one of the audio programs. For example, the program boundary metadata (of the encoded audio bitstream indicating the audio program) may include metadata indicating the position of the start of the program (e.g., the start of the "N"th frame of the bitstream, or the "M"th sample position in the "N"th frame of the bitstream), and additional metadata indicating the position of the end of the program (e.g., the start of the "J"th frame of the bitstream, or the "K"th sample position in the "J"th frame of the bitstream).
[0033] Throughout this disclosure, including the claims, the terms "coupled" or "is coupled" are used to mean a direct or indirect connection. Thus, if a first device is coupled to a second device, that connection may be by a direct connection, or by an indirect connection via other devices and connections. Detailed Description
[0034] A typical audio data stream includes both audio content (e.g., one or more channels of audio content) and metadata indicating at least one characteristic of the audio content. For example, in an AC-3 bitstream, there are several audio metadata parameters specifically intended to alter the sound of the program being delivered to the listening environment. One of the metadata parameters is the DIALNORM parameter, which is intended to indicate the average level of dialogue in an audio program and is used to determine the audio playback signal level.
[0035] During playback of a bitstream that includes a series of different audio program segments (each having a different DIALNORM parameter), the AC-3 decoder performs a type of loudness processing using the DIALNORM parameter of each segment, in which the AC-3 decoder modifies the playback level or loudness such that the perceived loudness of the dialogue of the series of segments is at a consistent level. Each encoded audio segment (item) in a series of encoded audio items will (typically) have a different DIALNORM parameter, and the decoder will scale the level of each item in the series such that the playback level or loudness of the dialogue of each item is the same or very similar, although this will require applying different amounts of gain to different items in the series during playback.
[0036] DIALNORM is typically set by the user rather than being automatically generated. However, there is a default DIALNORM value if the user does not set a value. For example, a content creator can use a device external to the AC-3 encoder to measure loudness and then transmit the result (indicating the loudness of the spoken dialogue of an audio program) to the encoder to set the DIALNORM value. Thus, it is dependent on the content creator to set the DIALNORM parameter correctly.
[0037] There are several different reasons why the DIALNORM parameter in an AC-3 bitstream can be incorrect. First, if the DIALNORM value is not set by the content creator, then each AC-3 encoder has a default DIALNORM value that is used during the generation of the bitstream. This default value may be significantly different from the actual dialogue loudness of the audio. Second, even if the content creator measures the loudness and sets the DIALNORM value accordingly, a loudness measurement algorithm or meter that does not conform to the recommended AC-3 loudness measurement method may have been used, resulting in an incorrect DIALNORM value. Third, even if an AC-3 bitstream has been created using a DIALNORM value that has been correctly measured and set by the content creator, the AC-3 bitstream may have been altered to an incorrect value during the transmission and / or storage of the bitstream. For example, this is not uncommon in television broadcast applications where an AC-3 bitstream is decoded, modified, and then re-encoded using incorrect DIALNORM metadata information. Thus, the DIALNORM value included in the AC-3 bitstream may be incorrect or inaccurate and may therefore have a negative impact on the quality of the listening experience.
[0038] In addition, the DIALNORM parameter does not indicate the loudness processing state of the corresponding audio data (e.g., what type of loudness processing has been performed on the audio data). Loudness processing state metadata (in the format in which it is provided in some embodiments of the present invention) facilitates the adaptive loudness processing of audio bitstreams and / or the verification of the loudness processing state and the validity of the loudness of audio content in a particularly efficient manner.
[0039] Although the present invention is not limited to using AC-3 bitstreams, E-AC-3 bitstreams, or Dolby E bitstreams, for convenience, it will be described in embodiments that generate, decode, or otherwise process such bitstreams.
[0040] An AC-3 encoded bitstream includes metadata and 1 to 6 channels of audio content. The audio content is audio data that has been compressed using perceptual audio coding. The metadata includes several audio metadata parameters that are intended to be used to alter the sound of the program being transmitted to the listening environment.
[0041] Each frame of an AC-3 encoded audio bitstream contains audio content and metadata for 1536 samples of digital audio. For a sampling rate of 48 kHz, this represents 32 milliseconds of digital audio or a rate of 31.25 frames per second of audio.
[0042] Depending on whether the frame contains 1, 2, 3, or 6 blocks of audio data respectively, each frame of an E-AC-3 encoded audio bitstream contains audio data and metadata for 256, 512, 768, or 1536 samples of digital audio. For a sampling rate of 48 kHz, this represents 5.333, 10.667, 16, or 32 milliseconds of digital audio respectively or rates of 189.9, 93.75, 62.5, or 31.25 frames per second of audio respectively.
[0043] As Figure 4 shown, each AC-3 frame is divided into parts (segments), including: a synchronization information (SI) part containing (as Figure 5 shown) a synchronization word (SW) and the first error correction word (CRC1) of two error correction words; a bitstream information (BSI) part containing most of the metadata; six audio blocks (AB0 to AB5) containing data-compressed audio content (and may also include metadata); a waste bits segment (W) (also known as the "skip field") containing any unused bits remaining after the compressed audio content; an auxiliary (AUX) information part that may contain more metadata; and the second error correction word (CRC2) of the two error correction words.
[0044] As Figure 7 shown, each E-AC-3 frame is divided into parts (segments), including: a synchronization information (SI) part containing (as Figure 5 shown) a synchronization word (SW); a bitstream information (BSI) part containing most of the metadata; six audio blocks (AB0 to AB5) containing data-compressed audio content (and may also include metadata); a waste bits segment (W) (also known as the "skip field") (although only one waste bits segment is shown, different waste bits or skip field segments can typically be after each audio block) containing any unused bits remaining after the compressed audio content; an auxiliary (AUX) information part that may contain more metadata; and an error correction word (CRC).
[0045] In an AC-3 (or E-AC-3) bitstream, there are several audio metadata parameters specifically intended to be used to alter the sound of the program being transmitted to the listening environment. One of the metadata parameters is the DIALNORM parameter, which is included in the BSI segment.
[0046] As Figure 6As shown, the BSI section of an AC-3 frame includes a 5-bit parameter ("DIALNORM") indicating the DIALNORM value of the program. If the audio coding mode ("acmod") of the AC-3 frame is 0, it includes a 5-bit parameter ("DIALNORM2") indicating the DIALNORM value of the 5-bit parameter of the second audio program carried in the same AC-3 frame, indicating the use of a dual mono-channel or "1+1" channel configuration.
[0047] The BSI section also includes a flag ("addbsie") indicating the presence (or absence) of additional bitstream information after the "addbsie" bit, a parameter ("addbsil") indicating the length of any additional bitstream information after the "addbsil" value, and up to 64 bits of additional bitstream information ("addbsi") after the "addbsil" value.
[0048] The BSI section includes other metadata values not specifically shown in Figure 6 .
[0049] According to one class of embodiments, the encoded bitstream indicates multiple substreams of audio content. In some cases, the substreams indicate the audio content of a multi-channel program, and each in the substream indicates one or more of the channels of the program. In other cases, multiple substreams of the encoded audio bitstream indicate the audio content of several audio programs - typically a "main" audio program (which can be a multi-channel program) and at least one other audio program (e.g., a program for commentary on the main audio program).
[0050] The encoded audio bitstream indicating at least one audio program needs to include at least one "independent" substream of the audio content. The independent substream indicates at least one channel of the audio program (e.g., the independent substream can indicate 5 full-range channels of a conventional 5.1-channel audio program). Herein, this audio program is referred to as the "main" program.
[0051] In some types of embodiments, the encoded audio bitstream indicates two or more audio programs (a "main" program and at least one other audio program). In such a case, the bitstream includes two or more independent substreams: a first independent substream indicating at least one channel of the main program; and at least one other independent substream indicating at least one channel of another audio program (a program different from the main program). Each independent substream can be decoded independently, and the decoder can operate to decode only a subset (not all) of the independent substreams of the encoded bitstream.
[0052] In a typical example of an encoded audio bitstream indicating two independent sub-streams, one of the independent sub-streams indicates the standard format speaker channels of a multi-channel main program (e.g., the left, right, center, left surround, and right surround full-range speaker channels of a 5.1-channel main program), while the other independent sub-stream indicates a mono-channel audio commentary regarding the main program (e.g., a director's commentary on a movie, where the main program is the soundtrack of the movie). In another example of an encoded audio bitstream indicating multiple independent sub-streams, one of the independent sub-streams indicates the standard format speaker channels of a multi-channel main program (e.g., a 5.1-channel main program) including dialogue in a first language (e.g., one of the speaker channels of the main program can indicate the dialogue), while each of the other independent sub-streams indicates a mono-channel translation of the dialogue (into a different language).
[0053] Optionally, an encoded audio bitstream indicating a main program (optionally also indicating at least one other audio program) includes at least one "dependent" sub-stream of audio content. Each dependent sub-stream is associated with an independent sub-stream of the bitstream and indicates at least one additional channel of the program (e.g., the main program) whose content is indicated by the associated independent sub-stream (i.e., the dependent sub-stream indicates at least one channel of the program that is not indicated by the associated independent sub-stream, while the associated independent sub-stream indicates at least one channel of the program).
[0054] In an example of an encoded bitstream including an independent sub-stream (indicating at least one channel of a main program), the bitstream also includes a dependent sub-stream (associated with the independent sub-stream) indicating one or more additional speaker channels of the main program. Such additional speaker channels are additional to the main program channels indicated by the independent sub-stream. For example, if the independent sub-stream indicates the left, right, center, left surround, and right surround full-range speaker channels of a 7.1-channel main program, then the dependent sub-stream can indicate the other two full-range speaker channels of the main program.
[0055] According to the E-AC-3 standard, an E-AC-3 bitstream must indicate at least one independent sub-stream (e.g., a single AC-3 bitstream) and can indicate up to 8 independent sub-streams. Each independent sub-stream of an E-AC-3 bitstream can be associated with up to 8 dependent sub-streams.
[0056] The E-AC-3 bitstream includes metadata indicating the substream structure of the bitstream. For example, the "chanmap" field in the bitstream information (BSI) portion of the E-AC-3 bitstream determines the channel mapping of the program channels indicated by the dependent substreams of the bitstream. However, the metadata indicating the substream structure is conventionally included in the E-AC-3 bitstream in a format that is convenient for access and use only by the E-AC-3 decoder (during decoding of the encoded E-AC-3 bitstream); it is not convenient for access and use after decoding (e.g., by a post-processor) or before decoding (e.g., by a processor configured to identify the metadata). Moreover, there is a risk that the decoder may incorrectly identify the substreams of a conventional E-AC-3 encoded bitstream using the conventionally included metadata, and prior to the present invention, it was not known how to include substream structure metadata in an encoded bitstream (e.g., an encoded E-AC-3 bitstream) in a format that allows for convenient and efficient detection and correction of errors in substream identification during decoding of the bitstream.
[0057] The E-AC-3 bitstream may also include metadata regarding the audio content of the audio program. For example, an E-AC-3 bitstream indicating an audio program includes metadata indicating the minimum frequency and maximum frequency at which spectral extension processing (and optionally also channel coupling coding) has been used to encode the content of the program. However, such metadata is typically included in the E-AC-3 bitstream in a format that is convenient for access and use only by the E-AC-3 decoder (during decoding of the encoded E-AC-3 bitstream); it is not convenient for access and use after decoding (e.g., by a post-processor) or before decoding (e.g., by a processor configured to identify the metadata). Moreover, such metadata is not included in the E-AC-3 bitstream in a format that allows for convenient and efficient error detection and error correction of the identification of such metadata during decoding of the bitstream.
[0058] According to typical embodiments of the present invention, PIM and / or SSM (and optionally also other metadata, such as loudness processing status metadata or "LPSM") are embedded in one or more reserved fields (or slots) of the metadata segment of an audio bitstream, which also includes audio data in other segments (audio data segments). Typically, at least one segment of each frame of the bitstream includes PIM or SSM, and at least one other segment of the frame includes the corresponding audio data (i.e., the audio data whose data structure is indicated by the SSM and / or at least one of whose characteristics or attributes is indicated by the PIM).
[0059] In one class of embodiments, each metadata segment is a data structure (sometimes referred to herein as a container) that can contain one or more metadata payloads. Each payload includes a header to provide an explicit indication of the type of metadata present in the payload, where the header includes a specific payload identifier (or payload configuration data). The order of the payloads within the container is not defined, such that the payloads can be stored in any order and an analyzer must be able to analyze the entire container to extract relevant payloads while ignoring irrelevant or unsupported payloads. Figure 8 Illustrates the structure of such a container and the payloads within the container (to be described below).
[0060] Communication metadata (e.g., SSM and / or PIM and / or LPSM) in an audio data processing chain is particularly useful when two or more audio processing units need to cooperate with each other throughout the processing chain (or content life cycle). In the case where metadata is not included in the audio bitstream, for example, when two or more audio codecs are utilized in the chain and single-ended volume is applied more than once during the bitstream path of a media consumption device (or the rendering point of the audio content of the bitstream), several media processing problems can occur, such as quality, level, and spatial degradation.
[0061] According to some embodiments of the present invention, loudness processing state metadata (LPSM) embedded in an audio bitstream can be authenticated and verified, for example, to enable a loudness adjustment entity to prove whether the loudness of a particular program is within a specified range and whether the corresponding audio data itself has not been modified (thereby ensuring compliance with applicable regulations). The loudness value included in a data block including the loudness processing state metadata can be read out to verify this without recalculating the loudness. In response to the LPSM, a management structure can determine that the corresponding audio content complies with loudness statutory and / or regulatory requirements (e.g., rules promulgated under the Commercial Advertisement Loudness Mitigation Act, also known as the "CALM" Act) as indicated by the LPSM without the need to calculate the loudness of the audio content.
[0062] Figure 1 Is a block diagram of an exemplary audio processing chain (audio data processing system) in which one or more of the elements of the system can be configured according to embodiments of the present invention. The system includes the following elements coupled together as shown: a preprocessing unit, an encoder, a signal analysis and metadata correction unit, a transcoder, a decoder, and a postprocessing unit. In a variant of the system shown, one or more of the elements are omitted, or additional audio data processing units are included.
[0063] In some implementations, Figure 1The preprocessing unit is configured to receive PCM (time domain) samples including audio content as input and output processed PCM samples. The encoder can be configured to receive the PCM samples as input and output an encoded (e.g., compressed) audio bitstream indicative of the audio content. The data of the bitstream indicative of the audio content is sometimes referred to herein as "audio data". If the encoder is configured according to an exemplary embodiment of the present invention, the audio bitstream output from the encoder includes PIM and / or SSM (optionally also including loudness processing status metadata and / or other metadata) as well as the audio data.
[0064] Figure 1 The signal analysis and metadata correction unit can receive one or more encoded audio bitstreams as input and determine (e.g., verify) whether the metadata (e.g., processing status metadata) in each encoded audio bitstream is correct by performing signal analysis (e.g., using the program boundary metadata in the encoded audio bitstream). If the signal analysis and metadata correction unit finds that the included metadata is invalid, the incorrect value is typically replaced with the correct value obtained from the signal analysis. Thus, each encoded audio bitstream output from the signal analysis and metadata correction unit can include the corrected (or uncorrected) processing status metadata as well as the encoded audio data.
[0065] Figure 1 The transcoder can receive an encoded audio bitstream as input and output a modified (e.g., differently encoded) audio bitstream in response (e.g., by decoding the input stream and re-encoding the decoded stream in a different encoding format). If the transcoder is configured according to an exemplary embodiment of the present invention, the audio bitstream output from the transcoder includes SSM and / or PIM (usually also including other metadata) as well as the encoded audio data. The metadata may already be included in the input bitstream.
[0066] Figure 1 The decoder can receive an encoded (e.g., compressed) audio bitstream as input and output (in response) a decoded PCM audio sample stream. If the decoder is configured according to an exemplary embodiment of the present invention, then in typical operation, the output of the decoder is or includes any one of the following:
[0067] An audio sample stream, and at least one corresponding stream of SSM and / or PIM (usually also other metadata) extracted from the input encoded bitstream; or
[0068] An audio sample stream, and a corresponding stream of control bits determined based on SSM and / or PIM (usually also other metadata, such as LPSM) extracted from the input encoded bitstream; or
[0069] An audio sample stream, but no corresponding stream of metadata or control bits determined from the metadata. In the last case, the decoder can extract the metadata from the input encoded bitstream and perform at least one operation (e.g., verification) on the extracted metadata, even if the extracted metadata or control bits determined from the metadata are not output.
[0070] By configuring according to an exemplary embodiment of the present invention Figure 1 a post - processing unit, which is configured to receive a decoded PCM audio sample stream and perform post - processing (e.g., volume leveling of the audio content) on it using the SSM and / or PIM received with the samples (usually other metadata such as LPSM as well), or control bits determined from the metadata received with the samples. The post - processing unit is also typically configured to render the post - processed audio content for playback by one or more speakers.
[0071] Exemplary embodiments of the present invention provide an enhanced audio processing chain, where audio processing units (e.g., encoders, decoders, transcoders, and pre - processing units and post - processing units) modify their respective processing to be applied to audio data according to the contemporaneous state of the media data indicated by metadata received separately by the audio processing units.
[0072] Input to Figure 1 any audio processing unit of the system (e.g., Figure 1 the encoder or transcoder of Figure 1 the system) can include SSM and / or PIM (optionally other metadata as well) and audio data (e.g., encoded audio data). The metadata can have been included in the input audio by Figure 1 another element of the system (or another source, not shown in
[0073] A typical implementation of the audio processing unit (or audio processor) of the present invention is configured to perform adaptive processing of audio data based on the status of the audio data indicated by metadata corresponding to the audio data. In some implementations, the adaptive processing is (or includes) loudness processing (if the metadata indicates that loudness processing or a process similar to loudness processing has not been performed on the audio data), and not (and does not include) loudness processing (if the metadata indicates that such loudness processing or a process similar to loudness processing has been performed on the audio data). In some implementations, the adaptive processing is or includes (e.g., performed in the metadata verification subunit) metadata verification to ensure that the audio processing unit performs other adaptive processing of the audio data based on the status of the audio data indicated by the metadata. In some implementations, this verification determines the reliability of the metadata associated with the audio data (e.g., included in a bitstream having the audio data). For example, if the verified metadata is reliable, then the results from a previously performed audio processing can be reused and new execution of the same type of audio processing can be avoided. On the other hand, if it is found that the metadata has been tampered with (or is otherwise unreliable), then a type of media processing that was allegedly previously performed (as indicated by the unreliable metadata) can be repeated by the audio processing unit, and / or other processing can be performed on the metadata and / or the audio data by the audio processing unit. If the unit determines that the metadata is valid (e.g., based on a match between an extracted encryption value and a reference encryption value), the audio processing unit can also be configured to signal to other audio processing units downstream in the enhanced media processing chain that the metadata (e.g., present in the media bitstream) is valid.
[0074] Figure 2 FIG. 4 is a block diagram of an encoder (100) as an implementation of the audio processing unit of the present invention. Any component or element of encoder 100 can be implemented as one or more processors and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware or software or a combination of hardware and software. Encoder 100 includes a frame buffer 110, an analyzer 111, a decoder 101, an audio status validator 102, a loudness processing stage 103, an audio stream selection stage 104, an encoder 105, a filler / formatter stage 107, a metadata generation stage 106, a dialogue loudness measurement subsystem 108, and a frame buffer 109, connected as shown. Encoder 100 generally also includes other processing elements (not shown).
[0075] An encoder 100 (which is a code converter) is configured to convert an input audio bitstream (e.g., which can be one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream) into an encoded output audio bitstream (e.g., which can be another one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream) by performing adaptive and automatic loudness processing by using loudness processing status metadata included in the input bitstream. For example, the encoder 100 can be configured to convert an input Dolby E bitstream (a format typically used in production and broadcast equipment but not in consumer equipment that receives audio programs that have been broadcast) into an encoded output audio bitstream in AC-3 or E-AC-3 format (suitable for broadcasting to consumer equipment).
[0076] Figure 2 The system further includes an encoded audio delivery subsystem 150 (which stores and / or delivers the encoded bitstream output from the encoder 100) and a decoder 152. The encoded audio bitstream output from the encoder 100 can be stored by the subsystem 150 (e.g., in DVD or Blu-ray Disc format), or transmitted by the subsystem 150 (which can implement a transmission line or network), or can be stored and transmitted by the subsystem 150. The decoder 152 is configured to decode the encoded audio bitstream (generated by the encoder 100) received via the subsystem 150 by extracting metadata (PIM and / or SSM, and optionally also loudness processing status metadata and / or other metadata) from each frame of the bitstream (and optionally also extracting program boundary metadata from the bitstream) and generating decoded audio data. Generally, the decoder 152 is configured to perform adaptive processing on the decoded audio data by using PIM and / or SSM and / or LPSM (optionally also using program boundary metadata), and / or forwarding the decoded audio data and metadata to a post-processor configured to perform adaptive processing on the decoded audio data by using the metadata. Generally, the decoder 152 includes a buffer that stores (e.g., in a non-transitory manner) the encoded audio bitstream received from the subsystem 150.
[0077] Various implementations of the encoder 100 and the decoder 152 are configured to perform different embodiments of the method of the present invention.
[0078] The frame buffer 110 is a buffer memory coupled to receive the encoded input audio bitstream. In operation, the buffer 110 stores (e.g., in a non-transitory manner) at least one frame of the encoded audio bitstream, and a sequence of frames of the encoded audio bitstream is set from the buffer 110 to the analyzer 111.
[0079] Couple the analyzer 111 and configure it to extract PIM and / or SSM, as well as loudness processing status metadata (LPSM), and optionally also program boundary metadata (and / or other metadata) from each frame of the encoded input audio including such metadata, and set at least the LPSM (and optionally also program boundary metadata and / or other metadata) to the audio state validator 102, the loudness processing stage 103, stage 106, and subsystem 108 to extract audio data from the encoded input audio and set the audio data to the decoder 101. The decoder 101 of the encoder 100 is configured to decode the audio data to generate decoded audio data and set the decoded audio data to the loudness processing stage 103, the audio stream selection stage 104, the subsystem 108, and generally also to the state validator 102.
[0080] The state validator 102 is configured to authenticate and verify the LPSM (optionally other metadata) set to it. In some embodiments, the LPSM is (or is included in) a data block that has been included in the input bitstream (e.g., according to an embodiment of the present invention). The block may include an encrypted hash (a hash-based message authentication code or “HMAC”) for processing the LPSM (optionally also other metadata) and / or the underlying audio data (provided from the decoder 101 to the validator 102). In these embodiments, the data block may be digitally marked such that downstream audio processing units can authenticate and verify the processing status metadata relatively easily.
[0081] For example, the HMAC is used to generate a digest, and the protection value included in the bitstream of the present invention may include this digest. The digest may be generated for an AC-3 frame as follows:
[0082] 1. After the AC-3 data and LPSM are encoded, the frame data bytes (connected frame data #1 and frame data #2) and the LPSM data bytes are used as inputs to the hash function HMAC. Other data that may be present in the auxiliary data field is not considered for calculating the digest. Such other data may be bytes that belong neither to the AC-3 data nor to the LPSM data. The protection bits included in the LPSM may not be considered for calculating the HMAC digest.
[0083] 2. After calculating the digest, it is written into the field reserved for the protection bits in the bitstream.
[0084] 3. The final step in generating the complete AC-3 frame is the calculation of the CRC checksum. This is written at the end of the frame and takes into account all the data belonging to the frame, including the LPSM bits.
[0085] Other encryption methods, including but not limited to any one of one or more non-HMAC encryption methods, can be used for the verification of LPSM and / or other metadata (e.g., in the validator 102) to ensure the secure transmission and reception of metadata and / or base audio data. For example, verification (using such an encryption method) can be performed in each audio processing unit of an embodiment that receives the audio bitstream of the present invention to determine whether the metadata included in the bitstream and the corresponding audio data have undergone (and / or have been generated) specific processing (indicated by the metadata) and have not been modified after such specific processing is performed.
[0086] The status validator 102 sets control data to the audio stream selection stage 104, the metadata generator 106, and the dialogue loudness measurement subsystem 108 to represent the result of the verification operation. In response to the control data, the stage 104 can select (and pass to the encoder 105):
[0087] The adaptively processed output of the loudness processing stage 103 (e.g., when the LPSM indicates that the audio data output from the decoder 101 has not undergone a specific type of loudness processing and the control bit from the validator 102 indicates that the LPSM is valid); or
[0088] The audio data output from the decoder 102 (e.g., when the LPSM indicates that the audio data output from the decoder 101 has undergone a specific type of loudness processing to be performed by the stage 103 and the control bit from the validator 102 indicates that the LPSM is valid).
[0089] The stage 103 of the encoder 100 is configured to perform adaptive loudness processing on the decoded audio data output from the decoder 101 based on one or more audio data characteristics indicated by the LPSM extracted by the decoder 101. The stage 103 can be an adaptive transform domain real-time loudness and dynamic range control processor. The stage 103 can receive user input (e.g., user target loudness / dynamic range value or dialogue normalization value), or other metadata input (e.g., one or more types of third-party data, track information, identifiers, ownership or standard information, user annotation data, user preference data, etc.) and / or other input (e.g., from fingerprint recognition processing), and use such input to process the decoded audio data output from the decoder 101. The stage 103 can perform adaptive loudness processing on the decoded audio data (output from the decoder 101) indicating a single audio program (represented by the program boundary metadata extracted by the analyzer 111), and can reset the loudness processing in response to receiving decoded audio data (output from the decoder 101) indicating different audio programs indicated by the program boundary metadata extracted by the analyzer 111.
[0090] When the control bit from validator 102 indicates that the LPSM is invalid, the dialogue loudness measurement subsystem 108 can operate to determine the loudness of segments of decoded audio (from decoder 101) representing dialogue (or other speech) using the LPSM (and / or other metadata) extracted by decoder 101. When the control bit from validator 102 indicates that the LPSM is valid, the operation of the dialogue loudness measurement subsystem 108 can be prohibited when the LPSM indicates the previously determined loudness of dialogue (or other speech) segments of the decoded audio (from decoder 101). Subsystem 108 can perform loudness measurements on decoded audio data representing a single audio program (indicated by program boundary metadata extracted by analyzer 111), and can reset the loudness processing in response to receiving decoded audio data representing different audio programs indicated by such program boundary metadata.
[0091] There are useful tools (e.g., Dolby LM100 loudness meter) for conveniently and easily measuring the level of dialogue in audio content. Some embodiments of the APU (e.g., stage 108 of encoder 100) of the present invention are implemented to include such a tool (or perform the functions of such a tool) to measure the average dialogue loudness of the audio content of an audio bitstream (e.g., a decoded AC-3 bitstream set from decoder 101 of encoder 100 to stage 108).
[0092] If stage 108 is implemented to measure the true average dialogue loudness of audio data, the measurement can include the step of separating segments of audio content that mainly contain speech. Then, the audio segments that are mainly speech are processed according to a loudness measurement algorithm. For audio data decoded from an AC-3 bitstream, the algorithm can be a standard K-weighted loudness measurement (according to the international standard ITU-R BS 1770). Alternatively, other loudness measurements (e.g., those based on loudness psychoacoustic models) can be used.
[0093] The separation of speech segments is not necessary for measuring the average dialogue loudness of audio data. However, it improves the accuracy of the measurement and generally provides more satisfactory results from the listener's perception. Since not all audio content contains dialogue (speech), the loudness measurement of the entire audio content can provide a sufficient approximation of the dialogue level of the audio where speech is present.
[0094] The metadata generator 106 generates (and / or passes to stage 107) metadata to be included by stage 107 in the encoded bitstream to be output from the encoder 100. The metadata generator 106 can pass LPSM (optionally also LIM and / or PIM and / or program boundary metadata and / or other metadata) extracted by the encoder 101 and / or the analyzer 111 to stage 107 (e.g., when the control bit from the validator 102 indicates that the LPSM and / or other metadata is valid), or generate new LIM and / or PIM and / or LPSM and / or program boundary metadata and / or other metadata and set the new metadata to stage 107 (e.g., when the control bit from the validator 102 indicates that the metadata extracted by the decoder 101 is invalid), or can set a combination of the metadata extracted by the decoder 101 and / or the analyzer 111 and the newly generated metadata to stage 107. The metadata generator 106 can include in the LPSM the loudness data generated by the subsystem 108 and at least one value indicating the type of loudness processing performed by the subsystem 108, and set the LPSM to stage 107 for inclusion in the encoded bitstream to be output from the encoder 100.
[0095] The metadata generator 106 can generate control bits (which can consist of or include a hash-based message authentication code or "HMAC") for at least one of decryption, authentication, or verification of the LPSM (optionally also other metadata) to be included in the encoded bitstream and / or the base audio data to be included in the encoded bitstream. The metadata generator 106 can provide such protection bits to stage 107 for inclusion in the encoded bitstream.
[0096] In typical operation, the dialogue loudness measurement subsystem 108 processes the audio data output from the decoder 101 to generate loudness values (e.g., gated and ungated dialogue loudness values) and dynamic range values in response to the audio data. In response to these values, the metadata generator 106 can generate loudness processing state metadata (LPSM) for inclusion (by the filler / formatter 107) in the encoded bitstream to be output from the encoder 100.
[0097] Additionally, optionally, or alternatively, the subsystems 106 and / or 108 of the encoder 100 can perform additional analysis of the audio data to generate metadata indicating at least one characteristic of the audio data for inclusion in the encoded bitstream to be output from stage 107.
[0098] The encoder 105 encodes the audio data output from the selection stage 104 (e.g., by performing compression on it), and sets the encoded audio to stage 107 for inclusion in the encoded bitstream to be output from stage 107.
[0099] Stage 107 multiplexes the encoded audio from encoder 105 and metadata (including PIM and / or SSM) from generator 106 to generate an encoded bitstream to be output from stage 107, preferably such that the encoded bitstream has a format specified by a preferred embodiment of the invention.
[0100] The frame buffer 109 is a buffer memory that stores (e.g., in a non-transitory manner) at least one frame of the encoded audio bitstream output from stage 107, and then a series of frames of the encoded audio bitstream are set from the buffer 109 as output from the encoder 100 to the transmission system 150.
[0101] The LPSM generated by metadata generator 106 and included in the encoded bitstream by stage 107 generally indicates a loudness processing state of the corresponding audio data (e.g., what type of loudness processing has been performed on the audio data) and the loudness of the corresponding audio data (e.g., measured dialogue loudness, gated and / or un-gated loudness, and / or dynamic range).
[0102] In this document, "gating" of loudness and / or level measurements performed on audio data refers to a certain level or loudness threshold above which calculated values exceeding the threshold are included in the final measurement (e.g., short-term loudness values below -60 dBFS are ignored in the final measured value). Gating for absolute values refers to a fixed level or loudness, while gating for relative values refers to a value that is dependent on the current "ungated" measurement value.
[0103] In some implementations of encoder 100, the encoded bitstream buffered in memory 109 (and output to delivery system 150) is an AC-3 bitstream or an E-AC-3 bitstream and includes audio data segments (e.g., Figure 4 ) and metadata segments, wherein the audio data segment indicates audio data, and each of at least some of the metadata segments includes PIM and / or SSM (and optionally other metadata). Stage 107 inserts the metadata segments (including metadata) into the bitstream in the following format. Each of the metadata segments including PIM and / or SSM is included in the useless bit segment of the bitstream (e.g., Figure 4 or Figure 7 ), or in the "addbsi" field of the bitstream information ("BSI") segment of a frame of a bitstream, or in the auxiliary data field at the end of a frame of a bitstream (e.g., Figure 4 or Figure 7 A frame of the bitstream may include one or two metadata segments, each metadata segment including metadata, and if a frame includes two metadata segments, one may be present in the addbsi field of the frame and the other in the AUX field of the frame.
[0104] In some embodiments, each metadata segment (sometimes referred to herein as a "container") inserted by stage 107 has a format that includes a metadata segment header (optionally also including other mandatory or "core" elements) and one or more metadata payloads following the metadata segment header. If present, the SIM is included in one of the payloads within the metadata payloads (identified by a payload header and typically having a first type of format). If present, the PIM is included in another of the payloads within the metadata payloads (identified by a payload header and typically having a second type of format). Similarly, each other type of metadata (if present) is included in another of the payloads within the metadata payloads (identified by a payload header and typically having a format for the type of metadata). The exemplary format enables access to SSM, PIM, and other metadata at times other than during decoding (e.g., by a post-processor after decoding, or by a processor configured to identify metadata without performing a full decode of the encoded bitstream), and allows for convenient and efficient error detection and correction during decoding of the bitstream (e.g., for substream identification). For example, without accessing the SSM in the exemplary format, the decoder may incorrectly identify the correct number of substreams associated with a program. One metadata payload within the metadata segment may include SSM, another metadata payload within the metadata segment may include PIM, and optionally, at least one other metadata payload within the metadata segment may include other metadata (e.g., loudness processing state metadata or "LPSM").
[0105] In some embodiments, the substream structure metadata (SSM) payload included in a frame of an encoded bitstream (e.g., an E-AC-3 bitstream indicating at least one audio program) by stage 107 includes SSM in the following format:
[0106] A payload header, typically including at least one identification value (e.g., a 2-bit value indicating the SSM format version, and optionally length, period, count, and substream associated values); and following the header:
[0107] Independent substream metadata indicating the number of independent substreams of the program indicated by the bitstream; and
[0108] Dependent substream metadata indicating whether each independent substream of the program has at least one associated dependent substream (i.e., whether at least one dependent substream is associated with each independent substream), and if so, the number of dependent substreams associated with each independent substream of the program.
[0109] It is expected that an independent sub-stream of the encoded bitstream can indicate the set of speaker channels of an audio program (e.g., the speaker channels of a 5.1 speaker channel audio program), and each of one or more dependent sub-streams (associated with the independent sub-stream and indicated by dependent sub-stream metadata) can indicate the target channels of the program. However, an independent bitstream of the encoded bitstream typically indicates the set of speaker channels of the program, and each dependent sub-stream (indicated by dependent sub-stream metadata) associated with the independent sub-stream indicates at least one additional speaker channel of the program.
[0110] In some embodiments, the program information metadata (PIM) payload included (by stage 107) in a frame of an encoded bitstream (e.g., an E-AC-3 bitstream indicating at least one audio program) has the following format:
[0111] A payload header, typically including at least one identification value (e.g., a value indicating the PIM format version, and optionally length, period, count, and sub-stream associated values); and PIM in the following format after the header:
[0112] Active channel metadata indicating each silent channel and each non-silent channel of the audio program (i.e., which channels of the program contain audio information and which channels (if any) contain only silence (typically with respect to the duration of the frame)). In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the active channel metadata in the frame of the bitstream can be combined with additional metadata of the bitstream (e.g., the audio coding mode (“acmod”) field of the frame, and, if present, the chanmap field in the frame or an associated dependent sub-stream frame) to determine which channels of the program contain audio information and which channels contain silence. The “acmod” field of an AC-3 or E-AC-3 frame indicates the number of full-range channels of the audio program indicated by the audio content of the frame (e.g., whether the program is a 1.0 channel mono program, a 2.0 channel stereo program, or a program including L, R, C, Ls, Rs full-range channels), or the frame indicates two independent 1.0 channel mono programs. The “chanmap” field of an E-AC-3 bitstream indicates the channel mapping of the dependent sub-stream indicated by the bitstream. The active channel metadata can assist in implementing upmixing (in a post-processor) downstream of the decoder, e.g., to add audio to channels containing silence at the output of the decoder;
[0113] Downmix processing status metadata indicating whether a program is downmixed (before or during encoding) and the type of downmix applied if the program is downmixed. The downmix processing status metadata can assist in performing upmixing downstream in a decoder (in a post-processor), e.g., to upmix the audio content of a program using parameters that best match the type of downmix applied. In implementations where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the downmix processing status metadata can be combined with the audio coding model ("acmod") field of a frame to determine the type of downmix (if any) applied to the channels of the program;
[0114] Upmix processing status metadata indicating whether a program is upmixed (e.g., from a smaller number of channels) before or during encoding and the type of upmix applied if the program is upmixed. The upmix processing status metadata can assist in performing downmixing downstream in a decoder (in a post-processor), e.g., to downmix the audio content of a program in a manner consistent with the type of upmix applied to the program (e.g., Dolby Pro Logic, or Dolby Pro Logic II Movie mode, or Dolby Pro Logic II Music mode, or Dolby Professional Upmixer). In implementations where the encoded bitstream is an E-AC-3 bitstream, the upmix processing status metadata can be combined with other metadata (e.g., the value of the "strmtyp" field of a frame) to determine the type of upmix (if any) applied to the channels of the program. The value of the "strmtyp" field (in the BSI field of a frame of an E-AC-3 bitstream) indicates whether the audio content of the frame belongs to an independent stream (which determines a program) or an independent substream (of a program that includes multiple substreams or is associated with multiple substreams) that can be encoded independently of any other substream indicated by the E-AC-3 bitstream, or whether the audio content of the frame belongs to a dependent substream (of a program that includes multiple substreams or is associated with multiple substreams) that must be decoded in combination with the independent substream with which it is associated; and
[0115] Preprocessing status metadata indicating: whether preprocessing has been performed on the audio content of a frame (before encoding the audio content that generates the encoded bitstream), and the type of preprocessing performed if preprocessing has been performed on the frame audio content.
[0116] In some implementations, the preprocessing status metadata indicates:
[0117] whether surround attenuation has been applied (e.g., before encoding, whether the surround channels of an audio program have been attenuated by 3 dB),
[0118] whether a 90° phase shift has been applied (e.g., before encoding, to the surround channels Ls and Rs of an audio program),
[0119] Before encoding, is a low-pass filter applied to the LFE channel of the audio program?
[0120] During generation, is the level of the LFE channel of the program monitored and, if the level of the LFE channel of the program is monitored, the monitored level of the LFE channel relative to the level of the full-range audio channels of the program?
[0121] Should dynamic range compression be performed (e.g., in the decoder) on each block of the decoded audio content of the program and, if dynamic range compression should be performed on each block of the decoded audio content of the program, the type (and / or parameters) of the dynamic range compression to be performed (e.g., the type of preprocessing status metadata can indicate which of the following compression profile types is assumed by the encoder to generate the dynamic range compression control value included in the encoded bitstream: movie standard, movie light, music standard, music light, or voice. Alternatively, the type of preprocessing status metadata can indicate that re-dynamic range compression (“compr” compression) should be performed on each frame of the decoded audio content of the program in a manner determined by the dynamic range compression control value included in the encoded bitstream).
[0122] Is spectral extension and / or channel coupling coding used to encode the program content in a specific frequency range and, if spectral extension and / or channel coupling coding is used to encode the program content in a specific frequency range, the minimum and maximum frequencies of the frequency components of the content on which spectral extension coding is performed, and the minimum and maximum frequencies of the frequency components of the content on which channel coupling coding is performed. This type of preprocessing status metadata information can assist in performing equalization (in the post-processor) downstream of the decoder. Both channel coupling information and spectral extension information assist in optimizing quality during transcoding operations and applications. For example, the encoder can optimize its behavior (including adaption of preprocessing steps such as headphone virtualization, upmixing, etc.) based on parameters such as the status of spectral extension and channel coupling information. Moreover, the encoder can dynamically modify its coupling parameters and spectral extension parameters based on the status of the incoming (and authenticated) metadata to match the optimal values and / or modify its coupling and spectral extension parameters to the optimal values, and
[0123] Is dialogue enhancement adjustment range data included in the encoded bitstream and, if the dialogue enhancement adjustment range data is included in the encoded bitstream, the range of adjustment available during the execution of dialogue enhancement processing (e.g., downstream of the decoder's post-processor) that adjusts the level of dialogue content relative to the level of non-dialogue content in the audio program?
[0124] In some implementations, additional pre - processing status metadata (e.g., metadata indicating parameters related to a headset) is included in the PIM payload of the encoded bitstream (by stage 107) to be output from encoder 100.
[0125] In some implementations, the LPSM payload included in a frame of an encoded bitstream (e.g., an E - AC - 3 bitstream indicating at least one audio program) by stage 107 includes an LPSM of the following format:
[0126] A header (generally including a sync word identifying the start of the LPSM payload, at least one identification value after the sync word, e.g., the LPSM format version, length, period, count, and sub - stream association values as represented in Table 2 below); and
[0127] Following the header:
[0128] At least one dialogue indication value (e.g., the parameter "dialogue channels" in Table 2) indicating whether the corresponding audio data indicates dialogue or not (e.g., which channels of the corresponding audio data indicate dialogue);
[0129] At least one loudness adjustment compliance value (e.g., the parameter "loudness adjustment type" in Table 2) indicating whether the corresponding audio content complies with the indicated set of loudness adjustments;
[0130] At least one loudness processing value of at least one type of loudness processing that has been performed on the corresponding audio data (e.g., one or more of the parameters "dialogue - gated loudness correction flag", "loudness correction type" in Table 2); and
[0131] At least one loudness value (e.g., one or more of the parameters "ITU relative - gated loudness", "ITU speech - gated loudness", "ITU (EBU 3341) short - term 3s loudness", and "true peak" in Table 2) indicating at least one loudness (e.g., peak or average loudness) characteristic of the corresponding audio data.
[0132] In some implementations, each metadata segment containing PIM and / or SSM (and optionally other metadata) contains a metadata segment header (and optionally additional core elements), and at least one metadata payload segment of the following format after the metadata segment header (or the metadata segment header and other core elements):
[0133] A payload header, generally including at least one identification value (e.g., SSM or PIM format version, length, period, count, and sub - stream association values), and
[0134] SSM or PIM (or another type of metadata) after the payload header.
[0135] In some implementations, each metadata segment (sometimes referred to herein as a "metadata container" or "container") in the padding bit segment / skip field segment (or "addbsi" field or auxiliary data field) of a frame inserted by stage 107 has the following format:
[0136] A metadata segment header (usually including a sync word indicating the start of the metadata segment, an identification value after the sync word, e.g., version, length, period, extended element count, and substream association values as represented in Table 1 below); and
[0137] At least one protection value (e.g., HMAC digest and audio fingerprint values of Table 1) after the metadata segment header that aids in at least one of decryption, authentication, or verification of the metadata segment or the corresponding audio data; and
[0138] A metadata payload identification ("ID") value and a payload configuration value that also follow the metadata segment header, identify the type of metadata in each of the following metadata payloads, and indicate at least one aspect of the configuration (e.g., size) of each such payload.
[0139] Each metadata payload follows the corresponding payload ID value and payload configuration value.
[0140] In some embodiments, each metadata segment in the padding bit segment (or auxiliary data field or "addbsi" field) of a frame has a three - level structure:
[0141] A high - level structure (e.g., a metadata segment header), including a flag indicating whether the padding (or auxiliary data or addbsi) field includes metadata, at least one ID value indicating what type of metadata is present, and typically also a value indicating how many bits of (e.g., each type of) metadata are present if the metadata is present. One type of metadata that can be present is PIM, another type of metadata that can be present is SSM, and other types of metadata that can be present are LPSM, and / or program boundary metadata, and / or media search metadata;
[0142] A middle - level structure, including data associated with each identified type of metadata (e.g., a metadata payload header, protection value, and payload ID value and payload configuration value for each identified type of metadata); and
[0143] Low-level structure, including a metadata payload of metadata about each identified type (e.g., if PIM is identified as being present, a series of PIM values, and / or if that other type of metadata is identified as being present, metadata values of another type (e.g., SSM or LPSM)).
[0144] Data values in such three-level structures can be nested. For example, the protection value of each payload identified by the high-level structure and the intermediate-level structure (e.g., each PIM, or SSM or other data payload) can be included after the payload (thus after the metadata payload header of the payload), or the protection values of all metadata payloads identified by the high-level structure and the intermediate-level structure can be included after the final metadata payload in the metadata segment (thus after the metadata payload headers of all payloads in the metadata segment).
[0145] In an example (to be described by the metadata segment or "container" with reference to Figure 8 ), the metadata segment header identifies 4 metadata payloads. As Figure 8 shown, the metadata segment header includes a container sync word (identified as "container sync") and version and key ID values. After the metadata segment header are 4 metadata payloads and protection bits. The payload ID value and payload configuration (e.g., payload size) value of the first payload (e.g., PIM payload) are after the metadata segment header, the first payload itself is after the ID and configuration values, the payload ID value and payload configuration (e.g., payload size) value of the second payload (e.g., SSM payload) are after the first payload, the second payload itself is after these ID and configuration values, the payload ID value and payload configuration (e.g., payload size) value of the third payload (e.g., LPSM payload) are after the second payload, the third payload itself is after these ID and configuration values, the payload ID value and payload configuration (e.g., payload size) value of the fourth payload are after the third payload, the fourth payload itself is after these ID and configuration values, and the protection value for all or some of the payloads in the payload (or for the high-level structure and the intermediate-level structure and all or some of the payloads in the payload) (identified as "protected data" in Figure 8 is after the last payload.
[0146] In some embodiments, if decoder 101 receives an audio bitstream with an encrypted hash generated according to an embodiment of the present invention, the decoder is configured to analyze and retrieve the encrypted hash based on data blocks determined from the bitstream, where the blocks include metadata. Validator 102 may use the encrypted hash to validate the received bitstream and / or associated metadata. For example, if validator 102 finds that the metadata is valid based on a match between a reference encrypted hash and the encrypted hash retrieved from the data blocks, then processor 103 may be prohibited from operating on the corresponding audio data, and selection stage 104 may be caused to pass the (unchanged) audio data. Additionally, optionally or alternatively, other types of encryption techniques may be used in place of the encrypted hash-based method.
[0147] Figure 2 Encoder 100 may determine (responsive to the LPSM extracted by decoder 101 and optionally also responsive to program boundary metadata) that a type of loudness processing has been performed on the audio data to be encoded by the post-processing / pre-processing unit (in elements 105, 106, and 107), and thus may create (in generator 106) loudness processing status metadata including and / or based on specific parameters for the previously performed loudness processing. In some implementations, as long as the encoder knows the type of processing that has been performed on the audio content, encoder 100 may create metadata indicating the processing history of the audio content (and include it in the encoded bitstream output from the encoder).
[0148] Figure 3 FIG. is a block diagram of a decoder (200) of an embodiment of the audio processing unit of the present invention and a post-processor (300) coupled to the decoder (200). The post-processor (300) is also an embodiment of the audio processing unit of the present invention. Any of the components or elements of encoder 200 and post-processor 300 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. Decoder 200 includes a frame buffer 201, an analyzer 205, an audio decoder 202, an audio status verification stage (validator) 203, and a control bit generation stage 204 connected as shown. Generally, decoder 200 also includes other processing elements (not shown).
[0149] Frame buffer 201 (buffer memory) stores (e.g., in a non-transitory manner) at least one frame of the encoded audio bitstream received by decoder 200. The sequence of frames of the encoded audio bitstream is set from buffer 201 to analyzer 205.
[0150] Couple to the analyzer 205 and configure it to extract PIM and / or SSM from each frame of the encoded input audio (optionally also extract other metadata, such as LPSM), set at least some of the metadata (e.g., LPSM and program boundary metadata, if either is extracted, and / or PIM and / or SSM) to the audio state validator 203 and stage 204, set the extracted metadata as the output (e.g., to the post-processor 300), extract audio data from the encoded input audio, and set the extracted audio data to the decoder 202.
[0151] The encoded audio bitstream input to the decoder 200 can be one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream.
[0152] Figure 3 The system further includes a post-processor 300. The post-processor 300 includes a frame buffer 301 and other processing elements (not shown) including at least one processing element coupled to the buffer 301. The frame buffer 301 stores (e.g., in a non-transitory manner) at least one frame of the decoded audio bitstream received by the post-processor 300 from the decoder 200. Couple the processing elements of the post-processor 300 and configure them to receive a series of frames of the decoded audio bitstream output from the buffer 301 and perform adaptive processing on them using the metadata output from the decoder 200 and / or the control bits output from stage 204 of the decoder 200. Generally, the post-processor 300 is configured to perform adaptive processing on the decoded audio data using the metadata from the decoder 200 (e.g., perform adaptive loudness processing on the decoded audio data using LPSM values and optionally also using program boundary metadata, where the adaptive processing can be based on the loudness processing state and / or one or more audio data characteristics indicated by the LPSM of the audio data indicating a single audio program).
[0153] Various implementations of the decoder 200 and the post-processor 300 are configured to perform different embodiments of the method of the present invention.
[0154] The audio decoder 202 of the decoder 200 is configured to decode the audio data extracted by the analyzer 205 to generate decoded audio data, and set the decoded audio data as the output (e.g., to the post-processor 300).
[0155] The status validator 203 is configured to authenticate and verify the metadata set thereto. In some embodiments, the metadata is (or is included in) a data block that has been included in an input bitstream (e.g., according to an embodiment of the present invention). The block may include an encrypted hash (a hash-based message authentication code or “HMAC”) for processing the metadata and / or the base audio data (provided from the analyzer 205 and / or the decoder 202 to the validator 203). The data block may be digitally tagged in these embodiments such that downstream audio processing units can relatively easily authenticate and verify the processing status metadata.
[0156] Other encryption methods, including but not limited to any one of one or more non-HMAC encryption methods, may be used for the verification of the metadata (e.g., in the validator 203) to ensure the secure transmission and reception of the metadata and / or the base audio data. For example, the verification (using such an encryption method) may be performed in each audio processing unit of an embodiment that receives the audio bitstream of the present invention to determine whether the metadata and the corresponding audio data included in the bitstream have undergone (and / or originated from) a specific processing (indicated by the metadata) and have not been modified after such specific processing is performed.
[0157] The status validator 203 sets control data to the control bit generator 204, and / or sets the control data as an output (e.g., sets to the post-processor 300) to indicate the result of the verification operation. In response to the control data (and optionally other metadata extracted from the input bitstream), stage 204 may generate (and set to the post-processor 300):
[0158] A control bit indicating that the decoded audio data output from the decoder 202 has undergone a specific type of loudness processing (when the LPSM indicates that the audio data output from the decoder 202 has undergone the specific type of loudness processing and the control bit from the validator 203 indicates that the LPSM is valid); or
[0159] A control bit indicating that the decoded audio data output from the decoder 202 should undergo a specific type of loudness processing (e.g., when the LPSM indicates that the audio data output from the decoder 202 has not undergone the specific type of loudness processing, or when the LPSM indicates that the audio data output from the decoder 202 has undergone the specific type of loudness processing but the control bit from the validator 203 indicates that the LPSM is invalid).
[0160] Alternatively, the decoder 200 sets the metadata extracted by the decoder 202 from the input bitstream and the metadata extracted by the analyzer 205 from the input bitstream to the post-processor 300, and the post-processor 300 performs adaptive processing on the decoded audio data using the metadata, or performs verification of the metadata, and then, if the verification indicates that the metadata is valid, performs adaptive processing on the decoded audio data using the metadata.
[0161] In some embodiments, if the decoder 200 receives an audio bitstream generated according to an embodiment of the present invention using an encrypted hash, the decoder is configured to analyze and retrieve the encrypted hash of data blocks determined from the bitstream, the blocks including loudness processing status metadata (LPSM). The validator 203 may use the encrypted hash to verify the received bitstream and / or associated metadata. For example, if the validator 203 finds the LPSM valid based on a match between a reference encrypted hash and the encrypted hash retrieved from the data block, then the downstream audio processing unit (e.g., the post-processor 300 which may be or include a volume leveling unit) may be signaled to pass the audio data of the (unchanged) bitstream. Additionally, optionally or alternatively, other types of encryption techniques may be used in place of the encrypted hash-based method.
[0162] In some implementations of the decoder 200, the received (and cached in the memory 201) encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and includes an audio data segment (e.g., Figure 4 segments AB0 to AB5 of the frames shown) and a metadata segment, where the audio data segment indicates audio data, and each of at least some of the metadata segments includes PIM or SSM (or other metadata). The decoder stage 202 (and / or the analyzer 205) is configured to extract metadata from the bitstream. Each metadata segment including PIM and / or SSM (optionally also including other metadata) in the metadata segment is included in the padding segment of the frame of the bitstream, or in the "addbsi" field of the bitstream information ("BSI") segment of the frame of the bitstream, or in the auxiliary data field at the end of the frame of the bitstream (e.g., Figure 4 the AUX segment shown). The frame of the bitstream may include one or two metadata segments, where each metadata segment includes metadata, and if the frame includes two metadata segments, one may be present in the addbsi field of the frame and the other in the AUX field of the frame.
[0163] In some embodiments, each metadata segment (sometimes referred to herein as a "container") of the bitstream cached in buffer 201 has a format that includes a metadata segment header (optionally also including other mandatory or "core" elements), and one or more metadata payloads following the metadata segment header. If present, the SIM is included in one of the payloads within the metadata payloads (identified by a payload header and typically having a first type of format). If present, the PIM is included in another of the payloads within the metadata payloads (identified by a payload header and typically having a second type of format). Similarly, other types of metadata (if present) are included in another of the payloads within the metadata payloads (identified by a payload header and typically having a format for the type of metadata). The exemplary format enables convenient access to SSM, PIM, and other metadata at times other than during decoding (e.g., by post-processor 300 after decoding, or by a processor configured to identify metadata without performing a full decode of the encoded bitstream), and allows for convenient and efficient error detection and correction during decoding of the bitstream (e.g., for substream identification). For example, without accessing the SSM in the exemplary format, decoder 200 may incorrectly identify the correct number of substreams associated with a program. One metadata payload within the metadata segment may include SSM, another metadata payload within the metadata segment may include PIM, and optionally, at least one other metadata payload within the metadata segment may include other metadata (e.g., loudness processing state metadata or "LPSM").
[0164] In some embodiments, the substream structure metadata (SSM) payload included in a frame of an encoded bitstream (e.g., an E-AC-3 bitstream indicating at least one audio program) cached in buffer 201 includes SSM in the following format:
[0165] A payload header, typically including at least one identification value (e.g., a 2-bit value indicating the SSM format version, and optionally length, period, count, and substream association values); and
[0166] Following the header:
[0167] Independent substream metadata indicating the number of independent substreams of the program indicated by the bitstream; and
[0168] Dependent substream metadata indicating: for each independent substream of the program, whether there is at least one dependent substream associated therewith, and if there is at least one dependent substream associated with each independent substream of the program, the number of dependent substreams associated with each independent substream of the program.
[0169] In some embodiments, the program information metadata (PIM) payload included in a frame of an encoded bitstream (e.g., an E-AC-3 bitstream indicating at least one audio program) cached in buffer 201 has the following format:
[0170] A payload header, typically including at least one identification value (e.g., a value indicating the PIM format version, and optionally length, period, count, and substream association values); and after the header, PIM in the following format:
[0171] Active channel metadata for each silent and each non-silent channel of the audio program (i.e., which channels of the program contain audio information and which channels (if any) contain only silence (typically with respect to the duration of the frame)). In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the active channel metadata in a frame of the bitstream can be combined with additional metadata of the bitstream (e.g., the audio coding mode (“acmod”) field of the frame, and if present, the chanmap field in the frame or associated dependent substream frame) to determine which channels of the program contain audio information and which channels contain silence;
[0172] Downmix processing status metadata, which indicates whether the program was downmixed (before or during encoding), and if the program was downmixed, the type of downmix applied. The downmix processing status metadata can assist in performing upmixing downstream in a decoder (in post-processor 300), e.g., to upmix the audio content of the program using parameters that best match the type of downmix applied. In embodiments where the encoded bitstream is an AC-3 or E-AC-3 bitstream, the downmix processing status metadata can be combined with the audio coding model (“acmod”) field of the frame to determine the type of downmix (if any) applied to the channels of the program;
[0173] Upmix processing status metadata that indicates whether a program is upmixed (e.g., from a smaller number of channels) before or during encoding, and if the program is upmixed, the type of upmix applied. The upmix processing status metadata can assist in performing a downmix downstream in a decoder (in a post-processor), e.g., to downmix the audio content of a program in a manner consistent with the type of upmix applied to the program (e.g., Dolby Pro Logic, or Dolby Pro Logic II movie mode, or Dolby Pro Logic II music mode, or Dolby Professional Upmixer). In an implementation where the encoded bitstream is an E-AC-3 bitstream, the upmix processing status metadata can be combined with other metadata (e.g., the value of the "strmtyp" field of a frame) to determine the type of upmix (if any) applied to the channels of a program. The value of the "strmtyp" field (in the BSI field of a frame of an E-AC-3 bitstream) indicates whether the audio content of the frame belongs to an independent stream (which determines a program) or an independent substream (of a program that includes multiple substreams or is associated with multiple substreams) and can thus be encoded independently of any other substream indicated by the E-AC-3 bitstream, or whether the audio content of the frame belongs to a dependent substream (of a program that includes multiple substreams or is associated with multiple substreams) and must thus be decoded in combination with the independent substream with which it is associated; and
[0174] Preprocessing status metadata that indicates whether preprocessing has been performed on the audio content of a frame (before encoding the audio content of the encoded bitstream), and if preprocessing has been performed on the frame audio content, the type of preprocessing performed.
[0175] In some implementations, the preprocessing status metadata indicates:
[0176] whether surround attenuation has been applied (e.g., whether the surround channels of an audio program have been attenuated by 3 dB before encoding),
[0177] whether a 90° phase shift has been applied (e.g., to the Ls and Rs channels of the surround channels of an audio program before encoding),
[0178] whether a low-pass filter has been applied to the LFE channel of an audio program before encoding,
[0179] whether the level of the LFE channel of a program has been monitored during generation, and if the level of the LFE channel of a program has been monitored, the monitored level of the LFE channel relative to the level of the full-range audio channels of the program,
[0180] Whether dynamic range compression should be performed (e.g., in a decoder) on each block of the decoded audio of a program, and if so, the type (and / or parameters) of dynamic range compression to be performed (e.g., this type of preprocessing status metadata can indicate which of the following compression profile types is assumed by the encoder to generate the dynamic range compression control values included in the encoded bitstream: movie standard, movie light, music standard, music light, or speech. Alternatively, this type of preprocessing status metadata can indicate that re-dynamic range compression (“compr” compression) should be performed on each frame of the decoded audio content of the program in a manner determined by the dynamic range compression control values included in the encoded bitstream),
[0181] Whether spectral extension and / or channel coupling coding is used to encode the content of a program in a specific frequency range, and if so, the minimum and maximum frequencies of the frequency components of the content on which spectral extension coding is performed, and the minimum and maximum frequencies of the frequency components of the content on which channel coupling coding is performed. This type of preprocessing status metadata information can assist in performing equalization (in a post-processor) downstream of the decoder. Both the channel coupling information and the spectral extension information also assist in optimizing quality during transcoding operations and applications. For example, the encoder can optimize its behavior (including adaptive preprocessing steps such as headphone virtualization, upmixing, etc.) based on the status of the parameters (e.g., spectral extension and channel coupling information). Moreover, the encoder can dynamically modify its coupling and spectral extension parameters based on the status of the incoming (and authenticated) metadata to match the optimal values and / or modify its coupling and spectral extension parameters to the optimal values, and
[0182] Whether dialogue enhancement adjustment range data is included in the encoded bitstream, and if so, the adjustment range available during the execution of dialogue enhancement processing (e.g., downstream of the decoder's post-processor) for adjusting the level of dialogue content relative to the level of non-dialogue content in the audio program.
[0183] In some embodiments, the LPSM payload included in a frame of the encoded bitstream (e.g., an E-AC-3 bitstream indicating at least one audio program) buffered in buffer 201 includes an LPSM of the following format:
[0184] A header (generally including a sync word identifying the start of the LPSM payload, and at least one identification value after the sync word, e.g., the LPSM format version, length, period, count, and substream association values indicated in Table 2 below); and
[0185] Following the header:
[0186] At least one dialogue representation value indicating whether the corresponding audio data indicates dialogue or not (e.g., which channels of the corresponding audio data indicate dialogue); (e.g., the parameter "dialogue channel" in Table 2).
[0187] At least one loudness adjustment compliance value indicating whether the corresponding audio content complies with the indicated set of loudness adjustments; (e.g., the parameter "loudness adjustment type" in Table 2).
[0188] At least one loudness processing value indicating at least one type of loudness processing that has been performed on the corresponding audio data; (e.g., one or more of the parameters "dialogue gating loudness correction flag", "loudness correction type" in Table 2); and
[0189] At least one loudness value indicating at least one loudness (e.g., peak or average loudness) characteristic of the corresponding audio data; (e.g., one or more of the parameters "ITU relative gating loudness", "ITU voice gating loudness", "ITU (EBU 3341) short-term 3s loudness", and "true peak" in Table 2).
[0190] In some implementations, the analyzer 205 (and / or decoder stage 202) is configured to extract each metadata segment having the following format from the padding bit segment, "addbsi" field, or auxiliary data segment of a frame of the bitstream:
[0191] A metadata segment header (typically including a sync word identifying the start of the metadata segment, and identification values after the sync word, such as version, length, period, extended element count, and substream association value); and
[0192] At least one protection value after the metadata segment header that contributes to at least one of decryption, authentication, or verification of at least one of the metadata segment or the metadata of the corresponding audio data; (e.g., the HMAC digest and audio fingerprint values in Table 1); and
[0193] A metadata payload identification ("ID") value and a payload configuration value that also come after the metadata segment header and identify the type of metadata in each of the following metadata payloads and represent at least one aspect of the configuration (e.g., size) of each such payload.
[0194] Each metadata payload segment (preferably having the format specified above) comes after the corresponding metadata payload ID value and metadata configuration value.
[0195] More generally, an encoded audio bitstream generated by a preferred embodiment of the present invention has a structure that provides a mechanism for labeling metadata elements and subelements as either core (mandatory) or extended (optional) elements or subelements. This enables the data rate of the bitstream (including its metadata) to be extended to a large number of applications. The core (mandatory) elements of the preferred bitstream syntax should also be able to signal the presence of extended (optional) elements associated with the audio content in (in-band) and / or remote locations (out-of-band).
[0196] The core elements are required to be present in each frame of the bitstream. Some of the subelements of the core elements are optional and can be present in any combination. The extended elements are not required to be present in each frame (to limit the overall bitrate overhead). Thus, the extended elements can be present in some frames and not in others. Some of the subelements of the extended elements are optional and can be present in any combination. However, some of the subelements of the extended elements can be mandatory (i.e., if the extended element is present in a frame of the bitstream).
[0197] In one class of embodiments, an encoded audio bitstream is generated (e.g., by an audio processing unit implementing the present invention) that includes a series of audio data segments and metadata segments. The audio data segments indicate audio data, each of at least some of the metadata segments includes PIM and / or SSM (and optionally at least one other type of metadata), and the audio data segments are time-division multiplexed with the metadata segments. In a preferred embodiment of this class, each of the metadata segments has a preferred format to be described herein.
[0198] In a preferred format, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each metadata segment that includes SSM and / or PIM is included (e.g., by stage 107 of a preferred implementation of encoder 100) as additional bitstream information in the "addbsi" field of the bitstream information ("BSI") segment of a frame of the bitstream Figure 6 (shown), or in the auxiliary data field of a frame of the bitstream, or in the padding bits segment of a frame of the bitstream.
[0199] In the preferred format, each in a frame includes a metadata segment (sometimes also referred to herein as a metadata container or container) in the padding bits segment (or addbsi field) of the frame. The metadata segment has the mandatory elements (collectively referred to as "core elements") shown in Table 1 below (and can include the optional elements shown in Table 1). At least some of the required elements shown in Table 1 are included in the metadata segment header of the metadata segment, but some can be included in other locations of the metadata segment:
[0200] Table 1
[0201]
[0202]
[0203] In a preferred format, each metadata segment containing SSM, PIM, or LPSM (in the padding bits segment or addbsi or auxiliary data field of a frame of the encoded bitstream) contains a metadata segment header (and optionally additional core elements), and one or more metadata payloads following the metadata segment header (or the metadata segment header and other core elements). Each metadata payload includes a metadata payload header included in the payload (indicating the specific type of metadata (e.g., SSM, PIM, or LPSM)), followed by the metadata of the specific type. Typically, the metadata payload header includes the following values (parameters):
[0204] A payload ID (identifying the type of metadata, e.g., SSM, PIM, or LPSM) following the metadata segment header (which may include values specified in Table 1);
[0205] A payload configuration value (usually indicating the size of the payload) following the payload ID;
[0206] And optionally further includes additional payload configuration values (e.g., an offset value indicating the number of audio samples from the start of the frame to the first audio sample involved in the payload, and a payload priority value, e.g., indicating the conditions under which the payload can be discarded).
[0207] Typically, the metadata of the payload has one of the following formats:
[0208] The metadata of the payload is SSM, including independent substream metadata indicating the number of independent substreams of the program indicated by the bitstream; and dependent substream metadata indicating whether each independent substream of the program has at least one dependent substream associated therewith, and if each independent substream of the program has at least one dependent substream associated therewith, the number of dependent substreams associated with each independent substream of the program;
[0209] The metadata of the payload is PIM, including active channel metadata indicating which channels of the audio program contain audio information and which channels (if any) contain only silence (typically regarding the duration of the frame); downmix processing status metadata indicating whether the program is downmixed (before or during encoding), and if the program is downmixed, the type of downmix applied; upmix processing status metadata indicating whether the program is upmixed (e.g., from a smaller number of channels) before or during encoding, and if the program is upmixed, the type of upmix applied; and preprocessing status metadata indicating whether preprocessing has been performed on the audio data of the frame (before encoding the audio content of the encoded bitstream), and if preprocessing has been performed on the audio data of the frame, the type of preprocessing performed; or
[0210] The metadata of the payload is LPSM, and the LPSM has a format as indicated in the following table (Table 2):
[0211] Table 2
[0212]
[0213]
[0214]
[0215]
[0216] In another preferred format of the encoded bitstream generated according to the present invention, the bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each metadata segment (e.g., level 107 of the preferred implementation of encoder 100) in the metadata segment including PIM and / or SSM (optionally further including at least one other type of metadata) is included in any of the following: the padding bit segment of the frame of the bitstream; or the "addbsi" field of the bitstream information ("BSI") segment of the frame of the bitstream ( Figure 6 as shown); or the auxiliary data field at the end of the frame of the bitstream (e.g., Figure 4the AUX segment shown in ). A frame may include one or two metadata segments, each of the metadata segments including PIM and / or SSM, and (in some embodiments) if the frame includes two metadata segments, one may be present in the addbsi field of the frame and the other in the AUX field of the frame. Each metadata segment preferably has the format specified above with reference to Table 1 above (i.e., including the core elements specified in Table 1, followed by a payload ID value (identifying the type of metadata in each payload of the metadata segment) and a payload configuration value, and each metadata payload). Each metadata segment including LPSM preferably has the format specified above with reference to Tables 1 and 2 above (i.e., including the core elements specified in Table 1, followed by a payload ID (identifying the metadata as LPSM) and a payload configuration value, followed by the payload (LPSM data having the format indicated in Table 2)).
[0217] In another preferred format, the encoded bitstream is a Dolby E bitstream, and each metadata segment including PIM and / or SSM (optionally also including other metadata) in the metadata segment is the first N sample positions of the Dolby E guard band interval. The Dolby E bitstream including such a metadata segment including LPSM preferably includes a value indicating the LPSM payload length signaled in the Pd word of the SMPTE 337M preamble (the SMPTE 337M Pa word repetition frequency preferably remains the same as the associated video frame rate).
[0218] In a preferred format, where the encoded bitstream is an E-AC-3 bitstream, each metadata segment including PIM and / or SSM (optionally also including LPSM and / or other metadata) in the metadata segment (e.g., stage 107 of the preferred implementation of encoder 100) is included as additional bitstream information in the "addbsi" field of the padding bit segment or the bitstream information ("BSI") segment of the frame that is the bitstream. The following additional aspects of encoding the E-AC-3 bitstream using LPSM in this preferred format are described next:
[0219] 1. During the generation of the E-AC-3 bitstream, although the E-AC-3 encoder (inserting LPSM values into the bitstream to be) is "active", for each generated frame (synchronization frame), the bitstream should include a metadata block (including LPSM) carried in the addbsi field (or padding bit segment) of the frame. The bits required to carry the metadata block should not increase the encoder bit rate (frame length);
[0220] 2. Each metadata block (containing LPSM) should contain the following information:
[0221] Loudness correction type flag: where "1" indicates that the loudness of the corresponding audio data is corrected upstream of the encoder, and "0" indicates that the loudness is corrected by a loudness corrector embedded in the encoder (e.g., Figure 2 the loudness processor 103 of encoder 100);
[0222] Speech channels: indicate which source channels contain speech (in the previous 0.5 seconds). If no speech is detected, this should be indicated;
[0223] Speech loudness: indicates the integrated speech loudness of each corresponding audio channel that includes speech (in the previous 0.5 seconds);
[0224] ITU loudness: indicates the integrated ITU BS.1770-3 loudness of each corresponding audio channel; and
[0225] Gain: the inverse loudness composite gain in the decoder (to indicate reversibility);
[0226] 3. When the E-AC-3 encoder (inserting LPSM values into the bitstream) is "active" and is receiving AC-3 frames with a "trusted" flag, the loudness controller in the encoder (e.g., Figure 2 the loudness processor 103 of encoder 100) should be bypassed. The "trusted" source dialogue normalization and DRC values should be passed (e.g., by the generator 106 of encoder 100) to the E-AC-3 encoder component (e.g., stage 107 of encoder 100). LPSM block generation continues, and the loudness correction type flag is set to "1". The loudness controller bypass sequence must be synchronized to the start of the decoded AC-3 frame where the "trusted" flag appears. The loudness controller bypass sequence should be implemented as follows: the flattener volume control decreases from value 9 to value 0 over 10 audio block periods (i.e., 53.3 milliseconds), and the flattener return end meter control is placed in the bypass mode (this operation should result in a seamless transition). The term "trusted" bypass of the regulator implies that the dialogue normalization value of the source bitstream is also reused at the output of the encoding. (For example, if the "trusted" source bitstream has a dialogue normalization value of -30, the output of the encoder should use -30 for the output dialogue normalization value);
[0227] 4. When the E-AC-3 encoder (inserting LPSM values into the bitstream) is "active" and is receiving AC-3 frames without a "trusted" flag, the loudness controller embedded in the encoder (e.g., Figure 2The loudness processor 103 of the encoder 100) shall be active. The LPSM block generation continues, and the loudness correction type flag is set to "0". The loudness controller activation sequence shall be synchronized to the start of the decoded AC-3 frame where the "trust" flag disappears. The loudness controller activation sequence shall be implemented as follows: the leveling meter control increases from value 0 to value 9 across 1 audio block period (e.g., 5.3 milliseconds), and the return end meter control is placed in the "active" mode (this operation shall result in a seamless transition and includes a return end meter comprehensive reset); and 5. During encoding, the graphical user interface (GUI) shall indicate the following parameters to the user: "Input audio program: [trusted / untrusted]" - the status of this parameter is based on the presence of the "trust" flag in the input signal; and "Real-time loudness correction: [enabled / disabled]" - the status of this parameter is based on whether the loudness controller embedded in the encoder is active.
[0228] When decoding an AC-3 or E-AC-3 bitstream that causes the LSPM (in a preferred format) to be included in the padding bit segment or skip field segment or bitstream information ("BSI") segment's "addbsi" field of each frame of the bitstream, the decoder shall analyze the LPSM block data (in the padding bit segment or addbsi field) and pass all the extracted LPSM values to the graphical user interface (GUI). Refresh the set of extracted LPSM values per frame.
[0229] In another preferred format of the encoded bitstream generated according to the present invention, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each metadata segment in the metadata segment that includes PIM and / or SSM (optionally also including LPSM and / or other metadata) (e.g., by stage 107 of a preferred implementation of the encoder 100) is included in the padding bit segment or AUX segment of the frame of the bitstream or as additional bitstream information in the "addbsi" field of the bitstream information ("BSI") segment ( Figure 6 as shown). In this format (which is a variant of the format described above with reference to Tables 1 and 2), each field in the addbsi (or AUX or padding) field that contains LPSM contains the following LPSM values:
[0230] The core elements specified in Table 1, followed by a payload ID (identifying the metadata as LPSM) and a payload value, followed by a payload (LPSM data) having the following format (similar to the mandatory elements shown in Table 2 above):
[0231] Version of the LPSM payload: A 2-bit field indicating the version of the LPSM payload;
[0232] dialchan: A 3-bit field indicating the left, right, and / or center channels of the corresponding audio data containing spoken dialogue. The bit assignment of the dialchan field can be as follows: Bit 0 indicating the presence of dialogue in the left channel is stored in the most significant bit of the dialchan field; and bit 2 indicating the presence of dialogue in the center channel is stored in the least significant bit of the dialchan field. If the corresponding channel contains spoken dialogue during the first 0.5 seconds of the program, each bit of the dialchan field is set to "1";
[0233] loudregtyp: A 4-bit field indicating which loudness adjustment standard the program loudness conforms to. Setting the "loudregtyp" field to "0000" indicates that LPSM does not indicate loudness adjustment compliance. For example, one value of this field (e.g., 0000) can indicate that compliance with the loudness adjustment standard is not indicated, another value of this field (e.g., 0001) can indicate that the audio data of the program conforms to the ATSC A / 85 standard, and another value of this field (e.g., 0010) can indicate that the audio data of the program conforms to the EBU R128 standard. In this example, if this field is set to any value other than "0000", the loudcorrdialgat and loudcorrtyp fields should subsequently be present in the payload;
[0234] loudcorrdialgat: A 1-bit field indicating whether dialogue gating correction has been applied. If the loudness of the program has been corrected using dialogue gating correction, the value of the loudcorrdialgat field is set to "1". Otherwise, it is set to "0";
[0235] loudcorrtyp: A 1-bit field indicating the type of loudness correction applied to the program. If the loudness of the program has been corrected using infinite-advance (file-based) loudness correction processing, the value of the loudcorrtyp field is set to "0". If the loudness of the program has been corrected using a combination of real-time loudness measurement and dynamic range control, the value of this field is set to "1";
[0236] loudrelgate: A 1-bit field indicating whether relative gated program loudness (ITU) exists. If the loudrelgate field is set to "1", the 7-bit ituloudrelgat field should subsequently be present in the payload;
[0237] loudrelgat: A 7-bit field indicating relative gated program loudness (ITU). This field indicates the integrated loudness of the audio program measured according to ITU-R BS.1770-3 without any gain adjustment due to the applied dialogue normalization and dynamic range compression (DRC). Values from 0 to 127 are interpreted as -58 LKFS to +5.5 LKFS in 0.5 LKFS steps;
[0238] loudspchgate: A 1-bit field indicating the presence of voice-gated loudness data (ITU). If the loudspchgate field is set to "1", then the payload shall subsequently be the 7-bit loudspchgat field;
[0239] loudspchgate: A 7-bit field indicating voice-gated program loudness. This field indicates the integrated loudness of the entire corresponding audio program measured according to formula (2) of ITU-R BS.1770-3 without any gain adjustment due to the applied dialogue normalization and dynamic range compression. Values from 0 to 127 are interpreted as -58 LKFS to +5.5 LKFS in 0.5 LKFS steps;
[0240] loudstrm3e: A 1-bit field indicating the presence of short-term (3-second) loudness data. If this field is set to "1", then the payload shall subsequently be the 7-bit loudstrm3s field;
[0241] loudstrm3s: A 7-bit field indicating the ungated loudness of the first 3 seconds of the corresponding audio program measured according to ITU-R BS.1771-1 without any gain adjustment due to the applied dialogue normalization and dynamic range compression. Values from 0 to 256 are interpreted as -116 LKFS to +11.5 LKFS in 0.5 LKFS steps;
[0242] truepke: A 1-bit field indicating the presence of true peak loudness data. If the truepke field is set to "1", then the payload shall subsequently be the 8-bit truepk field; and
[0243] truepk: An 8-bit field indicating the program true peak sample value measured according to Annex 2 of ITU-R BS.1770-3 without any gain adjustment due to the applied dialogue normalization and dynamic range compression. Values from 0 to 256 are interpreted as -116 LKFS to +11.5 LKFS in 0.5 LKFS steps.
[0244] In some embodiments, the core elements of the metadata segment in the padding bit section or auxiliary data (or "addbsi") field of a frame of an AC-3 bitstream or an E-AC-3 bitstream include a metadata segment header (which typically includes an identification value, e.g., a version), and, following the metadata segment header: a value indicating whether the metadata of the metadata segment includes fingerprint data (or other protection value), a value indicating the presence of external data (related to the audio data of the metadata corresponding to the metadata segment), a payload ID value and a payload configuration value for each type of metadata (e.g., PIM and / or SSM and / or LPSM and / or a type of metadata) identified by the core element, and a protection value for at least one type of metadata identified by the metadata segment header (or other core elements of the metadata segment). The metadata payload of the metadata segment follows the metadata segment header and (in some cases) is nested within the core elements of the metadata segment.
[0245] Embodiments of the present invention may be implemented in hardware, firmware, software, or a combination of hardware and software (e.g., as a programmable logic array). Unless otherwise specified, the algorithms or processes included as part of the present invention do not inherently involve any particular computer or other device. Specifically, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct a more specific device (e.g., an integrated circuit) to perform the required method steps. Thus, the present invention may be implemented as one or more computer programs executed on one or more programmable computer systems (e.g., Figure 1 the elements of, or Figure 2 the encoder 100 (or elements of the encoder), or Figure 3 the decoder (or elements of the decoder), or Figure 3 the post-processor (or elements of the post-processor)), each programmable computer system including at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The program code is applied to the input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.
[0246] Each such program may be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or object-oriented programming languages) to communicate with the computer system. In any case, the language may be a compiled language or an interpreted language.
[0247] For example, when implemented by a sequence of computer software instructions, the various functions and steps of the embodiments of the present invention can be implemented by a sequence of multithreaded software instructions running in appropriate digital signal processing hardware. In this case, the various apparatuses, steps, and functions of the embodiments can correspond to portions of the software instructions.
[0248] Each such computer program is preferably stored on or downloaded to a storage medium or device readable by a general or special purpose programmable computer (e.g., solid state memory or medium, magnetic medium, or optical medium). When the storage medium or device is read by a computer system to perform the processes described herein, it is used to configure and operate the computer. The system of the present invention can also be implemented as a computer-readable storage medium configured with (e.g., storing) a computer program, wherein the storage medium so configured causes the computer system to operate in a specific and predefined manner to perform the functions described herein.
[0249] Numerous embodiments of the present invention have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present invention. Given the above teachings, numerous modifications and variations of the present invention are possible. It should be understood that the present invention can be practiced differently within the scope of the appended claims than specifically described herein.
[0250] The present invention also includes the following solutions:
[0251] Solution 1. An audio processing unit, comprising:
[0252] A buffer memory; and
[0253] At least one processing subsystem coupled to the buffer memory, wherein the buffer memory stores at least one frame of an encoded audio bitstream, the frame including program information metadata or substream structure metadata in at least one metadata segment of at least one skip field of the frame and audio data in at least one other segment of the frame, wherein the processing subsystem is coupled and configured to perform at least one of generation of the bitstream, decoding of the bitstream, or adaptive processing of the audio data of the bitstream using the metadata of the bitstream, or perform at least one of authentication or verification of at least one of the audio data or metadata of the bitstream using the metadata of the bitstream,
[0254] wherein the metadata segment includes at least one metadata payload, the metadata payload including:
[0255] A header; and
[0256] At least a part of the program information metadata or at least a part of the substream structure metadata following the header.
[0257] Solution 2. The audio processing unit according to Solution 1, wherein the encoded audio bitstream indicates at least one audio program, and the metadata segment includes a program information metadata payload, and the program metadata payload includes:
[0258] A program information metadata header; and
[0259] After the program information metadata header, program information metadata indicating at least one attribute or characteristic of the audio content of the program, the program information metadata including active channel metadata indicating each non - silent channel and each silent channel of the program.
[0260] Solution 3. The audio processing unit according to Solution 2, wherein the program information metadata further includes one of the following:
[0261] Down - mixing processing status metadata, which indicates: whether the program has been down - mixed, and the type of down - mixing applied to the program in the case where the program has been down - mixed;
[0262] Up - mixing processing status metadata, which indicates: whether the program has been up - mixed, and the type of up - mixing applied to the program in the case where the program has been up - mixed;
[0263] Pre - processing status metadata, which indicates: whether pre - processing has been performed on the audio content of the frame, and the type of pre - processing performed on the audio content in the case where pre - processing has been performed on the audio content of the frame; or
[0264] Spectral expansion processing or channel coupling metadata, which indicates: whether spectral expansion processing or channel coupling has been applied to the program, and the frequency range to which spectral expansion or channel coupling is applied in the case where spectral expansion processing or channel coupling has been applied to the program.
[0265] Solution 4. The audio processing unit according to Solution 1, wherein the encoded audio bitstream indicates at least one audio program having at least one independent sub - stream with audio content, and the metadata segment includes a sub - stream structure metadata payload, and the sub - stream structure metadata payload includes:
[0266] A sub - stream structure metadata payload header; and
[0267] After the sub - stream structure metadata payload header, independent sub - stream metadata indicating the number of independent sub - streams of the program, and dependent sub - stream metadata indicating whether each independent sub - stream of the program has at least one associated dependent sub - stream.
[0268] Solution 5. The audio processing unit according to Solution 1, wherein the metadata segment includes:
[0269] A metadata segment header;
[0270] At least one protection value after the metadata segment header, which is used for at least one of decryption, authentication or verification of the program information metadata, or the sub-stream structure metadata, or the audio data corresponding to the program information metadata or the sub-stream structure metadata; and
[0271] A metadata payload identification value and a payload configuration value after the metadata segment header, wherein the metadata payload is after the metadata payload identification value and the payload configuration value.
[0272] Solution 6. The audio processing unit according to Solution 5, wherein the metadata segment header includes a synchronization word indicating the start of the metadata segment, and at least one identification value after the synchronization word, and the header of the metadata payload includes at least one identification value.
[0273] Solution 7. The audio processing unit according to Solution 1, wherein the encoded audio bitstream is an AC-3 bitstream or an E-AC-3 bitstream.
[0274] Solution 8. The audio processing unit according to Solution 1, wherein the buffer memory stores the frame in a non-transitory manner.
[0275] Solution 9. The audio processing unit according to Solution 1, wherein the audio processing unit is an encoder.
[0276] Solution 10. The audio processing unit according to Solution 9, wherein the processing subsystem includes:
[0277] A decoding subsystem configured to receive an input audio bitstream and extract input metadata and input audio data from the input audio bitstream;
[0278] An adaptive processing subsystem coupled and configured to perform adaptive processing on the input audio data using the input metadata, thereby generating processed audio data; and
[0279] An encoding subsystem coupled and configured to generate the encoded audio bitstream in response to the processed audio data, including by including the program information metadata or the sub-stream structure metadata in the encoded audio bitstream, and setting the encoded audio bitstream to the buffer memory.
[0280] Solution 11. The audio processing unit according to Solution 1, wherein the audio processing unit is a decoder.
[0281] Solution 12. The audio processing unit according to Solution 11, wherein the processing subsystem is a decoding subsystem coupled to the buffer memory and configured to extract the program information metadata or the substream structure metadata from the encoded audio bitstream.
[0282] Solution 13. The audio processing unit according to Solution 1, comprising:
[0283] A subsystem coupled to the buffer memory and configured to: extract the program information metadata or the substream structure metadata from the encoded audio bitstream, and extract the audio data from the encoded audio bitstream; and
[0284] A post-processor coupled to the subsystem and configured to perform adaptive processing on the audio data using at least one of the program information metadata or the substream structure metadata extracted from the encoded audio bitstream.
[0285] Solution 14. The audio processing unit according to Solution 1, wherein the audio processing unit is a digital signal processor.
[0286] Solution 15. The audio processing unit according to Solution 1, wherein the audio processing unit is a pre-processor configured to extract the program information metadata or the substream structure metadata and the audio data from the encoded audio bitstream, and perform adaptive processing on the audio data using at least one of the program information metadata or the substream structure metadata extracted from the encoded audio bitstream.
[0287] Solution 16. A method for decoding an encoded audio bitstream, the method comprising the steps of:
[0288] Receiving an encoded audio bitstream; and
[0289] Extracting metadata and audio data from the encoded audio bitstream, wherein the metadata is or includes program information metadata and substream structure metadata,
[0290] Wherein, the encoded audio bitstream includes a series of frames and indicates at least one audio program, the program information metadata and the substream structure metadata indicate the program, each of the frames includes at least one audio data segment, each of the audio data segments includes at least a part of the audio data, each frame in at least one subset of the frames includes a metadata segment, and each of the metadata segments includes at least a part of the program information metadata and at least a part of the substream structure metadata.
[0291] Solution 17. The method according to Solution 16, wherein the metadata segment includes a program information metadata payload, and the program information metadata payload includes:
[0292] A program information metadata header; and
[0293] Program information metadata indicating at least one attribute or characteristic of the audio content of the program after the program information metadata header, the program information metadata including active channel metadata indicating each non - silent channel and each silent channel of the program.
[0294] Solution 18. The method according to Solution 17, wherein the program information metadata further includes at least one of the following:
[0295] Down - mixing processing status metadata, which indicates: whether the program is down - mixed, and the type of down - mixing applied to the program in the case where the program is down - mixed;
[0296] Up - mixing processing status metadata, which indicates: whether the program is up - mixed, and the type of up - mixing applied to the program in the case where the program is up - mixed; or
[0297] Pre - processing status metadata, which indicates: whether pre - processing has been performed on the audio content of the frame, and the type of pre - processing performed on the audio content in the case where pre - processing has been performed on the audio content of the frame.
[0298] Solution 19. The method according to Solution 16, wherein the encoded audio bitstream indicates at least one audio program having at least one independent substream of audio content, and the metadata segment includes a substream structure metadata payload, and the substream structure metadata payload includes:
[0299] A substream structure metadata payload header; and
[0300] Independent sub-stream metadata indicating the number of independent sub-streams of the program and dependent sub-stream metadata indicating whether each independent sub-stream of the program has at least one associated dependent sub-stream, after the sub-stream structure metadata payload header.
[0301] Scheme 20. The method according to Scheme 16, wherein the metadata segment comprises:
[0302] A metadata segment header;
[0303] At least one protection value after the metadata segment header, for at least one of decryption, authentication or verification of the program information metadata or the sub-stream structure metadata or the audio data corresponding to the program information metadata and the sub-stream structure metadata; and
[0304] A metadata payload after the metadata segment header, comprising at least a part of the program information metadata and at least a part of the sub-stream structure metadata.
[0305] Scheme 21. The method according to Scheme 16, wherein the encoded audio bitstream is an AC-3 bitstream or an E-AC-3 bitstream.
[0306] Scheme 22. The method according to Scheme 16, further comprising the step of:
[0307] Performing adaptive processing on the audio data using at least one of the program information metadata or the sub-stream structure metadata extracted from the encoded audio bitstream.
Claims
1. An audio processing unit, comprising: One or more processors; A memory coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: Receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata includes loudness processing status metadata, and wherein the loudness processing status metadata includes metadata indicating the loudness of the audio program; Decoding the encoded audio data to obtain decoded audio data of the set of audio channels; Obtaining the loudness processing status metadata from the metadata of the encoded audio bitstream; and Performing adaptive loudness processing on the decoded audio data of the set of audio channels based on the loudness processing status metadata, wherein the metadata further includes program information metadata indicating a compression profile used to create dynamic range compression (DRC) data in the bitstream, and wherein the compression profile is a voice compression profile.
2. A method performed by an audio processing unit, comprising: Receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata includes loudness processing status metadata, and wherein the loudness processing status metadata includes metadata indicating the loudness of the audio program; Decoding the encoded audio data to obtain decoded audio data of the set of audio channels; Obtaining the loudness processing status metadata from the metadata of the encoded audio bitstream; And Performing adaptive loudness processing on the decoded audio data of the set of audio channels based on the loudness processing status metadata, wherein the metadata further includes program information metadata indicating a compression profile used to create dynamic range compression (DRC) data in the bitstream, and wherein the compression profile is a voice compression profile.
3. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: Receiving an encoded audio bitstream including an audio program, the encoded audio bitstream including encoded audio data of a set of one or more audio channels and metadata associated with the set of audio channels, wherein the metadata includes loudness processing status metadata, and wherein the loudness processing status metadata includes metadata indicating the loudness of the audio program; Decoding the encoded audio data to obtain decoded audio data of the set of audio channels; Obtaining the loudness processing status metadata from the metadata of the encoded audio bitstream; And Perform adaptive loudness processing on the decoded audio data of the set of audio channels based on the loudness processing status metadata, wherein the metadata further includes program information metadata that indicates a compression profile used to create dynamic range compression (DRC) data in the bitstream, and wherein the compression profile is a voice compression profile.
Citation Information
Patent Citations
Audio metadata verification
CN101160616A
System for combining loudness measurements in a single playback mode
CN102792588A