Adaptive processing by multiple media processing nodes
By using media processing state metadata to adaptively process media data, the enhanced media processing chain optimizes operations across distributed units, improving rendering quality and compliance.
Patent Information
- Application Number
- JP2025186032
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2011-11-10
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-10
AI Technical Summary
Media processing units operate in a blind manner, leading to unnecessary processing and degradation when distributed across a diverse network or cascade, as they do not consider the processing history of media data.
Media processing units in an enhanced chain adaptively process media data based on media processing state metadata, which includes information about previous processing, ensuring efficient and predictable media rendering by avoiding redundant operations.
The adaptive processing enhances media rendering quality by optimizing media processing across multiple units, maintaining compatibility with legacy systems, and ensuring regulatory compliance.
Smart Images

Figure 2026021488000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference to related applications and priority claims This application claims priority to U.S. Provisional Application No. 61 / 419,747, filed December 3, 2010, and U.S. Provisional Application No. 61 / 558,286, filed November 10, 2011, both of which are incorporated by reference in their entirety for all purposes.
[0002] technology The present invention relates generally to media processing systems, and more particularly to adaptively processing media data based on the media processing state of the media data. [Background technology]
[0003] Media processing units typically operate in a blind manner, paying no attention to the processing history of media data that occurred before the media data was received. This may work in a media processing framework in which a single entity performs all media processing and encoding for various target media rendering devices, which in turn perform all decoding and rendering of the encoded media data. However, this blind processing does not work well (or at all) in situations in which multiple media processing units are distributed across a diverse network or arranged in cascade (chain) and are expected to optimally perform each type of media processing. For example, some media data may be encoded for a high-performance media system and need to be converted to a reduced form suitable for mobile devices in the media processing chain. Thus, a media processing unit may unnecessarily perform a type of processing on the media data that has already been performed. For example, a volume leveling unit may perform processing on an input audio clip regardless of whether volume leveling has previously been performed on the input audio clip. As a result, the volume leveling unit performs leveling even when it is not necessary, and this unnecessary processing may result in the degradation and / or removal of certain characteristics when rendering the media content in the media data.
[0004] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise noted, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Likewise, it should not be assumed that problems identified with one or more approaches have been recognized by the prior art based on this section unless otherwise noted. Summary of the Invention [Problem to be solved by the invention]
[0005] Reduce or eliminate problems of the prior art. [Means for solving the problem]
[0006] This problem is solved by the means described in the claims. [Brief explanation of the drawings]
[0007] The present invention is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings in which like reference symbols refer to similar elements and in which: [Figure 1] FIG. 1 illustrates an exemplary media processing chain, in accordance with some possible embodiments of the present invention. [Figure 2] FIG. 1 illustrates an exemplary enhanced media processing chain in accordance with some possible embodiments of the present invention. [Figure 3] FIG. 1 illustrates an exemplary encoder / transcoder according to some possible embodiments of the present invention. [Figure 4] FIG. 2 illustrates an exemplary decoder according to some possible embodiments of the present invention. [Figure 5] FIG. 2 illustrates an exemplary post-processing unit, according to some possible embodiments of the present invention. [Figure 6] FIG. 2 illustrates an exemplary implementation of an encoder / transcoder in accordance with some possible embodiments of the present invention. [Figure 7] FIG. 10 illustrates an exemplary evolutionary decoder that controls the operation mode of a volume leveling unit based on the validity of loudness metadata in and / or associated with processing state metadata, according to some possible embodiments of the present invention. [Figure 8] FIG. 2 illustrates an exemplary configuration using data hiding to pass media processing information, according to some possible embodiments of the present invention. [Figure 9] 1A and 1B illustrate an exemplary process flow according to one possible embodiment of the present invention. [Figure 10] FIG. 1 illustrates an exemplary hardware platform upon which the computers or computing devices described herein may be implemented, in accordance with one possible embodiment of the present invention. [Figure 11] FIG. 1 illustrates a media frame along with processing state metadata associated with media data in the media frame, according to an example embodiment. [Figure 12A] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12B] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12C] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12D] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12E] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12F] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12G] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12H] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12I] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12J] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12K] 1 is a block diagram of an exemplary media processing node / device according to an embodiment of the present invention. [Figure 12L-1] 1 is a portion of a block diagram of an exemplary media processing node / device, according to an embodiment of the present invention. [Figure 12L-2] 1 is a portion of a block diagram of an exemplary media processing node / device, according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0008] Exemplary possible embodiments for adaptively processing media data based on media processing states of the media data are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without such specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily obscuring, burying, or obscuring the present invention.
[0009] Exemplary embodiments are now described according to the following outline.
[0010] 1. General Overview 2. Media Processing Chain 3. Media processing device or unit 4. Exemplary Adaptive Processing of Media Data 5. Data Hiding 6. Exemplary Process Flow 7. Implementation mechanism - Hardware overview 8. Numbering Example 9. Equivalents, Extensions, Substitutions, etc.
[0011] 1. General Overview This overview provides a basic description of some aspects of possible embodiments of the present invention. It should be noted that this overview is not a comprehensive or exhaustive summary of possible embodiments. Furthermore, it should be noted that this overview is not intended to identify any particularly significant aspects or elements of possible embodiments, nor is it intended to be understood as delineating the scope of possible embodiments or the present invention in particular in general. This overview merely presents some concepts related to exemplary possible embodiments in a simplified form, and should merely be understood as a conceptual prelude to the more detailed description of exemplary possible embodiments that is presented below.
[0012] Techniques for adaptive processing of media data based on a media processing state of the media data are described. In some possible embodiments, media processing units in an enhanced media processing chain are enabled to automatically obtain and examine media processing signals and / or processing state metadata, determine the state of the media data based on the media processing signals and / or processing state metadata, and adapt their processing based on the state of the media data. Media processing units in an enhanced media processing chain may include, but are not limited to, encoders, transcoders, decoders, pre-processing units, post-processing units, bitstream processing tools, Advanced Television Systems Committee (ATSC) codecs, Moving Picture Experts Group (MPEG) codecs, etc. The media processing units may be a media processing system or part of a media processing system.
[0013] As used herein, the term "processing state metadata" refers to data that is distinct from media data. Media data (e.g., video frames, perceptually coded audio frames, or PCM audio samples that comprise media content) refers to media sample data that represents media content and is used to render the media content as audio or video output. Processing state metadata is associated with media data and specifies what types of processing have already been performed on the media data. This association of processing state metadata with media data is time-synchronous. Thus, current processing state metadata indicates that the current media data contemporaneously contains the results of the indicated type of media processing and / or a description of the media features in the media data. In some possible embodiments, processing state metadata may include some or all of the processing history and / or parameters used in and / or derived from the indicated type of media processing. Additionally and / or optionally, the processing state metadata may include one or more different types of media features computed / extracted from the media data. Media features, as described herein, provide a semantic description of the media data and may include one or more of structural attributes, tonality including harmony and melody, timbre, rhythm, reference loudness, stereo mix or source of a certain amount of the media data, absence or presence of voices, repetition characteristics, melody, harmony, lyrics, timbre, perceptual features, digital media features, stereo parameters, voice recognition (e.g., what the speaker is saying), etc. The processing state metadata may also include other metadata that is not related to or derived from any processing of the media data.For example, third-party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, etc. may be added by a particular media processing unit to be passed to other media processing units. These independent types of metadata may be distributed around, verified, and used by media processing components in the media processing chain. The term "media processing signaling" refers to relatively lightweight control or status data (which may be a small amount of data compared to the processing state metadata) communicated between media processing units in a media bitstream. Media processing signaling may contain a subset or summary of the processing state metadata.
[0014] The media processing signal and / or processing state metadata may be embedded in one or more reserved fields (which may be, but are not limited to, currently unused fields), carried in a substream in the media bitstream, hidden in the media data, or provided with a separate media processing database. In some possible embodiments, the amount of media processing signal and / or processing state metadata may be small enough to be carried without affecting the bitrate allocated to carrying the media data (e.g., in reserved fields, or hidden in media samples using reversible data hiding techniques, or by calculating or obtaining media fingerprints from the media data and storing detailed processing state information in an external database). Communication of media processing signals and / or processing state metadata in an enhanced media processing chain is particularly useful when two or more media processing units need to cooperate with each other in a cascading manner throughout the media processing chain (or content lifecycle). Without media processing signals and / or processing state metadata, serious media processing problems such as quality, level, and spatial degradation may be likely to occur, for example, when more than one audio codec is used in the chain and single-ended volume leveling is applied more than once during the media content's journey to the media consumption device (or rendering point of the media content in the media data).
[0015] In contrast, the present techniques enhance the intelligence of any or all of the media processing units in the media processing chain (content lifecycle). Under the present techniques, any of these media processing units can "listen and adapt" and "announce" the state of the media data to downstream media processing units. Thus, under the present techniques, downstream media processing units may optimize their processing of media data based on knowledge of past processing of the media data performed by one or more upstream media processing units. Under the present techniques, media processing of the media data by the entire media processing chain becomes more efficient, more adaptive, and more predictable than would otherwise be the case. As a result, the overall rendering and handling of the media content in the media data is much improved.
[0016] Importantly, under the techniques described herein, the presence of the media data state indicated by the media processing signal and / or processing state metadata does not negatively impact legacy media processing units that may be present in the enhanced media processing chain. Legacy media processing units cannot proactively use the media data state to adaptively process the media data themselves. Furthermore, even if legacy media processing units in the media processing chain tend to tamper with the processing results of other upstream media processing devices, the processing state metadata described herein can be safely and securely passed to downstream media processing devices through secure communication methods that utilize cryptographic values, encryption, authentication, and data hiding. Examples of data hiding include both reversible and irreversible data hiding.
[0017] In some possible embodiments, to communicate the state of the media data to downstream media processing units, the techniques herein encapsulate and / or embed one or more processing subunits in the form of software, hardware, or both within the media processing unit so that the media processing unit can read, write, and / or verify the processing state metadata delivered with the media data.
[0018] In some possible embodiments, a media processing unit (e.g., encoder, decoder, leveler, etc.) may receive media data on which one or more types of media processing have previously been performed. However, 1) processing state metadata indicating the types of media processing previously performed may be absent and / or 2) the processing state metadata may be incorrect or incomplete. The types of media processing previously performed may include operations that may modify media samples (e.g., volume leveling) and operations that may not modify media samples (e.g., fingerprint extraction and / or feature extraction based on media samples). The media processing unit may be configured to automatically generate "correct" processing state metadata that reflects the "true" state of the media data and associate this state with the media data by communicating the generated processing state metadata to one or more downstream media processing units. Furthermore, the association of the media data with the processing state metadata may be performed in a manner that ensures that the resulting media bitstream is backward compatible with legacy media processing units, such as legacy decoders. As a result, legacy decoders that do not implement the techniques of this document may still correctly decode the media data while ignoring the associated processing state metadata that indicates the state of the media data, as the legacy decoder was designed to do. In some possible embodiments, the media processing unit of this document may also be configured with the parallel functionality of verifying the processing state metadata with the (source) media data via forensic analysis and / or verification of one or more embedded hash values (e.g., signatures).
[0019] Under the techniques described herein, adaptive processing of media data based on the contemporaneous state of the media data indicated by received processing state metadata may be performed at various points in the media processing chain. For example, if loudness metadata in the processing state metadata is valid, a volume leveling unit subsequent to the decoder may be notified by the decoder of the media processing signal and / or the processing state metadata, allowing the volume leveling unit to pass media data, such as audio, through unchanged.
[0020] In some embodiments, the processing state metadata includes media features extracted from the underlying media sample. The media features may provide a semantic description of the media sample and may be provided as part of the processing state metadata to indicate, for example, whether the media sample contains speech, music, someone singing in quiet or noisy conditions, singing in the presence of a crowd of people talking, whether dialogue is occurring, speech against a noisy background, or a combination of two or more of these. Adaptive processing of the media data may be performed at various points in the media processing chain based on the descriptions of the media features included in the processing state metadata.
[0021] Under the techniques described herein, processing state metadata embedded in a media bitstream along with the media data may be authenticated and verified. For example, this may be useful for a loudness regulatory entity to verify that the loudness of a particular program is already within a specified range and that the media data itself has not been modified (thereby ensuring regulatory compliance). To verify this, rather than recalculating the loudness, the loudness value contained in the data block containing the processing state metadata may be read.
[0022] Under the techniques described herein, data blocks containing processing state metadata may include additional reserved bytes for securely carrying third-party metadata. This feature may be used to enable a variety of applications. For example, a ratings agency (e.g., Nielsen Media Research) may choose to include content identification tags that can be used to identify specific programs viewed or listened to for purposes of calculating ratings, audience demographics, or audience demographics.
[0023] Significantly, the techniques described herein and variations of the techniques described herein can ensure that processing state metadata associated with media data is preserved throughout the media processing chain, from content creation to content consumption.
[0024] In some possible embodiments, the mechanisms described herein are part of a media processing system, including but not limited to handheld devices, game consoles, televisions, laptop computers, netbook computers, cellular wireless telephones, e-readers, point-of-sale terminals, desktop computers, computer workstations, computer kiosks, and various other types of terminals and media processing units.
[0025] Various modifications to the preferred embodiments and generic principles and features described herein will be readily apparent to those skilled in the art, and thus the present disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.
[0026] 2. Media Processing Chain FIG. 1 illustrates an exemplary media processing chain according to some possible embodiments of the present invention. The media processing chain may include, but is not limited to, an encoder, a decoder, pre-processing / post-processing units, a transcoder, and a signal analysis and metadata correction unit. These units in the metadata processing chain may be included in the same system or in different systems. In embodiments where the media processing chain spans different systems, these systems may be co-located or geographically distributed.
[0027] In some possible embodiments, the preprocessing unit of Figure 1 may accept as input PCM (time domain) samples containing media content and output processed PCM samples, and the encoder may accept PCM samples as input and output an encoded (e.g., compressed) media bitstream of the media content.
[0028] As used herein, data containing media content (e.g., carried in the main stream of a bitstream) is referred to as media data, while data separate from the media data that indicates the type of processing that has been performed on the media data at any given point in the media processing chain is referred to as processing state metadata.
[0029] The signal analysis and metadata correction unit may accept as input one or more encoded media bitstreams and verify whether included processing state metadata in the encoded media bitstreams is correct by performing signal analysis. If the signal analysis and metadata correction unit finds that the included metadata is invalid, the signal analysis and metadata correction unit replaces the incorrect values with correct values obtained from the signal analysis.
[0030] A transcoder may accept a media bitstream as input and output a modified media bitstream. A decoder may accept a compressed media bitstream as input and output a stream of decoded PCM samples. A post-processing unit may accept a stream of decoded PCM samples, perform any post-processing, such as volume leveling, on the media content therein, and render the media content in the decoded PCM samples on one or more speakers and / or display panels. Not all media processing units need to be capable of using processing state metadata to adapt the processing applied to the media data.
[0031] The techniques provided herein provide an improved media processing chain in which media processing units, such as encoders, decoders, transcoders, pre-processing and post-processing units, etc., adapt their respective processing to be applied to media data according to the synchronic state of the media data as indicated by the media processing signals and / or processing state metadata they respectively receive.
[0032] FIG. 2 illustrates an exemplary enhanced media processing chain having an encoder, decoder, pre-processing / post-processing units, transcoder, and signal analysis and metadata correction units according to some possible embodiments of the present invention. Some or all of the units in FIG. 2 may be modified to adapt the processing of media data based on the state of the media data. In some possible embodiments, each media processing unit in this exemplary enhanced media processing chain is configured to cooperate in performing non-redundant media processing and avoiding unnecessary and erroneous repetition of processing performed by upstream units. In some possible embodiments, the state of the media data at any point in the enhanced media processing chain, from content generation to content consumption, is understood by the current media processing unit at that point in the enhanced media processing chain.
[0033] 3. Media processing device or unit Figure 3 illustrates an exemplary (modified) encoder / transcoder according to some possible embodiments of the present invention. Unlike the encoder of Figure 1, the encoder / transcoder of Figure 3 may be configured to receive processing state metadata associated with input media data and to determine previous (pre / post) processing performed on the input media data (e.g., input audio) by one or more upstream units relative to the encoder / transcoder. The input media data is received by the modified encoder / transcoder from a logically upstream unit (e.g., the last upstream unit that performed processing on the input audio).
[0034] As used herein, the term "logically receiving" may mean that an intermediate unit may or may not be involved in communicating input media data from an upstream unit (e.g., the last upstream unit mentioned above) to a receiving unit, such as the encoder / transcoder unit in the present example.
[0035] In one example, an upstream unit that performed pre- / post-processing on the input media data may be in a different system than the system that the receiving unit is part of. The input media data may be a media bitstream output by the upstream unit and conveyed through an intermediate transmission unit, such as a network connection, USB, wide area network connection, wireless connection, optical connection, etc.
[0036] In another example, an upstream unit that performed pre- or post-processing on input media data may be located in the same system of which the receiving unit is a part. The input media data may be output by the upstream unit and transmitted over an internal connection through one or more internal units of the system. For example, the data may be physically delivered over an internal bus, crossbar connection, serial connection, etc. In either case, under the techniques herein, the receiving unit may logically receive the input media data from the upstream unit.
[0037] In some possible embodiments, the encoder / transcoder is configured to generate or modify processing state metadata associated with the media data, which may be a modified version of the input media data. The new or modified processing state metadata generated or modified by the encoder / transcoder may automatically and accurately capture the state of the media data output by the encoder / transcoder further along the media processing chain. For example, the processing state metadata may include whether certain processing (e.g., Dolby Volume, commercially available from Dolby Laboratories, upmixing) has been performed on the media data. Additionally and / or optionally, the processing state metadata may include parameters used in and / or derived from certain processing or any constituent operations of such processing. Additionally and / or optionally, the processing state metadata may include one or more fingerprints calculated / extracted from the media data. Additionally and / or optionally, the processing state metadata may include one or more different types of media features calculated / extracted from the media data. The media features described herein provide a semantic description of the media data and may include one or more of structural attributes, tonality including harmony and melody, timbre, rhythm, reference loudness, stereo mix or source of a certain amount of media data, absence or presence of voices, repetition characteristics, melody, harmony, lyrics, timbre, perceptual features, digital media features, stereo parameters, voice recognition (e.g., what the speaker is saying), etc. In some embodiments, the extracted media features are utilized to classify the underlying media data into one or more of a plurality of media data classes.The one or more media data classes may include, but are not limited to, any of a single overall / dominant "class" (e.g., class type) for the entire media and / or a single class (e.g., class subtype for a subset / subinterval of the entire work) representing a smaller time period such as a single media frame, a media data block, multiple media frames, multiple media data blocks, a fraction of a second, a second, multiple seconds, etc. For example, a class label may be calculated and inserted into the bitstream and / or hidden (using lossless or lossy data hiding techniques) every 32 msec of the bitstream. A class label may be used to indicate one or more class types and / or one or more class subtypes. In a media data frame, a class label may be inserted into a metadata structure preceding or alternatively following the media data block with which it is associated. This is shown in Figure 11. A media class may include, but is not limited to, any single class type, such as music, speech, noise, silence, or applause. The media processing devices described herein may also be configured to classify media data containing a mixture of media class types, such as speech over music. Additionally, alternatively, and optionally, the media processing devices described herein may be configured to carry an independent "likelihood" or probability value for the media class type or subtype indicated by the computed media class label. One or more such likelihood or probability values may be transmitted along with the media class label in the same metadata structure. The likelihood or probability value indicates the level of "trust" that the computed media class label has in relation to the media segment / block whose media class type or subtype is indicated by the computed media class label.The one or more likelihood or probability values in combination with the associated media class label may be utilized by a destination media processing device to adapt media processing in a manner that improves any of a wide variety of operations throughout the media processing chain, such as upmixing, encoding, decoding, transcoding, headphone virtualization, etc. The processing state metadata may include, but is not limited to, any of the media class type or subtype and the likelihood or probability value. Additionally, optionally, or alternatively, instead of passing the media class type / subtype and likelihood / probability value in a metadata structure inserted between media (audio) data blocks, some or all of the media class type / subtype and likelihood / probability value may be embedded in the media data (or samples) as hidden metadata and passed to the destination media processing node / device. In some embodiments, the results of content analysis of the media data included in the processing state metadata may include one or more indications of whether certain user-defined or system-defined keywords are spoken in any time segment of the media data. One or more applications may use such indicators to trigger the performance of related actions (eg, presenting contextually relevant advertisements for products and services related to the keyword).
[0038] In some embodiments, while processing media data with a first processor, the apparatus described herein may run a second processor in parallel to classify / extract media features from the media data. Media features may be extracted from a segment lasting over a period of time (one frame, multiple frames, one second, multiple seconds, one minute, multiple minutes, a user-defined period of time, etc.), or alternatively for a scene (based on detectable signal characteristic changes). The media features described by the processing state metadata may be used throughout the media processing chain. Downstream devices may adapt their own media processing of the media data based on one or more of the media features. Alternatively, downstream devices may choose to ignore the presence of any or all of the media features described in the processing state metadata.
[0039] An application on a device in the media processing chain may use media features in one or more of a variety of ways. For example, such an application may use media features to index the underlying media data. For a user who wants to go to the section where the judges talk about the performance, the application may skip other preceding sections. The media features described in the processing state metadata provide downstream devices with contextual information about the media data as an inherent part of the media data.
[0040] Two or more devices in a media processing chain may perform analysis to extract media features from the content of the media data, thereby avoiding downstream devices having to analyze the content of the media data.
[0041] In one possible embodiment, the generated or modified processing state metadata may be transmitted as part of a media bitstream (e.g., an audio bitstream with metadata about the audio state), and transmission rates may be on the order of 3-10 kbps. In some embodiments, the processing state metadata may be transmitted within the media data (e.g., PCM media samples) based on data hiding. A wide variety of data hiding techniques, which may alter the media data reversibly or irreversibly, may be used to hide some or all of the processing state metadata (including, but not limited to, authentication-related data) within the media samples. Data hiding may be achieved by altering / manipulating / modulating the signal characteristics (phase and / or amplitude in the frequency or time domain) of the underlying media sample signals. Data hiding may be implemented based on FSK, spread spectrum, or other available methods.
[0042] In some possible embodiments, pre-processing / post-processing units may perform processing of media data in a cooperative manner with an encoder / transcoder, and the processing performed by the cooperating pre- and post-processing units is also specified in processing state metadata that is communicated to downstream media processing units (e.g., via the audio bitstream).
[0043] In some possible embodiments, once a piece of processing state metadata (which may include a media fingerprint and any parameters used in or derived from one or more types of media processing) is derived, this piece of processing state metadata may be stored by media processing units in the media processing chain and passed on to all downstream units. Thus, in some possible embodiments, a piece of processing state metadata may be generated by the first media processing unit in the media processing chain (entire lifecycle) and passed on to the last media processing unit, either as embedded data within a media bitstream / substream or as data derivable from an external data source or media processing database.
[0044] 4 illustrates an exemplary decoder (e.g., an evolutionary decoder implementing the techniques of this document) according to some possible embodiments of the present invention. The decoder of some possible embodiments of the present invention may be configured to (1) parse and verify processing state metadata (e.g., processing history, media feature descriptions, etc.) and other metadata (e.g., third-party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, etc.) associated with incoming media data passed therethrough, and (2) determine a media processing state for the media data based on the verified processing state metadata. For example, by parsing and verifying the processing state metadata in a media bitstream carrying input media data and processing state metadata (e.g., an audio bitstream with metadata about the audio state), the decoder may determine that the loudness metadata (or media feature metadata) is valid, reliable, and generated by one of the enhancement content provider subunits (e.g., a Dolby media generator (DMG) commercially available from Dolby Laboratories) that implements the techniques described herein. In some possible embodiments, in response to determining that the received processing state metadata is valid and reliable, the decoder may then be configured to generate a media processing signal about the state of the media data using a lossy or lossy data hiding technique based, at least in part, on the received processing state metadata. The decoder may be configured to provide the media processing signal to a downstream media processing unit (e.g., a post-processing unit) in the media processing chain. This type of signal may be used, for example, when there is no dedicated (and synchronous) metadata path between the decoder and the downstream media processing unit.This situation may arise in some possible embodiments where the decoder and the downstream media processing unit exist as separate entities in a consumer electronics device, or in different subsystems or systems where a synchronous control or data path between the decoder and the subsequent processing unit is not available. In some possible embodiments, the media processing signal under the data hiding techniques herein may be transmitted as part of a media bitstream, and transmission rates may be on the order of 16 bits per second. A wide variety of data hiding techniques that can reversibly or irreversibly modify the media data may be used to hide some or all of the processing state metadata in the media samples. Data hiding techniques include, but are not limited to, perceptible or imperceptible secure communication channels, modifying / manipulating / modulating narrowband or spread-spectrum signal characteristics (phase and / or amplitude in the frequency or time domain) of one or more signals of the underlying media samples, or other available methods.
[0045] In some possible embodiments, the decoder may not attempt to pass on all received processing state metadata, but rather embed only enough information (e.g., within the limits of data hiding capacity) to change the operating mode of downstream media processing units based on the state of the media data.
[0046] In some possible embodiments, redundancy in audio or video signals in the media data may be exploited to carry the state of the media data. In some possible embodiments, some or all of the media processing signal and / or processing state metadata may be hidden in the least significant bits (LSBs) of multiple bytes in the media data or in a secure communication channel carried within the media data without causing audible or visible artifacts. The multiple bytes may be selected based on one or more factors or criteria, including whether the LSBs may cause perceptible or visible artifacts when the media sample with the hidden data is rendered by a legacy media processing unit. Other data hiding techniques that may reversibly or irreversibly alter the media data (e.g., perceptible or imperceptible secure communication channels, FSK-based data hiding techniques, etc.) may also be used to hide some or all of the processing state metadata in the media samples.
[0047] In some possible embodiments, data hiding techniques may be optional or not required, for example, if the downstream media processing unit is implemented as part of the decoder. For example, two or more media processing units may share a bus or other communication mechanism that allows metadata to be passed as an out-of-the-band signal from one media processing unit to another without hiding the data in the media samples.
[0048] FIG. 5 illustrates an exemplary post-processing unit (e.g., a Dolby Evolution post-processing unit) according to some possible embodiments of the present invention. The post-processing unit may be configured to first extract a media processing signal hidden in media data (e.g., PCM audio samples with embedded information) and determine the state of the media data indicated by the media processing signal. This may be done, for example, using an adjunct processing unit (e.g., an information extraction and audio restoration subunit in some possible embodiments in which the media data includes audio). In embodiments in which the media processing signal is hidden using a lossless data hiding technique, previous modifications performed on the media data by the data hiding technique (e.g., a decoder) to embed the media processing signal may be undone. In embodiments in which the media processing signal is hidden using an irreversible data hiding technique, previous modifications performed on the media data by the data hiding technique (e.g., a decoder) to embed the media processing signal may not be completely undone, but side effects on the quality of the media rendering may be minimized (e.g., minimal audio or visual artifacts). Then, based on the state of the media data indicated by the media processing signal, the post-processing unit may be configured to adapt its processing to be applied to the media data. In one example, volume processing may be turned off in response to a determination (from the media processing signal) that loudness metadata is valid and volume processing has been performed by an upstream unit. In another example, speech-recognized keywords may cause contextually relevant advertisements or messages to be presented or triggered.
[0049] In some possible embodiments, a signal analysis and metadata correction unit in a media processing system described herein may be configured to accept an encoded media bitstream as input and verify whether embedded metadata in the media bitstream is correct by performing signal analysis. After verifying that the embedded metadata in the media bitstream is valid or not, corrections may be applied, if necessary. In some possible embodiments, the signal analysis and metadata correction unit may be configured to perform analysis on media data or samples encoded in the input media bitstream in the time and / or frequency domain to determine media characteristics of the media data. After determining the media characteristics, corresponding processing state metadata (e.g., a description of one or more media characteristics) may be generated and provided to a device downstream of the signal analysis and metadata correction unit. In some possible embodiments, the signal analysis and metadata correction unit may be integrated with one or more other media processing units in one or more media processing systems. Additionally and / or optionally, the signal analysis and metadata correction unit may be configured to hide the media processing signal in the media data and signal to downstream units (encoders / transcoders / decoders) that the embedded metadata in the media data is valid and has been successfully verified. In some possible embodiments, the signaling data and / or processing state metadata associated with the media data may be generated and inserted into a compressed media bitstream carrying the media data.
[0050] Thus, the techniques described herein ensure that various processing blocks or media processing units (e.g., encoders, transcoders, decoders, pre-processing / post-processing units, etc.) in an enhanced media processing chain can determine the state of the media data. Each media processing unit can then adapt its processing according to the state of the media data indicated by the upstream unit. Furthermore, one or more lossless or lossy data hiding techniques may be used to ensure that signaling information about the state of the media data can be provided to downstream media processing units in an efficient manner that minimizes the required bitrate for transmitting the signaling information to the downstream media processing units. This is particularly useful when there is no metadata path between an upstream unit, such as a decoder, and a downstream unit, such as a post-processing unit, for example, when the post-processing unit is not part of the decoder.
[0051] In some possible embodiments, the encoder may be enhanced by or may include a preprocessing and metadata validation subunit. In some possible embodiments, the preprocessing and metadata validation subunit may be configured to ensure that the encoder performs adaptive processing of the media data based on the state of the media data indicated by the media processing signal and / or the processing state metadata. In some possible embodiments, through the preprocessing and metadata validation subunit, the encoder may be configured to validate processing state metadata associated with the media data (e.g., included in the media bitstream together with the media data). For example, if the metadata is verified to be authentic, results from a previously performed type of media processing may be reused, and new execution of that type of media processing may be avoided. On the other hand, if the metadata is found to be tampered with, the previously performed type of media processing may be repeated by the encoder. In some possible embodiments, once the processing state metadata (including metadata retrieval based on the media processing signal and fingerprints) is found to be unreliable, additional types of media processing may be performed on the metadata by the encoder.
[0052] If the processing state metadata is determined to be valid (e.g., based on a match between the extracted cryptographic value and the reference cryptographic value), the encoder may be configured to signal to other media processing units in the enhanced media processing chain, e.g., that the processing state metadata present in the media bitstream is valid. Any, some, or all of a variety of approaches may be implemented by the encoder.
[0053] Under the first approach, an encoder may insert a flag (e.g., an "evolution flag") into an encoded media bitstream to indicate that processing state metadata validation has already been performed on this encoded media bitstream. This flag may be inserted in a manner such that the presence of the flag does not affect "legacy" media processing units, such as decoders that are not configured to process and utilize processing state metadata as described herein. In one exemplary embodiment, an Audio Compression-3 (AC-3) encoder may be enhanced with a pre-processing and metadata validation sub-unit that sets an "evolution flag" in the xbsi2 field of the AC-3 media bitstream as specified in the ATSC standard (e.g., ATSC A / 52b). This "bit" may or may not be present in all coded frames carried in the AC-3 media bitstream. In some possible embodiments, the presence of this flag in the xbsi2 field does not affect deployed "legacy" decoders that are not configured to process and utilize processing state metadata as described herein.
[0054] Under the first approach above, there may be problems authenticating the information in the xbsi2 field: for example, a (e.g., malicious) upstream unit could turn the xbsi2 field "on" without actually verifying the processing state metadata, falsely signaling to other downstream units that the processing state metadata is valid.
[0055] To solve this problem, some embodiments of the present invention may use a second approach. To embed the "evolution flag," a secure data hiding method (including, but not limited to, any of several data hiding methods that create a secure communication channel within the media data itself, such as spread-spectrum-based methods, FSK-based methods, and methods based on other secure communication channels) may be used. This secure method is configured to prevent the "evolution flag" from being passed in the clear and therefore from being easily attacked, intentionally or unintentionally, by a unit or an intruder. Instead, under this second approach, the downstream unit may obtain the hidden data in encrypted form. Through a decryption and authentication subprocess, the downstream unit may verify the accuracy of the hidden data and trust the "evolution flag" in the hidden data. As a result, the downstream unit may determine that the processing state metadata in the media bitstream has previously been successfully verified. In various embodiments, any portion of the processing state metadata, such as the "evolution flag," may be delivered by an upstream device to a downstream device in any of one or more cryptographic methods (HMAC-based, non-HMAC-based).
[0056] In some possible embodiments, the media data may initially be a legacy media bitstream containing, for example, PCM samples. However, once the media data has been processed by one or more media processing units as described herein, processing state metadata generated by the one or more media processing units includes the state of the media data as well as relatively detailed information that can be used to decode the media data (including, but not limited to, any of one or more media characteristics determined from the media data). In some possible embodiments, the generated processing state metadata may include a media fingerprint, such as a video fingerprint, loudness metadata, dynamic range metadata, one or more hash-based message authentication codes (HMACs), one or more dialogue channels, an audio fingerprint, an enumerated processing history, audio loudness, dialogue loudness, true peak values, sample peak values, and / or any user- (third-party) specified metadata. The processing state metadata may also include an "evolution data block."
[0057] As used herein, the term "enhanced" refers to the ability of a media processing unit under the techniques described herein to cooperate with other media processing units or other media processing systems under the techniques described herein in a manner that can perform adaptive processing based on the state of the media data set by an upstream unit. The term "evolution" refers to the ability of a media processing unit under the techniques described herein to function in a manner compatible with legacy media processing units or legacy media processing systems, as well as the ability of a media processing unit under the techniques described herein to cooperate with other media processing units or other media processing systems under the techniques described herein in a manner that can perform adaptive processing based on the state of the media data set by an upstream unit.
[0058] In some possible embodiments, a media processing unit described herein may receive media data on which one or more types of media processing have been performed. However, there may be no or insufficient metadata associated with the media data indicating the one or more types of media processing. In some possible embodiments, such a media processing unit may be configured to generate processing state metadata indicating the one or more types of media processing performed by other units upstream from the media processing unit. Feature extraction not performed by the upstream device may also be performed and the processing state metadata may be forwarded to downstream devices. In some possible embodiments, a media processing unit (e.g., an evolutionary encoder / transcoder) may have a media forensic analysis subunit. The media forensic analysis subunit, such as an audio forensic subunit, may be configured to determine (without received metadata) whether certain types of processing have been performed on a piece of media content or media data. The analysis subunit may be configured to look for specific signal processing artifacts / signatures introduced and left behind by the certain types of processing. The media forensic subunit may be configured to determine whether certain types of feature extraction have been performed on a piece of media content or media data. The parsing subunit may be configured to look for the specific occurrence of feature-based metadata. For purposes of the present invention, the media forensic parsing subunit described herein may be implemented by any media processing unit in a media processing chain. Furthermore, processing state metadata generated by a media processing unit via the media forensic parsing subunit may therein be delivered to downstream units in the media processing chain.
[0059] In some possible embodiments, the processing state metadata described herein may include additional reserved bytes to support third-party applications. The additional reserved bytes may be guaranteed to be secure by assigning a separate encryption key to scramble any plaintext carried in one or more fields within those reserved bytes. Embodiments of the present invention support novel applications, including content identification and tracking. In one example, media bearing Nielsen ratings may carry a unique identifier for a program in the media bitstream. The Nielsen ratings may then use this unique identifier to calculate viewer or listener statistics for that program. In another example, the reserved bytes may carry keywords for a search engine such as Google. Google may then associate advertisements based on the keywords contained in one or more keyword-bearing fields in the reserved bytes. For purposes of the present invention, in applications such as those discussed herein, the techniques herein may be used to ensure that the reserved bytes are secure and cannot be decrypted by anyone other than a third party designated to use one or more fields of the reserved bytes.
[0060] The processing state metadata described herein may be associated with media data in any of a number of different ways. In some possible embodiments, the processing state metadata may be inserted into the outgoing compressed media bitstream carrying the media data. In some embodiments, the metadata is inserted in a manner that maintains backward compatibility with legacy decoders that are not configured to perform adaptive processing based on the processing state metadata described herein.
[0061] 4. Exemplary Adaptive Processing of Media Data FIG. 6 illustrates an exemplary implementation of an encoder / transcoder according to some possible embodiments of the present invention. Any of the depicted components may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and / or one or more integrated circuit (IC) circuits (including ASICs, FPGAs, etc.). The encoder / transcoder may also include several legacy subunits, such as front-end decode (FED), back-end decode (full mode) that may or may not perform dynamic range control / dialog norm (DRC / Dialnorm) processing based on whether such processing has already been performed, a DRC generator (DRC Gen), back-end encode (BEE), a stuffer, a CRC regeneration unit, etc. Using these legacy subunits, the encoder / transcoder can convert a bitstream (e.g., but not limited to, AC-3) into another bitstream containing the results of one or more types of media processing (e.g., but not limited to, AC-3 with adaptive and automated loudness processing). However, media processing (e.g., loudness processing) may be performed regardless of whether the loudness processing was previously performed and / or whether processing state metadata is present in the input bitstream. Thus, an encoder / transcoder with only legacy subunits may perform incorrect or unnecessary media processing.
[0062] Under the techniques described herein, in some possible embodiments shown in FIG. 6, the encoder / transcoder may include any of several new sub-units, such as a media data parser / verifier (which may be, for example, but is not limited to, an AC-3 flag parser and verifier), auxiliary processing units (e.g., adaptive transform-domain real-time loudness and dynamic range controllers, signal analysis, feature extraction, etc.), media fingerprint generation (e.g., audio fingerprint generation), metadata generators (e.g., evolutionary data generators and / or other metadata generators), media processing signal insertion (e.g., “add_bsi” insertion or insertion into ancillary data fields), HMAC generators (which may digitally sign one or more, or even all, frames to prevent tampering by malicious or legacy entities), one or more of other types of cryptographic processing units, one or more switches that operate based on processing state signals and / or processing state metadata (e.g., loudness flag “status” or flags for media features received from the flag parser & verifier), etc. Additionally, user input (e.g., user target loudness / dialnorm) and / or other input (e.g., from the video fingerprint generation process) and / or other metadata input (e.g., one or more types of third-party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, etc.) may be received by the encoder / transcoder. As shown, measured dialogue, gated and ungated loudness, and dynamic range values may also be inserted into the evolution data generator. Other media feature-related information may also be injected into the processing units described herein to generate part of the processing state metadata.
[0063] In one or more possible embodiments, the processing state metadata described herein is carried in the "add_bsi" field defined in the Enhanced AC-3 (E AC-3) syntax per ATSC A / 52b or in one or more ancillary data fields in the media bitstream described herein. In some possible embodiments, carrying the processing state metadata in these fields does not adversely affect the frame size and / or bitrate of the compressed media bitstream.
[0064] In some possible embodiments, processing state metadata may be included in an independent or dependent substream associated with the primary program media bitstream. An advantage of this approach is that the bitrate allocated for encoding the media data (carried by the primary program media bitstream) is not affected. If the processing state metadata is carried as part of the encoded frame, the bits allocated for encoding audio information may be reduced so that the frame size and / or bitrate of the compressed media bitstream may remain unchanged. For example, processing state metadata may have a reduced data rate representation, requiring a low data rate on the order of 10 kbps for transmission between media processing units. Thus, media data such as audio samples may be encoded at a rate reduced by 10 kbps to accommodate the processing state metadata.
[0065] In some possible embodiments, at least a portion of the processing state metadata may be embedded in the media data (or samples) via lossy or lossy data hiding techniques. An advantage of this approach is that the media samples and metadata may be received by downstream devices in the same bitstream.
[0066] In some possible embodiments, the processing state metadata may be linked to a fingerprint and stored in a media processing database. A media processing unit downstream from an upstream unit, such as an encoder / transcoder, that generates the processing state metadata may generate a fingerprint from the received media data and then use the fingerprint as a key to query the media processing database. After the processing state metadata is located in the database, a data block containing the processing state metadata associated with (or about) the received media data may be retrieved from the media processing database and made available to the downstream media processing unit. As used herein, a fingerprint may include, but is not limited to, any one or more media fingerprints generated to indicate media characteristics.
[0067] In some possible embodiments, the data block containing the processing state metadata includes a cryptographic hash (HMAC) of the processing state metadata and / or the underlying media data. The data block is assumed to be digitally signed in these embodiments, allowing downstream media processing units to authenticate and verify the processing state metadata relatively easily. Other cryptographic methods, including but not limited to any one or more non-HMAC-based cryptographic methods, may be used for secure transmission and reception of the processing state metadata and / or the underlying media data.
[0068] As previously described, a media processing unit, such as an encoder / transcoder described herein, may be configured to accept "legacy" media bitstreams and PCM samples. If the input media bitstream is a legacy media bitstream, the media processing unit may check for the presence of an evolution flag that may be present in the media bitstream or that may be hidden in the media data by one of the enhanced "legacy" encoders that include preprocessing and metadata validation logic as previously described. If the "evolution flag" is absent, the encoder is configured to perform adaptive processing and generate processing state metadata, as appropriate, in the output media bitstream or in a data block containing the processing state metadata. For example, as shown in FIG. 6, an exemplary unit such as a "transform domain real-time loudness and dynamic range controller" may adaptively process the audio content in the input media data it receives and automatically adjust the loudness and dynamic range if the "evolution flag" is absent in the input media data or source media stream. Additionally, optionally, or alternatively, another unit may utilize feature-based metadata to perform adaptive processing.
[0069] In the exemplary embodiment shown in FIG. 6, the encoder may know which post-processing / pre-processing units have performed certain types of media processing (e.g., loudness domain processing) and may therefore generate processing state metadata in the data block that includes specific parameters used in and / or derived from the loudness domain processing. In some possible embodiments, the encoder may generate processing state metadata that reflects the processing history of the content in the media data, so long as the encoder knows about the types of processing (e.g., loudness domain processing) that have been performed on the content in the media data. Additionally, optionally, or alternatively, the encoder may perform adaptive processing based on one or more media features described by the processing state metadata. Additionally, optionally, or alternatively, the encoder may perform analysis of the media data and generate a description of the media features as part of the processing state metadata to be provided to any other processing units.
[0070] In some possible embodiments, a decoder using the techniques described herein can understand the state of media data in the following scenarios:
[0071] Under the first scenario, if a decoder receives a media bitstream with an "evolution flag" set to indicate the validity of the processing state metadata in the media bitstream, the decoder may parse and / or extract the processing state metadata and signal it to downstream media processing units, such as appropriate post-processing units. On the other hand, if the "evolution flag" is absent, the decoder may signal to downstream media processing units that loudness metadata—which, for example, would in some possible embodiments be included in the processing state metadata if volume leveling had already been performed—is absent or cannot be trusted to be valid, and therefore volume leveling should still be performed.
[0072] In a second scenario, if a decoder receives a cryptographically hashed media bitstream generated by an upstream media processing unit, such as an evolutionary encoder, the decoder may parse and extract the cryptographic hash from a data block containing processing state metadata and use the cryptographic hash to validate the received media bitstream and associated metadata. For example, if the decoder finds that associated metadata (e.g., loudness metadata in the processing state metadata) is valid based on a match between a reference cryptographic hash and the cryptographic hash obtained from the data block, the decoder may signal a downstream media processing unit, such as a volume leveling unit, to pass the media data, such as audio, unchanged. Additionally, optionally, or alternatively, other types of cryptographic techniques may be used instead of cryptographic hash-based methods. Additionally, optionally, or alternatively, processing other than volume leveling may be performed based on one or more media characteristics of the media data described in the processing state metadata.
[0073] In a third scenario, a decoder receives a media bitstream generated by an upstream media processing unit, such as an evolutionary encoder, but if the media bitstream does not include a data block containing processing state metadata, that data block is stored in a media processing database. The decoder is configured to generate a fingerprint of the media data in the media stream, such as audio, and use the fingerprint to query the media processing database. The media processing database may return the appropriate data block associated with the received media data based on a fingerprint match. In some possible embodiments, the encoded media bitstream includes a simple universal resource locator (URL) to guide the decoder to send a fingerprint-based query to the media processing database, as discussed above.
[0074] In all of these scenarios, the decoder is configured to understand the media state and signal downstream media processing units to adapt their processing of the media data accordingly. In some possible embodiments, the media data herein may be re-encoded after being decoded. In some possible embodiments, a data block containing synchronic processing state information corresponding to the re-encoding may be passed to a downstream media processing unit, such as an encoder / converter after the decoder. For example, the data block may be included as associated metadata in the outgoing media bitstream from the decoder.
[0075] 7 illustrates an exemplary evolutionary decoder that controls the operation mode of a volume leveling unit based on the validity of loudness metadata in and / or associated with processing state metadata, according to some possible embodiments of the present invention. Other operations, such as feature-based processing, may also be handled. Any of the depicted components may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and / or one or more integrated circuit circuits (including ASICs and FPGAs). The decoder may have several legacy subunits, such as a frame information module (e.g., frame information module in AC-3, MPEG AAC, MPEG HE AAC, E AC-3, etc.), front-end decode (e.g., FED in AC-3, MPEG AAC, MPEG HE AAC, E AC-3, etc.), synchronization and transformation (e.g., synchronization and transformation module in AC-3, MPEG AAC, MPEG HE AAC, E AC-3, etc.), frame set buffer, back-end decode (e.g., BED (back end decode) in AC-3, MPEG AAC, MPEG HE AAC, E AC-3, etc.), back-end encoding (e.g., BEE in AC-3, MPEG AAC, MPEG HE AAC, E AC-3, etc.), CRC regeneration, media rendering (e.g., Dolby Volume), etc. Using these legacy subunits, the decoder can convey media content in the media data to downstream media processing units and / or render the media content. However, this decoder would not be able to communicate the state of the media data or provide media processing signals and / or processing state metadata in the output bitstream.
[0076] Under the techniques herein, in some possible embodiments, as shown in FIG. 7, the decoder may provide a number of features, such as metadata handling (evolutionary data and / or other metadata input including one or more types of third-party data, tracking information, identifiers, proprietary or standard information, user annotation data, user preference data, feature extraction, feature handling, etc.), secure (e.g., tamper-resistant) communication of processing state information (e.g., HMAC generators and signature verifiers, other cryptographic techniques), media fingerprint extraction (e.g., audio and video fingerprint extraction), adjunct media processing (e.g., speech channel(s) / loudness information, other The audio and video fingerprint extraction unit may also include any of several new subunits, such as: media feature extraction (type of media feature), data hiding (e.g., PCM data hiding, which can be destructive / irreversible or reversible), media processing signal insertion, HMAC generator (which may include, for example, "add_bsi" insertion(s) in one or more auxiliary data fields), other cryptographic techniques, hidden data recovery and verification (e.g., hidden PCM data recovery and verifier), data hiding "undo", one or more switches that operate based on processing state signals and / or processing state metadata (e.g., evolved data "valid" from the HMAC generator & signature verifier and data hiding insertion control), etc. As shown, the information extracted by the HMAC generator & signature verifier and audio & video fingerprint extraction unit may be output to or used for audio and video synchronization correction, rating, media rights, quality control, media location processing, feature-based processing, etc.
[0077] In some possible embodiments, a post-processing / pre-processing unit in a media processing chain does not operate independently. Rather, the post-processing / pre-processing unit may interact with an encoder or a decoder in the media processing chain. In the case of interaction with an encoder, the post-processing / pre-processing unit may help generate at least a portion of the processing state metadata about the state of the media data in the data block. In the case of interaction with a decoder, the post-processing / pre-processing unit is configured to determine the state of the media data and adapt its processing of the media data accordingly. As an example, in FIG. 7, an exemplary post-processing / pre-processing unit, such as a volume leveling unit, may obtain hidden data in PCM samples sent by an upstream decoder and, based on the hidden data, determine whether loudness metadata is valid. If the loudness metadata is valid, the input media data, such as audio, may be passed through the volume leveling unit unchanged. In another example, an exemplary post-processing / pre-processing unit may obtain hidden data in a PCM sample sent by an upstream decoder and, based on the hidden data, determine one or more types of media features previously determined from the content of the media sample. If a voice-recognized keyword is indicated, the post-processing unit may perform one or more specific actions related to the voice-recognized keyword.
[0078] 5. Data Hiding 8 shows an exemplary configuration using data hiding to pass media processing information according to some possible embodiments of the present invention. In some possible embodiments, data hiding may be used to enable signaling between an upstream processing unit, such as an evolutionary encoder or decoder (e.g., Audio Processing #1), and a downstream media processing unit, such as a post-processing / pre-processing unit (e.g., Audio Processing #2), when there is no metadata path between the upstream and downstream media processing units.
[0079] In some possible embodiments, lossless media data hiding (e.g., lossless audio data hiding) may be used to modify a media data sample (e.g., X) in the media data to become a modified media data sample (e.g., X') that carries the media processing signal and / or processing state metadata between two media processing units. In some possible embodiments, the modifications to the media data samples described herein are performed in a manner such that there is no perceptual degradation as a result of the modifications. Thus, even if there is no other media processing unit after media processing unit 1, no audible or visible artifacts may be perceived with respect to the modified media data sample. In other words, hiding the media processing signal and / or processing state metadata in a perceptually transparent manner does not cause any audible or visible artifacts when the audio and video of the modified media data sample are rendered.
[0080] In some possible embodiments, a media processing unit (e.g., audio processing unit #2 in FIG. 8 ) retrieves the embedded media processing signal and / or processing state metadata from the modified media data samples and restores the original media data samples by undoing the modifications. This may be done, for example, through subunits (e.g., information extraction and audio restoration). The retrieved embedded information may then serve as a signaling mechanism between two media processing units (e.g., audio processing units #1 and #2 in FIG. 8 ). The robustness of the data hiding techniques described herein may depend on the type of processing that may be performed by those media processing units. An example of media processing unit #1 may be a digital decoder in a set-top box, while an example of media processing unit #2 may be a volume leveling unit in the same set-top box. If the decoder determines that the loudness metadata is valid, the decoder may use a reversible data hiding technique to signal a subsequent volume leveling unit not to apply leveling.
[0081] In some possible embodiments, lossy media data hiding (e.g., a data hiding technique based on a secure communication channel) may be used to modify media data samples (e.g., X) in the media data to result in modified media data samples (e.g., X') that carry media processing signals and / or processing state metadata between two media processing units. In some possible embodiments, the modifications to the media data samples described herein are performed in a manner that minimizes perceptual degradation as a result of the modifications. Thus, minimal audible or visible artifacts may be perceived with respect to the modified media data samples. In other words, hiding the media processing signals and / or processing state metadata in a perceptually transparent manner will result in minimal audible or visible artifacts when the audio and video of the modified media data samples are rendered.
[0082] In some possible embodiments, modifications made to a modified media data sample through irreversible data hiding cannot be undone to restore the original media data sample.
[0083] 6. Exemplary Process Flow 9A and 9B illustrate an exemplary process flow according to some possible embodiments of the present invention. In some possible embodiments, this process flow may be performed by one or more computing devices or units in a media processing system.
[0084] In block 910 of FIG. 9A, a first device in a media processing chain (e.g., an enhanced media processing chain described herein) determines whether a type of media processing has been performed on an output version of media data. The first device may be part of or an entire media processing unit. In block 920, in response to determining that the type of media processing has been performed on the output version of media data, the first device may generate a media data state. In some possible embodiments, the media data state may specify a type of media processing, the results of which are included in the output version of the media data. A first device may communicate the output version of the media data and the media data state to a second device downstream in the media processing chain, for example, in an output media bitstream or in an auxiliary media bitstream associated with a separate media bitstream carrying the output version of the media data.
[0085] In some possible embodiments, the media data includes media content as one or more of audio content only, video content only, or both audio content and video content.
[0086] In some possible embodiments, the first device may provide the state of the media data to the second device as one or more of: (a) a media fingerprint; (b) processing state metadata; or (c) a media processing signal.
[0087] In some possible embodiments, the first device may store media processing data blocks in a media processing database, the media processing data blocks may include media processing metadata, and the media processing data blocks may be retrievable based on one or more media fingerprints associated with the media processing data blocks.
[0088] In some possible embodiments, the state of the media data includes a cryptographic hash value encrypted using the credential information, and the cryptographic hash value may be authenticated by the recipient device.
[0089] In some embodiments, at least a portion of the state of the media data includes one or more secure communication channels hidden in the media data, the one or more secure communication channels being authenticated by the recipient device. In one exemplary embodiment, the one or more secure communication channels may include at least one spread spectrum secure communication channel. In one exemplary embodiment, the one or more secure communication channels include at least one frequency shift keying secure communication channel.
[0090] In some possible embodiments, the state of the media data includes one or more sets of parameters used in and / or derived from media processes of said type.
[0091] In some possible embodiments, at least one of the first device or the second device includes one or more of a pre-processing unit, an encoder, a media processing subunit, a transcoder, a decoder, a post-processing unit, or a media content rendering subunit. In one example embodiment, the first device is an encoder (e.g., an AVC encoder), while the second device is a decoder (e.g., an AVC decoder).
[0092] In some possible embodiments, the type of processing is performed by the first device, and in other possible embodiments, the type of processing is instead performed by a device upstream from the first device in the media processing chain.
[0093] In some possible embodiments, a first device may receive an input version of media data, the input version of the media data including any state of the media data that indicates the type of media processing. In these embodiments, the first device may analyze the input version of the media data to determine the type of media processing that has already been performed on the input version of the media data.
[0094] In some possible embodiments, the first device encodes loudness and dynamic range in the state of the media data.
[0095] In some possible embodiments, a first device may adaptively avoid performing media processing of a type performed by an upstream device. However, even when media processing of that type has been performed, the first device may receive a command to override the media processing of that type performed by the upstream device. Instead, the first device may be commanded to still perform media processing of that type, e.g., using the same or different parameters. The media data state communicated from the first device to a second device downstream in the media processing chain may include an output version of the media data containing the results of the media processing of that type performed by the first device under the command, and a media data state indicating that media processing of that type has already been performed in the output version of the media data. In various possible embodiments, the first device may receive the command from one of: (a) user input, (b) a system configuration setting of the first device, (c) a signal from a device external to the first device, or (d) a signal from a subunit within the first device.
[0096] In some embodiments, the state of the media data includes at least a portion of the state metadata that is hidden in one or more secure communication channels.
[0097] In some embodiments, the first device modifies a number of bytes in the media data to store at least a portion of the state of the media data.
[0098] In some embodiments, at least one of the first device and the second device includes one or more of an Advanced Television Systems Committee (ATSC) codec, a Motion Picture Experts Group (MPEG) codec, an Audio Codec 3 (AC-3) codec, and an Enhanced AC-3 codec.
[0099] In some embodiments, the media processing chain includes: a pre-processing unit configured to accept as input time-domain samples comprising media content and to output processed time-domain samples; an encoder configured to output a compressed media bitstream of the media content based on the processed time-domain samples; a signal analysis and metadata correction unit configured to verify processing state metadata in the compressed media bitstream; a transcoder configured to modify the compressed media bitstream; a decoder configured to output decoded time-domain samples based on the compressed media bitstream; and a post-processing unit configured to perform post-processing of the media content in the decoded time-domain samples. In some embodiments, at least one of the first device and the second device includes one or more of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder, and the post-processing unit. In some embodiments, at least one of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder, and the post-processing unit performs adaptive processing of the media content based on processing metadata received from an upstream device.
[0100] In some embodiments, the first device determines one or more media features from the media data and includes a description of the one or more media features in the state of the media data. The one or more media features may include at least one media feature determined from one or more of frames, seconds, minutes, a user-definable time interval, a scene, a song, a piece of music, and a recording. The one or more media features include a semantic description of the media data. In various embodiments, the one or more media features include one or more of structural attributes, tonality including harmony and melody, timbre, rhythm, loudness, stereo mix, a quantity of audio source of the media data, absence or presence of voices, repetition characteristics, melody, harmony, lyrics, timbre, perceptual features, digital media features, stereo parameters, and one or more portions of speech content.
[0101] In block 950 of FIG. 9B, a first device in a media processing chain (e.g., an enhanced media processing chain described herein) determines whether some type of media processing has already been performed on the input version of the media data.
[0102] In block 960, in response to determining that the type of media processing is already being performed on the input version of the media data, the first device adapts its processing of the media data to disable performance of the type of media processing on the first device. In some possible embodiments, the first device may turn off one or more types of media processing based on the input state of the media data.
[0103] In some possible embodiments, a first device may communicate to a downstream second device in the media processing chain an output version of the media data and a state of the media data indicating that media processing of the type has already been performed on the output version of the media data.
[0104] In some possible embodiments, the first device may encode loudness and dynamic range in the state of the media data. In some possible embodiments, the first device may automatically perform one or more of adapting corrective loudness or dynamics audio processing based at least in part on whether said type of processing has already been performed on an input version of the media data.
[0105] In some possible embodiments, a first device may perform a second, different type of media processing on media data, and may communicate to a second device downstream in the media processing chain an output version of the media data and a state of the media data indicating that the second, different type of media processing and the second, different type of media processing have been performed on the output version of the media data.
[0106] In some possible embodiments, the first device may obtain a media data input state associated with the input version of the media data. In some possible embodiments, the media data input state is carried together with the input version of the media data in an input media bitstream. In some possible embodiments, the first device may extract the media data input state from data units encoding media content in the media data.
[0107] In some possible embodiments, the first device may restore a version of the data unit that does not include the input state of the media data and render the media content based on the restored version of the data unit.
[0108] In some possible embodiments, the first device may authenticate the input state of the media data by verifying a cryptographic hash value associated with the input state of the media data.
[0109] In some embodiments, the first device may authenticate the input state of the media data by verifying one or more fingerprints associated with the input state of the media data, where at least one of the one or more fingerprints is generated based on at least a portion of the media data.
[0110] In some embodiments, the first device may verify the input state of the media data by verifying one or more fingerprints associated with the input state of the media data, where at least one of the one or more fingerprints is generated based on at least a portion of the media data.
[0111] In some possible embodiments, a first device may receive input states of media data described by processing state metadata. The first device may generate a media processing signal based, at least in part, on the processing state metadata. The media processing signal may indicate the input states of the media data, even if it requires a smaller amount of data and / or a lower bit rate than the processing state metadata. The first device may transmit the media processing signal to a media processing device downstream of the first device in the media processing chain. In some possible embodiments, the media processing signal is hidden in one or more data units in the output version of the media data using a reversible data hiding technique such that one or more modifications to the media data can be removed by a receiving device. In some embodiments, at least one of the one or more modifications to the media data is hidden in one or more data units in the output version of the media data using an irreversible data hiding technique such that it cannot be removed by a receiving device.
[0112] In some possible embodiments, the first device determines one or more media features based on a description of the one or more media features in the state of the media data. The one or more media features may include at least one media feature determined from one or more of frames, seconds, minutes, a user-definable time interval, a scene, a song, a piece of music, and a recording. The one or more media features include a semantic description of the media data. In some embodiments, the first device performs one or more specified actions in response to determining the one or more media features.
[0113] In some possible embodiments, a method is provided that is performed by one or more computing devices, comprising: computing, by a first device in a media processing chain, one or more data rate reduced representations of source frames of media data; and simultaneously and securely conveying the one or more data rate reduced representations to a second device in the media processing chain within the state of the media data itself.
[0114] In some possible embodiments, the one or more data rate reduced representations are carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients.
[0115] In some possible embodiments, the one or more reduced data rate representations include synchronization data used to synchronize audio and video delivered within the media data.
[0116] In some possible embodiments, the one or more data rate reduced representations comprise media fingerprints (a) generated by a media processing unit and (b) embedded in the media data for one or more of quality monitoring, media rating, media tracking, or content retrieval.
[0117] In some possible embodiments, the method further includes calculating, by at least one of the one or more computing devices in the media processing chain, a cryptographic hash value based on the media data and / or a state of the media data, and transmitting it within one or more encoded bitstreams carrying the media data.
[0118] In some possible embodiments, the method further includes authenticating, by the receiving device, the cryptographic hash value; signaling, by the receiving device to one or more downstream media processing units, a determination of whether the state of the media data is valid; and, in response to determining that the state of the media data is valid, signaling, by the receiving device to the one or more downstream media processing units, the state of the media data.
[0119] In some possible embodiments, the cryptographic hash value representing the media state and / or media data is carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients.
[0120] In some possible embodiments, there is provided a method comprising: adaptively processing, by one or more computing devices in a media processing chain including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder or a stream processor, an input version of the media data based on a past history of loudness processing of the media data by one or more upstream media processing units as indicated by a state of the media data; and normalizing the loudness and / or dynamic range of an output version of the media data at the end of the media processing chain to consistent loudness and / or dynamic range values.
[0121] In some possible embodiments, the consistent loudness value includes a loudness value that is (1) controlled or selected by a user or (2) adaptively signaled by conditions within the input version of the media data.
[0122] In some possible embodiments, the loudness value is calculated for a dialogue portion of the media data.
[0123] In some possible embodiments, the loudness values are calculated for absolute, relative and / or ungated portions of the media data.
[0124] In some possible embodiments, the consistent dynamic range values include dynamic range values that are (1) controlled or selected by a user or (2) adaptively signaled by conditions within the input version of the media data.
[0125] In some possible embodiments, the dynamic range values are calculated for dialogue portions of the media data.
[0126] In some possible embodiments, the dynamic range values are calculated for absolute, relative and / or ungated portions of the media data.
[0127] In some possible embodiments, the method further includes: calculating one or more loudness and / or dynamic range gain control values for normalizing the output version of the media data to a consistent loudness value and a consistent dynamic range; and simultaneously conveying said one or more loudness and / or dynamic range gain control values within the state of the output version of the media data at the end of the media processing chain, wherein said one or more loudness and / or dynamic range gain control values are usable by another device to reverse-apply said one or more loudness and / or dynamic range gain control values to restore the original loudness value and original dynamic range in the input version of the media data.
[0128] In some possible embodiments, the one or more loudness and / or dynamic range control values representing the state of the output version of the media data are carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients.
[0129] In some possible embodiments, there is provided a method comprising performing, by one or more computing devices in a media processing chain including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder or a stream processor, one of inserting, extracting or editing related and unrelated media data positions and / or the state of related and unrelated media data positions in one or more encoded bitstreams.
[0130] In some possible embodiments, the one or more related and unrelated media data positions and / or the status of related and unrelated media data positions in the encoded bitstream are carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients.
[0131] In some possible embodiments, there is provided a method comprising performing one or more of insertion, extraction or editing of related and unrelated media data and / or state of related and unrelated media data in one or more encoded bitstreams by one or more computing devices in a media processing chain including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder or a stream processor.
[0132] In some possible embodiments, the one or more related and unrelated media data and / or the state of the related and unrelated media data in the encoded bitstream is carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients.
[0133] In some possible embodiments, the media processing system is configured to compute, by one or more computing devices in a media processing chain including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder or a stream processor, a cryptographic hash value based on the media data and / or a state of the media data, and convey it in one or more encoded bitstreams.
[0134] As used in this document, the term "related and unrelated media data locations" can refer to information that may include media resource locators such as absolute paths, relative paths and / or URLs that indicate the location of related media (e.g., copies of media in different bitstream formats) or absolute paths, relative paths and / or URLs that indicate the location of unrelated media or other types of information that are not directly related to the essence or bitstream in which the media data locations are found (e.g., the location of new pieces of media such as commercials, advertisements, web pages, etc.).
[0135] As used herein, the term "state of related and unrelated media data locations" can refer to the validity of the related and unrelated media locations (as the locations can be edited / updated throughout the lifecycle of the bitstream in which they are carried).
[0136] As used herein, the term "related media data" can refer to the carriage of related media data in the form of a secondary media data bitstream that is highly correlated with the primary media that the bitstream represents (e.g., the carriage of a copy of the media data in a second (independent) bitstream format). In the context of unrelated media data, this information can refer to the carriage of a secondary media data bitstream that is independent of the primary media data.
[0137] As used herein, "state" for related media data can refer to any signaling information (processing history, updated target loudness, etc...) and / or metadata as well as the validity of the related media data. "State" for unrelated media data can refer to independent signaling and / or metadata, including validity information, that can be carried separately (independently) from the state of the "related" media data. The state of unrelated media data refers to media data that is "unrelated" to the media data bitstream in which this information is found (as this information can be edited / updated independently throughout the lifecycle of the bitstream in which it is carried).
[0138] As used in this paper, the terms "absolute, relative, and / or ungated portions of media data" refer to gating of loudness and / or level measurements performed on media data. Gating refers to a specific level or loudness threshold, where calculated values above the threshold are included in the final measurement (e.g., ignoring short-term loudness values below -60 dBFS in the final measurement). Gating on an absolute value refers to a fixed level or loudness, while gating on a relative value refers to a value that is dependent on the current "ungated" measurement.
[0139] 12A-12L further illustrate block diagrams of some exemplary media processing nodes / devices according to some possible embodiments of the present invention.
[0140] As shown in FIG. 12A , a signal processor (which may be node 1 of the N nodes) is configured to receive an input signal that may include audio PCM samples. The audio PCM samples may or may not include processing state metadata (or media state metadata) hidden between the audio PCM samples. The signal processor of FIG. 12A may include a media state metadata extractor configured to decode, extract, and / or interpret processing state metadata from the audio PCM samples, provided by one or more media processing units preceding the signal processor of FIG. 12A . At least a portion of the processing state metadata may be provided to an audio encoder in the signal processor of FIG. 12A to adapt processing parameters for the audio encoder. In parallel, an audio analysis unit in the signal processor of FIG. 12A may analyze media content passed in the input signal. Feature extraction, media classification, loudness estimation, fingerprint generation, etc. may be implemented as part of the analysis performed by the audio analysis unit. At least a portion of the results of this analysis may be provided to an audio encoder in the signal processor of FIG. 12A to adapt processing parameters for the audio encoder. The audio encoder encodes audio PCM samples in the input signal based on processing parameters into a coded bitstream in the output signal. The coded bitstream analysis unit in the signal processor of Figure 12A may be configured to determine whether media data or samples in the coded bitstream to be transmitted in the output signal of the signal processor of Figure 12A have room to store at least a portion of processing state metadata. The new processing state metadata to be transmitted by the signal processor of Figure 12A includes some or all of the processing state metadata extracted by the media state metadata extractor, the processing state metadata generated by the audio analysis unit and the media state metadata generator of the signal processor of Figure 12A, and / or any third-party data.If it is determined that the media data or samples in the encoded bitstream have room to store at least some of the processing state metadata, some or all of the new processing state metadata may be stored as hidden data in the media data or samples in the output signal. Additionally, optionally, or alternatively, some or all of the new processing state metadata may be stored in a metadata structure separate from the media data and samples in the output signal. Thus, the output signal may include an encoded bitstream including the new processing state (or "media state") metadata carried within and / or between media samples (essences) via a secure, covert or non-covert communication channel.
[0141] As shown in FIG. 12B , a signal processor (which may be node 1 of the N nodes) is configured to receive an input signal that may include audio PCM samples. The audio PCM samples may or may not include processing state metadata (or media state metadata) hidden between the audio PCM samples. The signal processor of FIG. 12B may include a media state metadata extractor configured to decode, extract, and / or interpret processing state metadata from the audio PCM samples provided by one or more media processing units preceding the signal processor of FIG. 12B. At least a portion of the processing state metadata may be provided to a PCM audio sample processor in the signal processor of FIG. 12B to adapt processing parameters for the PCM audio sample processor. In parallel, an audio analysis unit in the signal processor of FIG. 12B may analyze the media content passed in the input signal. Feature extraction, media classification, loudness estimation, fingerprint generation, etc. may be implemented as part of the analysis performed by the audio analysis unit. At least a portion of the results of this analysis may be provided to an audio encoder in the signal processor of Figure 12B to adapt processing parameters for the PCM audio sample processor. The PCM audio sample processor processes the audio PCM samples in the input signal based on the processing parameters into an encoded PCM audio (sample) bitstream in the output signal. The PCM audio analysis unit in the signal processor of Figure 12B may be configured to determine whether media data or samples in the PCM audio bitstream to be transmitted in the output signal of the signal processor of Figure 12B have room to store at least a portion of the processing state metadata.The new processing state metadata to be transmitted by the signal processor of FIG. 12B may include some or all of the processing state metadata extracted by the media state metadata extractor, the processing state metadata generated by the audio analysis unit and the media state metadata generator of the signal processor of FIG. 12B, and / or any third-party data. If it is determined that the media data or samples in the PCM audio bitstream have room to store at least some of the processing state metadata, some or all of the new processing state metadata may be stored as hidden data in the media data or samples in the output signal. Additionally, optionally, or alternatively, some or all of the new processing state metadata may be stored in a metadata structure separate from the media data and samples in the output signal. Thus, the output signal may include a PCM audio bitstream including the new processing state (or "media state") metadata carried within and / or between media samples (essences) via a secure, covert or non-covert communication channel.
[0142] As shown in FIG. 12C , a signal processor (which may be node 1 of the N nodes) is configured to receive an input signal that may include a PCM audio (sample) bitstream. The PCM audio bitstream may include processing state metadata (or media state metadata) carried within and / or between media samples (essence) in the PCM audio bitstream via a secure, covert or non-covert communication channel. The signal processor of FIG. 12C may include a media state metadata extractor configured to decode, extract, and / or interpret the processing state metadata from the PCM audio bitstream. At least a portion of the processing state metadata may be provided to a PCM audio sample processor in the signal processor of FIG. 12C to adapt processing parameters for the PCM audio sample processor. The processing state metadata may include descriptions of media features, media class types or subtypes, or likelihood / probability values determined by one or more media processing units prior to the signal processor of FIG. 12C , which the signal processor of FIG. 12C may be configured to utilize without performing its own media content analysis. Additionally, optionally, or alternatively, the media state metadata extractor may be configured to extract third-party data from the input signal and send the third-party data to a downstream processing node / entity / device. In one embodiment, the PCM audio sample processor processes the PCM audio bitstream into audio PCM samples for the output signal based on processing parameters set based on the processing state metadata provided by the one or more media processing units prior to the signal processor of FIG. 12C .
[0143] As shown in FIG. 12D , a signal processor (which may be node 1 of the N nodes) is configured to receive an input signal that may include an encoded audio bitstream. The encoded audio bitstream includes processing state metadata (or media state metadata) carried within and / or hidden among media samples via a secure, covert or non-covert communication channel. The signal processor of FIG. 12D may include a media state metadata extractor configured to decode, extract, and / or interpret the processing state metadata from the encoded bitstream, as provided by one or more media processing units preceding the signal processor of FIG. 12D . At least a portion of the processing state metadata may be provided to an audio decoder in the signal processor of FIG. 12D to adapt processing parameters for the audio decoder. In parallel, an audio analysis unit in the signal processor of FIG. 12D may analyze the media content passed in the input signal. Feature extraction, media classification, loudness estimation, fingerprint generation, etc. may be implemented as part of the analysis performed by the audio analysis unit. At least a portion of the results of this analysis may be provided to an audio decoder in the signal processor of Figure 12D for adapting processing parameters for the audio decoder. The audio decoder converts the encoded audio bitstream in the input signal into a PCM audio bitstream in the output signal based on the processing parameters. The PCM audio analysis unit in the signal processor of Figure 12D may be configured to determine whether media data or samples in the PCM audio bitstream have room to store at least a portion of the processing state metadata.The new processing state metadata to be transmitted by the signal processor of FIG. 12D may include some or all of the processing state metadata extracted by the media state metadata extractor, the processing state metadata generated by the audio analysis unit and the media state metadata generator of the signal processor of FIG. 12D, and / or any third-party data. If it is determined that the media data or samples in the PCM audio bitstream have room to store at least some of the processing state metadata, some or all of the new processing state metadata may be stored as hidden data in the media data or samples in the output signal. Additionally, optionally, or alternatively, some or all of the new processing state metadata may be stored in a metadata structure separate from the media data and samples in the output signal. Thus, the output signal may include a PCM audio (sample) bitstream including processing state (or "media state") metadata carried within and / or between the media data / samples (essence) via a secure, covert or non-covert communication channel.
[0144] As shown in FIG. 12E, a signal processor (which may be node 1 of the N nodes) is configured to receive an input signal that may include an encoded audio bitstream. The encoded audio bitstream may include processing state metadata (or media state metadata) carried within and / or between media samples (essences) in the encoded audio bitstream via a secure, covert or non-covert communication channel. The signal processor of FIG. 12E may have a media state metadata extractor configured to decode, extract, and / or interpret the processing state metadata from the encoded audio bitstream. At least a portion of the processing state metadata may be provided to an audio decoder in the signal processor of FIG. 12E to adapt processing parameters for the audio decoder. The processing state metadata may include descriptions of media features, media class types or subtypes, or likelihood / probability values determined by one or more media processing units prior to the signal processor of FIG. 12E, which the signal processor of FIG. 12E may be configured to utilize without performing its own media content analysis. Additionally, optionally, or alternatively, the media state metadata extractor may be configured to extract third-party data from the input signal and transmit the third-party data to a downstream processing node / entity / device. In one embodiment, the audio decoder processes the encoded audio bitstream into audio PCM samples of the output signal based on processing parameters set based on the processing state metadata provided by the one or more media processing units prior to the signal processor of Figure 12E.
[0145] As shown in FIG. 12F, a signal processor (which may be node 1 of the N nodes) is configured to receive an input signal that may include an encoded audio bitstream. The encoded audio bitstream includes processing state metadata (or media state metadata) carried within and / or hidden among media samples via a secure, covert or non-covert communication channel. The signal processor of FIG. 12F may include a media state metadata extractor configured to decode, extract, and / or interpret the processing state metadata from the encoded bitstream, as provided by one or more media processing units preceding the signal processor of FIG. 12F. At least a portion of the processing state metadata may be provided to a bitstream transcoder (or encoded audio bitstream processor) in the signal processor of FIG. 12F to adapt processing parameters for the bitstream transcoder. In parallel, an audio analysis unit in the signal processor of FIG. 12F may analyze the media content passed in the input signal. Feature extraction, media classification, loudness estimation, fingerprint generation, etc. may be implemented as part of the analysis performed by the audio analysis unit. At least a portion of the results of this analysis may be provided to a bitstream transcoder in the signal processor of FIG. 12F to adapt processing parameters for the bitstream transcoder. The bitstream transcoder converts an encoded audio bitstream in the input signal into an encoded audio bitstream in the output signal based on the processing parameters. The encoded bitstream analysis unit in the signal processor of FIG. 12F may be configured to determine whether media data or samples in the encoded audio bitstream have room to store at least a portion of the processing state metadata.The new processing state metadata to be transmitted by the signal processor of FIG. 12F may include some or all of the processing state metadata extracted by the media state metadata extractor, the processing state metadata generated by the audio analysis unit and the media state metadata generator of the signal processor of FIG. 12F, and / or any third-party data. If it is determined that the media data or samples in the encoded audio bitstream have room to store at least some of the processing state metadata, some or all of the new processing state metadata may be stored as hidden data in the media data or samples in the output signal. Additionally, optionally, or alternatively, some or all of the new processing state metadata may be stored in a metadata structure separate from the media data in the output signal. Thus, the output signal may include an encoded audio bitstream including processing state (or "media state") metadata carried within and / or between media data / samples (essence) via a secure, covert, or non-covert communication channel.
[0146] FIG. 12G illustrates an exemplary configuration partially similar to FIG. 12A. Additionally, optionally, or alternatively, the signal processor of FIG. 12G may include a media state metadata extractor configured to query a local and / or external media state metadata database, which may be operatively linked to the signal processor of FIG. 12G via an intranet and / or the Internet. The query sent to the database by the signal processor of FIG. 12G may include one or more fingerprints associated with the media data, one or more names associated with the media data (e.g., song title, movie title), or any other type of identifying information associated with the media data. Based on the information in the query, matching media state metadata stored in the database may be located and provided to the signal processor of FIG. 12G. The media state metadata may be included in processing state metadata provided by the media state metadata extractor to a downstream processing node / entity, such as an audio encoder. Additionally, optionally, or alternatively, the signal processor of Figure 12G may include a media state metadata generator configured to provide any generated media state metadata and / or associated identification information, such as fingerprints, names, and / or other types of identification information, to a local and / or external media state metadata database as shown in Figure 12G. Additionally, optionally, or alternatively, one or more portions of the media state metadata stored in the database may be provided to the signal processor of Figure 12G for communication to downstream media processing nodes / devices within and / or between media samples (essences) via secure, covered or uncovered communication channels.
[0147] FIG. 12H illustrates an exemplary configuration partially similar to FIG. 12B. Additionally, optionally, or alternatively, the signal processor of FIG. 12H may include a media state metadata extractor configured to query a local and / or external media state metadata database, which may be operatively linked to the signal processor of FIG. 12H via an intranet and / or the Internet. The query sent to the database by the signal processor of FIG. 12H may include one or more fingerprints associated with the media data, one or more names associated with the media data (e.g., song title, movie title), or any other type of identifying information associated with the media data. Based on the information in the query, matching media state metadata stored in the database may be located and provided to the signal processor of FIG. 12H. The media state metadata may be included in processing state metadata provided by the media state metadata extractor to a downstream processing node / entity, such as a PCM audio sample processor. Additionally, optionally, or alternatively, the signal processor of Figure 12H may include a media state metadata generator configured to provide any generated media state metadata and / or associated identification information, such as fingerprints, names, and / or other types of identification information, to a local and / or external media state metadata database as shown in Figure 12H. Additionally, optionally, or alternatively, one or more portions of the media state metadata stored in the database may be provided to the signal processor of Figure 12H for communication to downstream media processing nodes / devices within and / or between media samples (essences) via secure, covert, or non-covert communication channels.
[0148] FIG. 12I illustrates an exemplary configuration partially similar to FIG. 12C. Additionally, optionally, or alternatively, the signal processor of FIG. 12I may include a media state metadata extractor configured to query a local and / or external media state metadata database, which may be operatively linked to the signal processor of FIG. 12I via an intranet and / or the Internet. The query sent to the database by the signal processor of FIG. 12I may include one or more fingerprints associated with the media data, one or more names associated with the media data (e.g., song title, movie title), or any other type of identifying information associated with the media data. Based on the information in the query, matching media state metadata stored in the database may be located and provided to the signal processor of FIG. 12I. The media state metadata may be provided to a downstream processing node / entity, such as a PCM audio sample processor.
[0149] FIG. 12J illustrates an exemplary configuration partially similar to FIG. 12D. Additionally, optionally, or alternatively, the signal processor of FIG. 12J may include a media state metadata extractor configured to query a local and / or external media state metadata database, which may be operatively linked to the signal processor of FIG. 12J via an intranet and / or the Internet. The query sent to the database by the signal processor of FIG. 12J may include one or more fingerprints associated with the media data, one or more names associated with the media data (e.g., song title, movie title), or any other type of identifying information associated with the media data. Based on the information in the query, matching media state metadata stored in the database may be located and provided to the signal processor of FIG. 12J. The media state metadata from the database may be included in the processing state metadata provided to a downstream processing node / entity, such as an audio decoder. Additionally, optionally, or alternatively, the signal processor of Figure 12J may have an audio analysis unit configured to provide any generated media state metadata and / or associated identification information, such as fingerprints, names, and / or other types of identification information, to a local and / or external media state metadata database as shown in Figure 12J. Additionally, optionally, or alternatively, one or more portions of the media state metadata stored in the database may be provided to the signal processor of Figure 12J to be communicated to downstream media processing nodes / devices within and / or between media samples (essences) via secure, covered or uncovered communication channels.
[0150] FIG. 12K illustrates an exemplary configuration partially similar to FIG. 12F. Additionally, optionally, or alternatively, the signal processor of FIG. 12K may include a media state metadata extractor configured to query a local and / or external media state metadata database, which may be operatively linked to the signal processor of FIG. 12K via an intranet and / or the Internet. The query sent to the database by the signal processor of FIG. 12K may include one or more fingerprints associated with the media data, one or more names associated with the media data (e.g., song title, movie title), or any other type of identifying information associated with the media data. Based on the information in the query, matching media state metadata stored in the database may be located and provided to the signal processor of FIG. 12K. The media state metadata from the database may be included in processing state metadata provided to a downstream processing node / entity, such as a bitstream transcoder or encoded audio bitstream processor. Additionally, optionally or alternatively, one or more portions of the media state metadata stored in the database may be provided to the signal processor of FIG. 12K for communication to downstream media processing nodes / devices within and / or between media samples (essences) via secure covered or uncovered communication channels.
[0151] FIG. 12L illustrates signal processor node 1 and signal processor node 2 according to an example embodiment. Signal processor node 1 and signal processor node 2 may be part of an overall media processing chain. In some embodiments, signal processor node 1 adapts its media processing based on processing state metadata received by signal processor node 2. Meanwhile, signal processor node 2 adapts its media processing based on processing state metadata received by signal processor node 2. The processing state metadata received by signal processor node 2 may include processing state metadata and / or media state metadata added by signal processor node 1 after signal processor node 1 analyzes the content of the media data. As a result, signal processor node 2 can directly utilize the metadata provided by signal processor node 1 in its media processing without repeating some or all of the analysis previously performed by signal processor node 1.
[0152] 7. Implementation mechanism - Hardware overview According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may include digital electronic devices, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), that are hardwired to execute the techniques or that are persistently programmed to execute the techniques, or may include one or more general-purpose hardware processors that are programmed to execute the techniques according to program instructions in firmware memory, other storage, or a combination thereof. Such special-purpose computing devices may combine custom hardwired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices, or any other devices incorporating hardwired and / or program logic to implement the techniques.
[0153] 10 is a block diagram illustrating a computer system 1000 in which embodiments of the present invention may be implemented. The computer system 1000 includes a bus 1002 or other communication mechanism for communicating information, and a hardware processor 1004 coupled to the bus 1002 for processing information. The hardware processor 1004 may be, for example, a general-purpose microprocessor.
[0154] Computer system 1000 also includes a main memory 1006, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 1002 for storing information and instructions to be executed by processor 1004. Main memory 1006 may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 1004. Such instructions, when stored on a non-transitory storage medium accessible to processor 1004, render computer system 1000 a special-purpose machine customized to perform the operations specified in the instructions.
[0155] Computer system 1000 includes a read-only memory (ROM) 1008 or other static storage device, coupled to bus 1002, for storing static information and instructions for processor 1004. A storage device 1010, such as a magnetic disk or optical disk, is provided and coupled to bus 1002 for storing information and instructions.
[0156] Computer system 1000 may be coupled via bus 1002 to a display 1012, such as a cathode ray tube (CRT), for displaying information to a computer user. An input device 1014, including alphanumeric and other keys, is coupled to bus 1002 for communicating information and command selections to processor 1004. Another type of user input device is a cursor control 1016, such as a mouse, trackball, or cursor direction keys, for communicating directional information and command selections to processor 1004 and for controlling cursor movement on display 1012. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allow the device to specify a position in a plane.
[0157] Computer system 1000 may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, configures or programs computer system 1000 as a special-purpose machine. According to one embodiment, the techniques described herein are performed by computer system 1000 in response to processor 1004 executing one or more sequences of one or more instructions contained in main memory 1006. Such instructions may be read into main memory 1006 from another storage medium, such as storage device 1010. Execution of the sequences of instructions contained in main memory 1006 causes processor 1004 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.
[0158] As used herein, the term "storage media" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific fashion. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 1010. Volatile media include dynamic memory, such as main memory 1006. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, flash EPROMs, NVRAM, and any other memory chip or cartridge.
[0159] Storage media is distinct from, but may be used in the context of, transmission media. Transmission media participates in transferring information between storage media. For example, transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise bus 1002. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0160] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 1004 for execution. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 1000 can receive the data on the telephone line and convert the data to an infrared signal using an infrared transmitter. An infrared detector can receive the data carried in the infrared signal and appropriate circuitry can place the data on bus 1002. Bus 1002 carries the data to main memory 1006, from which processor 1004 retrieves and executes the instructions. The instructions received by main memory 1006 may optionally be stored on storage device 1010 either before or after execution by processor 1004.
[0161] Computer system 1000 also includes a communication interface 1018 coupled to bus 1002. The communication interface 1018 provides a two-way data communication coupling to a network link 1020 that is connected to a local network 1022. For example, communication interface 1018 may be an integrated services digital network (ISDN) card, cable modem, satellite modem, or modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 1018 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 1018 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0162] Network link 1020 typically provides data communication through one or more networks to other data devices. For example, network link 1020 may provide a connection through local network 1022 to a host computer 1024 or to data equipment operated by an Internet Service Provider (ISP) 1026. ISP 1026 provides data communication services through the world-wide packet data communication network now commonly referred to as the "Internet" 1028. Local network 1022 and the Internet 1028 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 1020 and through communication interface 1018, which carry the digital data to and from computer system 1000, are exemplary forms of transmission media.
[0163] Computer system 1000 can send messages and receive data, including program code, through the network(s), network link 1020 and communication interface 1018. In the Internet example, a server 1030 might transmit a requested code for an application program through Internet 1028, ISP 1026, local network 1022 and communication interface 1018.
[0164] The received code may be executed by processor 1004 as it is received, and / or stored in storage device 1010, or other non-volatile storage for later execution.
[0165] 8. Numbering Example Accordingly, embodiments of the present invention may relate to one or more of the numbered examples below. Each numbered example is an example, and, like any other related discussion provided above, should not be construed as limiting any claim, whether as currently written or as subsequently amended, replaced, or added. Likewise, these examples should not be considered limiting with respect to any claim of any related patent and / or patent application (including foreign or international counterparts and / or patents, divisionals, continuations, reissues, etc.). [Numbered Example 1] 1. A method comprising: determining, by a first device in a media processing chain, whether a type of media processing has been performed on an output version of media data; and in response to determining that the type of media processing has been performed on the output version of the media data by a device of the first device: (a) generating a media data state specifying the type of media processing that has been performed on the output version of the media data by the first device; and (b) communicating the output version of the media data and the media data state from the first device to a second device downstream in the media processing chain. [Numbered Example 2] 2. The method of numbered embodiment 1, wherein the metadata includes media content as one or more of: audio content only, video content only, or both audio content and video content. [Numbered Example 3] The method of numbered embodiment 1, further comprising providing to the second device the state of the media data as one or more of: (a) a media fingerprint, (b) processing state metadata, (c) extracted media feature values, (d) media class type or subtype descriptions and / or values, (e) media feature class and / or subclass probability values, (f) cryptographic hash values, or (f) media processing signals. [Numbered Example 4] 10. The method of claim 1, further comprising storing a media processing data block in a media processing database, the media processing data block including media processing metadata, the media processing data block being retrievable based on one or more media fingerprints associated with the media processing data block. [Numbered Example 5] 10. The method of claim 1, wherein the state of the media data includes a cryptographic hash value encrypted using credential information, and the cryptographic hash value is authenticated by the recipient device. [Numbered Example 6] The method of numbered Example 1, wherein at least a portion of the state of the media data includes one or more secure communication channels hidden in the media data, and the one or more secure communication channels are authenticated by a recipient device. [Numbered Example 7] 7. The method of embodiment 6, wherein the one or more secure communication channels include at least one spread spectrum secure communication channel. [Numbered Example 8] 7. The method of embodiment 6, wherein the one or more secure communication channels include at least one frequency shift keying secure communication channel. Numbered Example 9 2. The method of numbered embodiment 1, wherein the state of the media data is carried along with an output version of the media data in an output media bitstream. Numbered Example 10 2. The method of numbered Example 1, wherein the state of the media data is carried in an auxiliary metadata bitstream associated with a separate media bitstream carrying an output version of the media data. Numbered Example 11 2. The method of numbered embodiment 1, wherein the media data state includes one or more sets of parameters related to media processing of the type. Numbered Example 12 The method of numbered embodiment 1, wherein at least one of the first device or the second device includes one or more of: a pre-processing unit, an encoder, a media processing subunit, a transcoder, a decoder, a post-processing unit, or a media content rendering subunit. Numbered Example 13 2. The method of claim 1, wherein the first device is an encoder and the second device is a decoder. Numbered Example 14 2. The method of embodiment 1, further comprising performing, by the first device, a media process of the type. [Numbered Example 15] said type of media processing being performed by a device upstream relative to said first device in said media processing chain, said method further comprising: receiving, by the first device, an input version of the media data, the input version of the media data including any state of the media data that indicates the type of media processing; analyzing the input version of the media data to determine the type of media processing that has already been performed on the input version of the media data; Method described in numbered Example 1. [Numbered Example 16] 2. The method of Example 1, further comprising encoding loudness and dynamic range values in the media data. Numbered Example 17 said type of media processing having previously been performed by a device upstream to said first device in said media processing chain, said method further comprising: receiving, by said first device, a command to override previously performed media processing of said type; performing, by said first device, media processing of said type; communicating from the first device to a second device downstream in the media processing chain an output version of the media data and a state of the media data indicating that media processing of the type has been performed on the output version of the media data; Method described in numbered Example 1. Numbered Example 18 The method of numbered Example 17 further includes receiving the command from one of (a) user input, (b) system configuration settings of the first device, (c) a signal transmission from a device external to the first device, or (d) a signal transmission from a subunit within the first device. Numbered Example 19 2. The method of claim 1, further comprising communicating one or more types of metadata from the first device to the second device downstream in the media processing chain, the types of metadata being independent of the state of the metadata. [Numbered Example 20] 2. The method of numbered embodiment 1, wherein the state of the media data includes at least a portion of state metadata that is hidden in one or more secure communication channels. [Numbered Example 21] 2. The method of embodiment 1, further comprising modifying a plurality of bytes in the media data to store at least a portion of a state of the media data. Numbered Example 22 The method of numbered embodiment 1, wherein at least one of the first device and the second device includes one or more of an Advanced Television Systems Committee (ATSC) codec, a Motion Picture Experts Group (MPEG) codec, an Audio Codec 3 (AC-3) codec, and an Enhanced AC-3 codec. [Numbered Example 23] The media processing chain includes: a pre-processing unit configured to accept as input time-domain samples comprising media content and to output processed time-domain samples; an encoder configured to output a compressed media bitstream of the media content based on the processed time-domain samples; a signal analysis and metadata correction unit configured to verify processing state metadata within the compressed media bitstream; a transcoder configured to modify the compressed media bitstream; a decoder configured to output decoded time-domain samples based on the compressed media bitstream; a post-processing unit configured to perform post-processing of the media content within the decoded time-domain samples. Method described in numbered Example 1. [Numbered Example 24] The method described in numbered embodiment 23, wherein at least one of the first device and the second device includes at least one of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder, and the post-processing unit. [Numbered Example 25] The method of numbered Example 23, wherein at least one of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder, and the post-processing unit performs adaptive processing of the media content based on processing metadata received from an upstream device. [Numbered Example 26] determining one or more media characteristics from the media data; and including a description of the one or more media features in the state of the media data. Method described in numbered Example 1. [Numbered Example 27] The method of numbered Example 26, wherein the one or more media features include at least one media feature determined from one or more of frames, seconds, minutes, a user-definable time interval, a scene, a song, a piece of music, and a recording. [Numbered Example 28] 27. The method of numbered Example 26, wherein the one or more media features include a semantic description of the media data. [Numbered Example 29] The method of numbered Example 26, wherein the one or more media features include one or more of structural attributes, sound quality including harmony and melody, timbre, rhythm, loudness, stereo mix, amount of audio from the media data, absence or presence of voices, repetition characteristics, melody, harmony, lyrics, timbre, perceptual features, digital media features, stereo parameters, and one or more portions of speech content. [Numbered Example 30] 27. The method of numbered embodiment 26, further comprising classifying the media data into one or more media data classes among a plurality of media data classes using the one or more media features. Numbered Example 31 The method of numbered embodiment 30, wherein the one or more media data classes include a single overall / dominant media data class for the entire media or a single class representing a time period shorter than the entire media. Numbered Example 32 32. The method of numbered embodiment 31, wherein the shorter time period represents a single media frame, a single media data block, multiple media frames, multiple media data blocks, a fraction of a second, one second, or multiple seconds. [Numbered Example 33] 31. The method of numbered embodiment 30, wherein one or more media data class labels representing the one or more media data classes are calculated and inserted into the bitstream. Numbered Example 34 The method of numbered embodiment 30, wherein one or more media data class labels representing the one or more media data classes are calculated and signaled to a recipient media processing node as hidden data embedded in the media data. [Numbered Example 35] The method of numbered embodiment 30, wherein one or more media data class labels representing the one or more media data classes are calculated and signaled to a recipient media processing node in a separate metadata structure between the blocks of media data. [Numbered Example 36] The method of numbered Example 31, wherein the single overall / dominant media data class represents one or more of a single class type such as music, speech, noise, silence, applause, or a mixed class type such as speech over music, speech over noise, or other mixture of media data types. Numbered Example 37 The method of numbered embodiment 30 further includes associating one or more likelihood or probability values with the one or more media data class labels, the likelihood or probability values representing a level of confidence that the calculated media class label has with the media segment / block with which the calculated media class label is associated. Numbered Example 38 The method of numbered embodiment 37, wherein the likelihood or probability value is used by a recipient media processing node in the media processing chain to adapt processing to improve one or more operations such as upmixing, encoding, decoding, transcoding or headphone virtualization. Numbered Example 39 The method of numbered Example 38, wherein at least one of the one or more operations eliminates the need for pre-configured processing parameters, reduces the complexity of processing units throughout the media chain, or extends battery life, thereby avoiding complex analytical operations to classify media data by a recipient media processing node. [Numbered Example 40] determining whether a type of media processing has already been performed on an input version of the media data by a first device in the media processing chain; and in response to determining that media processing of the type has already been performed on the input version of the media data by the first device, performing an adaptation of processing of the media data to disable performance of media processing of the type on the first device. method. Numbered Example 41 The method of numbered embodiment 40 further includes communicating from the first device to a second device downstream in the media processing chain an output version of the media data and a state of the media data indicating that the type of media processing has already been performed on the output version of the media data. [Numbered Example 42] The method of Example 41, further comprising encoding loudness and dynamic range values in the media data. [Numbered Example 43] performing, by the first device, a second type of media processing on the media data that is different from the first type of media processing; communicating from the first device to a second device downstream in the media processing chain an output version of the media data and a state of the media data indicating that the second type of media processing has been performed on the output version of the media data. Method described in numbered Example 40. [Numbered Example 44] The method of numbered Example 40, further comprising automatically performing one or more of adapting corrective loudness or dynamics audio processing based at least in part on whether that type of processing has already been performed on the input version of the media data. [Numbered Example 45] The method of numbered embodiment 40, further comprising extracting an input state of the media data from a data unit in the media data that encodes the media content. [Numbered Example 46] The method of numbered embodiment 45 further comprises restoring a version of the data unit that does not include the input state of the media data, and rendering the media content based on the restored version of the data unit. [Numbered Example 47] 47. The method of embodiment 46, further comprising obtaining an input state of the media data associated with the input version of the media data. Numbered Example 48 48. The method of numbered embodiment 47, further comprising authenticating the input state of the media data by verifying a cryptographic hash value associated with the input state of the media data. [Numbered Example 49] The method of numbered Example 47, further comprising authenticating the input state of the media data by verifying one or more fingerprints associated with the input state of the media data, at least one of the one or more fingerprints being generated based on at least a portion of the media data. [Numbered Example 50] 48. The method of numbered embodiment 47, further comprising verifying the media data by verifying one or more fingerprints associated with the input state of the media data. Numbered Example 51 48. The method of numbered embodiment 47, wherein the input state of the media data is carried along with the input version of the media data in an input media bitstream. Numbered Example 52 The method of Example 47, further comprising turning off one or more types of media processing based on the input state of the media data. [Numbered Example 53] The input state of the media data is described with processing state metadata, and the method further comprises: generating a media processing signal indicative of the input state of the media data based at least in part on the processing state metadata; transmitting the media processing signal to a media processing device downstream of the first device in the media processing chain. Method described in numbered Example 47. [Numbered Example 54] 54. The method of numbered embodiment 53, wherein the media processing signal is hidden in one or more data units in the output version of the media data. [Numbered Example 55] The method of numbered embodiment 54, wherein the transmission of the media processing signal is performed using a reversible data hiding technique such that one or more modifications to the media data can be removed by a receiving device. [Numbered Example 56] The method of numbered embodiment 54, wherein the transmission of the media processing signal is performed using an irreversible data hiding technique such that at least one of one or more modifications to the media data cannot be removed by a receiving device. Numbered Example 57 The method of numbered embodiment 46 further includes receiving one or more types of metadata from an upstream device in the media processing chain, the metadata being independent of any past media processing performed on the media data. [Numbered Example 58] 48. The method of numbered embodiment 47, wherein the state of the media data includes at least a portion of state metadata hidden in one or more secure communication channels. [Numbered Example 59] 47. The method of embodiment 46, further comprising modifying a plurality of bytes of the media data to store at least a portion of the state of the media data. [Numbered Example 60] The method of numbered embodiment 46, wherein the first device includes one or more of an Advanced Television Systems Committee (ATSC) codec, a Motion Picture Experts Group (MPEG) codec, an Audio Codec 3 (AC-3) codec, and an Enhanced AC-3 codec. Numbered Example 61 The media processing chain includes: a pre-processing unit configured to accept as input time-domain samples comprising media content and to output processed time-domain samples; an encoder configured to output a compressed media bitstream of the media content based on the processed time-domain samples; a signal analysis and metadata correction unit configured to verify processing state metadata within the compressed media bitstream; a transcoder configured to modify the compressed media bitstream; a decoder configured to output decoded time-domain samples based on the compressed media bitstream; a post-processing unit configured to perform post-processing of the media content within the decoded time-domain samples. Method described in numbered Example 46. Numbered Example 62 The method of numbered embodiment 61, wherein the first device includes one or more of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder, and the post-processing unit. Numbered Example 63 The method of numbered embodiment 61, wherein at least one of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder and the post-processing unit performs adaptive processing of the media content based on processing metadata received from an upstream device. Numbered Example 64 The method of Example 47, further comprising determining one or more media features based on a description of the one or more media features in the state of the media data. Numbered Example 65 The method of numbered Example 64, wherein the one or more media features include at least one media feature determined from one or more of frames, seconds, minutes, a user-definable time interval, a scene, a song, a piece of music, and a recording. [Numbered Example 66] 65. The method of numbered Example 64, wherein the one or more media features include a semantic description of the media data. Numbered Example 67 The method of numbered embodiment 64, further comprising performing one or more specific actions in response to determining the one or more media characteristics. Numbered Example 68 The method of numbered Example 43, further comprising providing to the second device in the media processing chain the state of the media data as one or more of: (a) a media fingerprint, (b) processing state metadata, (c) extracted media feature values, (d) media class type or subtype descriptions and / or values, (e) media feature class and / or subclass probability values, (f) cryptographic hash values, or (f) media processing signals. Numbered Example 69 computing, by a first device in the media processing chain, one or more data rate reduced representations of source frames of media data; and simultaneously and securely conveying the one or more reduced data rate representations within the media data itself to a second device in the media processing chain; A method performed by one or more computing devices. [Numbered Example 70] The method described in Example 69, wherein the one or more data rate reduced representations are carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients. [Numbered Example 71] 69. The method of claim 69, wherein the one or more data rate reduced representations include synchronization data used to synchronize audio and video delivered within the media data. [Numbered Example 72] The method of numbered Example 69, wherein the one or more data rate reduced representations include media fingerprints (a) generated by a media processing unit and (b) embedded in the media data for one or more of quality monitoring, media rating, media tracking, or content search. [Numbered Example 73] 70. The method of numbered Example 69, wherein the one or more data rate reduced representations include at least a portion of state metadata hidden in one or more secure communication channels. [Numbered Example 74] The method of numbered embodiment 69, further comprising modifying a plurality of bytes of the media data to store at least a portion of one of the one or more data rate reduced representations. [Numbered Example 75] The method of numbered embodiment 69, wherein at least one of the first device and the second device includes one or more of an Advanced Television Systems Committee (ATSC) codec, a Motion Picture Experts Group (MPEG) codec, an Audio Codec 3 (AC-3) codec, and an Enhanced AC-3 codec. [Numbered Example 76] The media processing chain includes: a pre-processing unit configured to accept as input time-domain samples comprising media content and to output processed time-domain samples; an encoder configured to output a compressed media bitstream of the media content based on the processed time-domain samples; a signal analysis and metadata correction unit configured to verify processing state metadata within the compressed media bitstream; a transcoder configured to modify the compressed media bitstream; a decoder configured to output decoded time-domain samples based on the compressed media bitstream; a post-processing unit configured to perform post-processing of the media content within the decoded time-domain samples. Method described in numbered Example 69. [Numbered Example 77] The method of numbered embodiment 76, wherein at least one of the first device and the second device includes one or more of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder, and the post-processing unit. [Numbered Example 78] The method of numbered Example 76, wherein at least one of the pre-processing unit, the signal analysis and metadata correction unit, the transcoder, the decoder, and the post-processing unit performs adaptive processing of the media content based on processing metadata received from an upstream device. [Numbered Example 79] The method of numbered Example 69, further comprising providing the state of the media data to the second device as one or more of: (a) a media fingerprint, (b) processing state metadata, (c) extracted media feature values, (d) media class type or subtype descriptions and / or values, (e) media feature class and / or subclass probability values, (f) cryptographic hash values, or (f) media processing signals. [Numbered Example 80] adaptively processing, by one or more computing devices in a media processing chain including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder or a stream processor, an input version of the media data based on a past history of loudness processing of the media data by one or more upstream media processing units as indicated by a state of the media data; normalizing the loudness and / or dynamic range of an output version of said media data at the end of said media processing chain to consistent loudness and / or dynamic range values; method. [Numbered Example 81] The method of numbered Example 80, wherein the consistent loudness values include loudness values that are (1) controlled or selected by a user or (2) adaptively signaled by conditions within the input version of the media data. [Numbered Example 82] 81. The method of numbered embodiment 80, wherein the loudness value is calculated for a dialogue portion of the media data. [Numbered Example 83] 81. The method of numbered Example 80, wherein the loudness values are calculated for absolute, relative and / or ungated portions of the media data. [Numbered Example 84] The method of numbered example 80, wherein the consistent dynamic range values include dynamic range values that are (1) controlled or selected by a user or (2) adaptively signaled by conditions within the input version of the media data. [Numbered Example 85] 85. The method of Example 84, wherein the dynamic range values are calculated for a dialogue (speech) portion of the media data. [Numbered Example 86] 85. The method of numbered Example 84, wherein the dynamic range values are calculated for absolute, relative and / or ungated portions of the media data. [Numbered Example 87] calculating one or more loudness and / or dynamic range gain control values for normalizing the output version of the media data to a consistent loudness value and a consistent dynamic range; and simultaneously conveying the one or more loudness and / or dynamic range gain control values within the state of the output version of the media data at the end of the media processing chain, wherein the one or more loudness and / or dynamic range gain control values are usable by another device to reverse-apply the one or more loudness and / or dynamic range gain control values to restore the original loudness values and original dynamic range in the input version of the media data. Method described in numbered Example 80. [Numbered Example 88] The method described in Example 87, wherein the one or more loudness and / or dynamic range control values representing the state of the output version of the media data are carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients. [Numbered Example 89] The method of numbered embodiment 80 further comprises calculating, by at least one of the one or more computing devices in the media processing chain, a cryptographic hash value based on the media data and / or the state of the media data and transmitting it within one or more encoded bitstreams carrying the media data. [Numbered Example 90] authenticating, by a recipient device, the cryptographic hash value; signaling by the recipient device to one or more downstream media processing units a determination of whether the state of the media data is valid; signaling, by the recipient device to the one or more downstream media processing units, the state of the media data in response to determining that the state of the media data is valid. Method described in numbered Example 80. [Numbered Example 91] The method of numbered embodiment 89, wherein the cryptographic hash value representing the media state and / or the media data is carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields or one or more transform coefficients. [Numbered Example 92] The method of numbered embodiment 80, wherein the state of the media data includes one or more of: (a) a media fingerprint, (b) processing state metadata, (c) extracted media feature values, (d) media class type or subtype descriptions and / or values, (e) media feature class and / or subclass probability values, (f) cryptographic hash values, or (f) media processing signals. [Numbered Example 93] 1. A method comprising performing, by one or more computing devices in a media processing chain including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder or a stream processor, one of inserting, extracting or editing related and unrelated media data positions and / or the state of related and unrelated media data positions in one or more encoded bitstreams. [Numbered Example 94] The method of numbered embodiment 93, wherein the one or more related and unrelated media data positions and / or the status of the related and unrelated media data positions in the encoded bitstream are carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields, or one or more transform coefficients. [Numbered Example 95] 1. A method comprising performing, by one or more computing devices in a media processing chain including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder or a stream processor, one or more of: insertion, extraction or editing of related and unrelated media data and / or states of related and unrelated media data in one or more encoded bitstreams. [Numbered Example 96] The method of numbered embodiment 95, wherein the one or more related and unrelated media data and / or the state of the related and unrelated media data in the encoded bitstream is carried in at least one of a substream, one or more reserved fields, an add_bsi field, one or more auxiliary data fields, or one or more transform coefficients. [Numbered Example 97] The method of numbered Example 93, further comprising providing from an upstream media processing device to a downstream media processing device the state of the media data as one or more of: (a) a media fingerprint, (b) processing state metadata, (c) extracted media feature values, (d) media class type or subtype descriptions and / or values, (e) media feature class and / or subclass probability values, (f) cryptographic hash values, or (f) media processing signals. [Numbered Example 98] A media processing system is provided that is configured to compute cryptographic hash values based on media data and / or a state of the media data by one or more computing devices in a media processing chain, including one or more of a psychoacoustic unit, a transform, a waveform / spatial audio coding unit, an encoder, a decoder, a transcoder, or a stream processor, for conveyance in one or more encoded bitstreams. [Numbered Example 99] The system of numbered Example 98, wherein the state of the media data includes one or more of: (a) a media fingerprint, (b) processing state metadata, (c) extracted media feature values, (d) media class type or subtype descriptions and / or values, (e) media feature class and / or subclass probability values, (f) cryptographic hash values, or (f) media processing signals. [Numbered Example 100] A media processing system configured to adaptively process media data based on a state of the media data received from one or more secure communication channels. [Numbered Example 101] A media processing system as described in Example 100, including one or more processing nodes, wherein the processing nodes include a media delivery system, a media distribution system, and a media rendering system. [Numbered Example 102] The media processing system of numbered embodiment 101, wherein the one or more secure communication channels include at least one secure communication channel traversing two or more of the compressed / encoded bitstream and PCM processing nodes. [Numbered Example 103] 102. The media processing system of embodiment 101, wherein the one or more secure communication channels include at least one secure communication channel spanning two separate media processing devices. [Numbered Example 104] 102. The media processing system of embodiment 101, wherein the one or more secure communication channels include at least one secure communication channel spanning two media processing nodes within a single media processing device. [Numbered Example 105] A media processing system as described in numbered embodiment 100, configured to perform autonomous media processing operations independently of the order of media processing systems in the media processing chain of which the media processing system is a part. [Numbered Example 106] The media processing system of numbered embodiment 100, wherein the state of the media data includes one or more of: (a) a media fingerprint, (b) processing state metadata, (c) extracted media feature values, (d) media class type or subtype descriptions and / or values, (e) media feature class and / or subclass probability values, (f) cryptographic hash values, or (f) media processing signals. [Numbered Example 107] 10. A media processing system configured to perform the method recited in any one of numbered embodiments 1-99. [Numbered Example 108] An apparatus having a processor and configured to perform the method recited in any one of numbered Examples 1-99. [Numbered Example 109] A computer-readable storage medium comprising software instructions that, when executed by one or more processors, cause the method described in any one of numbered Examples 1 to 99 to be performed.
[0166] 9. Equivalents, Extensions, Substitutions, etc. The foregoing specification describes possible embodiments of the invention, with reference to numerous specific details that may vary depending on the implementation. Accordingly, the only indication of what is, and is intended by applicant to be, the invention is the set of claims issuing from this application, in the specific form in which such claims are issued, including any subsequent amendments. The definitions expressly set forth herein for terms contained in such claims govern the meaning of such terms as used in such claims. Accordingly, no limitation, element, attribute, feature, advantage, or characteristic not expressly recited in a claim should in any way limit the scope of such claim. Accordingly, the specification and drawings are to be regarded in an illustrative, and not a restrictive, sense.
Claims
1. 1. A method of audio decoding comprising: receiving an encoded bitstream, the encoded bitstream including encoded input audio data and processing state metadata including loudness values and sample peak values; decoding the encoded input audio data; receiving signaling data indicating whether loudness processing should be performed on the decoded audio data; When the signaling data indicates that loudness processing should be performed on the decoded audio data: obtaining the loudness value from the processing state metadata; normalizing the loudness of the decoded input audio data to a consistent loudness according to the loudness value to provide output audio data; method.
2. The processing state metadata further indicates whether dynamic range processing should be performed, and the method further comprises: When the processing state metadata further indicates that dynamic range processing is to be performed, obtaining a dynamic range value from the processing state metadata; normalizing the dynamic range of the decoded input audio data in accordance with the dynamic range value. The method of claim 1.
3. The method of claim 1 , wherein the loudness value is calculated for a dialogue portion of the audio data.
4. receiving an encoded bitstream, the encoded bitstream including encoded input audio data and processing state metadata including loudness values and sample peak values; a decoder configured to perform the steps of: receiving signaling data indicating whether loudness processing should be performed on the decoded audio data; When the signaling data indicates that loudness processing should be performed on the decoded audio data: obtaining the loudness value from the processing state metadata; and normalizing the loudness of the decoded input audio data to a consistent loudness according to the loudness value to provide output audio data. Audio decoding system.
5. The processing state metadata further indicates whether dynamic range processing should be performed, and the post-processing unit further: When the processing state metadata further indicates that dynamic range processing is to be performed, obtaining a dynamic range value from the processing state metadata; configured to normalize a dynamic range of the decoded input audio data according to the dynamic range value. The system of claim 4.
6. 5. The system of claim 4, wherein the loudness value is calculated for a dialogue portion of the audio data.
7. A computer program for causing a computer to carry out the method according to any one of claims 1 to 3.
8. A storage medium having a software program adapted for execution on a processor to perform the method of any one of claims 1 to 3 when executed on a computing device.