Provenance determination and authentication of content using watermarks
Patent Information
- Application Number
- PCT/US2024/054814
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2024-11-06
- Publication Date
- 2025-08-07
AI Technical Summary
Existing technologies face challenges in authenticating and determining the provenance of linear broadcast streams and on-demand video assets, especially after undergoing various distribution routes, encoding, and transcoding processes.
The use of watermarks and cryptographic metadata to authenticate and verify the provenance of broadcast content, including the application of watermarks to pre-recorded assets and continuous service watermarking for linear playout streams, allows for the retrieval of canonical representations and validation of media objects.
This approach effectively authenticates and verifies the provenance of broadcast content, ensuring its integrity and origin, even after multiple distribution and processing stages, thereby combating misinformation and ensuring trust in digital media.
Smart Images

Figure US2024054814_07082025_PF_FP_ABST
Abstract
Description
PROVENANCE DETERMINATION AND AUTHENTICATION OF CONTENT USING WATERMARKSFIELD OF INVENTION
[0001] The present invention generally relates to provenance determination and authentication of content.BACKGROUND
[0002] This section is intended to provide a background or context to the disclosed embodiments that are recited in the claims. The description herein may include concepts that could be pursued but are not necessarily ones that have been previously conceived or pursued. Therefore, unless otherwise indicated herein, what is described in this section is not prior art to the description and claims in this application and is not admitted to be prior art by inclusion in this section.
[0003] Linear broadcast streams as well as on-demand video assets are subject to a variety of distribution routes and various encoding and transcoding processes. This presents challenges when attempting to authenticate and determine the provenance of such content downstream of these activities.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] FIG. 1 is a block diagram of metadata and watermarking architecture showing the relationship between the registered content distributors, media objects, their embedded watermarks, associated cryptographic metadata, and the canonical representations of the media object itself in accordance with an exemplary embodiment.
[0005] FIG. 2 illustrates an exemplary production flow in which watermarks are applied to enable provenance across multiple distribution paths, with asset watermarksapplied to pre-recorded assets and service watermarking continuously applied to a linear playout stream in accordance with an exemplary embodiment.
[0006] FIG. 3 illustrates a media object validation scenario in accordance with an exemplary embodiment.
[0007] FIG. 4 illustrates a media object canonical processing scenario in accordance with an exemplary embodiment.
[0008] FIG. 5 shows data hash segments, cryptographic metadata, and watermarks in the scenario of a live news broadcast with a watermark consisting of a constant service / asset identifier component and a time- varying index code in accordance with an exemplary embodiment.
[0009] FIG. 6 shows the creation of a canonical media object where the watermark in a live recording is time -varying in accordance with an exemplary embodiment.
[0010] FIG. 7 is a diagram of Broadcast OTT content provenance categories in accordance with an exemplary embodiment.
[0011] FIG. 8 illustrates a variety of OTT files which are canonical OTT candidates in accordance with an exemplary embodiment.
[0012] FIG. 9 illustrates a sequence diagram for creating the fMP4 replica and canonical OTT files in accordance with an exemplary embodiment.
[0013] FIG 10 illustrates the relationship of Server Code, Interval Code and fMP4Replica in accordance with an exemplary embodiment.
[0014] FIG. 11 illustrates the relationship of fMP4 replicas to broadcast channel in accordance with an exemplary embodiment.
[0015] FIG. 12 illustrates general validation steps when there is no embedded manifest in accordance with an exemplary embodiment.
[0016] FIG. 13 is a table showing OTT content watermark categories in accordance with an exemplary embodiment.
[0017] FIG. 14 illustrates an example of a single channel Canonical Candidate (4), enclosed within a single Data Hash Segment in accordance with an exemplary embodiment.
[0018] FIG. 15 illustrates an example of a single channel Canonical Candidate spanning a boundary between two Data Hash Segments in accordance with an exemplary embodiment.
[0019] FIG. 16 illustrates an example of a Multichannel Canonical Candidate in accordance with an exemplary embodiment.
[0020] FIG. 17 illustrates the Single Channel, single Data Hash Segment version of a Stitched Canonical Candidate in accordance with an exemplary embodiment.
[0021] FIG. 18 illustrates a single channel, single Data Hash Segment example of a Partial Canonical Candidate.
[0022] FIG. 19 illustrates the general Canonical OTT Substitution steps in accordance with an exemplary embodiment.
[0023] FIG. 20 shows the system architecture for the watermark reference model for C2PA in accordance with an embodiment of the invention.
[0024] FIG. 21 show a process sequence diagram of the watermark reference model for C2PA in accordance with an embodiment of the invention.
[0025] FIG. 22 shows screen shots of the step of a user searching for broadcast station news in accordance with an exemplary embodiment.
[0026] FIG. 23 shows a screen shot of the launch of a Broadcast News Authentication application, which is enabled by a watermark in accordance with an exemplary embodiment.
[0027] FIG. 24 shows the presentation of a QR code linked to a secure local broadcaster website prompted by the user accepting a call-to-action in accordance with an exemplary embodiment.
[0028] FIG. 25 shows the activation of the user’s mobile phone to capture an image of the QR code in accordance with an exemplary embodiment.
[0029] FIG. 26 shows the step of the user accessing a link to a secure broadcaster website through the mobile phone in accordance with an exemplary embodiment.
[0030] FIG. 27 shows the engagement with the secure local broadcaster website using the mobile phone to view authentic media and transcripts thereof directly from the local station.
[0031] FIG. 28 illustrates a block diagram of a device that can be used for implementing various disclosed embodiments.SUMMARY OF THE INVENTION
[0032] This section is intended to provide a summary of certain exemplary embodiments and is not intended to limit the scope of the embodiments that are disclosed in this application.
[0033] Systems and methods for provenance determination and authentication of broadcast content, such as OTT files, using watermarks. The method includes determining if the OTT file is linear broadcast content, and if so, determining if the OTTfile is one of the following categories: i) approved C2PA encoding with manifest removed; ii) approved C2PA encoding with endorsed C2PA transcoding with manifest removed; iii) approved C2PA encoding with unendorsed C2PA transcoding with manifest removed; or iv) approved C2PA encoding with unapproved mp4 transcoding. If the OTT file is one of these categories, the OTT file is validated. If the OTT file is not one of categories i), ii), iii), or iv), it is determined if the OTT file is one of the following categories: v) unapproved C2PA encoding with manifest removed; or vi) unapproved mp4 encoding. If the OTT file is one of these categories, a broadcast canonical OTT substitution is prepared for the OTT file.
[0034] Systems and methods for provenance determination and authentication of broadcast content, using watermarks. The method includes determining whether received media contains a manifest, and if so, determining if the manifest is trustworthy and validates the media, and if so, declaring that the media is validated. If the received media does not contain a manifest, or is not trustworthy, or does not validate the media, determining if the media contains a watermark. If the media does not contain a watermark, declaring that the media is unvalidated. If the media does contain a watermark, retrieving a manifest referenced by the watermark. Determining if the manifest is trustworthy and if not, declaring that the media is unvalidated. If the manifest is trustworthy, determining if the media validates the media and if so declaring that the media is validated. If the manifest is not trustworthy, notifying a user of a referenced asset, wherein the referenced asset is referenced by the retrieved manifest. If a request for the referenced asset is received, retrieving the referenced asset and determining if the manifest validates the retrieved asset, and if the manifest validates the retrieved asset, declaring that the retrieved asset is validated.
[0035] These and other advantages and features of disclosed embodiments, together with the organization and manner of operation thereof, will become apparent from the following detailed description when taken in conjunction with the accompanying drawings.DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS
[0036] In the following description, for purposes of explanation and not limitation, details and descriptions are set forth in order to provide a thorough understanding of the disclosed embodiments. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments that depart from these details and descriptions.
[0037] Additionally, in the subject description, the word “exemplary” is used to mean serving as an example, instance, or illustration. Any embodiment or design described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or designs. Rather, use of the word exemplary is intended to present concepts in a concrete manner.
[0038] (Final IBC paper) Interoperable Provenance Authentication of Broadcast Media using Open Standards-based Metadata, Watermarking and Cryptography
[0039] The spread of false and misleading information is receiving significant attention from legislative and regulatory bodies. Consumers place trust in specific sources of information, so a scalable, interoperable method for determining the provenance and authenticity of information is needed. The disclosed embodiments address the posting of broadcast news content to a social media platform, the role of open standards, the interplay of cryptographic metadata and watermarks when validating provenance, and likely success and failure scenarios. The open standards for cryptographically authenticated metadata developed by the Coalition for Provenance and Authenticity (C2PA) and for audio and video watermarking developed by the Advanced Television Systems Committee (ATSC) are well suited to address broadcast provenance. The disclosed embodiments teach methods for using these standards for optimal success.
[0040] Introduction
[0041] In our interconnected world, information flows ceaselessly, shaping opinions, policies, and societies. Within this digital torrent false and misleading information often obscures the truth.
[0042] False information may take the form of misinformation, spread when well- intentioned individuals share what they found online, neglecting to verify what they found. Or it may be disinformation, false or misleading information intentionally created and spread to deceive.
[0043] Both forms of false information are harmful, and both thrive in the global digital ecosystem. Social media platforms amplify their reach, turning falsehoods into viral storms. A rumor, a manipulated video, a fabricated statistic — these can cascade across screens, eroding public discourse.
[0044] Provenance and Authenticity
[0045] Any attempt to address false information on the web must proceed from an understanding of how people come to place trust in information.
[0046] The prevalence of information ‘bubbles’ demonstrates that people primarily place trust in specific sources of information. If information appears unaltered and from a trusted source, we often consider that information to be factual.
[0047] In other words, most of us judge what is factual based on the provenance and authenticity of the information, where provenance refers to the origin, history, and chain of custody of a piece of audio-video content, and authenticity refers to whether the content has been manipulated or altered in a way out of the control of the trusted source of the information.
[0048] The Role of Standards
[0049] There are two general methods for conveying provenance and authenticity metadata in association with audio-video content. Metadata can be cryptographically bound to the audio-video content, perhaps stored at the audio-video container level. Metadata can also be embedded as a watermark in the audio-video elementary stream.
[0050] For practical reasons described in herein these two metadata approaches are interdependent. Both cryptographic and watermarking provenance and authenticity methods should provide a reasonable degree of provenance assurance.
[0051] A critical issue to address is the impact of adopting proprietary solutions on interoperability and scalability, an issue often encountered. For example, fifteen years ago, Digital Rights Management (DRM) on the web had not yet been standardized. Prior to the 1SO / 1EC Common Encryption standard playback devices would have to support every major variety of digital rights management software, and there would be as many versions of the audio-video content as there were DRM systems. Had this continued it would have resulted in a combinatorial explosion, an effective barrier to large scale growth of commercial web media. It is no wonder that Netflix was one of the first companies to recognize the value of the common encryption standard.
[0052] It is reasonable to expect that the same will hold true for provenance and authenticity. For scalability and interoperability, the cryptographic metadata bound to the audio-video container and the watermark metadata embedded in the audio-video elementary streams must include open standard options.
[0053] A solution for provenance and authenticity for broadcast content distributed on social media platforms is described, utilizing metadata, watermarking and cryptographic standards. The disclosed embodiments demonstrate this can be used with broadcast news content while pointing out several important implementation considerations.
[0054] Provenance and Authenticity Success Scenarios
[0055] A provenance and authenticity use case can encompass multiple scenarios, including success scenarios, where everything goes roughly as intended and various exception scenarios, which lead to undesirable outcomes. All these scenarios should be describable as discrete programmatic steps to uncover the functional requirements for addressing provenance and authenticity in practice.
[0056] The following disclosure will examine the details of one specific, provenance and authenticity use case - the posting of what appears to be broadcast news content to a social media platform. What is particularly interesting is the interplay between provenance validation using tamper-evident cryptographic bindings and metadata retrieval using elementary stream watermarks, with a focus on the constituent ‘success’ and ‘exception’ scenarios.
[0057] A broadcaster produces content for linear distribution by an affiliate / network / platform operator. This content consists of a series of audio-video programs comprising a single linear broadcast TV channel.
[0058] There are a variety of scenarios where some of the broadcaster content finds its way into Internet distribution and is uploaded to a social media platform. At a minimum it will then be transcoded into a platform’s preferred framerates, resolutions, bitrates, codecs, and container formats. It may also be truncated to meet the platform’s maximum size limits.
[0059] Verifying Authenticity
[0060] Before being posted to a social media platform, broadcast news content may be altered such that there are observable, meaningful differences between what was depicted in the original broadcast and the posted video. This manipulation could be done for artistic, creative, or deceptive reasons, depending on the intention of the editor.
[0061] One way to characterize these differences is to ask whether the posted video is an authentic representation of the original, whether it is true to the original, without anyjudgement as to whether the original itself depicted what transpired in front of the camera lens and microphone.
[0062] In this definition, an authentic representation of the broadcaster content may not be bit-wise identical to the original, it may be an unaltered clip from the original, or it may be a transcoding of the original, but it may nonetheless accurately represent what was depicted by the original.
[0063] The C2PA standard provides a mechanism for editing the original - e.g., transcoding or clipping - and a means to cryptographically verify the authenticity of the result, but this requires that the tool used to alter the broadcaster content supports the use of this standard. We believe it is highly unlikely that in the near-temi social media platforms will reject content that was edited using a tool that does not implement cryptographic metadata standards.
[0064] Verifying Provenance
[0065] When consuming video in a linear TV receiver, consumers quite reasonably believe that the network / platform operator is accurately identifying the channel and content creator.
[0066] Video content on the Internet might misrepresent the identity of the original content creator, the identity of who subsequently transcoded the video and what authorized or unauthorized changes were made.
[0067] One way to characterize this history is to use the term “provenance,” meaning, the identifiable source of the content and an accurate history of the content’s transformation from that source.
[0068] There are cryptographic methods for verifying the provenance of video posted to a social media platform using the C2PA standards, but again it is unlikely in the short-term that social media platforms will reject videos that do not enable the use of these methods to identify the source of the original content.
[0069] Canonical Representation of a Media Object
[0070] A tamper-evident cryptographic binding to an audio-video media object which contains provenance information can be used to validate the provenance and authenticity of that object. It is surely a successful outcome if the provenance is validated, but what should happen if the provenance and authenticity fail to be verified? What constitutes success in this scenario?
[0071] There are multiple scenarios where the content may have been innocently modified by the user when preparing to post to a social media platform, since even a single bit change to the content will invalidate a cryptographic binding. Generating numerous alerts for innocent alterations to the media could lead to “security alert fatigue,” diminishing user trust in the alert’s salience. Doing nothing is also an unattractive option because it leaves users blind to the “trust signal” conveyed by the presence of a provenance assertion.
[0072] Should the cryptographic verification of the provenance and authenticity of a media object fail a successful outcome is for the social media platform to use an embedded watermark to retrieve the authoritative version - the canonical representation of the media object. This can be done by first using the watermark to retrieve the cryptographic metadata associated with the original media object as distributed and then use that trusted metadata to retrieve the media object’s canonical representation.
[0073] An Approach to Authentication using Metadata and Watermarking
[0074] Architecture
[0075] The provenance and authenticity approach in this paper builds on the relationship between the registered content distributors, media objects, their embeddedwatermarks, associated cryptographic metadata, and the canonical representations of the media object itself. This relationship is shown in FIG. 1.
[0076] Security Model
[0077] If the tamper evident cryptographic metadata associated with a media object is stored as a component of the media object container, it is relatively easy to remove. A durable embedded watermark can enable cryptographic metadata to be brought back into association with the media object.
[0078] Watermark security is typically maintained by making the watermark difficult to remove or alter by keeping the watermark technology secret. This approach works against availability and interoperability by demanding hardened implementations and strict access controls. It can also provide only weak security assurances because its secrecy impedes comprehensive security assessment. Because recorded broadcast content has a long lifespan on the Internet, the security of “closed” watermarks requires successful long-term protection of the associated secrets. And furthermore, recent advances in attacks on watemiarking have demonstrated that advances in artificial intelligence render even robust, secret watermarks automatically removable, further diminishing their potential advantages.
[0079] This motivates a security approach that does not treat the watermark as a root of trust. Instead, we assume that they are durable, i.e. that they survive content processing that causes traditional metadata formats to be lost, but that they are otherwise as mutable as traditional metadata and can be modified or removed by any intermediary. Tike traditional metadata, data conveyed via watermarking is treated as untrusted and must be validated using cryptographic methods.
[0080] This same approach was advocated by England et al. in their foundational work on media provenance authentication. See: England, P. et al. 2021. AMP:Authentication of Media via Provenance. 12th ACM Multimedia Systems Conference. July 2021.
[0081] That work, however, assumed the presence of a signature in the watermark payload. We view that signature as unnecessary and assume that the watermark carries only a URL and media timeline. The root of trust is a manifest that has been retrieved using the watermark and cryptographically validated using an appropriate trust list.
[0082] Watermarking Audio-Video Content
[0083] Our success scenario demands a path to validated content regardless of distribution source, which for broadcasters must encompass both linear and on-demand delivery. To achieve this, it must be possible to apply watermarking in asset-based digital publishing as well as within the live production chain.
[0084] FIG. 2 illustrates an exemplary production flow in which watermarks are applied to enable provenance across multiple distribution paths, with asset watermarks applied to pre-recorded assets and service watermarking continuously applied to a linear playout stream.
[0085] The broadcaster may produce content to be published on their website (2-1).They would apply an asset watermark (2-2), generate cryptographic metadata or a “manifest” for that content (2-3), store the asset (2-10) and distribute the asset to their website (2-4).
[0086] The broadcaster may also want to take live and third-party assets (2-5) and prepare them for linear playout (2-6). They would apply a service watermark with a time varying component (2-7), distribute the content (2-8) and periodically generate cryptographic metadata or “manifest” information for that broadcast (2-9).
[0087] A watermark can be used to retrieve the associated, static cryptographic metadata and canonical content. And the time- varying service watermarks can be used toretrieve the associated, time-varying cryptographic metadata and canonical content (2- 10).
[0088] Validating Content Authenticity using Watermarking
[0089] FIG. 3 and FIG. 4 summarize how watermarks can be integrated into the content validation process. FIG. 3 illustrates a media object validation scenario. FIG. 4 illustrates a media object canonical processing scenario.
[0090] Media validation uses the cryptographic metadata which may be stored in the media object, distributed with the media object and / or retrievable from the cloud. This object is referred to as a ‘manifest’ in the C2PA standard referenced above.
[0091] Media object validation scenarios
[0092] If the content can be validated by a contained or retrieved manifest, then a successful outcome does not require utilizing canonical content.
[0093] If the media contains a manifest (3-1), the manifest corresponds to a registered distributor (3-2), and the manifest validates the media object (3-3), validation is achieved without reference to a watermark. We would view this as a success scenario.
[0094] Otherwise, if the media object does not contain a watermark (3-4), the media remains unvalidated, this is an exception scenario.
[0095] If the media does contain a watermark (3-4) then the manifest is retrieved from the manifest cloud store (3-5). If this retrieval fails, for example if the URI Authority field provided in the watermark does not correspond to a registered broadcaster, the media object is not validated, an exception scenario.
[0096] If the retrieved manifest’s digital signature does not correspond to an approved broadcaster (3-6), then the media object cannot be validated. Another exception scenario.
[0097] Otherwise, if the manifest’s digital signature is trustworthy (3-6) and the manifest validates the content (3-7), validation is achieved by using the watermark. A success scenario.
[0098] Media object canonical representation scenarios
[0099] If a retrieved manifest (3-5) is trustworthy (3-6) but it does not validate the content (3-7), then it is the view of this paper that the only success scenarios involve canonical processing.
[0100] The decision to perform canonical processing (4-8) can be made by the user posting the content or by the platform supporting the validation logic, depending on the policy being adhered to by the social media platform.
[0101] If the decision is to perform canonical process (4-9), the validator retrieves the canonical content (4-10).
[0102] The previously retrieved manifest (3-5) should always validate the canonical content (4-11). If it does not, it is an error and an exception scenario.
[0103] We view validation of the retrieved canonical content as optional because its retrieval location has been established as trusted through validation of the manifest that contains it (3-6) (3-7).
[0104] Media object canonical processing
[0105] The availability of a canonical version of the media object presents the social media platform with additional success scenario opportunities, including one or more of the following:
[0106] a) Posting the uploaded content together with the retrieved asset, or a link to it.
[0107] b) Providing the uploader with a choice between which version of the content should be posted and posting that version with an appropriate label.
[0108] c) Performing an automated comparison of the uploaded and reference asset to determine the nature and amount of difference between the two.
[0109] d) Automatically replacing the uploaded content with the valid asset content.
[0110] 3) Forwarding the uploaded content and the retrieved asset to an internal content moderation process.
[0111] The Above Approach Applied to Live Broadcast
[0112] Low Latency Considerations
[0113] Real-time broadcast and live streaming, often referred to as “glass to glass,” is a process where content is captured through a camera lens and transmitted to a viewer’s screen with minimal delay. Although it is a real-time transmission, it always involves some degree of delay or latency, incidental and / or intentional.
[0114] Live scenarios may be categorized by the degree of latency required. This depends on the content’s nature and the desired viewer experience. Real-time, low latency live is essential for live sports and breaking news, where timely viewing is important. Higher latency can be introduced for any number of reasons. For example, content is often recorded, edited, or processed before broadcast.
[0115] Using digital signatures to protect provenance metadata for ‘glass to glass’ real-time live streaming scenarios is technically challenging. The primary difficulty is that performing a digital signing operation on a Content Delivery Network (CDN) edge server is not adequately secure and performing that operation in a Hardware Security Module (HSM) is unlikely to achieve the low latency desired.
[0116] However, any live scenario where the content is captured downstream, edited, and subsequently posted will introduce an inherent latency sufficient to allow the use of an HSM for provenance metadata protection.
[0117] Live Broadcast News Content Posted to a Social Media Platform
[0118] A 30-minute evening news program is broadcast. The live broadcast is captured and recorded on a device downstream of an HDMI port. A 20-second clip of the news broadcast is created as an MP4 file and posted to a social media platform.
[0119] No manifest can be present with the content since only the elementary stream will make its way beyond the HDMI port. The social media platform can examine the posted video for a watermark, but what would be the success scenario?
[0120] The fragmented MP4 broadcast replica
[0121] The following approach provides a reasonable degree of provenance and authenticity assurance for broadcast news content posted to social media platforms.
[0122] The live news program is broadcast with a watermark consisting of a constant service / asset identifier component and a time-varying index code. See FIG. 5, which shows data hash segments, cryptographic metadata, and watermarks in the scenario of a live news broadcast with a watermark consisting of a constant service / asset identifier component and a time- varying index code. This watermark approach is in use today for delivering metadata for interactive television services [1][2][3] and can readily support the retrieval of provenance metadata without the need to modify the watermark itself. See the previously mentioned ATSC 3.0 specifications A / 334, A / 335, and A / 336.
[0123] The broadcaster or the network / platform operator on behalf of the broadcaster produces a secure transcoding of specified portions of the live linear broadcast into a fragmented MP4 format - an ‘IMP4 Replica’ of the portion of the linear broadcast for which provenance and authenticity is to be established.
[0124] Periodically a C2PA manifest is produced for this fMP4 Replica. In this design the portion of the Replica that each manifest corresponds to is defined as the Data Hash Segment (DHS). The real-time duration of a DHS defines a minimum lag time behind the linear live edge for the availability of DHS Replica Manifests.
[0125] The Replica itself consists of a sequence of fragmented MP4 segments or chunks for each track, adequate to cover the length of the Data Hash Segment. Each segment or chunk includes auxiliary ‘c2pa’ boxes (defined in the previously mentioned C2PA specification) which can be used by a C2PA validator to validate any portion of the DHS Replica, as described below.
[0126] Apart from the addition of c2pa-specific ISOBMFF boxes, the Replica format is identical to the format in common use for adaptive bitrate streaming, the Common Media Application Format or CMAF. See: ISO / IEC 23000-19:2020, “Information technology - Multimedia application format (MPEG- A) - Part 19: Common media application format (CMAF) for segmented media”.
[0127] The fragmented MP4 replica C2PA manifest
[0128] The fMP4 Replica Manifest is constructed in the exact same way as a C2PA Manifest for audio-video streaming. See section 9.2.3 of the C2PA specification.
[0129] Before the manifest is generated, a DHS initialization segment is produced for the content stored in the DHS Replica. The cryptographic metadata stored in this initialization segment is identical to that specified by C2PA for adaptive bitrate delivery. See Section 9.2.3 of the C2PA specification.
[0130] The c2pa-specific box in each track’s initialization segment will contain the C2PA manifest, which, as is the case for adaptive delivery, must be identical across tracks. The Manifest’s c2pa.bmff.hash assertion will contain CBOR with an array of Merkle rows, one per track.
[0131] In the C2PA specification for adaptive delivery provenance validation, the Merkle tree associated with the entire video stream enables piecewise validation of individual fragment components of the stream without access to the entire stream [4] . The same mechanism enables piecewise validation of arbitrary portions of the DHS using a single DHS manifest.
[0132] Producing a canonical live recording
[0133] If the watermark in the live recording is time-varying, it can be used create a canonical live recording, as shown in FIG. 6.
[0134] The time-varying watermark is used to derive the manifest recovery URL. OTT BINX and OTT EINX are the time indices (Interval Code) corresponding to the start and end of the posted video, respectively.
[0135] A recovery request (6-1) is sent. The DHS Manifest is provided in a recovery response (6-2).
[0136] This recovery request response is identical to the method used today for interactive television. The only change is in the payload of the response from the provenance-authenticity server.
[0137] The DHS corresponding to the manifest is accessed from the Asset Reference Assertion in the DHS Manifest (6-3).
[0138] The IMP4 Replica is used as an authenticated mezzanine format, to produce a canonical MP4 representation of the posted live content. Any portion of the Data Hash Segment can be validated with the DHS Manifest.
[0139] Beginning with the OTT BINX (6-4), the algorithm walks the Data Hash Segments provided in the Recovery Response (6-7) until the EIDX of the posted live recording is reached (6-5).
[0140] The Present Approach Applied to Web Published Content
[0141] Differences without a Distinction
[0142] During validation of content posted to a social media platform, even the slightest alteration to the content can cause it to be flagged as inauthentic. There are multiple scenarios where the content may have been innocently modified by a user, making their edits from a provenance perspective a ‘difference without a distinction’.
[0143] As discussed, we believe the successful outcome for a validation failure to be for the social media platform to recover the original content and use it in one of the ways we outlined. There are cases, however, where this too will result in an unsuccessful outcome.
[0144] Clipped Web Published News Content Posted to Social Media
[0145] One of the most likely such cases is what we are calling “the clipped news segment” scenario.
[0146] Consider the following example. The broadcaster publishes a 30-minute evening news program to their website. The published video file includes cryptographic metadata, and it is watermarked. A user wishes to share a 20-second clip from that 30- minute program. They download the broadcaster published video file, edit it to produce a 20-second clip, and attempt to post it to a social media platform.
[0147] If the editing tool the user used removed the manifest, the social media platform can recover the metadata using the watermark, as described above. Regardless, the file will be flagged as inauthentic. And recovering the canonical version of the content will result in a 30-minute post.
[0148] Producing the canonical news clip
[0149] If the watermark in the clipped news content is time-varying, it is used to derive the manifest recovery URL, a recovery request (1) is sent where BINX is the time index corresponding to the start of the clip. The retrieved DHS Manifest in a recovery response (2) can be used to produce a canonical version of the news clip, as shown in FIG. 6, following the same steps as producing a canonical live recording.
[0150] Since social media platforms transcode posted video into a multitude of targeted formats, it is likely that they would treat the DHS as a canonical mezzanine format to produce a wide variety of device targeted formats. Using fragmented MP4 as a mezzanine format is commonly done. In addition, the MPEG DASH specification provides support to access segments of presentations at a specified media time through the use of an MPD Anchor, using a query parameter "t=" that a client can append to an MPD URL with either an NPT or UTC time. This could be used to access portions of the fMP4 Replica.
[0151] Broadcast Provenance and Authenticity Problem Statements
[0152] How to Link Provenance and Authenticity from Linear to OTT Files
[0153] There are two general TV distribution models: linear (satellite direct-to-home, terrestrial cable / IP / DTT) and Over-the-Top (OTT). Satellite and terrestrial network distributed content is extremely difficult to maliciously alter before it arrives at a receiver. Consequently, when consuming content in a linear TV receiver, consumers quite reasonably expect that the network / platform operator is accurately identifying the provenance (broadcaster) and that the content being rendered at their linear receiver is authentic, viz., unmodified from what was broadcast.
[0154] The C2PA specifications described above provide a means for representing provenance / authenticity for OTT video content. The specifications describe how to manage that provenance / authenticity through a series of content modifications,transcoding, and edits, in a provenance chain from the active C2PA manifest to each preceding manifest in the manifest store.
[0155] Although there is an implicit provenance / authenticity assurance in the linear TV distribution upstream from the receiver, the linear distribution does not include a C2PA manifest.
[0156] How to Validate Provenance and Authenticity for broadcast OTT Files
[0157] Once we have a way of linking provenance and authenticity between linear and OTT content, we then need a way to validate provenance and authenticity for content which to an end user might appear to be broadcaster OTT content.
[0158] Provenance Categories for Broadcast OTT Files
[0159] From the perspective of provenance, there are ten types of broadcaster OTT content which a C2PA validator may encounter. All of this content begins with authentic segments of the linear broadcast stream.
[0160] Note that this “omniscient” view of the provenance of the OTT content is not what a validator is assumed to possess. It is the job of the validator to determine the provenance / authenticity of the content.
[0161] FIG. 7 is a diagram of Broadcast OTT content provenance categories in accordance with an exemplary embodiment.
[0162] With reference to FIG. 7:
[0163] 1) ‘ appro ved / unapproved C2PA encoding’ means an mp4 encoding of the linear broadcast content which is c2pa-compliant, performed with / without the approval of the broadcaster.
[0164] 2) ‘unapproved mp4 encoding’ means an mp4 encoding of the linear broadcast content which is not c2pa-compliant (all broadcaster-approved encoding is assumed to be c2pa-compliant).
[0165] 3) ‘unapproved mp4 transcoding’ means an mp4 transcoding which is not c2pa-compliant (all broadcaster-approved transcoding is assumed to be c2pa-compliant).
[0166] 4) ‘ endorsed / unendorsed C2PA transcoding’ takes the specific meaning of transcoding endorsement as given in the C2PA specifications.
[0167] 5) ‘ C2PA compliance’ provenance categories are provenance threats which should be C2PA signing key compliance rule violations; viz., it should be a C2PA signing key compliance rule violation to sign a C2PA claim for an h) ‘unapproved’ C2PA encoding of a broadcaster’s content or e) an ‘unendorsed’ C2PA transcoding of linear broadcast content.
[0168] a) Approved c2pa-compliant encoding of linear broadcast content. Compliant OTT content. This is a C2PA encoding from linear with the approval of the broadcaster which is C2PA compliant. Approval is by a method that is out-of-band from the C2PA specification.
[0169] Example: the broadcaster might have a business arrangement with the network / platform operator giving them permission to create a c2pa-compliant mp4 encoding of all or a portion of the broadcast content.
[0170] b) Manifest removed from approved c2pa-compliant encoding. A non- compliant state. This is category a) with the manifest removed.
[0171] Example 1 : a malicious actor has removed the manifest of an approved C2PA encoding of a news broadcast to produce disinformation.
[0172] Example 2 : The manifest is lost as the result of a legacy multiplexer implementation that discards metadata that it thinks is superfluous.
[0173] c) Endorsed C2PA transcoding of approved c2pa-compliant encoding. Compliant OTT content. This is category a) subsequently transcoded to another format with an endorsement assertion in the C2PA manifest claim.
[0174] Example: the broadcaster endorses a Facebook transcoding an approved C2PA encoding to that social media platform’s native format.
[0175] d) Manifest removed from endorsed C2PA transcoding of approved c2pa- compliant encoding. A non-compliant state. This is category c) with the manifest removed.
[0176] Example 1 : a malicious actor has downloaded an endorsed C2PA transcoding of a news broadcast and removed the manifest in order to produce disinformation.
[0177] Example 2: a broadcaster tools / infrastructure limitation results in the manifest being lost.
[0178] e) Unendorsed C2PA transcoding of an approved c2pa-compliant encoding. Compliant with the C2PA specifications, but not compliant with this proposed specification. A C2PA signing key compliance issue. This is category a) subsequently transcoded to another format without an endorsement assertion in the C2PA manifest claim.
[0179] Example: the broadcaster does not endorse a TikTok transcoding of an approved C2PA encoding to that social media platform’s native format. TikTok has a C2PA signing certificate but permits uploading the broadcaster’s content without an endorsement assertion.
[0180] f) Manifest removed from an unendorsed transcoding of an approved C2PA encoding. A non-compliant state and a C2PA signing key compliance issue. This is category e) content with the manifest removed.
[0181] Example: the broadcaster does not endorse a Reddit transcoding of an approved C2PA encoding to that social media platform’s native format. Reddit has a C2PA signing certificate but permits uploading the broadcaster’s content without an endorsement assertion. A Reddit user downloads the broadcaster content and removes the manifest to produce disinformation.
[0182] g) An unapproved mp4 transcoding of an approved C2PA encoding. A non- compliant state. This is category a) transcoded to an mp4 format which is not c2pa- compliant.
[0183] Example: a social media platform does not support c2pa. They transcode an approved c2pa-compliant encoding removing the C2PA metadata.
[0184] h) Unapproved c2pa-compliant encoding of linear broadcast content.Compliant with the C2PA specification but not with this proposed specification. A C2PA signing key compliance issue. This is a C2PA encoding from linear without the approval of the broadcaster, but which is otherwise C2PA compliant.
[0185] Example: A network / platform operator with a C2PA signing key decides to produce c2pa-compliant encodings directly from the linear broadcaster content, without an agreement to do so with the broadcaster.
[0186] i) Manifest removed from an unapproved, c2pa-compliant encoding. A non- compliant state and a C2PA signing key compliance issue. This is category h) content with the manifest removed.
[0187] Example: the broadcaster does not endorse Comcast encoding their linear content into a c2pa-compliant format. Comcast has a C2PA signing certificate andencodes the broadcaster content for distribution to Comcast subscribers. A Comcast user downloads the broadcaster content and removes the manifest to produce disinformation.
[0188] j) Unapproved MP4 encoding of linear content. A non-compliant state. Broadcaster content encoded without the approval of the broadcaster, and which is not c2pa-compliant.
[0189] Example: an analog capture of the broadcast, encapsulated as an MP4 file and posted to a social media site.
[0190] Validation Functional Requirements
[0191] The functional requirements for validating the provenance and authenticity of OTT which are C2PA compliant with an intact C2PA manifest are documented in the C2PA specifications (see section 15. Validation of the C2PA specification). It is the steps required of a validator when there is no C2PA manifest that are important to this specification.
[0192] The requirements when processing an OTT File without a C2PA manifest are:
[0193] 1. Determine whether this is linear broadcast content.
[0194] 2. If it is linear broadcast content, determine its category (see FIG. 7).
[0195] 3. If it is category b), d), I), or g) - attempt to validate the OTT File (see theBroadcast Provenance Validation Section) and optionally prepare a Broadcast Canonical OTT Substitution for the File (see the Broadcast Canonical OTT Substitution Section).
[0196] 4. If it is category i), or j), prepare a Broadcast Canonical OTT Substitution for the File.
[0197] Broadcast Provenance and Authenticity Services
[0198] The following solution is an exemplary embodiment.
[0199] Overview
[0200] The broadcaster or the network / platform operator on behalf of the broadcaster produces a secure transcoding of defined portions of the linear broadcast (see FIG. 9) into a fragmented MP4 format - an 1MP4 Replica of the linear broadcast.
[0201] The fMP4 Replica or simply “Replica” is intended to never be moved from secure storage.
[0202] The Replica has no boundary and corresponds to the entire linear broadcast. Periodically a Replica Manifest is produced for a defined range of the Replica. The Replica Manifest contains the provenance and authenticity C2PA metadata needed to validate a broadcast-originated OTT file.
[0203] The portion of the Replica that each Replica Manifest spans is the Data Hash Segment (DHS). A DHS may correspond to a single linear broadcast program, multiple programs, or a portion of a program. This is a broadcaster decision. One consideration will be that the real-time duration of a DHS defines a minimum lag time behind the linear live edge for the construction of Canonical OTT files and the availability of Replica Manifests to be provided to C2PA validators.
[0204] Canonical OTT files are mp4 files produced from the fMP4 Replica. They correspond to category a) in FIG. 7. There is a single canonical OTT file corresponding to each Data Hash Segment. Additional Canonical Files corresponding to multiple Data Hash Segments, or a portion of a single Data Hash Segment can be produced, all derived from the Replica.
[0205] OTT Files which are produced by clipping Canonical OTT files are Clipped Canonical OTT Files. They can be validated using the same Replica Manifest that corresponds to the entire FIG. 7, OTT Files produced from Canonical OTT Files fall into categories b) through g). C2PA Compliant files with an intact manifest - categories a), c)and e) can be validated without reference to watermarks, so it is only categories b), d), f) and g) which concern us in this specification.
[0206] When a C2PA validator encounters an OTT File without a manifest (see FIG. 7):
[0207] If the OTT File is a Canonical Candidate - categories b, d), f) or g) - the validator can attempt to validate the content using the appropriate algorithm as described in the Broadcast Provenance Validation section.
[0208] In all categories absent a C2PA manifest - b), d), f), g), i) and j) - the C2PA validator can provide a Canonical OTT version of the linear broadcast as a substitute for the OTT File, including for Clipped Canonical Files, following the process described in the Broadcast Canonical OTT Substitution section.
[0209] The variety of Canonical Candidates are shown in FIG 8. Whether an OTT File is a Canonical Candidate can be determined by examination of the VP1 watermark, using the logic found in Table 1 shown in FIG. 13.
[0210] The following types of OTT Files found in the wild are discussed in both the Broadcast Provenance Validation and the Broadcast Canonical OTT Substitution sections, below.
[0211] OTT Files which have a VP1 Server Code throughout and a consistent, monotonically increasing VP1 Interval Code are Single Channel Canonical Candidates (6.3) As used herein, the term “Server Code” means a value conveyed in a watermark embedded in media content that identifies a remote resource on a network that provides metadata associated with the content. The term “Interval Code” means a value conveyed in a watermark embedded in a specific region of a media content item that can be used in combination with metadata retrieved from a remote resource, to obtain metadata associated with the region of media content. For additional details regarding ServerCodes and Interval Codes reference is made to U.S. Patent No. 9, 596,521, owned by Verance corporation and incorporated herein by reference.
[0212] Stitched Canonical Candidates containing content from multiple channels are Stitched Multichannel Canonical Candidates (6.4).
[0213] OTT Files which have a single VP1 Server Code throughout but gaps in the Interval Code are Stitched Single Channel Canonical Candidates. If validated, they may prove to be Clipped Canonical OTT Files (6.5).
[0214] OTT Files with watermark gaps but contain regions which would otherwise be considered Canonical Candidates are Partial Single Channel Canonical Candidates (6.6).
[0215] Hybrid variations are also possible. For example, Partial Multichannel Canonical Candidates, Stitched, Partial Single Channel Canonical Candidates, and so forth. The important point is that since they are Canonical Candidates, it is possible to retrieve DHS Manifests to validate them.
[0216] In summary, there are two broad categories of Broadcast Linear OTT Files without a manifest, Canonical OTT Candidates - b), d), f) and g) - and Substitution Candidates - i) and j). Canonical OTT Candidates can be validated and / or replaced, but Substitution Candidates can only be replaced.
[0217] Generating the fMP4 Replica
[0218] The IMP4 Replica of the linear broadcast is created piecewise, corresponding to a series of linear data hash segments. As each data hash segment processing is completed, a C2PA manifest is produced for the data hash segment, including a newly defined IMP4 Replica hard binding assertion (5.3) and a VP1 soft binding assertion (5.4).
[0219] Once this Replica manifest has been produced, the Replica can be used to construct arbitrary c2pa-compliant MP4 canonical OTT files, based on time ranges asindicated by VP1 Server Code / Interval Code values (5.6), and the Broadcast Provenance Service (5.5) can assist C2PA validators (6) to validate broadcast OTT content which are missing a manifest.
[0220] FIG. 9 shows a sequence diagram for creating the fMP4 replica and canonical OTT files. In FIG. 9, the following steps are performed to produce the fMP4 Replica:
[0221] 1. Using the broadcaster preference, the automation system selects the begin and end time for each of the data hash segments, using the broadcaster’s preferred timeline (e.g., PTS, TEMI, composite time, etc.)
[0222] 2. The automation system signals the watermark embedder to stripe the media essence with VP1 field 1 = the Server Code and with field 2 = Interval Code, a monotonically increasing value. For this Data Hash Segment, the first value of field 2 is defined as BINX, the last value is defined as EIDX, and the Watermark Media Time (WMT) derived from the broadcaster’s preferred timeline.
[0223] 3. The watermark embedder then signals the Prov / Auth Server to transcode the data hash segment into C2PA compliant IMP4 segments for each track, each segment consists of a moof box and an mdat box, where the moof box contains metadata about the precise byte range locations of the mdat segments, and the mdat box contains the actual media data.
[0224] 4. Once completing the transcoding of the entire Data Hash Segment - fromInterval Code = BINX to Interval Code EIDX, the Prov / Auth Server produces a Canonical OTT File, an MP4 consisting of the moof and mdat boxes of the MP4 Replica which correspond to one Data Hash Segment.
[0225] 5. The Prov / Auth Server next constructs a C2PA Replica Manifest for thatData Hash Segment, including fMP4 Replica Hard Binding (5.3) and VP1 Soft Binding (5.4) assertions.
[0226] 6. The Prov / Auth Server then stores the Replica Manifest and its corresponding Canonical OTT File, retrievable from the VP1 URL derived from Field 1 = Server Node and BINX <= Field 2 <= EIDX.
[0227] Note that while the example described here synchronizes the locations of the VP1 Payload and Data Hash Segments boundaries, in some broadcast production environments such synchronization may be difficult to achieve. This synchronization is not required, so long as the location of the VP 1 Payload and Data Hash Segment boundaries on a common media presentation timeline are both recorded in the Canonical OTT file index. This index can be recovered and used, when a segment of watermarked media content is discovered and its extent on the media timeline is determined, to: (a) determine a set of Data Hash Segments that fully lie within the watermark segment and attempt to validate that portion of the media file; or (b) identify a set of Data Hash Segments within a Canonical OTT file that fully contain the watermarked segment and can be validated.
[0228] The fMP4 Replica Hard Binding
[0229] The fMP4 Replica hard binding assertion is constructed in the exact same way as a C2PA manifest for audio-video streaming (see 11.3.2, “Embedding manifests into BMFF-based assets”).
[0230] The fMP4 Replica consists of a large sequence of fragmented MP4 segments or chunks for each track, adequate to cover the length of the Data Hash Segment. Each segment or chunk includes auxiliary ‘c2pa’ boxes (see 11.3.2.3. Auxiliary 'c2pa' Boxes for Large and Fragmented Files) which are used by a C2PA validator to validate the asset.
[0231] The C2PA streaming and the Replica manifest and hard binding assertions are constructed the same, but the Replica manifest differs in two important ways - how it isused to validate OTT files (6.2), and how it is used to produce canonical OTT files (5.6) and perform Broadcast Canonical OTT Substitution (7).
[0232] The VP1 Soft Binding
[0233] Advanced Television systems Committee (ATSC) A / 336 standard specifies an open, extensible standard by which broadcasters can apply VP 1 or similar watermarks to their linear services, publish time-based metadata associated those services, and enable clients equipped with watermark detectors to acquire and use that time-based metadata from versions of the linear service received from broadband, broadcast, or pay TV sources. For further information regarding VP1 watermarks, see US patent Pub.20150261753, owned by Verance Corporation, which is incorporated herein by reference.
[0234] The VP 1 watermarks can be conveyed using openly specified audio and / or video watermarks, as specified in ATSC A / 334 and A / 335 respectively and remain present and readable from media content following all common distribution processing (e.g., transcoding, frame rate or resolution conversion, spatial rendering, etc.). ATSC standards A / 334, A / 335 and A / 336 may be found at https: / / www.atsc.org / atsc- documents / type / 3-0-standards /
[0235] The timed metadata delivered by A / 336 already supports identification of the broadcast service, the media presentation timeline, and program and advertising content identifiers. It uses the MIME format to support delivery of any asset with an IANA- registered DTD, so C2PA simply needs to register a DTD for whatever information it wishes to deliver to a validator that discovers a watermarked asset and wishes to use the watermark to acquire a manifest or canonical file.
[0236] Broadcast Provenance Service
[0237] When the IMP4 Replica was constructed the manifest for each data hash segment is stored in a database, along with the c2pa-compliant IMP4 encoding, retrievable from the VP1 URL (Server Code, Interval Code).
[0238] FIG. 10 illustrates the relationship of Server code, interval code and fMP4 Replica. In FIG 10:
[0239] 1) The VP1 Field 1 is the Server Code for the broadcast. It is a constant for the broadcast channel.
[0240] 2) The VP1 Field 2 is an Interval Code which increments sequentially, and both identifies a metadata resource on the web server and identifies a time on the media timeline.
[0241] 3) The Data Hash Segments are the range of Interval Codes which are c2pa- compliant MP4 encoded to produce an IMP4 Replica. The IMP4 Replica correspond to a Program Group, a single Program, or a Portion of a Program, depending on the broadcaster preference.
[0242] 4) In addition to any other supported, time-varying metadata, a Canonical OTTFile and the corresponding Replica Manifest are retrievable from the Broadcast Provenance Service using the VP1 Server Code and Interval Code.
[0243] For example, the Canonical OTT File #M and its corresponding C2PA Replica Manifest #M are retrievable using the VP1 Server Code and an Index Code between BINX(M) and EINX(M), inclusive. The term Index Code and Interval Code are interchangeable.
[0244] FIG. 11 illustrates the relationship of IMP4 replicas to broadcast channel. FIG 11 represents a portion of the linear broadcast between ‘begin WMT(M)’ and ‘end WMT(R)’. The categories of OTT MP4 content that a C2PA validator may encounter (see FIG. 7), include:
[0245] The OTT MP4 content encountered may correspond precisely to one specific Data Hash Segment of the 1MP4 Replica. For example, DHS #P with VP1 Field 1 =Server Code, the first Index Code = BINX(P) and last Index Code = EINX(P). In which case the Manifest #P and Canonical OTT File #P would be retrievable.
[0246] On the other hand, the OTT MP4 content encountered may correspond to an arbitrary range of presentation time stamps, as indicated by the VP1 Interval Code. For example, (see FIG. 10) from VP1 Field 2 M+2 to VP1 Field 2 N+l. In this case a Canonical OTT could be constructed from the fragmented MP4 Replicas, and a new Manifest created based on the that MP4 file; but as we will see, this is unnecessary, since provenance validation and canonical OTT substitution can take place using the original DHS manifests and Canonical OTT Files.
[0247] Canonical OTT Files
[0248] The creation of the IMP4 Replica, DHS Canonical OTT Files and DHS Manifests establishes a provenance anchor from OTT Files to linear broadcast content; but this has other benefits.
[0249] The creation of Canonical OTT ecosystem provides a means for the broadcaster to promote the use of authenticated content on the Internet in general and on social media platforms in particular, enabling end-users to post authenticated content to social media applications directly, and to provide immediate access to authenticated versions of broadcast content when attempting to upload a version which may have been modified without authority (see the Broadcast Canonical OTT Substitution section).
[0250] Broadcast Provenance Validation
[0251] Overview
[0252] As explained previously (see the Broadcast Provenance and Authentication Services section and FIG. 7), there are two kinds of Broadcast Linear OTT Files without a manifest, Canonical OTT Candidates - categories b), d), I) and g) - and SubstitutionCandidates - categories i) and j). Canonical OTT Candidates can be validated and / or replaced, but Substitution Candidates can only be replaced.
[0253] This section describes the process for Validating Canonical OTT Candidates. Broadcast Canonical OTT Substitution is described in the Broadcast Canonical OTT Substitution section.
[0254] The general steps for C2PA asset validation are: locate the active manifest (14.1) and the claim (14.2); validate the signature (14.3), timestamp (14.4), and credential revocation information (14.5); validate the assertions (14.6); validate the integrity of ingredients (14.7), and validate the asset’s content (14.9).
[0255] Using the fMP4 Replica for OTT Validation
[0256] As constructed, the fMP4 Replica manifest is identical to a C2PA manifest created for audio-video streaming. How the Replica manifest is used to validate OTT content is entirely different.
[0257] For streaming validation, the C2PA validator does not have access to the entire video stream. The auxiliary C2PA boxes are used to validate the fragmented MP4 chunk without having access to the rest of the chunks in the video track. This also means subsets of the video stream can be validated in isolation. This feature is useful for validating OTT content which may have been abstracted from a canonical OTT file. For further background information see the Section 8.2.2 of the C2PA specification at: https: / / c2pa.Org / specifications / specifications / l.0 / specs / C2PA_Specification.html#_hashin g a bmff formatted asset
[0258] For MP4 Replica validation, the validator has full access to an OTT object. Using the VP1 watermark, the C2PA validator can use recovery requests to receive prov / auth metadata sufficient to validate the OTT asset.
[0259] FIG. 12 illustrates general validation steps when there is no embedded manifest. FIG. 12 describes the general validation method to be followed when the OTT File does not have an embedded manifest. All of the cases listed below - single and multichannel canonical candidates, stitched canonical candidates and partial canonical candidates follow this same general validation process. The detailed differences for each OTT type can be found below.
[0260] FIG. 13 shows Table 1, which shows the OTT content watermark categories. Table 1 describes the 4 categories of VP1 watermark states the validator might encounter in OTT content. Validation steps for each of these 4 categories are listed below, to enable a validator to determine the OTT content provenance category from those identified in FIG. 7.
[0261] Note that Canonical Candidates are restricted to be OTT Files without a manifest, at least potentially derived from one or more Canonical OTT Files, category a) in FIG. 7. In other words, categories b), d), f), or g) in FIG 7.
[0262] 1. Single and multi-channel canonical candidates - (4), (5)
[0263] 2. Single and multi-channel stitched canonical candidates - (1), (6)
[0264] 3. Single and multi-channel partial canonical candidates - (2), (3), (7), (8)
[0265] 4. No detected watermark - (9)
[0266] Single Channel Canonical Candidates
[0267] (4) A single channel is present throughout the content with a consistent, monotonically increasing Interval Code.
[0268] This simplest case may be the result of an OTT file produced from a single channel Canonical OTT File. The OTT file might correspond to a single Data Hash Segment, or multiple Data Hash Segment.
[0269] FIG. 14 illustrates an example of a single channel Canonical Candidate (4), enclosed within a single Data Hash Segment. FIG. 15 shows an example of a single channel Canonical Candidate spanning a boundary between two Data Hash Segments. The model is generalizable to any number of contiguous Data Hash Segments.
[0270] This is distinct from the single channel stitched OTT file (6.5) in that there are no gaps in the timeline and that is no content included in the OTT file without a watermark, as there are with categories (2) and (3).
[0271] Validation Steps:
[0272] 1) Construct a recovery request VP1 URL from the first field- 1 / field-2 (e.g., server code / interval code) pair (OTT BINX).
[0273] 2) Set INDEX = OTT BINX
[0274] 3) Retrieve the DHS BINX, DHS EINX for the corresponding DHS from the recovery response.
[0275] 4) Retrieve the DHS Replica Manifest from the recovery response.
[0276] 5) Scan the OTT file and find the OTT EINX
[0277] 6) If the OTT EINX <= DHS EINX, then
[0278] a) use the DHS Replica Manifest to complete validation of the OTT File (
[0279] For further background information see the Section 8.2.2 of the C2PA specification at: https: / / c2pa.Org / specifications / specifications / l.0 / specs / C2PA_Specification.html#_hashin g_a_bmff_fomiatted_asset
[0280] b) Exit
[0281] 7) Else (see FIG. 15)
[0282] a) Use the DHS Replica Manifest to validate the content from INDEX toDHS EINX
[0283] b) Use the recovery response to find the next DHS in the IMP4 Replica
[0284] c) Retrieve the DHS BINX, DHS EINX for the corresponding DHS
[0285] d) Retrieve the DHS Replica Manifest
[0286] e) Set INDEX = DHS BINX
[0287] f) Go to Step 6
[0288] 8) End
[0289] Multichannel Canonical Candidates
[0290] (5) Multiple channels are present in the OTT content with a consistent, monotonically increasing Interval Code for each channel.
[0291] Similar to the single channel canonical candidate, the OTT file may have been produced from multiple channel’s Canonical OTT Files. Again, the OTT file might correspond to a single Data Hash Segment, or multiple Data Hash Segments.
[0292] This is distinct from the multi-channel stitched OTT file in that there are no gaps in the timeline and that is no content included in the OTT file without a watemiark, as there are with categories (7) and (8).
[0293] FIG. 16 shows an example of a Multichannel Canonical Candidate. Of course, either channel could involve multiple Data Hash Segments as shown in FIG. 15.
[0294] Validation Steps are the same as for a single channel canonical candidate, only performed for each Server Code (channel) encountered in the OTT file.
[0295] Single and Multichannel Stitched Canonical Candidates (1), (6)
[0296] A single channel (1) or multichannel (6) found throughout the OTT content, but the time order is not consistent. One or more Data Hash Segments may be involved.
[0297] FIG. 17 shows the Single Channel, single Data Hash Segment version of a Stitched Canonical Candidate. The distinguishing feature for all Stitched Canonical Candidates is that watermarks are present throughout the OTT File, but there are gaps in the time order.
[0298] Validation Steps are the same as for Canonical Candidates, since the C2PA algorithm allows for processing subsets of the Merkle tree.
[0299] Single and Multichannel Partial Canonical Candidates (2), (3), (7), (8)
[0300] Single or Multichannel watermarks are present in the OTT File, but there are regions of the file which contain no watermarks.
[0301] FIG. 18 shows a single channel, single Data Hash Segment example of a Partial Canonical Candidate. The defining characteristic is the presence of content which has no watermark. Otherwise, the OTT File can be similar to Canonical Candidates and Stitched Canonical Candidates, single and multichannel, single and multiple Data Hash Segments.
[0302] Validation follows the same steps as with canonical candidates, except that the unknown content needs to be shown as unvalidated during rendering to the end user.
[0303] No Detected Watermark (9)
[0304] No watermark detected anywhere in the OTT content.
[0305] When no watermark is present the validation cannot proceed.
[0306] The C2PA Decoupled Use Case
[0307] Assume that the broadcaster provides a mechanism for end-users to produce a Canonical OTT files from the Replica (5.6). This would be a smart move, since once it is propagated to network / platform operators, the ease of use would make fraudulent OTT content stand out.
[0308] Using DHS Manifests to validate decoupled broadcast OTT flies
[0309] If the content corresponds to a single Data Hash Segment, then the produced MP4 files could be managed as a single channel canonical files (6.3). The manifest used would correspond to the entire Data Hash Segment but could be used for any portion of the Data Hash Segment.
[0310] If the content corresponds to multiple Data Hash Segments, the algorithm given in (6.3) for handling the case where OTT EIDX > DHS EIDX could be used.
[0311] Using Interval-Range specific Manifests to validate decoupled broadcast OTT files
[0312] Alternatively, the user-facing application could create a Canonical OTT from the fMP4 Replicas, and a new Manifest created based on the that MP4 file.
[0313] If URL(Server Code, BINX, EIDX) already exists, the service would provide the corresponding Canonical OTT file to the end-user.
[0314] Else the service would create a canonical OTT file from the Replica, sign the claim in the manifest, populate the URL (Server Code, BNDX, ENDX) and then return the canonical OTT file to the end user.
[0315] Broadcast Canonical OTT Substitution
[0316] Overview
[0317] As explained previously (see the Broadcast Provenance and Authentication Services section and FIG. 7), there are two kinds of Broadcast Linear OTT Files without a manifest, Canonical OTT Candidates - categories b), d), f) and g) - and Substitution Candidates - categories i) and j). Canonical OTT Candidates can be validated and / or replaced, but Substitution Candidates can only be replaced.
[0318] This section describes the process for Broadcast Canonical OTT Substitution. Validating Canonical OTT Candidates is described in the Broadcast Provenance Validation section.
[0319] Using the fMP4 Replica for Canonical OTT Substitution
[0320] The process for producing a Canonical OTT Substitution mirrors the process for validation. The principle difference is that an OTT Playlist is constructed which indicates what range(s) of the Data Hash Segment(s) are played in order to reproduce the timeline of the OTT File encountered.
[0321] The OTT Playlist is constructed by walking through the OTT File from OTT BINX to OTT EINX, keeping track of any jumps in the VP1 Interval Code and / or Server Code. The playlist it constructed as a tuple - Server Code, DHS ID, Start Interval Code, End Interval).
[0322] Server Code, DHS ID, (Start Index, Stop Index), (Start Index), (Stop Index), ...
[0323] Whenever the end of a DHS is reached INDX > DHS EIDX, the next DHS in the sequence is retrieved from the Recovery Response.
[0324] When this process is completed, the (Start Index, End Index) values are converted to the Timeline designation that the broadcaster is using in the Canonical OTT Files.
[0325] Since the Data Hash Segment Manifest works for any subset of the Data Hash Segment, no new manifest need be created. Provenance Service constructs an MP4 file from the Canonical OTT Files using the DHS Manifest.
[0326] FIG. 19 describes the general Canonical OTT Substitution steps. To determine which of these categories are appropriate for any particular OTT File, see Table 1 in FIG. 13.
[0327] Substituting Canonical Candidates
[0328] (4) A single channel is present throughout the content with a consistent, monotonically increasing Interval Code.
[0329] (5) Multiple channels are present in the OTT content with a consistent, monotonically increasing Interval Code for each channel.
[0330] (1), (6) A single channel (1) or multichannel (6) found throughout the OTT content, but the time order is not consistent. One or more Data Hash Segments may be involved.
[0331] (2), (3), (7), (8) Single or Multichannel watermarks are present in the OTTFile, but there are regions of the file which contain no watermarks.
[0332] All of the categories from FIG. 7 which are downstream from the Canonical OTT File, and which have no manifest - categories b), d), f), and g) - can be validated against the Canonical OTT File; but they can also be substituted.
[0333] FIG. 14 provides an example of a single channel Canonical Candidate, enclosed within a single Data Hash Segment. FIG. 15 shows an example of a single channel Canonical Candidate spanning a boundary between two Data Hash Segments. The model is generalizable to any number of contiguous Data Hash Segments.
[0334] Canonical OTT Substitution Steps:
[0335] 1) Construct a recovery request VP1 URL from the first field- 1 / field-2 pair(OTT BINX).
[0336] 2) Set INDX = OTT BINX
[0337] 3) Retrieve the DHS BINX, DHS EINX for the corresponding DHS from the recovery response.
[0338] 4) Retrieve the DHS Canonical OTT File from the recovery response.
[0339] 5) Retrieve the DHS Replica Manifest from the recovery response.
[0340] 6) Scan the OTT file and find the OTT EINX
[0341] 7) If the OTT EINX <= DHS EINX, then
[0342] a) use the DHS OTT Canonical OTT File to finish the Server Code, Interval Code Playlist
[0343] i) If there are gaps in the Interval Code, Create (start, stop) pairs
[0344] b) add the DHS Replica Manifest to the Canonical OTT Substitute Manifest List
[0345] c) Convert the Server Code, Interval Code Playlist to a Server Code, Timeline Playlist
[0346] d) Construct the Canonical OTT Substitute from the Server Code, Timeline Playlist and DHS Replica Manifest List.
[0347] e) Exit
[0348] 8) Else
[0349] a) Use the DHS OTT Canonical OTT File to build the Server Code, Interval Code Playlist from INDEX to DHS EINX
[0350] i) If there are gaps in the Interval Code, Create (start, stop) pairs
[0351] b) Add the DHS Replica Manifest to the Canonical OTT Substitute Manifest List
[0352] c)Use the recovery response to find the next DHS in the fMP4 Replica
[0353] d) Retrieve the DHS BINX, DHS EINX for the corresponding DHS
[0354] e) Retrieve the DHS Replica Manifest
[0355] I) Set INDEX = DHS BINX
[0356] g) Go to Step 6
[0357] 9) End
[0358] A Soft binding Reference Model for C2PA
[0359] There is a need to get the C2PA soft binding reference model crisply defined independent of any specific soft binding technologies; viz., defined soft binding use cases with corresponding requirements, application / session logic, an abstract payload definition, and a reference model.
[0360] A reference model acts as a blueprint that shapes the way use cases, scenarios, and functional requirements are defined and documented. It ensures that all aspects of system design and development are aligned with a common set of principles and standards, facilitating interoperability, and integration among different system components.
[0361] For all C2PA use cases, there is an ‘exception’ scenario which includes a soft binding functional step, using watermarking or fingerprinting. The role of soft binding when validating content should be crisply defined in C2PA, including the role in recovering verified content whether the content was broadcaster published to a website or recorded from broadcast without authority.
[0362] System architecture
[0363] Entities
[0364] FIG. 20 shows the system architecture for the watermark reference model for C2PA, including:
[0365] 1) Content Creator: An actor who uses a Claim Generator to produce Assets as defined by the C2PA Specification.
[0366] 2) Claim Generator: A function that produces Watermarked Assets as defined by the C2PA Specification that employs a Soft Binding Watermark Generator.
[0367] 3) Soft Binding Watermark Generator: A function for producing WatermarkedContent with functionality as defined in the Process Sequence Diagram section of this document.
[0368] 4) Manifest Service: A networked store of decoupled manifests referenced byURIs.
[0369] 5) Manifest: Either a Manifest Store or an Active Manifest as defined by theC2PA Specification.
[0370] 6) Alteration: A modification applied to an Asset that will cause a validation using only hard bindings to fail.
[0371] 7) Verifier: An actor who uses a Validator to validate Watermarked Content.
[0372] 8) Validator: A function for authenticating Assets as defined by the C2PASpecification that employs a Soft Binding Watermark Reader.
[0373] 9) Soft Binding Watermark Reader: A function for reading watermarks fromWatermarked Content with functionality as defined in the Process Sequence Diagram section of this document.
[0374] Data Objects
[0375] 1) Content: Media content, which may or may not be an Asset.
[0376] 2) Watermarked Asset: An Asset as defined by C2PA Specification that contains a soft-binding watermark.
[0377] 3) Watermarked Content: Content that contains a soft-binding watermark.
[0378] 4) Soft Binding Watermark: Information embedded in content that provides a durable association between an asset and a decoupled manifest URL
[0379] 5) Manifest URL A URI that reference a C2PA Manifest stored in a ManifestService.
[0380] 6) Manifest: Metadata conveying Asset assertions and authentication information as defined in the C2PA Specification.
[0381] 7) Manifest Reference: A data object describing a manifest association recovered from a soft binding watermark in content, including the Manifest URI and Regions associated with the watermark.
[0382] 8) Region: A region map that specifies a spatial, textual, or temporal range in content as defined in the C2PA Specification.
[0383] 9) Soft Binding Assertion: A metadata record in the Manifest that is associated with a soft binding watermark instance. Its contents, format and use are specific to the soft binding watermark technology.
[0384] 10) Validation Result: Information produced by a Validator for use by aVerifier in making trust decisions. This may include a Validator status codes, referenced Manifests, Manifest References from Soft Binding Watermark Readers, and other information.
[0385] Process Sequence Diagram
[0386] FIG. 21 shows a process sequence diagram of the watermark reference model for C2PA.
[0387] Soft Binding Watermark Embedder (GenerateWatermarkedContent)
[0388] Inputs
[0389] Content. The source of content that is to be watermarked.
[0390] a) Instances of the function are expected to support only a limited set of media types.
[0391] b) In the case of generative Al, Content may be of a different media type than the WatermarkedContent (i.e., a prompt).
[0392] ManifestURI'. A URI associated with the manifest to be referenced by the generated soft binding watermark.
[0393] Note that a particular soft-binding technology may restrict ManifestURI support in a variety of ways, including by quantity, protocol types, or endpoint identity
[0394] Outputs
[0395] a) WatermarkedContent'. Media content (which may be of static or stream format) derived from the input Content that contains a generated soft binding watermark.
[0396] b) SoftBindingAssertion'. A soft binding assertion associated with the generated soft binding watermark for inclusion by the Claim Generator in the Manifest for the Asset containing the WatermarkedContent.
[0397] c) Status'. Description of the outcome of the soft binding watermark generation function; i.e. success or error information.
[0398] Description
[0399] This function is used to generate WatermarkedContent that includes a soft binding watermark and an associated SoftBindingAssertion.
[0400] An instance of the GenerateWatermarkedContent function is associated with a particular soft binding watermark technology. The methods used by a soft binding watermark technology are out of scope of C2PA.
[0401] When Content is input in a streamed format, the function’s output data is also streamed. That is, WatermarkedContent, SoftBindingAssertion, and Status are output piecewise as Content input progresses, potentially via asynchronous callbacks.
[0402] For non-ML / AI use cases (traditional media), WatermarkedContent is produced by embedding a soft binding watermark into the input Content. For ML / Al use cases, the GenerateWatermarkedContent function may also include content generation, in which case Content could be a prompt with a different media type (e.g., text vs. image) from WatermarkedContent.
[0403] The behavior of the GenerateWatermarkedContent function when presented with Content input that contains a watermark of the same technology is technology dependent. The function may either replace the prior watermark, embed an overlappingwatermark, refuse to embed a watermark in regions that would overlap with a preexisting watermark, or refuse to embed any watermark at all (i.e. even in non-overlapping regions).
[0404] WatermarkedContent contains a soft binding watermark described by SoftBindingAssertion subject to any error conditions indicated by Status.
[0405] SoftBindingAssertion contains a soft binding assertion associated with the embedded soft binding watermark for the Claim Generator to include in the Manifest of the WatermarkedContent.
[0406] Status identifies any applicable success or error conditions.
[0407] Soft Binding Watermark Reader RecoverWatermark)
[0408] Inputs
[0409] Content. The source of content from which soft binding watermarks are desired to be recovered.
[0410] Outputs
[0411] ManifestReference'. A compound data object representing a soft binding watermark referencing a manifest.
[0412] a) ManifestURT. A URI associated with the manifest referenced by a soft binding watermark embedded in the content.
[0413] b) ContentRegion'. A description of the region (i.e., C2PA “region-map”) of the Content in which the soft binding watermark was recovered.
[0414] c) AssetRegion'. A description of the region (i.e., C2PA “region-map”) of the asset in which the recovered soft binding watermark was embedded.
[0415] Description
[0416] This function is used to recover manifest references from a soft binding watermark in Content.
[0417] A ManifestReference object is provided for each region of the Content in which a soft binding watermark is recovered. ManifestURI has the values of the corresponding data element provided to the GenerateWatermarkedContent function when the soft binding watermark was generated.
[0418] ContentRegion has values corresponding to the region in Content where the soft binding watermark was found. This may describe a subset of the Content in the event that the Content is a composition of the watermarked asset with other content (e.g. a collage or mash-up).
[0419] AssetRegion has values corresponding to the region of the watermarked asset in which the recovered soft binding watermark was embedded. This will be a subset of the WatermarkedContent if the Content contains only a portion (e.g. a crop or clip) of the watermarked asset.
[0420] Multiple ManifestReference objects, each referencing different manifests, may be returned for Content that contains multiple soft binding watemiarks. One use case in which this might occur is if the Content is composed of content derived from different watermarked assets.
[0421] ManifestReference objects may be output piecewise as Content input progresses via asynchronous callbacks.
[0422] Multimedia Content Validation
[0423] Background
[0424] In our interconnected world there is a flood of information, and it is often difficult to discriminate what is true and what is not. Typically, users look for the source of the information to decide whether they trust it. In other words, most of us judge what is factual based on the provenance and authenticity of the information, where provenance refers to the origin, history, and chain of custody of a piece of audio-video content, and authenticity refers to whether the content has been manipulated or altered in a way out of the control of the trusted source of the information.
[0425] The C2PA addresses the prevalence of misleading information online through the development of technical standards for certifying the source and history (or provenance) of media content. C2PA is a Joint Development Foundation project, formed through an alliance between Adobe, Arm, Intel, Microsoft and Truepic.
[0426] There are two types of bindings supported by C2PA - hard bindings and soft bindings. The hard binding uses a cryptographic hashing algorithm over some or all of the bytes of an asset. It can be used to detect tampering. Soft bindings may be a perceptual hash computed from the digital content (a.k.a. fingerprint), or a watermark embedded in the digital content. C2PA metadata, including content hashes and provenance data, is stored in so called “manifests” that are embedded into content containers.
[0427] However, the distribution channels for multimedia assets typically involve content processing such as cropping, transcoding, adjustment of frame rates, resolutions, bitrates, audio channel counts, container formats etc. Any processing would require new hard binding calculation and creation of a new C2PA manifest. The current distribution channels and most distribution channels in near future are unlikely to be able to create new C2PA manifests. To recover the provenance and authenticity information in this scenario C2PA introduces soft binding which is also known as Automatic Content Recognition (ACR) techniques. ACR can be done based on either fingerprints or embedded watermarks.
[0428] According to C2PA standards whenever a hard binding data and a manifest are created a canonical representation of associated media object is saved in a “golden store” and made available for retrieving based on ACR data. Typically, multimedia content, such as news show, or sport events, are segmented, and each segment is saved in the “golden store” together with its manifest.
[0429] In principle the ACR data may lead to retrieving C2PA manifest only, and this manifest may be successful in validating the received content. But in most cases the distribution channel processing will prevent manifest based validation and retrieving canonical representation is recommended for media object validation.
[0430] This disclosure teaches how to use ACR technology to validate reliably and conveniently received content in the absence of valid hard binding.
[0431] Detailed description of Multimedia Content Validation
[0432] One of the problem scenarios appears when users upload a news clip to a social media claiming that it was taken from a reputable source, say a TV network. Typically, the clip is short, and its boundaries will not coincide with segment boundaries of the canonical representation. Also, typically the clip’s bit rate is lower than in the original broadcast, and the number of audio channels is reduced from 5.1 to stereo (to reduce bandwidth requirements). So, in this scenario the C2PA manifest will not be able to validate the uploaded content, and the social media platform may try to use ACR technology to try content validation.
[0433] As discussed above, the Advanced Television Systems Committee (ATSC) has specified digital watermarking that has been used already in many broadcast stations which would make it an obvious choice for the ACR technology. The watermarking technique labeled VP1 has the payload that includes Server Code (SC), which points to a server that contains the content metadata, as well as the Interval Code (IC) that detemiines the media timeline with precision of about one millisecond. The metadataretrieved from servers pointed by SC can be used to retrieve manifests and canonical representation, and IC from uploaded content and canonical representation can be used to align uploaded content and canonical representation.
[0434] However, it should be clear to those skilled in the art that other watermarking technologies can lead to the server containing metadata as well as establish the uploaded clip location within canonical representation.
[0435] On the other hand, fingerprinting technology has a predefined server, or servers, with fingerprinting database, while fingerprints themselves, after finding a match in the database, would point to the clip’s location within canonical representation, although typically with less precision than watermarks. It should be noted also that fingerprints have scalability problem, and that an increase in number of fingerprints in the database reduces reliability of fingerprint matches and increases processing. Further, the reliability of fingerprint matching decreases with reduced size of the clip and depends strongly on content characteristics. Finally, the same clip may originate at different broadcast stations, which could cause incorrect provenance determination.
[0436] The first solution to our problem scenario is to retrieve canonical representation corresponding to detected set of watermarks, cut from it the clip that has the same pattern of watermarks as uploaded clip and replace uploaded clip with this newly created clip. This solution has a few problems. First, the user may not approve replacing his clip with something else. Secondly, there are likely to be some constraints on what can be posted on the social media platform, in terms of frame rates, resolutions, bitrates, audio channel counts, container formats etc. which may not be matched by the canonical representation, and the conversion may not be trivial. Finally, the canonical representation segment may be much larger than uploaded clip interval, and as a result, retrieving canonical representation, manifest and performing validation may be burdensome.
[0437] An alternate solution is to try to establish if uploaded content and the content located in canonical representation based on watermarks are substantially the same. In other words, if the difference between the two contents can be fully explained by standard processing such as transcoding, Dynamic Range Compression (DRC), bit rate reduction etc., the uploaded clip may be validated, while user editing should be red flagged, and result in a failed validation.
[0438] One way to achieve this is by fingerprinting, i.e., by comparing the perceptual hashes of the two contents. For example, the fingerprints of the canonical segments can be calculated at the source and cryptographically linked to provenance data within manifests. Those manifests can be retrieved by social media platforms, and selected portion of them compared to fingerprints of the uploaded clips. If they are the same, the uploaded clip is validated. However, even if the fingerprints aren’t the same, but the difference between them is recognized to be caused by standard processing, the uploaded clip may be validated.
[0439] It should be noted that this use of fingerprints doesn’t suffer the same drawbacks as ACR usage discussed above. In ACR usage the fingerprints need to be matched in database to many distinct fingerprints of different contents and of unknown location within the content. Here we have to compare fingerprints of two precisely aligned contents, which is much easier task, and thus has much more favorable tradeoff between false positives and false negatives.
[0440] Many different fingerprints could serve this purpose, and in order to simplify adoption probably some standardized fingerprints, such as SMPTE 2064, which specifies both audio and video fingerprints, can be used. If both of them match, the upload can be validated. For details of SMPTE 2064-1 and SMPTE 2064-2.
[0441] It should be noted that multiple fingerprinting algorithms could be used simultaneously. Some of them could be based on Artificial Intelligence (Al). Recent developments on Al based speech to text conversion, speaker voice recognition, and facerecognition would help minimize danger that attacker could introduce only subtle changes of the spoken words or speaker substitutions.
[0442] However, at this moment the Al fingerprinting algorithms are in R&D phase, and it is expected that in near future they would keep improving. Similarly, it is expected that Al could be used to introduce new attacks on deployed fingerprinting algorithms. Having this in mind, it may be better for social media platform to download canonical representation of the segment(s) together with manifest(s) and perform cryptographical validation. After the cryptographical validation of canonical representation the platform may calculate the fingerprint of the uploaded content and compare it to fingerprint of the clip taken from canonical representation at location determined based on watermarks, and if they match, the upload is validated. In this scenario there is no need to standardize the fingerprint calculation or no need for fingerprint coordination between the content source and the media platforms, and novel fingerprinting technologies, parameter adjustments etc., could be introduced based on new insights about attacks observed in the field.
[0443] Once the canonical representation is retrieved and fingerprints calculated, the resulting fingerprints could be stored in a database and linked to the corresponding set of watermarks. Next time if watermarks of uploaded clip point to the canonical representation that was already retrieved and processed for fingerprint extraction, the content retrieving step could be omitted, and the fingerprints from the database could be used.
[0444] Furthermore, the fingerprint database may also contain the hash of the uploaded content as well as the result of the content validation. Subsequent uploads of the same content can be recognized by comparing the content hash with previously created hashes, and if a match is found then the validation outcome can be reapplied to newly uploaded content, and thus some processing would be saved.
[0445] If the fingerprint database described above is created by a trusted entity, it could be made available to consumer devices too. For example, Apple devices could bereading watermarks, generating fingerprints, and pulling canonical fingerprints from an Apple server using Apple fingerprint technology. Similarly, Microsoft browsers or Google browsers could use fingerprint databases stored on Microsoft or Google servers.
[0446] An attack to consider is where watermark spoofing leads to the “Rickrolling” type of mischief. See: https: / / en.wikipedia.org / wiki / Rickrolling.
[0447] The objective is not to misinform but to misdirect user’s browser to unexpected site. At worse it would annoy users, and perhaps make them disable C2PA validation based on watermarks. The solution to this attack would again be to compare fingerprints of uploaded clip with fingerprints taken from the clip located in canonical representation based on watermarks and reject the upload if they do not match.
[0448] In an alternative embodiment, the association between the watermark and fingerprint of the canonical representation are obtained previously through an alternative means such as broadcast monitoring, upload through trusted portal, web spidering, etc.).
[0449] Finally, note that the two approaches described above, one where fingerprints of canonical representation are included in the manifest, and the other where they are not included, are not mutually exclusive. For example, the standardized fingerprints could be part of the manifest, but some social media platforms may choose to improve validation process by introducing additional, proprietary, fingerprint-based validation.
[0450] Broadcast News Authentication Embodiment
[0451] FIGS. 22-27 illustrate screenshots of an exemplary embodiment for the broadcast news use case. In general, FIGS. 22-27 illustrate a user experience in the scenario where (a) a television set that includes a validator as described in this disclosure receives a video stream with a VP1 watermark; (b) the validator determines that there is no C2PA manifest present in the received video stream (c) the validator detects the presence of the VP1 watermark in the video stream; (d) the validator employs the VP1 watermark using the A / 336 content recovery protocol to retrieve a C2PA manifest; (e) thevalidator determines that the received video stream cannot be validated using the recovered C2PA manifest; (f) the validator determines that the C2PA manifest contains a valid link to a canonical instance of the content; (g) the television set offers the viewer the ability to access the canonical instance; (h) the viewer accesses to the canonical instance of the content using the valid link contained in the recovered C2PA manifest.
[0452] FIG. 22 shows screen shots of the step of a user searching for broadcast station news. The screen on the left shows a social media platform search and the screen on the right shows a search of antenna or pay TV using a program guide. FIG. 23 shows the launch of a Broadcast News Authentication application. This application is enabled by the VP1 watermark discussed above. The screen displays a notice indicating that “news authentication” is available, which also prompts the user to take an action.
[0453] FIG. 24 shows the presentation of a link to a secure broadcaster website prompted by the user accepting a call-to-action. FIG. 25 shows the presentation of a QR code, and the activation of the user’s mobile phone. FIG. 26 shows the step of the user accessing the link through the mobile phone. FIG. 27 shows the engagement with the local broadcast news station using the mobile phone to view authentic media and transcripts thereof directly from the local station.
[0454] In one exemplary embodiment, a method of validating the provenance of an OTT file is disclosed comprising determining if the OTT file is linear broadcast content, and if so, determining if the OTT file is one of the following categories: i) approved C2PA encoding with manifest removed, ii) approved C2PA encoding with endorsed C2PA transcoding with manifest removed, iii) approved C2PA encoding with unendorsed C2PA transcoding with manifest removed, or iv) approved C2PA encoding with unapproved mp4 transcoding, and if so, validating the OTT file; andif the OTT file is not one of categories i), ii), iii), or iv), determining if the OTT file is one of the following categories: v) unapproved C2PA encoding with manifest removed or vi) unapproved mp4 encoding, and if so, preparing a broadcast canonical OTT substitution for the OTT file.
[0455] In another exemplary embodiment, A method of validating the provenance of digital media comprises: determining whether received media contains a manifest, and if so, determining if the manifest is trustworthy and validates the media, and if so, declaring that the media is validated; if the received media does not contain a manifest, or is not trustworthy, or does not validate the media, determining if the media contains a watermark; if the media does not contain a watermark, declaring that the media is unvalidated; if the media does contain a watermark, retrieving a manifest referenced by the watermark; determining if the manifest is trustworthy and if not, declaring that the media is unvalidated; if the manifest is trustworthy, determining if the media validates the media and if so declaring that the media is validated; if the manifest is not trustworthy, notifying a user of a referenced asset, wherein the referenced asset is referenced by the retrieved manifest; if a request for the referenced asset is received, retrieving the referenced asset and determining if the manifest validates the retrieved asset; and if the manifest validates the retrieved asset, declaring that the retrieved asset is validated.
[0456] According to another embodiment, a method for providing provenance authentication of broadcast media comprises: receiving a media object from a broadcast; determining if the received media object contains a manifest corresponding to a registered distributor, and if it does, determining if the manifest validates the media object, and if it does, declaring the media validated; if the media object does not contain a manifest or if the manifest is not validated, determining if the media object contains a watermark, and if so, using the watemiark to retrieve the manifest; determining if the retrieved manifest has a digital signature that corresponds to an approved broadcaster, and if so, determining if the manifest validates the content, and if so declaring the mediaobject validated; if the manifest does not validate the content, initiating a canonical process to retrieve a canonical representation of the media object.
[0457] According to another embodiment, a method for validating media content comprises: receiving a media object; detecting a watermark in the media object; calculating a first fingerprint based on the media object; calculating a second fingerprint based on a canonical representation of the media object; and determining whether the first and second fingerprints match, and if they match, declaring the received media object validated.
[0458] According to another embodiment, a method for validating media content comprises: receiving a media object; detecting a watermark in the media object; determining a first fingerprint based on a canonical representation of the media object containing the detected watermark; calculating a second fingerprint based on the media object; and determining whether the first and second fingerprints match, and if they match, declaring the received media object validated.
[0459] It is understood that the various embodiments of the present invention may be implemented individually, or collectively, in devices comprised of various hardware and / or software modules and components. These devices, for example, may comprise a processor, a memory unit, an interface that are communicatively connected to each other, and may range from desktop and / or laptop computers, to consumer electronic devices such as media players, mobile devices, and the like. For example, FIG. 28 illustrates a block diagram of a device 1000 within which the various disclosed embodiments may be implemented. The device 1000 comprises at least one processor 1002 and / or controller, at least one memory 1004 unit that is in communication with the processor 1002, and at least one communication unit 1006 that enables the exchange of data and information, directly or indirectly, through the communication link 1008 with other entities, devices and networks. The communication unit 1006 may provide wired and / or wireless communication capabilities in accordance with one or more communication protocols,and therefore it may comprise the proper transmitter / receiver antennas, circuitry and ports, as well as the encoding / decoding capabilities that may be necessary for proper transmission and / or reception of data and other information.
[0460] Referring back to FIG. 28 the device 1000 and the like may be implemented in software, hardware, firmware, or combinations thereof. Similarly, the various components or sub-components within each module may be implemented in software, hardware, or firmware. The connectivity between the modules and / or components within the modules may be provided using any one of the connectivity methods and media that is known in the art, including, but not limited to, communications over the Internet, wired, or wireless networks using the appropriate protocols.
[0461] Various embodiments described herein are described in the general context of methods or processes, which may be implemented in one embodiment by a computer program product, embodied in a computer-readable medium, including computerexecutable instructions, such as program code, executed by computers in networked environments. A computer-readable medium may include removable and non-removable storage devices including, but not limited to, Read Only Memory (ROM), Random Access Memory (RAM), compact discs (CDs), digital versatile discs (DVD), etc. Therefore, the computer-readable media that is described in the present application comprises non-transitory storage media. Generally, program modules may include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of program code for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps or processes.
[0462] The foregoing description of embodiments has been presented for purposes of illustration and description. The foregoing description is not intended to be exhaustive orto limit embodiments of the present invention to the precise form disclosed, and modifications and variations are possible in light of the above teachings or may be acquired from practice of various embodiments. The embodiments discussed herein were chosen and described in order to explain the principles and the nature of various embodiments and its practical application to enable one skilled in the art to utilize the present invention in various embodiments and with various modifications as are suited to the particular use contemplated. The features of the embodiments described herein may be combined in all possible combinations of methods, apparatus, modules, systems, and computer program products.
Claims
WHAT IS CLAIMED IS:
1. A method of validating the provenance of an OTT file comprising: determining if the OTT file is linear broadcast content, and if so, determining if the OTT file is one of the following categories: i) approved c3pa encoding with manifest removed, ii) approved C2PA encoding with endorsed C2PA transcoding with manifest removed, iii) approved C2PA encoding with unendorsed C2PA transcoding with manifest removed, or iv) approved C2PA encoding with unapproved mp4 transcoding, and if so, validating the OTT file; and if the OTT file is not one of categories i), ii), iii), or iv), determining if the OTT file is one of the following categories: v) unapproved C2PA encoding with manifest removed or vi) unapproved mp4 encoding, and if so, preparing a broadcast canonical OTT substitution for the OTT file.
2. The method of claim 1 further comprising: after validating the OTT file, preparing a broadcast canonical OTT Substitution for the OTT file.
3. The method of claim 1 wherein the step of validating the OTT file further comprises:1) Construct a recovery request VP1 URL from the first field- 1 / field-2 pair (OTT BINX).2) Set INDEX = OTT BINX3) Retrieve the DHS BINX, DHS EINX for the corresponding DHS from the recovery response.4) Retrieve the DHS Replica Manifest from the recovery response.5) Scan the OTT file and find the OTT EINX6) Ifthe OTT EINX <= DHS EINX, then a) use the DHS Replica Manifest to complete validation of the OTT File b) Exit7) Else a) Use the DHS Replica Manifest to validate the content from INDEX to DHS EINX b) Use the recovery response to find the next DHS in the fMP4 Replica c) Retrieve the DHS BINX, DHS EINX for the corresponding DHS d) Retrieve the DHS Replica Manifest e) Set INDEX = DHS BINX f) Go to Step 68) End4. The method of claim 1 wherein the step of preparing a broadcast canonical OTT substitution further comprises:1) Construct a recovery request VP1 URL from the first field- 1 / field-2 pair (OTT BINX).2) Set INDX = OTT BINX3) Retrieve the DHS BINX, DHS EINX for the corresponding DHS from the recovery response.4) Retrieve the DHS Canonical OTT File from the recovery response.5) Retrieve the DHS Replica Manifest from the recovery response.6) Scan the OTT file and find the OTT EINX7) Ifthe OTT EINX <= DHS EINX, then a) use the DHS OTT Canonical OTT File to finish the Server Code, Interval Code Playlisti) If there are gaps in the Interval Code, Create (start, stop) pairs b) add the DHS Replica Manifest to the Canonical OTT Substitute Manifest List c) Convert the Server Code, Interval Code Playlist to a Server Code, TimelinePlaylist d) Construct the Canonical OTT Substitute from the Server Code, Timeline Playlist and DHS Replica Manifest List. e) Exit8) Else a) Use the DHS OTT Canonical OTT File to build the Server Code, Interval CodePlaylist from INDEX to DHS EINX i) If there are gaps in the Interval Code, Create (start, stop) pairs b) Add the DHS Replica Manifest to the Canonical OTT Substitute Manifest List c) Use the recovery response to find the next DHS in the IMP4 Replica d) Retrieve the DHS BINX, DHS EINX for the corresponding DHS e) Retrieve the DHS Replica Manifest f) Set INDEX = DHS BINX g) Go to Step 69) End5. A method of validating the provenance of digital media comprising: determining whether received media contains a manifest, and if so, determining if the manifest is trustworthy and validates the media, and if so, declaring that the media is validated; if the received media does not contain a manifest, or is not trustworthy, or does not validate the media, determining if the media contains a watermark; if the media does not contain a watermark, declaring that the media is unvalidated; if the media does contain a watermark, retrieving a manifest referenced by the watermark;determining if the manifest is trustworthy and if not, declaring that the media is unvalidated; if the manifest is trustworthy, determining if the media validates the media and if so declaring that the media is validated; if the manifest is not trustworthy, notifying a user of a referenced asset, wherein the referenced asset is referenced by the retrieved manifest; if a request for the referenced asset is received, retrieving the referenced asset and determining if the manifest validates the retrieved asset; and if the manifest validates the retrieved asset, declaring that the retrieved asset is validated.
6. A method for providing provenance authentication of broadcast media comprising: receiving a media object from a broadcast; determining if the received media object contains a manifest corresponding to a registered distributor, and if it does, determining if the manifest validates the media object, and if it does, declaring the media validated; if the media object does not contain a manifest or if the manifest is not validated, determining if the media object contains a watermark, and if so, using the watermark to retrieve the manifest; determining if the retrieved manifest has a digital signature that corresponds to an approved broadcaster, and if so, determining if the manifest validates the content, and if so declaring the media object validated; if the manifest does not validate the content, initiating a canonical process to retrieve a canonical representation of the media object.
7. The method of claim 6 wherein the canonical process comprises: making a decision to perform canonical processing; retrieving a canonical representation of the media object; validating the canonical representation using the retrieved manifest; andmaking the canonical representation of the media object available to a user.
8. The method of claim 7 wherein the step of making the canonical representation of the media object available to a user includes at least one of: posting the received content together with the canonical representation or a link to it; providing the user with a choice between which version of the content to post and posting that version with an appropriate label; performing an automated comparison of the received media object and the canonical representation to determine the nature and amount of difference between the two; automatically replacing the received media object with the canonical representation; and forwarding the received media object and canonical representation to an internal content moderation process.
9. A method for validating media content comprising: receiving a media object; detecting a watermark in the media object; calculating a first fingerprint based on the media object; calculating a second fingerprint based on a canonical representation of the media object; and determining whether the first and second fingerprints match, and if they match, declaring the received media object validated.
10. The method of claim 9 wherein the media object is retrieved from a broadcast;12. The method of claim 9 wherein the watermark contains a server code that points to a server containing metadata for the media object.
13. The method of claim 9 wherein the watermark contains an interval code, and further comprising using the interval code to align the uploaded media object with the canonical representation.
14. The method of claim 9 further comprising calculating fingerprints of canonical segments at its source; and cryptographically linking the fingerprints of canonical segments to provenance data within a manifest attached to the media object.
15. The method of claim 9 wherein a social media platform downloads the canonical representation of the media object and a manifest and performs cryptographical validation.
16. The method of claim 9 wherein a social media platforms steps of calculating the first and second fingerprints, determining whether they match, and declaring the received media object validated.
17. The method of claim 9 further comprising: storing the first and second fingerprints in a database and linking them to a corresponding set of watemiarks; and receiving a second media object that points to the canonical representation associated with the received media object; and using the stored fingerprint in the database associated with the canonical representation for validation of the second media object.
18. The method of claim 17 further comprising: storing the fingerprint of second media object in the data base along with the results of the validation of the second media object; receiving a third media object that has a fingerprint that matches the second media object; and using the stored validation for the second media object to validate the third media object.
19. The method of claim 17 further comprising: making the database available to a consumer device, wherein the consumer device performs the steps of receiving, detecting, retrieving, calculating a first fingerprint, determining and declaring a received media object validated.
20. The method of claim 9 further comprising: retrieving a canonical representation of the media object using at least one of: the detected watermark; broadcast monitoring; upload through a trusted portal; and web spidering.
21. A method for validating media content comprising: receiving a media object; detecting a watermark in the media object; determining a first fingerprint based on a canonical representation of the media object containing the detected watermark; calculating a second fingerprint based on the media object; and determining whether the first and second fingerprints match, and if they match, declaring the received media object validated.
22. The method of claim 21 wherein the media object is retrieved from a broadcast;23. The method of claim 21 wherein the step of determining a first fingerprint based on a canonical representation of the media object comprises: the watermark containing a server code that points to a server containing metadata for the media object.
24. The method of claim 21 wherein the watermark contains an interval code, and further comprising using the interval code to align the uploaded media object with the canonical representation.
25. The method of claim 21 wherein the step of determining a first fingerprint based on a canonical representation of the media object comprises: calculating fingerprints of canonical segments at its source; and cryptographically linking the fingerprints of canonical segments to provenance data within a manifest attached to the media object.
26. The method of claim 21 wherein a social media platform downloads the canonical representation of the media object and a manifest and performs cryptographical validation.
27. The method of claim 21 wherein a social media platforms steps of calculating the first and second fingerprints, determining whether they match, and declaring the received media object validated.
28. The method of claim 21 further comprising: storing the first and second fingerprints in a database and linking them to a corresponding set of watermarks; and receiving a second media object that points to the canonical representation associated with the received media object; and using the stored fingerprint in the database associated with the canonical representation for validation of the second media object.
29. The method of claim 28 further comprising: storing the fingerprint of second media object in the data base along with the results of the validation of the second media object;receiving a third media object that has a fingerprint that matches the second media object; and using the stored validation for the second media object to validate the third media object.
30. The method of claim 28 further comprising: making the database available to a consumer device, wherein the consumer device performs the steps of receiving, detecting, retrieving, calculating a first fingerprint, determining and declaring a received media object validated.
Citation Information
Patent Citations
Methods, systems and computer program products for providing a media file to a designated set-top box
US20160156982A1
Protected multimedia content transport and playback system
US20190246149A1
Method and apparatus for single-signature content integrity and provenance validation for named data networking
US20210105281A1
Publishing a Disparate Live Media Output Stream using Pre-Encoded Media Assets
US20210211750A1
Media Distribution And Management Platform
US20220232268A1