Video data stream reliability

A unique identifier and digital signature system within the video data stream, combined with an external editor track, addresses the limitations of existing authenticity checks, providing adaptable and efficient reliability verification across various streaming scenarios.

JP2026016324APending Publication Date: 2026-02-03FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025113501
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-05
Filing Date
2025-07-04
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing video data stream authenticity checking methods lack adaptability to various application scenarios and compatibility with video codecs, particularly in streaming environments, and require inefficient bit-rates for reliability verification.

Method used

Incorporating a unique identifier and digital signature system within the video data stream, utilizing hash functions to verify authenticity, and maintaining checkability through an external editor track, allowing for adaptable and efficient authenticity checks across different streaming scenarios.

Benefits of technology

Enhances adaptability and compatibility with video codecs while reducing the required bit-rate for authenticity checks, ensuring reliable verification of video data streams in diverse applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016324000001_ABST
    Figure 2026016324000001_ABST
Patent Text Reader

Abstract

Aspects of a reliability check of a video data stream are described.SOLUTION: According to a first aspect, a unique identifier identifying the media asset to which the portion of the video data stream to be authenticity checked belongs is included in the authenticity check. According to a second aspect, the certificate of the content provider for performing the authenticity check is obtained from a track of the editor stored in the external resource. A third aspect provides a method for identifying a portion of a video data stream to be checked for authenticity using a digital signature. According to a fourth aspect, a digital signature for checking a portion of a video data stream is obtained from an external resource.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of the present invention relate to an apparatus for checking a video data stream for reliability, an apparatus for rendering a video data stream in which the video is coded in a way that is checkable for reliability, a video decoder, a video encoder, a method for checking a video data stream for reliability, a method for rendering a video data stream in which the video is coded in a way that is checkable for reliability, a method for decoding video, a method for encoding video, a video, and / or a video data stream. [Background technology]

[0002] Content authentication is important to avoid media manipulation. Rapid advances in AI have given rise to the creation of sophisticated deepfakes, blurring the line between real and fake content and raising serious cybersecurity and copyright concerns. Therefore, being able to verify the authenticity of media has become increasingly important in recent years. An exemplary method for performing such authentication includes or consists of digitally signing the media by first hashing the media asset and then signing it with the content generator's private key, so that on the client side, given the content generator's public key, the client can compare the provided signature with the value of the hash that it computed based on the media asset it received. If the values ​​match, the client can safely assume that the media has not been tampered with. Summary of the Invention [Problem to be solved by the invention]

[0003] Existing concepts for authenticity checking of video data streams still leave room for improvement, for example in terms of adaptability to application scenarios, e.g. usefulness in streaming scenarios, as well as compatibility with the structure of video data streams. [Means for solving the problem]

[0004] An object of embodiments of the present invention is to provide a concept for authenticity checking of video data streams, which offers an improved trade-off between the low bit-rate within the video data stream required to provide authenticity checkability, a high degree of adaptability to video codecs, and a high suitability to a number of application scenarios, such as streaming scenarios, e.g., when enabling the extraction of sub-streams of the video data stream.

[0005]

[0013] An embodiment of the first aspect of the present invention relies on the idea of ​​performing a check of the authenticity of a portion of a video data stream by including a unique identifier in the authenticity check, where the unique identifier uniquely identifies the media asset to which the portion being checked belongs. In particular, the authenticity check can be performed by subjecting the portion to a hash function to obtain a hash value and checking whether the combination of the value and the unique identifier matches a digital signature for checking that portion of the video data stream. In other words, the digital signature can be used to verify the combination of the portion and the unique identifier. For this purpose, for example, a digital signature can be obtained by jointly signing the combination of the hash value derived by hashing the portion and the unique identifier. When using a similar approach for other components of a media asset, such as audio or subtitles, a client can verify the combination of the media components it processes.

[0006] An embodiment according to a first aspect of the present invention provides an apparatus for checking a video data stream in which video is encoded for authenticity. The apparatus is configured to apply a predetermined portion of the video data stream, or data derived therefrom, to a hash function to obtain a hash value, obtain a unique identifier (e.g., from the video data stream or a reference, e.g., using a URI) that uniquely identifies a media asset to which the predetermined portion belongs, obtain a digital signature based on the video data stream (e.g., from the video data stream, e.g., from a twcv_signature, or from a metadata file or manifest file, e.g., a C2PA file), and check whether a combination of the hash value and the unique identifier (e.g., a combination of multiple information including the hash value and the unique identifier) ​​matches the digital signature to determine whether the video data stream is authentic.

[0007] A further embodiment according to a first aspect of the present invention provides an apparatus for decoding a video-encoded video data stream, the apparatus being configured to: decode a syntax structure from the video data stream; check the authenticity of the video data stream based on a predetermined portion of the video data stream, the predetermined portion being subjected to a hash function or used to derive data to be subjected to the hash function, derive a hash value operative to check the authenticity of the video data stream; decode a unique identifier or a reference pointing to a unique identifier from the video data stream, the unique identifier uniquely identifying a media asset to which the predetermined portion belongs; and decode an indication of a digital signature from the video data stream (e.g., from the video data stream, e.g., twcv_signature, or from a metadata file or manifest file, e.g., C2PA file), the digital signature being based on a combination of the hash value and the unique identifier (e.g., a combination of multiple information including the hash value and the unique identifier).

[0008] A further embodiment according to the first aspect of the present invention provides an apparatus for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the apparatus being configured to apply a predetermined portion of the video data stream, or data from which the video data stream is derived, to a hash function to obtain a hash value, assign a unique identifier to the predetermined portion that uniquely identifies the media asset to which the predetermined portion belongs, and sign a combination of the hash value and the unique identifier (e.g., a combination of multiple information including the hash value and the unique identifier) ​​to obtain a digital signature.

[0009] A method for checking a video data stream in which video is encoded for authenticity, the method including: subjecting a predetermined portion of the video data stream, or data derived therefrom, to a hash function to obtain a hash value; obtaining a unique identifier (e.g., from the video data stream, or a reference, e.g., using a URI) that uniquely identifies the media asset to which the predetermined portion belongs; obtaining a digital signature based on the video data stream (e.g., from the video data stream, e.g., from a twcv_signature, or from a metadata file or manifest file, e.g., a C2PA file); and checking whether a combination of the hash value and the unique identifier (e.g., a combination of multiple information including the hash value and the unique identifier) ​​matches the digital signature to determine whether the video data stream is authentic.

[0010] A method for decoding a video data stream in which video is encoded, the method including: decoding a syntax structure from the video data stream, including information for checking the authenticity of the video data stream based on a predetermined portion of the video data stream, the predetermined portion being subjected to a hash function or used to derive data to be subjected to a hash function, deriving a hash value that functions to check the authenticity of the video data stream; decoding a unique identifier or a reference pointing to a unique identifier from the video data stream, the unique identifier uniquely identifying a media asset to which the predetermined portion belongs; and decoding an indication of a digital signature from the video data stream (e.g., from the video data stream, e.g., twcv_signature, or from a metadata file or manifest file, e.g., a C2PA file), the digital signature being based on a combination of the hash value and the unique identifier (e.g., a combination of multiple information including the hash value and the unique identifier).

[0011] A method for rendering a video data stream in which the video is encoded in a manner that can be checked for authenticity, the method comprising: subjecting a predetermined portion of the video data stream, or data from which the video data stream is derived, to a hash function to obtain a hash value; assigning a unique identifier to the predetermined portion that uniquely identifies the media asset to which the predetermined portion belongs; and signing a combination of the hash value and the unique identifier (e.g., a combination of multiple information including the hash value and the unique identifier) ​​to obtain a digital signature.

[0012] An embodiment according to the second aspect of the present invention relies on the idea of ​​providing a concept that allows modification of a video data stream while maintaining checkability of authenticity. To this end, in an embodiment of the second aspect of the present invention, a video data stream whose authenticity is to be checked may include an indication of an external resource that keeps a track of the editors of the video data stream. To check the authenticity of a video data stream, an entity can query the editor's track on the external resource for the certificate of the content provider who was the last editor of the video data stream, e.g., the most recent editor, and derive from the external resource a key of this last editor that can be used to perform an authenticity check of the video data stream. For example, the editor's track may include the tracks of all editors who contributed to the video data stream, e.g., from the editor who originally generated the video data stream to any editors who performed changes to the video data stream. Thus, the editor's track may provide a seamless track of changes, each of which is verifiable, e.g., by their respective endorsement certificates. This concept may enable, for example, a trusted transcoder to extract portions of the video data stream, e.g., by selecting one or more substreams from the video data stream. For example, a video data stream may include multiple sub-streams, each of which may represent video at a particular resolution and / or frame rate. A further parameter of data stream scalability may be the extraction of a number of pictures, or, in the case of a multi-view data stream, the extraction of a certain view. A transcoder may, for example, extract sub-streams from a video data stream on behalf of a client requesting a video data stream at a particular bitrate. A reliable transcoder may check any incoming video data stream for reliability, extract the required parts of the video data stream, and render the extracted video data stream checkable for reliability.The trusted transcoder can then complete a certificate for tracking changes to the video data stream so that a receiver of the extracted video data stream can verify the extracted video data stream using the certificate of the trusted transcoder.

[0013] An embodiment according to a second aspect of the present invention provides an apparatus for checking a video data stream in which video is encoded for authenticity, the apparatus being configured to apply a predetermined portion of the video data stream, or data derived therefrom, to a hash function to obtain a hash value, check whether the hash value matches a digital signature (e.g., derived from the video data stream and derived from a criterion indicated in the video data stream) to determine whether the video data stream is authentic, decrypt the digital signature using a public key of an asymmetric decryption scheme to obtain a check value, and check whether the hash value matches the check value, the apparatus being configured to check whether the video data stream includes an indication of an external resource (e.g., a metadata structure in the external resource, such as a manifest file) that includes a track of an editor of the video data stream (e.g., a record of edits or modifications from the generation of the video data stream to its current version, and / or the identities of the corresponding editors), and if the video data stream includes an indication of an external resource that includes a track of an editor of the video data stream, query (or retrieve) the editor's track for a certificate of a content provider that was the last editor of the video data stream, and derive a public key based on the content provider's certificate.

[0014] A further embodiment according to a second aspect of the present invention provides an apparatus for transcoding a video data stream in which video has been encoded with respect to authenticity, the apparatus being configured to receive an input video data stream, check the authenticity of the input video data stream, transcode the input video data stream to generate an output data stream, apply a hash function to a predetermined portion of the output video data stream or of the data from which the output video data stream is derived to obtain a hash value, sign the hash value using a private key of an asymmetric cryptography scheme to obtain a digital signature, provide an editor's track of the output video data stream (e.g., a record of edits or modifications from the generation of the video data stream to its current version and / or the identity of the corresponding editor), to the editor's track being provided in an external resource (e.g., a metadata structure, e.g., a manifest file, at the external resource), a certificate of the content provider (e.g., identifying the apparatus), the certificate including or pointing to the public key of the asymmetric cryptography scheme, and provide the digital signature to the output video data stream (e.g., in an SEI message) or to the external resource (e.g., insert the digital signature into the metadata structure or a further metadata structure and provide the same to the external resource) or the further external resource.

[0015] A further embodiment according to a second aspect of the present invention provides an apparatus for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the apparatus being configured to: hash a predetermined portion of the video data stream, or data from which the video data stream is derived, to obtain a hash value, sign the hash value using a private key of an asymmetric cryptography scheme to obtain a digital signature, provide an editor's track of the video data stream (e.g. a record of edits or modifications from the generation of the video data stream to its current version and / or the identity of the corresponding editor), to the editor's track being provided in an external resource (e.g. a metadata structure, e.g. a manifest file, in the external resource), a certificate of the content provider (e.g. identifying the apparatus), the certificate including or pointing to the public key of the asymmetric cryptography scheme, and provide the digital signature to the video data stream (e.g. in an SEI message) or to the external resource (e.g. inserting the digital signature in the metadata structure or a further metadata structure and providing the same to the external resource) or to the further external resource.

[0016] a public key for the external resource that contains a track of an editor of the video data stream (e.g., a record of edits or modifications from the generation of the video data stream to a current version, and / or the identity of the corresponding editor); and if the video data stream contains an indication of an external resource that contains a track of an editor of the video data stream (e.g., a record of edits or modifications from the generation of the video data stream to a current version, and / or the identity of the corresponding editor); and if the video data stream contains an indication of an external resource that contains a track of an editor of the video data stream, querying (or retrieving) the editor's track for a certificate of a content provider that was the last editor of the video data stream, and deriving a public key based on the content provider's certificate.

[0017] 1. A method for transcoding a video data stream in which video is encoded, the method comprising: receiving an input video data stream; checking authenticity of the input video data stream; transcoding the input video data stream to generate an output data stream; subjecting a predetermined portion of the output video data stream, or of the data from which the output video data stream is derived, to a hash function to obtain a hash value; signing the hash value using a private key of an asymmetric cryptography scheme to obtain a digital signature; providing an editor's track of the output video data stream (e.g., a record of edits or modifications from generation of the video data stream to a current version, and / or the identity of the corresponding editor), the editor's track being provided in an external resource (e.g., a metadata structure, e.g., a manifest file, at the external resource), a certificate of the content provider (e.g., identifying a device), the certificate including or pointing to the public key of the asymmetric cryptography scheme; and providing the digital signature to the output video data stream (e.g., in an SEI message) or the external resource (e.g., inserting the digital signature in a metadata structure or a further metadata structure and providing the same to the external resource) or the further external resource.

[0018] 1. A method for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the method comprising: subjecting a predetermined portion of the video data stream, or data from which the video data stream is derived, to a hash function to obtain a hash value; signing the hash value using a private key of an asymmetric cryptography scheme to obtain a digital signature; providing an editor's track of the video data stream (e.g., a record of edits or modifications from the generation of the video data stream to a current version and / or the identity of the corresponding editor); the editor's track being provided in an external resource (e.g., a metadata structure, e.g., a manifest file) at the external resource; a certificate of the content provider (e.g., identifying the device), the certificate including or pointing to the public key of the asymmetric cryptography scheme; and providing the digital signature to the output video data stream (e.g., in an SEI message) or the external resource (e.g., inserting the digital signature in the metadata structure or a further metadata structure and providing the same to the external resource) or the further external resource.

[0019] An embodiment according to a third aspect of the invention provides a concept for determining the part of a video data stream that is checked for authenticity, which part a digital signature for performing the authenticity check refers to, which results in an improved trade-off between the bit rate required to signal an indication identifying a part in a video data stream and a high degree of adaptability of the concept to the structure of the video data stream.

[0020] An embodiment according to a first type of the third aspect of the invention is based on the idea that the identification of the part of the video data stream to which the digital signature refers for performing the authenticity check is performed based on one or more syntax elements defining the structure of the video data stream, in particular based on one or more of a temporal layer identifier, one or more layer identifiers, a combination of temporal layer identifier and layer identifier, a time frame identifier, a priority level identifier and the AVC nal_ref_id.

[0021] By utilizing syntax elements that define the structure of the video data stream by assigning units of the video data stream, such as pictures, to specific sub-portions of the video data stream, such as temporal layers, tiers, time frames, or priority levels, identification of the portions used for reliability checks is possible without requiring any additional association between the units of the video data stream and the portions used for reliability checks.

[0022] An embodiment according to a first type of the third aspect provides an apparatus for checking a video data stream in which video is encoded for authenticity, the apparatus being configured to apply a predetermined portion of the video data stream, or data derived therefrom, to a hash function to obtain a hash value, and to check whether the hash value matches a digital signature (e.g., derived from the video data stream and derived from criteria indicated within the video data stream) to determine whether the video data stream can be trusted. The device is configured to determine the predetermined portion based on one or more of: a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream that identifies a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream (e.g., layer_id in HEVC / VVC, dependency_id and / or quality_id in AVC) that identify the layer of the video data stream to which the respective picture belongs (e.g., the video data stream is a layered video data stream including multiple layers, such as a base layer and one or more enhancement layers, that represent video at different resolutions or from different viewpoints); a combination of the temporal layer identifier and the layer identifier; a temporal frame identifier (e.g., picture order count, POC); a priority level identifier indicating a priority level of the picture (e.g., priority_id in AVC); and nal_ref_id in AVC.

[0023] A further embodiment according to the first type of the third aspect provides an apparatus for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the apparatus being configured to subject a predetermined portion of the video data stream, or a predetermined portion of data from which a further portion of the video data stream is derived, to a hash function to obtain a hash value, and to sign the hash value (e.g., by using a private key in an asymmetric cryptography scheme) to obtain a digital signature. The device is configured to determine the predetermined portion based on one or more of: a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream that identifies a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream (e.g., layer_id in HEVC / VVC, dependency_id and / or quality_id in AVC) that identify the layer of the video data stream to which the respective picture belongs (e.g., the video data stream is a layered video data stream including multiple layers, such as a base layer and one or more enhancement layers, that represent video at different resolutions or from different viewpoints); a combination of the temporal layer identifier and the layer identifier; a temporal frame identifier (e.g., picture order count, POC); a priority level identifier indicating a priority level of the picture (e.g., priority_id in AVC); and nal_ref_id in AVC.

[0024] A further embodiment according to the first type of the third aspect is a method for checking a video data stream in which video is encoded for authenticity, the method applying a predetermined portion of the video data stream, or data derived therefrom, to a hash function to obtain a hash value, and checking whether the hash value matches a digital signature (e.g. derived from the video data stream and derived from criteria indicated in the video data stream) to determine whether the video data stream is authentic, and the method comprising: a temporal layer (e.g. temporal_id) identifier associated with pictures of the video data stream, which identifies a subset of temporal frames of the video data stream to which each picture belongs; the method includes determining the predetermined portion based on one or more of a temporal layer identifier, one or more layer identifiers associated with pictures of the video data stream (e.g., layer_id in HEVC / VVC, dependency_id and / or quality_id in AVC) that identify the layer of the video data stream to which the respective picture belongs (e.g., the video data stream is a layered video data stream including multiple layers, e.g., a base layer and one or more enhancement layers, that represent video at different resolutions or from different viewpoints), one or more layer identifiers, a combination of the temporal layer identifier and the layer identifier, a time frame identifier (e.g., picture order count, POC), a priority level identifier indicating a priority level of the picture (e.g., AVC priority_id), and nal_ref_id in AVC.

[0025] A further embodiment according to the first type of the third aspect is a method for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the method comprising subjecting a predetermined portion of the video data stream, or a predetermined portion of data from which a further portion of the video data stream is derived, to a hash function to obtain a hash value, and signing the hash value (e.g., by using a private key in an asymmetric cryptography scheme) to obtain a digital signature. The method includes determining the predetermined portion based on one or more of: a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream that identifies a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream (e.g., layer_id in HEVC / VVC, dependency_id and / or quality_id in AVC) that identify a layer of the video data stream to which the respective picture belongs (e.g., the video data stream is a layered video data stream including multiple layers, e.g., a base layer and one or more enhancement layers, that represent video at different resolutions or from different viewpoints); a combination of the temporal layer identifier and the layer identifier; a temporal frame identifier (e.g., picture order count, POC); a priority level identifier indicating a priority level of the picture (e.g., priority_id in AVC); and nal_ref_id in AVC.

[0026] An embodiment according to the second type of the third aspect of the present invention relies on the idea of ​​providing an indication within the syntax structure of the video data stream, which indicates a manner in which to determine the portions of the video data stream on which the reliability check is performed. By signaling the indication, a high degree of flexibility in defining the portions for the reliability check is achieved. For example, the indication can distinguish between different modes of determining the portions for the reliability check, which may include a mode using one or more syntax elements defining the structure of the video data stream as described with respect to the first type of the third aspect of the present invention, or a mode applying dedicated instructions within the video data stream, and assigning units of the video data stream to the portions for the reliability check. Therefore, providing an indication indicating a manner in which to determine the portions for the reliability check, such as by using one or more syntax elements defining the structure of the video data stream, improves the trade-off between a low bit rate for identifying portions, such as by using one or more syntax elements defining the structure of the video data stream, and a high degree of flexibility in defining the portions for the reliability check, such as by providing dedicated instructions associating units of the video data stream with portions.

[0027] A second type of embodiment of the third aspect of the present invention provides an apparatus for checking a video data stream in which video is encoded for authenticity, the apparatus being configured to apply a predetermined portion of the video data stream, or data derived therefrom, to a hash function to obtain a hash value, check whether the hash value matches a digital signature (e.g., derived from the video data stream and derived from criteria indicated in the video data stream) to determine whether the video data stream is authentic, and derive an indication from the video data stream, the indication indicating how the predetermined portion is determined. A further embodiment according to a second type of the third aspect of the present invention provides an apparatus for decoding a further video data stream in which video is encoded. The apparatus is configured to derive from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion of the video data stream (e.g., to be applied with a hash function or used to derive data to be applied with a hash function to derive a hash value that functions to check the authenticity of the video data stream). The syntax structure includes instructions indicating how to determine the predetermined portion of the video data stream.

[0028] A further embodiment according to a second type of the third aspect of the present invention provides an apparatus for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the apparatus being configured to: apply a predetermined portion of the video data stream, or a predetermined portion of data from which a further portion of the video data stream is derived, to a hash function to obtain a hash value, sign the hash value (e.g. by using a private key in an asymmetric cryptography scheme) to obtain a digital signature, and insert instructions into the video data stream, the instructions indicating how to determine the predetermined portion. A further embodiment according to a second type of the third aspect of the present invention provides a method for checking a video data stream in which video is encoded for authenticity, the method comprising: subjecting a predetermined portion of the video data stream, or data derived therefrom, to a hash function to obtain a hash value; checking whether the hash value matches a digital signature (e.g., derived from the video data stream and derived from criteria indicated in the video data stream) to determine whether the video data stream is authentic; and deriving an indication from the video data stream, the indication indicating how the predetermined portion is determined.

[0029] A further embodiment according to a second type of the third aspect of the present invention provides a method for decoding a video data stream in which video is encoded, the method comprising deriving from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion of the video data stream (e.g., data to be applied with a hash function or used to derive data to be applied with a hash function to derive a hash value that functions to check the authenticity of the video data stream), the syntax structure including instructions indicating how to determine the predetermined portion of the video data stream.

[0030] A further embodiment according to a second type of the third aspect of the present invention provides a method for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the method comprising subjecting a predetermined portion of the video data stream, or a predetermined portion of data from which a further portion of the video data stream is derived, to a hash function to obtain a hash value, signing the hash value (e.g. by using a private key in an asymmetric encryption scheme) to obtain a digital signature, and inserting instructions into the video data stream, the instructions indicating a manner in which the predetermined portion is determined.

[0031] An embodiment according to the fourth aspect of the present invention relies on the idea of ​​storing a digital signature for verifying a video data stream in an external resource, for example, instead of signaling the digital signature within the video data stream. For example, the digital signature stored in the external resource can provide verification of the temporal consistency of multiple portions of the video data stream. For example, the digital signature may be obtained by signing a combination of multiple hashes obtained from respective portions of the video data stream. To check the authenticity of the video data stream, a client can obtain the digital signature from the external resource and check whether a hash value obtained from a portion of the video data stream matches the digital signature, for example, by comparing the hash value with a check value whose authenticity is guaranteed by the digital signature. For example, the check value may be part of a check value obtained by decrypting the digital signature or may be verifiable by the digital signature.

[0032] An embodiment according to a fourth aspect of the present invention provides an apparatus for checking a video data stream in which video is encoded for authenticity, wherein the apparatus is configured to apply a predetermined portion of the video data stream (e.g., video data associated with an access unit, e.g., a time frame) or data derived therefrom to a hash function to obtain a hash value, derive a digital signature associated with the predetermined portion from an external resource (e.g., a server), and check whether the hash value matches the digital signature to determine whether the video data stream is authentic.

[0033] A further embodiment according to a fourth aspect of the present invention provides an apparatus for decoding a video-encoded video data stream, the apparatus being configured to derive from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion of the video data stream (e.g., to be subjected to a hash function or used to derive data to be subjected to a hash function to derive a hash value that is functional for checking the authenticity of the video data stream), the syntax structure including a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a digital signature associated with the predetermined portion.

[0034] A further embodiment according to a fourth aspect of the present invention provides an apparatus for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the apparatus being configured to apply a predetermined portion of the video data stream, or data from which the video data stream is derived, to a hash function to obtain a hash value, sign the hash value (e.g., by using a private key in an asymmetric cryptography scheme) to obtain a digital signature, provide the digital signature to an external resource, and insert an indication of the external resource (e.g., a reference to the digital signature of the external resource) into the video data stream (e.g., a URI of the external resource or of the digital signature).

[0035] A further embodiment according to a fourth aspect of the present invention provides a method for checking a video data stream in which video is encoded for authenticity, the method comprising subjecting a predetermined portion of the video data stream, or data derived therefrom (e.g., video data associated with an access unit, e.g., a time frame), to a hash function to obtain a hash value, deriving a digital signature associated with the predetermined portion from an external resource (e.g., a server), and checking whether the hash value matches the digital signature to determine whether the video data stream can be trusted.

[0036] A further embodiment according to a fourth aspect of the present invention provides a method for decoding a video data stream in which video has been encoded. The method includes deriving from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion of the video data stream (e.g., to be subjected to a hash function or used to derive data to be subjected to a hash function to derive a hash value that functions to check the authenticity of the video data stream). The syntax structure includes a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a digital signature associated with the predetermined portion.

[0037] A further embodiment according to a fourth aspect of the present invention provides a method for rendering a video data stream in which the video is encoded in a way that can be checked for authenticity, the method comprising: subjecting a predetermined portion of the video data stream, or data from which the video data stream is derived, to a hash function to obtain a hash value; signing the hash value (e.g., by using a private key in an asymmetric cryptography scheme) to obtain a digital signature; providing the digital signature to an external resource; and inserting an indication of the external resource (e.g., a reference to the digital signature of the external resource) into the video data stream (e.g., a URI of the external resource or of the digital signature).

[0038] A further embodiment of the present invention provides a video data stream, for example stored on a non-transitory digital storage medium, comprising a video data stream obtained by any of the methods described above. Advantageous embodiments are defined by the subject matter of the dependent claims. Embodiments of the present disclosure are described in more detail below with reference to the drawings. [Brief explanation of the drawings]

[0039] [Figure 1]1 illustrates an apparatus for checking the authenticity of a video data stream according to an embodiment; [Figure 2] 1 illustrates an apparatus for decoding video according to an embodiment. [Figure 3] 1 shows an apparatus for checking the authenticity of a video data stream according to an embodiment of the first aspect; [Figure 4] 1 illustrates a verification module according to an embodiment. [Figure 5] 1 illustrates an apparatus for authenticity-checkable rendering of a video data stream according to an embodiment; [Figure 6] 1 shows an apparatus for authenticity checkable rendering of a video data stream according to an embodiment of the first aspect; [Figure 7] 4 shows an apparatus for checking the authenticity of a video data stream according to an embodiment of the second aspect; [Figure 8] 4 shows an apparatus for authenticity checkable rendering of a video data stream according to an embodiment of the second aspect; [Figure 9] 4 illustrates a transcoder according to an embodiment of the second aspect; [Figure 10] 4 shows an apparatus for checking the authenticity of a video data stream according to an embodiment of the fourth aspect; [Figure 11] 10 shows an apparatus for authenticity checkable rendering of a video data stream according to an embodiment of the fourth aspect; [Figure 12] 1 illustrates a video encoder according to an embodiment; [Figure 13] 1 illustrates a video decoder according to an embodiment; [Figure 14] 1 illustrates block partitions of a picture of a video, according to an embodiment. [Figure 15] 10 illustrates the construction of an identification string IdString according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0040]

[0023] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings, in which the same or similar elements, or elements having the same or similar functions, are assigned the same reference numerals or are identified by the same names. In the following description, numerous details are set forth to provide a thorough description of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that other embodiments may be practiced without these specific details. Furthermore, features of different embodiments described herein can be combined with each other unless otherwise specified.

[0041] The description begins with a description of an apparatus for checking a video data stream for authenticity with reference to Figure 1 and a decoder with reference to Figure 2, with further optional details being described with reference to Figure 3. Figure 4 describes an apparatus for rendering a video data stream in an authenticity-checkable manner. The apparatuses of Figures 1 and 2 and 4 may provide a framework within which aspects of the present invention may be implemented. In other words, any of the features and functions described with reference to Figures 1 to 4 may optionally be applied to the embodiments described below, and the features described with reference to Figures 1 to 4 may optionally be combined individually or together with any of the subsequent embodiments.

[0042] FIG. 1 shows an apparatus 16 for checking the authenticity of a video data stream 14. For example, authenticity may mean that the content and / or content provider of the data stream or a predetermined portion are verified to be authentic. The video data stream 14 contains encoded video. The apparatus 16 is configured to apply a predetermined portion 13 of the video data stream 14 to a hash function 31 to obtain a hash value 33. Alternatively, instead of applying the predetermined portion 13 to the hash function, the apparatus 16 may apply data 62 derived from the predetermined portion 13 to the hash function 31 to obtain the hash value 33. The latter option is exemplarily visualized in FIG. 1 by an optional block 61 that can derive data 62 to be subjected to the hash function 31 from the predetermined portion 13. The apparatus 16 comprises an extractor 21 that extracts the predetermined portion 13 from the video data stream 14.

[0043] The device 16 further comprises a verification information derivation unit 51 that obtains the digital signature 43 based on the video data stream 14. For example, the digital signature may be included in the data stream 14, or the data stream 14 may contain a reference to the digital signature. The device 16 further comprises a verification module 41 which checks whether the hash value 33 matches the digital signature 43 to determine whether the data stream 14 can be trusted.

[0044] For example, the extractor 21 can extract information from the data stream 14 to be used in the verification process 41, such as verification information 53 that can be used by the verification information derivation unit 51 to derive a digital signature 43 or a reference to the digital signature 43. For example, validation information 53 may include or consist of one or more syntax elements and / or one or more syntax structures. For example, validation information 53 may include one or more SEI messages.

[0045] For example, video data stream 14 may include multiple payload packets, e.g., referred to as network abstraction layer (NAL) units, e.g., in H.264, H.265, or H.266. The payload packets may include coded video payload packets, e.g., referred to as video coding layer (VCL) NAL units, as well as supplemental information payload packets, e.g., referred to as supplemental enhancement information (SEI) NAL units, which carry information about the coded video data and / or for the decoding process of the coded video data and / or for coding options for decoding the coded video data. The supplemental information payload packets may include one or more supplemental information messages, e.g., referred to as SEI messages. For example, the verification information derivation unit 51 may derive the digital signature 43 from the data stream 14, e.g., from a syntax element that carries the digital signature 43. Alternatively, the verification information 53 may indicate a reference to a metadata file or a manifest file, e.g., a C2PA file, and the verification information derivation unit 51 may derive the digital signature from that reference.

[0046] According to an embodiment, device 16 is configured to derive digital signature 43 from video data stream 14, e.g., from payload packets interspersed in the video data stream between video payload packets carrying encoded video data. For example, device 16 may derive digital signature 43 from an SEI message of the video data stream, e.g., a trustworthy_content_verification SEI message. According to an alternative embodiment, device 16 is configured to derive an indication of an external resource, e.g., a URI, from video data stream 14, e.g., from an SEI message, e.g., trustworthy_content_verification SEI message, of the video data stream. Device 16 can derive digital signature 43 from the external resource. That is, according to an embodiment, the indication of the external resource is a uniform resource identifier that points to a manifest file stored on the server.

[0047] According to an embodiment, the hash value 33 depends on all bits of a predetermined portion 13 of the video data stream. According to an embodiment, the hash value 33 depends on all bits of a given portion 13 of the video data stream in the coded domain (eg in a domain where at least a portion of the video data stream is entropy coded). According to an embodiment, the predetermined portion 13 of the video data stream extends across multiple access units (or time frames) of the video data stream, such that the hash value 33 depends on bits from more than one access unit. Alternatively, the predetermined portion 13 includes video data from only one access unit (or time frame).

[0048] As far as block 61 is concerned, block 61 may, for example, comprise a reconstruction of a portion of a video, which portion is represented by predetermined portion 13. In other words, according to an embodiment, when device 16 subjects predetermined portion 13 of a video data stream, or data derived therefrom, to hash function 31 to obtain hash value 33, it may reconstruct the video with respect to predetermined portion 13 to obtain a reconstructed portion of the video, and subject the reconstructed portion to hash function 31.

[0049] According to an embodiment, device 16 may be a decoder for decoding video data stream 14, e.g., for reconstructing video encoded in the video data stream. For example, device 16 may reconstruct predetermined portion 13 to obtain a reconstructed portion of the video. For example, the reconstruction of the predetermined portion may be part of block 61, which provides data 62 to be subjected to hash function 31. For example, data 62 may correspond to a reconstructed portion reconstructed based on predetermined portion 13. Alternatively, block 61 may derive data 62 from the reconstructed portion of predetermined portion 13. In other words, data 62 derived from predetermined portion 13 may be a reconstructed portion of the video, or even data derived from the reconstructed portion.

[0050] For example, the extractor 21 may comprise a decoding module for decoding the digital signature indications from the video data stream, in other words, the device 16 may be a decoder. According to an embodiment, device 16 is a decoder (e.g., the decoder is configured to decode video from the video data stream by block-based predictive decoding and transform-based residual decoding) for decoding a video data stream (e.g., a decoder compliant with H.264 / AVC or H.265 / HEVC or H.266 / VVC).

[0051] 2 shows an apparatus 20 for decoding a video-encoded video data stream 14 according to an embodiment. The apparatus 20 may be referred to as a decoder 20. The decoder 20 comprises a decoding module 21, which may optionally correspond to the extractor 21 of FIG. 1. The decoding module 21 is configured to decode a syntax structure 52 from the video data stream, which contains information for checking the authenticity of the video data stream based on a predetermined portion 13 of the video data stream that is subjected to a hash function 31, or which is used to derive data that is subjected to the hash function 31 to derive a hash value 33 that serves to check the authenticity of the video data stream. The decoding module 21 is further configured to decode an indication 44 of a digital signature 43 from the video data stream (e.g., from the video data stream, e.g., twcv_signature, or from a metadata file or manifest file, e.g., C2PA file). For example, syntax structure 52 and instructions 44 may be part of validation information 53 of FIG.

[0052] For example, syntax structure 52 may be included in or correspond to a payload packet interspersed among coded video data payload packets, e.g., a supplemental information message as described with respect to FIG. 1, e.g., a first payload packet, which may hereinafter be referred to as a first SEI message. For example, instruction 44 may be included in a further payload packet interspersed among coded video data payload packets, e.g., a supplemental information message, e.g., a twsc_content_verification SEI message. In other words, syntax structure 52 and instruction 44 may be included in different payload packets of an SEI message. The decoder 20 may optionally include the functionality of the device 60 of Figure 1. Furthermore, the decoder 20 comprises a decoding module 63 for decoding the video 11, in particular the predetermined portion 13.

[0053] 3 shows an example of a verification module 41 according to an embodiment. According to this embodiment, the verification module 41 comprises a decryption module 46 that decrypts the digital signature 43 to obtain a check value 47. The verification module 41 according to this embodiment further comprises a verification block 49 that checks whether the hash value 33 matches the check value 47. For example, decryption block 46 may use an asymmetric decryption scheme to decrypt digital signature 43. For example, decryption block 46 may use the public key of the asymmetric decryption scheme to decrypt digital signature 43 to obtain check value 47.

[0054] According to an embodiment, verification block 49 performs a check whether the hash value matches the check value by forming a verification string based on hash value 33 and based on further information. For example, as described below, according to an embodiment of the first aspect of the invention, the further information includes unique identifier 45. Verification block 49 then compares the verification string with check value 47. In an example, comparing the verification string with check value 47 may include a further hash of the verification string, as described in more detail below. In other words, according to an embodiment, the verification module 41 performs a check as to whether the hash value 33 matches the digital signature 43 by forming a verification string based on the hash value and based on further information and comparing the verification string with the digital signature 43 using the public key (comparing the verification string with the digital signature may include decryption performed by the decryption block 46).

[0055] For example, the generation of the digital signature 43 may be performed on the encoder side by forming a verification string and signing it using the private key of an asymmetric cryptography scheme. For example, the signature may comprise a further hash, i.e. hashing the verification string with a further hash function to obtain a further hash value and signing the further hash value. In this example, it may not be possible to reconstruct the verification string from the digital signature 43 at the decoder side, but instead it is only possible to check whether a check value formed using the hash value 33 matches the digital signature, e.g. by forming a verification string and hashing the verification string with a further hash function. In other words, in this case, verification by verification module 41 may comprise hashing the verification string with a further hash function to obtain a further hash value, checking whether the further hash value matches the digital signature, e.g. by decrypting the digital signature using a public key, and checking whether the resulting check value is equal to the further hash value.

[0056] In other words, according to an embodiment, checking whether the hash value 33 matches or matches the check value 47 may involve forming a verification string using the hash value 33, e.g., by concatenating the hash value 33 with further information, such as a further hash value or a hash function identifier as described below, and hashing the verification string, e.g., using the further hash function. The verification module 41 may then check whether the hashed verification string is equal to the check value 47 decrypted from the digital signature. On the encoder side, according to this embodiment, the digital signature may be generated by forming a verification string in the same way as on the decoder side, hashing it using the further hash function, and signing the hashed verification string to obtain the digital signature 43.

[0057] According to alternative embodiments, check value 47 may correspond to a verification string, e.g., hash value 33, or a concatenation of hash value 33 with further information, e.g., a further hash value or a hash function identifier. In other words, decrypting the digital signature in this case may yield hash value 33 as part of check value 47 (or the entire check value 47). In this case, the digital signature may be larger due to the omission of the further hash. For example, whether one or the other of the above alternatives is used may depend on the hash function selected. According to an embodiment, the device 16 derives an indication of an external resource, e.g., a URI, for obtaining the public key from the video data stream 14. In this embodiment, the verification information derivation unit 51 obtains the public key from the external resource indicated in the video data stream 14.

[0058] FIG. 4 illustrates an apparatus 15 according to an embodiment. The apparatus 15 is for rendering a video data stream 14 that can be checked for authenticity of the video encoding. The apparatus 15 is configured to apply a predetermined portion 13 of the video data stream 14, or a predetermined portion 13 of data 62 from which the video data stream 14 is derived, to a hash function 31 to obtain a hash value 33. For example, the data 62 is the data from which the predetermined portion 13 is derived. In this regard, the description of the apparatus 16 in FIG. 1 applies in an equivalent manner as described, for example, with respect to optional block 61. In particular, the data subjected to the hash function 31 to obtain the hash value 33 may be the same as that used by the apparatus 16 to derive the hash value 33. The apparatus 15 comprises a signing module 71 configured to determine a digital signature 43 based on the hash value 33. For this purpose, the signing module 71 can sign the hash value 33 individually or in combination with further data. In other words, the signing module 71 can sign a combination, e.g., a concatenation, of one or more pieces of information including the hash value 33. The apparatus 15 further comprises an inserter 77 configured to insert an indication of the digital signature 43 into the data stream 14, for example in the form of or as part of the verification information 51 described with reference to Figure 1. In other words, the inserter 77 may insert, into the video data stream 14, an indication of the digital signature 43, for example by encoding the digital signature 43, or an indication of criteria from which the digital signature 43 can be derived, for example by encoding the digital signature 43.

[0059] Any description of device 16 may optionally equally apply to device 15, in the sense that device 16 derives information from data stream 14 that can be inserted into data stream 14 by device 15. Furthermore, any hash function, such as hash function 31, used by device 15 may be equivalent to a corresponding hash function used by device 16. The same is true for the input of a corresponding hash function, such as hash function 31, used to derive hash value 33. For example, signing 71 to obtain digital signature 43 and verifying 41 digital signature 43, performed by device 15 and device 16, respectively, may be part of an asymmetric encryption / decryption scheme, each performed by a private / public key pair, where at least the private key is used for signing and the public key is used for decryption. As far as block 61 is concerned, according to an embodiment, device 15 reconstructs the video with respect to the previously determined portion 13 to obtain a reconstructed portion of the video, and data 62 to which hash function 31 is applied may correspond to or be derived from the reconstructed portion.

[0060] According to an embodiment, the device 15 is an encoder configured to encode video into a data stream 14 . According to an embodiment, the allocation module 71 forms a verification string based on the hash value 33 and based on one or more further pieces of information. According to this embodiment, the signing module 71 further signs the verification string using a private key, for example using a signing algorithm, to obtain a digital signature 43. With respect to device 16 embodiments, depending on which device 16 derives information from an external resource or reference, device 15 may be configured to provide this information to the external resource or reference.

[0061] An embodiment of the first aspect of the present invention is described below. Figure 5 shows an apparatus 16 for checking a video data stream 14 for authenticity according to an embodiment of the first aspect of the invention. The apparatus 16 of Figure 5 may optionally correspond to the apparatus 16 of Figure 1, i.e. the apparatus 16 of Figure 5 may be based on any of the embodiments described in relation to Figure 1. Furthermore, the embodiments described below may optionally be combined with any of the embodiments described in relation to the first aspect. The device 16 of FIG. 5 obtains a unique identifier 45 that uniquely identifies the media asset to which a given portion 13 belongs.

[0062] For example, the verification information derivation unit 51 may derive the unique identifier 45 from the video data stream 14, for example from a syntax element, for example a dedicated syntax element carrying the unique identifier, for example having a value corresponding to the unique identifier. Alternatively, the verification information derivation unit 51 may derive the unique identifier from a reference indicated in the video data stream 14, for example by a unique resource identifier (URI). In other words, the verification information 53 may comprise an indication of a reference, for example a URI, from which the device 16 can derive the unique identifier. In other words, according to an embodiment, device 16 derives a unique identifier from video data stream 14 .

[0063] According to an embodiment, device 16 derives the unique identifier 45 from a payload packet signaled in video data stream 14, for example an SEI message. For example, the SEI message may further include one or more of an indication of a hash function, an indication of some portions of the video data stream for which a digital signature is available to verify the authenticity of the video data stream, and an indication of how to obtain a public key to check whether the combination of the hash value and the unique identifier matches the digital signature.

[0064] According to the embodiment of FIG. 5, the verification module 41 checks whether the combination of the hash value 33 and the unique identifier 45 matches the digital signature 43 to determine whether the data stream 14 can be trusted. In other words, according to an embodiment of the first aspect of the present invention, the device 16 obtains a unique identifier that uniquely identifies the media asset to which the given portion 13 belongs. Furthermore, the verification module 41 checks whether the combination of the hash value and the unique identifier 45 matches the digital signature 43 to determine whether the video data stream is authentic.

[0065] In other words, according to the embodiment of the first aspect, the unique identifier of the media asset to which a given portion belongs is included in the verification of authenticity. This verifies not only the authenticity of the given portion 13 itself, but also its association with the media asset. Therefore, the combination of media belonging to a media asset can be verified as the combination of media provided by the content provider that provided the digital signature. Therefore, the embodiment of the first aspect makes it possible to verify the authenticity of the combination of different media substreams of a media asset, thereby discovering, for example, whether a video is combined with an audio stream different from the one provided by the content provider. Furthermore, the use of a unique identifier for the purpose of verifying the combination of media components of a media asset enables verification even when only a subset of the components of the media asset is available, for example, when only one of multiple available audio streams is streamed together with the video stream. When video and audio are assigned together to obtain a digital signature, it is not possible to remove individual components, such as individual audio streams, resulting in the need to always stream the entire media asset or to provide various combinations of different components of the media asset that are signed together. Instead, the use of a unique identifier enables individual verification that a video data stream belongs to a media asset. A similar process can be performed for any further components of the media asset, such as one or more audio streams and / or subtitles. According to an embodiment, the device 16 is configured to check whether a combination of multiple pieces of information, including the hash value 33, the unique identifier 45, and instructions for the hash function 31, matches the digital signature 43 to determine whether the video data stream is trustworthy, and for example the verification module 41 can use this information to construct a verification string.

[0066] According to an embodiment of the first aspect, the verification block 49 of FIG. 3 checks whether the combination of the hash value 33 and the unique identifier 45 matches or coincides with the check value 47 . For example, verification block 49 may form a verification string based on hash value 33 and unique identifier 45, and verification module 41 may use the public key to compare the verification string to digital signature 43. For example, comparing the verification string to the digital signature may include decrypting digital signature 43, e.g., as described with respect to decryption block 46.

[0067] An example of the construction of the verification string is shown in Figure 15, whereby the verification string includes a hash value 33, optionally a unique identifier 45, and further a hash value of a previous part of the video data stream for verifying temporal consistency, and an identifier of the hash function 31. As already mentioned above, the device 16 is able to derive instructions for an external resource to obtain the public key from the video data stream and to derive the public key from the external resource. According to an embodiment of the first aspect, the device 16, e.g. the verification information derivation unit 51, may derive the unique identifier 45 from an external resource, e.g. the same external resource from which the public key is derived. For example, verification information derivation unit 51 may derive the public key and the unique identifier based on the same information derived from video data stream 14. For example, verification information 53 may include an indication of an external resource from which verification information derivation unit 51 may derive unique identifier 45 and the public key.

[0068] According to an embodiment, device 16 checks whether unique identifier 45 matches a unique identifier associated with one or more further media components, such as audio or subtitles. For example, the further media components may be signaled in a data stream that includes a video data stream. For example, checking whether the unique identifier matches a unique identifier associated with one or more further media components may be performed by checking whether the unique identifier is equal to a unique identifier associated with one or more further video components.

[0069] According to an embodiment, device 16 performs a check of the authenticity of video data stream 14 sequentially on multiple portions of the video data stream. The multiple portions may include predetermined portion 13. According to this embodiment, device 16 applies hash function 31 to predetermined portion 13, or data 62 derived therefrom, to obtain a hash value 33. Furthermore, device 16 applies hash function 31 to a further portion of video data stream 14, or further data derived from the further portion of the video data stream, to obtain a further hash value. For example, the further portion is a portion preceding predetermined portion 13, e.g., a portion of the video data stream that precedes the predetermined portion. According to this embodiment, verification module 41 checks whether a combination of the hash value, the further hash value, and the unique identifier matches digital signature 43. In other words, the combination of the multiple pieces of information may include a hash value and a unique identifier. Optionally, the combination may include further information, such as a hash function identifier, as described below. In other words, by way of example, the verification string that may be formed by verification block 49 may include a further hash value derived by subjecting a further portion of the video data stream to hash function 31 .

[0070] 2, according to an embodiment of the first aspect, the decoder 20 decodes from the video data stream a unique identifier 45, or a reference pointing to the unique identifier 45, that uniquely identifies the media asset to which the given portion 13 belongs. Furthermore, the digital signature 43 decoded by the decoder 20 may be based on a combination of the hash value 33 and the unique identifier 45 (e.g., a combination of multiple pieces of information including the hash value and the unique identifier 45). According to an embodiment, the unique identifier 45 is signaled in a syntax structure 52 .

[0071] FIG. 6 illustrates an apparatus 15 for rendering a data stream 14 checkably for authenticity according to an embodiment of the first aspect of the present invention. The apparatus 15 of FIG. 6 may correspond to the apparatus 15 of FIG. 4 . That is, the apparatus 15 of FIG. 6 may be implemented based on any of the embodiments described with reference to FIG. 4 . According to an embodiment of the first aspect, the apparatus 15 comprises a media asset identification module 120 that assigns a unique identifier to a given portion 13. The unique identifier 45 uniquely identifies the media asset to which the given portion 13 belongs. According to this embodiment, a signing module 71 signs a combination of the hash value and the unique identifier to obtain a digital signature 43. For example, the combination includes a plurality of pieces of information including the hash value and the unique identifier, and optionally one or more further pieces of information, such as an identifier of a hash function 31 and / or one or more further hash values. According to an embodiment, the signing module 71 forms a verification string based on the hash value 33 and the unique identifier 45, and optionally based on one or more of the further information already mentioned above.

[0072] An embodiment of the second aspect of the present invention is described below. Figure 7 shows an apparatus 16 for checking a video data stream 14 for authenticity according to an embodiment of the second aspect of the invention. The apparatus 16 of Figure 7 may optionally correspond to the apparatus 16 of Figure 1, i.e. the apparatus 16 of Figure 7 may be based on any of the embodiments described in relation to Figure 1. Furthermore, the embodiments described below may optionally be combined with any of the embodiments described in relation to the first aspect.

[0073] According to the embodiment described with respect to FIG. 7 , the verification module 41 uses the public key 57 of the asymmetric decryption scheme to decrypt 46 the digital signal channel 43 to obtain a check value 47, and checks whether the hash value 33 matches the digital signature 43 by checking whether the hash value 33 matches or coincides with the check value 47 to determine whether the data stream 14 can be trusted. For example, the validation module 41 may be implemented as described with respect to FIG.

[0074] 7, device 16 checks whether video data stream 14 includes an indication 55 of an external resource 280 that includes an editor's track 231 of video data stream 14. If video data stream 14 includes an indication 55 of an external resource that includes an editor's track 231 of video data stream 14, device 16 queries or searches editor track 231 for a certificate 233 of the content provider who is the last editor of the video data stream, for example, according to editor track 231. If video data stream 14 includes indication 55, device 16 derives public key 57 based on certificate 231 of the content provider who is the last editor of the video data stream. For example, the instruction 55 may be part of the verification information 53. For example, a check may be performed by the verification information derivation unit 53 as to whether the instruction 55 is included in the video data stream 14.

[0075] According to an embodiment, the device 16 performs the checking whether the video data stream 14 includes an indication 55 of an external resource including an editor's track 231 of the video data stream 14 by deriving a syntax element from the video data stream, the syntax element being: 1) The video data stream is A) a URI that points directly to the content provider's certificate, or B) a URI pointing to a register of certificates of content providers and an index to a register pointing to content providers of video data streams 14; or 2) Indicate whether or not the video data stream 14 contains indications of external resources, including editor tracks, or distinguish between them.

[0076] In other words, the syntax element can indicate or distinguish between two cases: a first case in which the video data stream includes a URI that leads to the content provider of the video data stream 14, either directly or by pointing to a register of the content provider, and a second case in which the video data stream 14 includes an indication 55 of an external resource, including an editor's track 55. Thus, for example, the same syntax element can be used to signal a URI or to signal an instruction 55, and the syntax element that indicates or distinguishes between those two cases indicates how to read the syntax element that signals the URI or instruction 55 of an external resource, including the editor's track.

[0077] According to an embodiment, the syntax elements and, if present, the indication of external resources 55 including the editor's track are alternatively sent in an SEI message of the video data stream, for example the trustworthy_content_initialization SEI message described below. According to an embodiment, the indication of the external resource 55 containing the editor's track, if present, is sent in an SEI message of the video data stream, e.g., a first SEI message, e.g., called trustworthy_content_initialization SEI message. Furthermore, according to this embodiment, the device 16, e.g., the verification information derivation unit 53, is configured to derive a further digital signature from the external resource 280 if the indication of the external resource 55 is present, and to check whether the payload of the SEI message, i.e., the first SEI message, or a predetermined part thereof, matches the further digital signature.

[0078] For example, external resource 280 may be a metadata structure, e.g., a manifest file, in an external resource; in other words, the external resource may include a metadata structure, e.g., a manifest file. For example, the metadata structure may include information about a content provider or multiple content providers that added video data stream 14. For example, editor track 231 may be a track of editor recordings or modifications made from the generation of video data stream 14 and / or the corresponding editor identifier to the current version. For example, the metadata structure may include, for each editor, information about the identity of the editor and, optionally, further information, such as a location relative to the video data stream and / or a time at which the video data stream was edited.

[0079] By checking whether the payload of the first SEI message, or a predetermined portion thereof, matches the further digital signature, it is possible to not only verify that the video data stream originates from the content provider identified by the certificate of public key 57, but also to verify that metadata indicated in the external resource, which extends beyond identifying the content provider, such as the location and time of editing the video data stream, is associated with video data stream 14. For example, the SEI message, or a predetermined portion thereof, that is checked against the further digital signal signature may be unique, e.g., unique with respect to a further video data stream, e.g., a further data stream from the same content provider, and optionally, unique with respect to a further content provider. Thus, performing the further digital signature on the payload of the first SEI message, or a predetermined portion thereof, prevents erroneous metadata from being associated with the video data stream, e.g., by swapping the external resource's indication to point to another resource from the same content provider.

[0080] According to an embodiment, the first SEI message or a predetermined portion of the payload of the first SEI message includes a unique identifier, e.g., a payload portion that is a designation for a video data stream. Checking the predetermined portion of the SEI message, or the SEI message, with a further digital signature, an external resource, and thereby information, e.g., a manifest file stored in the external resource, is securely associated with a particular video data stream. According to an embodiment, a predetermined portion of the payload of the SEI message excludes indications of external resources, including tracks for the editor. By excluding the external resource indication 55 from a predetermined portion of the payload of the first SEI message, it becomes possible to change the location at which the external resource is provided without having to recalculate further digital signatures.

[0081] According to an embodiment, the syntax structure further includes a media component identifier. For example, the media component identifier identifies the video data stream 14 among multiple media components of the media message. According to this embodiment, the device 16 is configured to use the media component identifier to select a further digital signature from a set of one or more digital signatures included in the external resource 280. For example, each of the one or more digital signatures is associated with a media component, e.g., audio, video, subtitles. According to an embodiment, the syntax structure further comprises one or more of an indication of a hash function 31 and an indication of several parts of the video data stream for which a digital signature is available to verify the authenticity of the video data stream 14, e.g., the syntax structure is a trustworth_content_initialization SEI message.

[0082] FIG. 8 illustrates an apparatus 15 for rendering a video data stream 14 in a manner that can be checked for authenticity, according to an embodiment of the second aspect of the present invention. The apparatus 15 of FIG. 8 may optionally correspond to the apparatus 15 of FIG. 4, i.e., the apparatus 15 of FIG. 8 may be based on any of the embodiments described with reference to FIG. 4. According to the embodiment described with reference to FIG. 8, the apparatus 15 signs the hash value 33 using a private key 58 of an asymmetric cryptography scheme to obtain a digital signature 43. The apparatus 15 according to this embodiment provides, in an editor's track 231 of the video data stream 14, a certificate 233 that identifies, for example, the apparatus 15, a content provider, such as the content provider of the video data stream 14, for example, the editor's track provided on an external resource 280. The certificate 233 includes or points to a public key 57 for the asymmetric cryptography scheme. According to this embodiment, the apparatus 15 provides the digital signature 43 in the video data stream 14, for example, by inserting the digital signature 43 into a metadata structure or further metadata structure, and provides the digital signature 43 to the external resource 280. Further alternatively, device 15 may provide digital signature 43 to a further external resource different from external resource 280 .

[0083] 9 shows an apparatus 17 for transcoding a video data stream according to an embodiment of the second aspect of the present invention. The apparatus 17 is configured to receive an input video data stream 14' and to check the authenticity of the input video data stream 14'. To this end, the apparatus 17 may comprise an apparatus 15' for checking the video data stream for authenticity. The apparatus 15' may correspond to any of the apparatuses 15 described herein. The apparatus 17 further comprises a transcoder 12 for transcoding the input data stream 14'. In doing so, the transcoder 12 results in a data stream 14", based on which the device 17 derives the output data stream 14. For example, when transcoding the input data stream 14', the transcoder 12 can extract portions of the input data stream 14' to be forwarded in the data stream 14. For example, the transcoder 12 can selectively forward portions to be forwarded in the data stream 14. That is, the transcoder 12 can drop portions of the input data stream 14'. For example, the transcoder 12 can select one or more sub-streams of the input video data stream 14' to be forwarded in the data stream 14. Furthermore, the transcoder 12 can adapt information to be forwarded, such as supplemental enhancement information. For example, the transcoder 12 can adapt verification information 53. However, in an alternative example, the input video data stream 14' is not necessarily verifiable. Therefore, in an example, the device 17 can add verification information 53 in the output video data stream 14.

[0084] The device 17 renders the output video data stream 14 authentically checkable, for example, as described with respect to any of the devices 15 according to the embodiment of the second aspect described with respect to FIG. 8 . In other words, the device 17 applies a predetermined portion 13 of the output video data stream 14, or data 62 from which the output data stream 14 is derived, to a hash function 31 to obtain a hash value 33. The device 17 signs the hash value 33 using a private key 58 of the asymmetric cryptography to obtain a digital signature 43. The device 17 further provides a content provider certificate 233 in an editor's track of the output video data stream 14, the edit track being presented in an external resource 280, and the certificate 233 includes or points to a public key of the asymmetric cryptography. An inserter 77 of the device 17 provides the digital signature 43 in the output video data stream 14, or in the extension resource 280 or a further external resource. Any optional features and details described with respect to device 15 of Figure 8 may optionally be applied to device 17 of Figure 9. In particular, like reference numerals in Figures 8 and 9 may represent like functions and features.

[0085] Embodiments of the third aspect of the invention are described below with respect to Figures 1, 2 and 4. Embodiments of the third aspect of the invention may be combined with any of the features and details described with respect to any of the previous embodiments of Figures 1-9. Referring to FIG. 1, according to a first type of embodiment of the third aspect of the present invention, the device 16 comprises: a temporal layer identifier, such as temporal_ID, associated with the pictures of the video data stream. The temporal layer identifier identifies the subset of temporal frames of the video data stream to which the respective subset of pictures belongs. - One or more layer identifiers associated with the picture of the video data stream. - a combination of a time layer identifier and a layer identifier; - time frame identifier, a priority level identifier indicating the priority level of the picture, - The H.264 / AVC syntax element nal_ref_id. The predetermined portion 13 is configured to determine the predetermined portion 13 based on one or more of the following:

[0086] For example, a temporal layer of a video data stream may include a subset of the temporal frames of the video data stream, with the temporal frames of different temporal layers being interleaved with each other in the temporal order of the pictures of the video data stream, e.g., the inter-picture presentation order. Thus, for example, a single temporal layer may represent video at a first frame rate, while a combination of multiple temporal layers may represent video at a second frame rate that is higher than the first frame rate. In other words, pictures of two temporal layers of a video data stream may be interleaved in the temporal order of the pictures of the video data stream. As far as a layer identifier is concerned, the layer identifier identifies the layer of the video data stream to which the picture associated with the layer identifier belongs.

[0087] For example, a video data stream may include, for a timestamp, multiple pictures associated with different layers of the video data stream, e.g., within one access unit, where the pictures of the multiple layers represent the picture at the timestamp at different resolutions or qualities, or provide different perspectives on the timestamp, or provide different types of information.

[0088] For example, the video data stream may be a layered video data stream including multiple layers, including, for example, a base layer and one or more enhancement layers representing video at different resolutions, or multiple layers representing video from different viewpoints. For example, the layer identifier may refer to one or a combination of the syntax element layer_id in HAV / VVC and the two syntax elements dependency_id and quality_id in AVC. The above-mentioned time frame identifier may, for example, indicate the position of the picture with which it is associated within a defined temporal order among pictures, for example a presentation order called a Picture Order Count POC. The above priority level identifiers can refer to the AVC syntax element priority_id.

[0089] According to an embodiment, the device 16 is configured to derive instructions from the video data stream 14, the instructions indicating how to determine the predetermined portion 13. For example, the instructions may be part of the verification information 53 . For example, the indication may be signaled in a syntax structure, such as the first syntax structure. For example, the indication may be signaled in a sample extended information message.

[0090] According to an embodiment, the instructions indicating how to determine the predetermined portion 13 include: an indication associated with a time frame, such as an access unit, of the video data stream, the indication indicating whether the time frame belongs to a given portion 13; - a time layer identifier, one or more layer identifiers, - a combination of a time layer identifier and a layer identifier; - time frame identifier, - priority level identifier, - NAL_REF_ID for AVC. Distinguish one or more of the following.

[0091] In other words, the instruction indicating the manner in which the predetermined portion is determined may indicate which of the above syntax elements or instructions is used to determine the predetermined portion. For example, the indication associated with a time frame of the video data stream may refer to dedicated signaling of the predetermined portion, e.g., provided by one or more SEI messages signaled within the video data stream, the SEI messages indicating which portions of the video data stream belong to the predetermined portion 13. For example, the indication indicating where a time frame belongs to a predetermined portion may be provided by a trustworthy_content_initialization SEI message and / or a trustworthy_content_selection SEI message, e.g., as described below.

[0092] For example, device 16 may perform verification of video data stream 14 in portions called substreams, such as verification substreams, etc. To this end, device 16 may determine, for each verification substream, the portion of the video data stream that is used to verify the respective verification substream, i.e., the portion of the video data stream that is part of the portion that is applied with hash function 31 to verify the respective verification substream. For example, device 16 may determine, for each picture in the video data stream or a subset of pictures in the video data stream, which of one or more verification substreams each picture belongs to, and include the picture, e.g., the coded video payload packet of the picture, in the portion that is subjected to hash function 31.

[0093] According to an embodiment, device 16 performs the determination of which verification sub-stream a picture should be associated with in response to one of the syntax elements / indications mentioned above. According to an embodiment, the device 16 determines the predetermined portion 13 based on a temporal layer identifier, a layer identifier or a time frame identifier. According to this embodiment, the device 16 derives a range of values ​​from the video data stream, the range of values ​​indicating the value of the respective identifier, i.e., the temporal layer identifier, the layer identifier or the time frame identifier, and the value is associated with the predetermined portion 13. That is, for example, pictures for which the respective identifier has a value within the range of signal values ​​belong to the predetermined portion 13.

[0094] According to a further embodiment of the first type of the third aspect of the present invention, the device 15 of Fig. 4 may determine or select the predetermined portion 13 to be multiplied by the hash function 31 based on the same criteria as described with respect to the device 16. As far as the indication indicating how to determine the predetermined portion 13 is concerned, the device 15 may provide this indication in the video data stream 14, for example in an SEI message, for example in the first syntax structure.

[0095] According to a second type of embodiment of the third aspect of the invention, the device 16 derives an indication from the video data stream (for example, twci_substream_selection_idc, as described below), which indicates how to determine the predetermined portion 13. In other words, according to an embodiment, device 16 is capable of determining predetermined portion 13 of video data stream 14 in response to instructions indicating how to determine the predetermined portion. According to an embodiment, the instructions indicating how to determine the predetermined portion 13 include: an indication associated with a time frame, such as an access unit, of the video data stream, indicating whether the time frame belongs to a given portion 13; - a time layer identifier, one or more layer identifiers, - a combination of a time layer identifier and a layer identifier; - time frame identifier, - priority level identifier, -nal_ref_id of AVC. Distinguish one or more of the following.

[0096] Details regarding these indications and identifiers, as well as the manner of signaling the indications indicating the manner of determining the predetermined portion described with respect to the first type of embodiment of the third aspect of the invention, may optionally be applied in an equivalent manner to the second type of embodiment of the third aspect of the invention. In other words, according to an embodiment, the indication of how to determine the predetermined portion n is a syntax element that distinguishes between multiple modes of determining the predetermined portion.

[0097] For example, the plurality of modes may include a first mode and / or a second mode. In one embodiment, the plurality of modes may consist of a first mode and a second mode. Thus, for example, a syntax element indicating how to derive a given portion may be a flag with exactly two states. For example, the first mode may be a mode in which, for a given picture, the decision of whether to include the picture in a given portion depends on the assignment of the picture to one of the layers. In other words, the decision may depend on the layer to which the given picture belongs, for example, the layer index. For example, a given portion may be associated with one of multiple layers of a video data stream, and if the layer to which the given picture belongs corresponds to the layer associated with the given portion, the given picture is included in the given portion, and if not, it is not included in the given portion.

[0098] For example, the reliability of a video data stream may be checked in portions called verification substreams, each of which is identified using a substream id, as in the example syntax presented below. For example, in a first mode, a given picture of the video data stream may be associated with one of the verification substreams or portions depending on one of the above-mentioned attributes of the given picture, such as the layer index or layer identifier of the layer to which the picture belongs and / or the index or identifier of the temporal layer to which the picture belongs. For example, in the first mode, a given picture may be assigned to the verification substream associated with the layer and / or temporal layer to which the given picture belongs.

[0099] According to an embodiment, in the second mode, a given picture may be assigned a predetermined one of the multiple validation substreams, for example a default one, for example a substream equal to 0, which may be the predetermined portion 13. As already mentioned above and as will be explained in more detail below, the video data stream may further include a dedicated indication for a picture, indicating the verification substream to which the picture belongs, see for example the content selection SEI message. According to an embodiment, the manner of determining the predetermined portion may depend on the presence of a dedicated indication, such as an SEI message for the picture, in addition to the mode indicated by the indication of the manner of determining the predetermined portion. For example, if such a dedicated indication is present for a picture, the picture may be assigned to the verification substream indicated in the dedicated indication, e.g., the substream ID indicated in the content selection SEI message present for the picture; if such a dedicated indication is not present, the assignment of the picture to one of the verification substreams may be performed according to the mode indicated by the indication of the manner of determining the predetermined portion, e.g., according to the first mode or the second mode. In other words, for example, the above-mentioned predetermined picture may be a picture for which the dedicated identification information of the verification sub-stream to which the picture is associated is not signaled in the data stream.

[0100] According to a further embodiment of the second type of the third aspect of the invention, the device 15 inserts an indication into the video data stream 14, the indication indicating how to determine the predetermined portion 13. For example, the device 15 may determine or select the predetermined portion according to any of the criteria described with respect to the third aspect of the present invention, and the device 15 may indicate the scheme used to determine the predetermined portion 13 within the video data stream 14.

[0101] Embodiments of the fourth aspect of the present invention are described below. Figure 10 shows an apparatus 16 according to an embodiment of the fourth aspect of the present invention. The apparatus 16 of Figure 10 may optionally correspond to the apparatus 16 of Figure 12, i.e. any of the features and details described in relation to Figure 12 may optionally be applied to the apparatus 16 of Figure 10. The apparatus 16 according to Fig. 10 is configured to apply a predetermined portion 13 of a video data stream, or data 62 derived therefrom, to a hash function 31 to obtain a hash value 33. For example, according to an embodiment of the fourth aspect, the predetermined portion 13 may be a time frame, e.g., a coded video associated with a certain time frame of the video data stream, e.g., the predetermined portion 13 may be an access unit. The apparatus 16 according to Fig. 10 is configured to derive a digital signature 43 from an external resource 480, such as a server. The digital signature 43 may be associated with the predetermined portion 13, i.e., for example, the digital signature 43 may include or be derived based on a hash value derived from the predetermined portion. The apparatus 16 according to Fig. 10 is configured to check whether the hash value 33 matches the digital signature 43 to determine whether the video data stream is authentic.

[0102] According to an embodiment, device 16 is configured to derive a reference to an external resource from the video data stream. For example, device 16 may derive a reference to external resource 480 from a syntax structure, such as a first syntax structure, such as an SEI message of the video data stream, such as a trustworthy_content_initialization SEI message. According to an embodiment, device 16 decrypts digital signature 43 using the public key of the asymmetric decryption scheme to obtain check value 47, for example, as described with respect to decryption block 46 of Figure 14. Device 16 can check whether the hash value also matches or coincides with check value 47, for example, as described with respect to verification block 49.

[0103] For example, the check value 47 here refers to a portion of a value obtained by decrypting the digital signature 43. For example, the check value may be a portion of a value obtained by decrypting a digital signature associated with a given portion 13 of the video data stream 14. In other words, according to an embodiment, decrypting the digital signature 43 may generate a value that is a combination or instantiation of multiple hash values ​​obtained from multiple portions of the video data stream. In other words, according to an embodiment of the fourth aspect, the digital signature 43 stored in the external resource may be a signed version of a combination or instantiation of multiple hash values, each associated with a respective portion of the video data stream. For example, signing of the combination or instantiation of hash values ​​may be performed without further hashing, so that decryption results in the combination or instantiation of the originally signed hashes.

[0104] According to an embodiment, device 16 derives a portion identifier from video data stream 14, the portion identifier being associated with a given portion 13. For example, the portion identifier is a hash identifier or hash index, e.g., twcs_associated_hash_IDX, that identifies a portion of a digital signature associated with the given portion. For example, the portion identifier may be associated with the given portion in the sense that the portion identifier is signaled for the given portion. For example, the portion identifier may be signaled in a syntax structure, e.g., an SEI message, that is signaled before the given portion, e.g., access unit, to which the portion identifier refers. According to an embodiment, device 16 uses a portion identifier to identify a portion of the check value. In checking whether a hash value matches the check value, device 16 can check whether hash value 33 matches the portion of the check value identified by the portion identifier.

[0105] According to an embodiment, device 16 derives a media component identifier, e.g., twcs_associated_hash_group_ID, from video data stream 14, where the media component identifier indicates the media type of a given portion 13. For example, the media type may be one of video, audio, or subtitles. Device 16 can use the media component identifier to identify a portion of the check value, i.e., a portion to be compared with hash value 33 when checking whether the hash value matches the check value. For example, device 16 can use the media component identifier in addition to the portion identifier to identify a portion of the check value to be compared with hash value 33 to check whether hash value 33 matches the check value.

[0106] Figure 11 shows an apparatus 15 according to an embodiment of the fourth aspect of the present invention. The apparatus 15 of Figure 11 may correspond to the apparatus 15 of Figure 4, i.e. any of the features and details described in relation to Figure 4 may optionally be applied to the apparatus 15 of Figure 11. The device 15 of Figure 11 is configured to apply a predetermined portion 13 of the video data stream, or data 62 from which the predetermined portion 13 of the video data stream is derived, to a hash function 31 to obtain a hash value 33. The device 15 of Figure 11 signs the hash value 33 to obtain a digital signature 43, for example, by using a private key in an asymmetric cryptography scheme. The device 16 provides the digital signature 43 to an external resource 480. The device 16 further asserts an indication of the external resource 480 in the video data stream. For example, the device 16 can insert a reference to the digital signature 43 on the external resource 480, for example, a URI of the external resource or the digital signature, into the video data stream 14. Note that signing of the hash value by device 16 to obtain digital signature 43 may be optional. Alternatively, device 16 may provide hash value 33 to an external resource, and signing may be performed at the external resource, e.g., a server.

[0107] According to an embodiment of the fourth aspect, the signing of hash value 33 may be performed in combination with further hash values, i.e., a combination or instantiation of hash values ​​of multiple portions of video data stream 14 may be formed and the combination may be assigned together to give digital signature 43. In this way, the combination of portions of the video data stream may be verified. This aspect applies regardless of whether the signing is performed by device 15 or by external resource 480.

[0108] The further details and aspects described in relation to the fourth aspect of the invention in relation to device 16 may optionally be applied in a corresponding manner to device 15, for example in the sense that device 15 inserts information derived from video data stream 14 by device 16 into video data stream 14.

[0109] Below, aspects of the invention are again described in different terms, and specific implementations and further embodiments of the invention are described. The embodiments described with respect to Figures 1-11 can be considered generalizations of the embodiments described below, but the following description can further include additional embodiments of the invention. Embodiments of the first aspect of the present invention may rely on the discovery that the first problem arises due to the following facts: 1) A media asset may consist of several components: audio, video, and subtitles. 2) Each of these components may be available in several bit rates, resolutions, and languages.

[0110] Given this, the hash or signature cannot be applied jointly to the entire content consumed, i.e., across the different combinations that each different receiver may get (e.g., receiver 1 may consume 4K+English audio, receiver 2 8K+German, etc.) It is important to verify these different components together, as otherwise the audio and video of different videos may be mixed, which could lead to fake media.

[0111] As a first embodiment, a solution to this problem that does not involve hashing / signing different components jointly consists of adding a unique identifier to the SEI message (twci_content_uuid of the Trusted Content Initialization SEI message in the example below) that is the unique identifier of the specific media asset (same for each component, such as audio, video, subtitles, etc.) used during hashing / signing. For example, once a hash value is calculated for a particular set of coded pictures, the hash + unique identifier are signed together. Alternatively, the hash of a previous or dependent set of coded pictures, together with the hash value of the current set of coded pictures and the hash method type value and unique identifier, are composed into a string signed with the content provider's key. Further media components use similar unique identifiers as well. Alternatively, instead of adding a unique identifier to the SEI message, the unique identifier can be included in a reference containing a public key (a pointer or metadata containing the certificate used for signing), which is used to compute a hash or digital signature as described above.

[0112] Embodiments of the second aspect of the invention may rely on the discovery that problems that arise when transmitting media streams may require modifications to the transmission chain. For example, if there is insufficient bandwidth in the network, a trusted transcoder may need to change the video resolution or video bitrate and re-encode it. When this occurs, authentication of the original media stream cannot occur because it may have been altered. However, if each entity in the chain is trusted, each entity authenticates the incoming data and digitally signs the outgoing data while still providing metadata that tracks the changes, and the receiver can track all changes and verify the data with the last entity's key while still ensuring that the data has not been tampered with but that only "authorized" changes have been made (e.g., bitrate reduction through transcoding). In the following embodiments, a URI is provided that identifies metadata (e.g., a C2PA Manifest) that indicates the changes and further provides the certificate of the last signing entity within that metadata. However, a "man in the middle" could take the stream and link an incorrect, unauthorized link to such a metadata file. Further embodiments generate an SEI message pointing to that metadata-URI with an additional payload that makes such an SEI unique using a hash / digitally signed value included in the metadata calculated by the unique payload of such an SEI. Note that the linking achieved by hashing the unique SEI payload contained in the indicated metadata may be optional and may be indicated by an additional flag in the SEI message (although not present in the example below, a trusted content initialization SEI message may include the syntax element twci_payload_hash_in_c2pa_flag).

[0113] In the following, exemplary syntax for implementing embodiments of the first and second aspects of the present invention is described, which are illustrated with examples of a common syntax, although embodiments of the first and second aspects may be implemented independently of each other. 1.1 Reliable Content Initialization SEI Message 1.1.1 Reliable Content Initialization SEI Message Syntax JPEG2026016324000002.jpg74170 For example, if mode_idc is 0, then twci_content_uuid_present_flag should be 1. For example, twci_key_retrieval_mode_idc is used to distinguish between the mode when the certificate is in a C2PA Manifest Store identified by a URI and the mode when the URI (+idx) directly identifies the certificate. The unique identifier 45 may be signaled using twci_content_uuid. In other words, compared to previous implementations, the embodiment of the first aspect may introduce the syntax elements twci_content_uuid and, optionally, twci_content_uuid_present_flag.

[0114] The indication 55 described with respect to the second aspect may be signaled by twci_key_source_uri, optionally in combination with the further syntax element twci_c2pa_hash_idx (see below), if twci_key_retrieval_mode_idc = 0. Thus, in the above implementation, an embodiment of the second aspect may introduce the syntax element twci_key_retrieval_mode compared to the previous implementation. Thus, lines 5-7 and 12 of the above syntax may represent modifications to the previous implementation.

[0115] 1.1.2 Reliable Content Initialization SEI Message Semantics The Trusted Content Initialization SEI message, the Trusted Content Selection SEI message, and the Trusted Content Verification SEI message provide mechanisms for verifying that the coded video was generated by a trusted content provider. The Trusted Content Initialization SEI message provides information about the secure hash algorithm used to compute a message digest, which is used together with the digital signature present in the Trusted Content Verification SEI message to verify the authenticity of the VCL NAL units present in the coded video sequence. The Trusted Content Initialization SEI message further provides information about the digital signature algorithm used and the content provider's public key. The Trusted Content Initialization SEI message can provide the content provider's public key by providing a URI that identifies a C2PA Manifest Store that contains a certificate with the content provider's public key, or by providing a URI that directly identifies the certificate.

[0116] If any reliable content initialization SEI message, reliable content selection SEI message, or reliable content verification SEI message is present in a coded video sequence, it is a requirement for bitstream conformance that a reliable content initialization SEI message be present in all access units of the coded video sequence, including IDR access units and CRA pictures. It is a requirement for bitstream conformance that any reliable content selection and reliable content verification SEI messages of an access unit be preceded by a reliable content initialization SEI message.

[0117] The reliable content initialization SEI message applies to the current coded picture and all subsequent coded pictures until one or more of the following conditions become true: - The bitstream ends. - A new coded video sequence begins. A new reliable content initialization SEI message is received. twci_hash_method_type indicates the secure hash algorithm used to compute a message digest of a subset of the VCL NAL units of a coded video sequence. Based on these message digests and the digital signature present in the trusted content verification SEI message, a decoder can verify that the coded video was generated by the content originator indicated by the syntax elements twci_use_key_register_idx_flag, twci_key_source_uri, and, if twci_key_register_idx flag is equal to 1, twci_key_register_idx. Supported values ​​of the syntax element twci_hash_method_type, the block size used to compute the message digest, and the size of the computed message digest are specified in Table 1. Values ​​of twci_hash_method_type not listed in the table are reserved for future use by ITU-T|ISO / IEC and shall not be present in payload data conforming to this version of the specification. Decoders shall ignore Trusted Initialization SEI messages that contain reserved values ​​of twci_hash_method_type. The secure hash algorithms listed in Table 1 are specified in the "Secure Hash Standard," FIPS PUB 180-4. Table 1 - Supported values ​​for twci_hash_method_type JPEG2026016324000003.jpg53170 twci_num_verification_substreams_minus1 plus 1 indicates the number of substreams over which the message digest is calculated and signatures may be present in the following trusted content verification SEI messages. The variable NumVerificationSubstream is derived as follows: NumVerificationSubstream=twci_num_verification_substreams_minus1+1. twci_use_key_register_idx_flag equal to 1 indicates that the URI contained in twci_key_source_uri specifies a certificate register and the syntax element twci_key_register_idx is present in the SEI message. twci_use_key_register_idx_flag equal to 0 indicates that the URI contained in twci_key_source_uri specifies a certificate and the syntax element twci_key_register_idx is not present in the SEI message. twci_content_uuid_present_flag equal to 1 specifies that the syntax element twci_content_uuid is present. twci_content_uuid_present_flag equal to 0 specifies that the syntax element twci_content_uuid is not present.

[0118] twci_key_source_uri contains a URI with syntax and semantics specified in IETF Internet Standard 66. If twci_use_key_register_idx_flag is equal to 0, the URI identifies the content provider's certificate that can be used to verify the signature present in subsequent trusted verification SEI messages. Otherwise (twci_use_key_register_idx_flag is equal to 1), the URI identifies the certificate register and content provider's certificate that can be used to verify the signature present in subsequent trusted verification SEI messages indicated by twci_key_register_idx. twci_key_retrieval_mode_idc equal to 0 indicates that the URI contained in twci_key_source_uri specifies a C2PA Manifest Store as specified in the C2PA Technical Specification. twci_key_retrieval_mode_idc equal to 1 indicates the URI contained in twci_key_source_uri, if present, and twci_key_register_idx specifies a certificate.

[0119] twci_c2pa_hash_idx, if present, contains an index specifying an entry in c2pa.hash.data of the Active Manifest, as defined in the C2PA Technical Specification, associated with the current Trusted Content Initialization SEI message. For twci_key_retrieval_mode_idc equal to 0, the media asset for which the Active Manifest brings content bindings is a Trusted Content Initialization SEI message, as specified in the C2PA Technical Specification. The following constraints apply to the C2PA Manifest Store identified by twci_key_source_uri:

[0120] -The Active Manifest shall contain exactly one c2pa.hash.data, which is strongly binding on the content assertion, as specified in the C2PA Technical Specification. For example, this makes it possible to verify the UUID of the SEI used to sign a VCL NAL unit. The exclusion range indicated in -c2pa.hash.data shall match the twci_key_source_uri byte in the Trusted Content Initialization SEI message. For example, by not hashing the URI, you can change the location of the manifest without having to modify the manifest. twci_key_register_idx contains an index specifying the content provider's certificate in the certificate register indicated by twci_key_source_uri, which can be used to verify the signature present in the following trusted verification SEI message.

[0121] The certificate indicated by the syntax elements twci_key_retrieval_mode_idc, twci_use_key_register_idx_flag, twci_key_source_uri, and twci_key_register_idx if twci_use_key_register_idx_flag is equal to 1, shall specify the digital signature method together with associated parameters (if applicable) and the content provider's public key. The format in which this information is provided when twci_key_retrieval_mode_idc is equal to 1 is outside the scope of this specification. It is suggested to use a digital signature algorithm that complies with FIPS 186-5, the "Digital Signature Standard". twci_content_uuid, if present, indicates the identifier of the video content and shall have a value specified as a UUID according to the procedures of ISO / IEC 11578:1996, Annex A.

[0122] When a reliable content initialization SEI message is received, the calculation of the NumVerificationSubstream message digests is initialized according to the FIPS PUB 180-4 specification for the specified twci_hash_method_type. Each VCL NAL unit following the reliable content initialization SEI message is associated with one of the NumVerificationSubstream message digests. The verification substream id is either indicated by the reliable content selection SEI message or inferred to be equal to 0 if no reliable content selection SEI message is present for the coded picture. The message used to calculate the kth message digest, where k is in the range 0 to twci_num_verification_substreams_minus1, inclusive, is obtained by concatenating all VCL NAL units associated with the kth verification substream. The message digest calculation is performed on a block-by-block basis, with the block size specified in Table 1 depending on the value of twci_hash_method_type. For each VCL NAL unit, the associated message digest is updated according to the algorithm specified in FIPS PUB 180-4 for the specified twci_hash_method_type. Note that because the message digest is calculated for the concatenation of all VCL NAL units for the verification substream, some of the processing blocks typically span two or more consecutive VCL NAL units.

[0123] 1.2 Trusted Content Selection SEI Message 1.2.1 Reliable Content Selection SEI Message Syntax JPEG2026016324000004.jpg261701.2.2 Reliable Content Selection SEI Message Semantics The trusted content selection SEI message provides a mechanism for associating a coded picture with one of the verification substreams indicated in the trusted content initialization SEI message. It is a bitstream conformance requirement that any reliable content selection SEI message be preceded by a reliable content initialization SEI message within the same coded video sequence.

[0124] twcs_verification_substream_id indicates the verification substream to which the VCL NAL units of the current coded picture are assigned. If a reliable content initialization SEI message was present in the current coded video sequence but no reliable content selection SEI message was present in the coded picture, the value of twcs_verification_substream_id is inferred to be equal to 0. The value of twcs_verification_substream_id shall be in the range from 0 to twci_num_verification_substream_minus1, inclusive. As specified in Section 0, the message digest of the verification substream with id equal to twcs_verification_substream_id is updated with the VCL NAL unit of the current coded picture according to the twci_hash_method_type specified in the preceding Reliable Content Initialization SEI message.

[0125] 1.3 Trusted Content Validation SEI Messages 1.3.1 Reliable Content Validation SEI Message Syntax JPEG2026016324000005.jpg37170 1.3.2 Reliable Content Validation SEI Message Semantics The trusted content verification SEI message provides a mechanism for verifying the authenticity of video content. It is a bitstream conformance requirement that any reliable content verification SEI message be preceded by a reliable content initialization SEI message within the same coded video sequence. When a coded video sequence includes a reliable content initialization SEI message, it is a bitstream conformance requirement that the last coded picture of the verification substream in the coded video sequence be associated with the reliable content verification SEI message. twcs_verification_substream_id indicates the verification substream to which the SEI message applies. twcv_signature_length_in_octets_minus1 plus 1 specifies the length of the syntax element twcv_signature in octets (an octet consists of 8 bits).

[0126] twcv_signature contains the digital signature of the verification substream indicated by twcs_verification_substream_id, which is sent in the trusted content selection SEI message preceding the trusted content verification SEI message of the same access unit, or is inferred to be equal to 0. If VerificationSubstreamId is the value of twcs_verification_substream_id associated with a trusted content verification SEI message, verification consists of the following ordered steps:

[0127] 1. The calculation of a message digest, called CurrDigest, is determined as follows: The concatenation of VCL NAL units for the verification substream with id equal to VerificationSubstreamId is padded according to the specifications of FIPS PUB 180-4. Note that it is sufficient to pad the last VCL NAL unit of the verification substream. The calculation of the message digest CurrDigest is determined according to the FIPS PUB 180-4 specification. The length of the message digest (in bits) is shown in Table 1. 2. The reference message digest, RefDigest, is determined as follows: If VerificationSubstreamId is greater than 0, the reference message digest RefDigest is the last calculated message digest of the verification substream with id equal to VerificationSubstreamId-1. It is a bitstream conformance requirement that any trusted content verification SEI associated with verification substream id equal to VerificationSubstreamId-1 precedes any trusted content verification SEI message with verification substream id equal to VerificationSubstreamId. Otherwise, if the current trusted content verification SEI message is the first trusted content verification SEI message in the coded video sequence with verification ID equal to 0 and the preceding coded video sequence did not contain any trusted content initialization SEI messages (this includes the case where the current coded video sequence is the first coded video sequence in the bitstream), then RefDigest is set equal to a bit string consisting of DigestSize bits equal to 1, where DigestSize is the size of the message digest as specified in Table 1. Otherwise, the reference message digest RefDigest is the last computed message digest of the validation substream with id equal to 0.

[0128] 3. The identification string IdString is constructed by concatenating the reference message digest RefDigest, the current message digest, and the binary representation of twci_hash_method_type and, if present, twci_content_uuid, as shown in FIG. The number of bits in RefDigest is determined by the value of twci_hash_method_type that was in effect when calculating the value of RefDigest, the number of bits in CurrDigest is determined by the current value of twci_hash_method_type, the value of twci_hash_method_type is represented in 8 bits, and the value of twci_content_uuid, if present, is represented in 128 bits.

[0129] 4. The identification string IdString represents the message used to verify the signature. The signature verification algorithm and the public key used for signature verification are indicated by the syntax elements twci_use_key_register_idx_flag, twci_key_source_uri, and, if twci_use_key_register_idx_flag is equal to 1, twci_key_register_idx.

[0130] NOTE 1 - Because the bitstream used for signature verification includes the RefDigest, it is possible not only to verify that the VCL NAL units used to compute the current message digest are correct, but also to further verify that no additional VCL NAL units were added to the bitstream and no VCL NAL units were removed from the bitstream. NOTE 2 - When a decoder tunes to the bitstream, it cannot correctly calculate the value of RefDigest and therefore cannot verify the IdString constructed for the first trusted content verification SEI message. However, it can verify the signature starting from the second trusted content verification SEI message. After verification, the message digest of the verification substream with id equal to VerificationSubstreamId is reinitialized according to the FIPS PUB 180-4 specification for the specified twci_hash_method_type.

[0131] Embodiments of the third aspect of the present invention rely on the discovery that a third problem is how to identify which coded pictures are used for a particular hash / digital signature value. Identifying the coded pictures used for a particular hash / digital signature can require high overhead if an indication needs to be sent for each picture. The association of NAL units to hashed / digitally signed substreams can be achieved by: Default: If no association is indicated, a preselected substream is used, for example the substream with ID 0. - Prefix SEI: A "Substream Selection" SEI containing the ID of the substream to be used is signaled for each access unit.

[0132] When a single bitstream is used, a single substream is used, and therefore no indication needs to be sent. However, when a substream is generated for each temporal layer or scalable layer, many coded pictures require substream indication. As a further embodiment, a more compact indication that does not require sending an indication for each picture can be performed by sending an indication to combine a temporal layer, a scalable layer, or a combination thereof into a substream for a group of pictures (e.g., all pictures in the CVS, or all pictures of the current picture up to the new indication).

[0133] For example, below is an idc directive: JPEG2026016324000006.jpg57170 twci_substream_selection_idc indicates how VCL NAL units are associated with substreams. Table A - Supported values ​​for twci_substream_selection_idc JPEG2026016324000007.jpg67170 For H.264 dependency_id, DQId, and priority_id, these syntax elements and variables are defined in the NAL unit header SVC extension. If the NAL unit header SVC extension is not available in the bitstream, the default substream (e.g., ID 0) is used.

[0134] Further values ​​can be used: Table A - Supported values ​​for twci_substream_selection_idc JPEG2026016324000008.jpg92170 -Temporal layer ID (temporal_id) for temporal scalability The substreams are inferred directly from the signaled temporal sublayer, e.g., the value of temporal_id is used. -Layer ID (layer_id) for spatial scalability, multiview, or SNR scalability The substreams are inferred directly from the signaled (spatial, SNR, multiview, 3D) layer, e.g., the value of layer_id is used. - Layer ID and time ID combination The substreams are inferred directly from the signaled temporal sublayers and (spatial, SNR, multiview, 3D) layers, e.g., the value N*layer_id+temporal_id is used, where N is the maximum allowed number of temporal sublayers. -H.264 / AVC Dependency ID / DQId The substream is inferred directly from the Dependency ID signaled in H.264 / AVC or the calculated DQId value. If the NAL unit header SVC extension is not available in the bitstream, a default substream (e.g., ID 0) is used. -nal_ref_idc for H.264 / AVC The substreams are inferred directly from the nal_ref_idc values ​​signaled in H.264 / AVC. -H.264 / AVC priority_id The substream is inferred directly from the priority_id value signaled in H.264 / AVC. If the NAL unit header SVC extension is not available in the bitstream, the default substream (e.g., ID 0) is used.

[0135] Alternatively, ranges can be provided for each different substream id of the mode (shown only for temporal id and layer id, the same applies for further modes and their combinations for a particular idc-s). JPEG2026016324000009.jpg111170 A similar problem arises when indicating the span of a substream in the time domain, i.e., the number of pictures used for a particular segment / chunk used to calculate the hash / digital signature. Instead of giving such an indication for each picture, in a further embodiment, the Content Initialization SEI message can indicate a mode for defining which pictures are used to calculate it. An example is given below, where POC is used for this purpose based on the previous semantics: JPEG2026016324000010.jpg58170 twci_substream_selection_idc indicates how VCL NAL units are associated with substreams. Table A - Supported values ​​for twci_substream_selection_idc If JPEG2026016324000011.jpg98170 POC mapping is used, a list is signaled in the "Initialize" SEI message. The list contains POC ranges and their association with substreams. An example syntax is shown below: JPEG2026016324000012.jpg100170 The start and end POC may be transmitted as absolute values ​​or as relative differences, for example: JPEG2026016324000013.jpg102170 An alternative signaling is as follows:

[0136] For each list entry, only the end POC is signaled. The start POC is inferred to be equal to the end POC of the previous list entry plus 1. For the first list entry, the start POC is inferred to be equal to 0. JPEG2026016324000014.jpg95170Embodiments of the fourth aspect of the present invention rely on the discovery that a further problem is that digitally signing two "segments" together to avoid removing portions of media or adding additional media leads to problems in adaptive bitrate streaming.

[0137] Adaptive bitrate streaming (e.g., DASH) is currently performed by encoding several versions of the content and letting the receiver decide, for each segment, which version to download. If the encoded media already contains a signature that spans multiple segments (for each of the versions), when the client changes from one version to another, the hash / signature across the switch for the two segments will not match the one calculated at the client.

[0138] An alternative to jointly signing hashes of segments to prevent segment removal or insertion or reordering consists of computing a hash for each segment separately and storing them externally in some metadata (e.g., as C2PA does in the C2PA manifest). These hashes are then signed within the manifest and can be used to compare the computed hashes with the corresponding values ​​in the manifest.

[0139] However, since video coding streams are not files, identifying the corresponding hashes requires some mapping. In a further embodiment, information is included in the video stream to assign a value that represents the index of the hash value of the additional metadata with which the NAL unit is associated. JPEG2026016324000015.jpg26170 In some cases, hashes for different types of media (e.g., video, audio) may be stored in the same metadata, and therefore some identifier is required to identify which hash should be used (twcs_associated_hash_group_id in the example below - e.g., if hashes for both streams are stored in the same C2PA manifest, the value for audio will be 0 and the value for video will be 1). An example is shown below. JPEG2026016324000016.jpg32170 Note that in a live scenario, it is not possible to store all the hashes as the content is encoded and transmitted and then reference them in external metadata as they are being calculated, so this alternative does not require sending the signature in the stream in the SEI message since it is stored in external metadata, and can only be done in video-on-demand.

[0140] The following describes video coding schemes in which embodiments of the present invention may be optionally implemented. In other words, the decoder 20 of Figure 14 may be optionally implemented according to any of the embodiments of the decoder 20 described below. Similarly, the device 15 may optionally be an encoder according to any of the embodiments of the encoder 10 described below.

[0141] The following description of the figures begins with the presentation of a description of an encoder and decoder of a block-based predictive codec for coding pictures of video to form an example of a coding framework into which embodiments of the present invention may be incorporated. The respective encoders and decoders are described with reference to Figures 12, 13, and 14. Below, descriptions of embodiments of the inventive concepts are presented along with an explanation of how such concepts may be incorporated into the respective encoders and decoders of Figures 12 and 13, although the embodiments described in subsequent figures and in subsequent matters may also be used to form encoders and decoders that do not operate according to the coding framework underlying the encoders and decoders of Figures 12, 13, and 14.

[0142] FIG. 12 illustrates an apparatus for predictively coding picture 12 into data stream 14, illustratively using transform-based residual coding. The apparatus or encoder is designated using reference numeral 10. FIG. 13 illustrates a corresponding decoder 20, i.e., apparatus 20 configured to predictively decode picture 12′ from data stream 14, also using transform-based residual decoding, with an apostrophe used to indicate that picture 12′ reconstructed by decoder 20 deviates from picture 12 originally encoded by apparatus 10 in terms of coding loss introduced by quantization of the predictive residual signal. While FIGS. 12 and 13 illustratively use transform-based predictive residual coding, embodiments of the present application are not limited to this type of predictive residual coding. This also applies to other details described with respect to FIGS. 12 and 13, as outlined below.

[0143] The encoder 10 is configured to perform a spatial-spectral transform of the prediction residual signal and to encode the prediction residual signal thus obtained into a data stream 14. Similarly, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and to perform a spectral-spatial transform of the prediction residual signal thus obtained. Internally, the encoder 10 may comprise a prediction residual signal former 22 that generates a prediction residual 24 to measure the deviation of a prediction signal 26 from an original signal, i.e., picture 12. The prediction residual signal former 22 may, for example, be a subtractor that subtracts the prediction signal from the original signal, i.e., from picture 12. The encoder 10 then further comprises a transformer 28 that performs a spatial-spectral transform on the prediction residual signal 24 to obtain a spectral-domain prediction residual signal 24′, which is then quantized by a quantizer 32 also included in the encoder 10. The prediction residual signal 24″ thus quantized is coded into the bitstream 14. For this purpose, the encoder 10 may optionally comprise an entropy coder 34 that entropy codes the prediction residual signal that has been transformed and quantized into the data stream 14. The prediction signal 26 is generated by a prediction stage 36 of the encoder 10 based on the prediction residual signal 24″ that has been coded into the data stream 14 and that can be decoded therefrom. For this purpose, as shown in FIG. 12 , the prediction stage 36 may internally comprise an inverse quantizer 38 that inversely quantizes the prediction residual signal 24″ to obtain a spectral-domain prediction residual signal 24″ that corresponds to the signal 24′ without the quantization losses, followed by an inverse transformer 40 that subjects the latter prediction residual signal 24′′ to an inverse transform, i.e., a spectral-spatial transform, to obtain a prediction residual signal 24″″ that corresponds to the original prediction residual signal 24 without the quantization losses. A combiner 42 of the prediction stage 36 then recombines the prediction signal 26 and the prediction residual signal 24″″, for example by addition, to obtain a reconstructed signal 46, i.e., a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to the signal 12′. A prediction module 44 of the prediction stage 36 then generates a prediction signal 26 based on the signal 46, for example by using spatial prediction, i.e., intra-picture prediction, and / or temporal prediction, i.e., inter-picture prediction.

[0144] Similarly, decoder 20 may be internally configured from components corresponding to prediction stage 36 and interconnected in a manner corresponding to prediction stage 36, as shown in Figure 13. In particular, entropy decoder 50 of decoder 20 may entropy decode quantized spectral domain prediction residual signal 24" from the data stream, with inverse quantizer 52, inverse transformer 54, combiner 56, and prediction module 58 interconnected and cooperating in the manner described above with respect to the modules of prediction stage 36 to recover a reconstructed signal based on prediction residual signal 24" such that the output of combiner 56 provides a reconstructed signal, i.e., picture 12', as shown in Figure 13.

[0145] Although not specifically described above, it is readily apparent that the encoder 10 can set several coding parameters, including, for example, prediction modes, motion parameters, etc., according to several optimization schemes, such as, for example, schemes that optimize several rate- and distortion-related criteria, i.e., coding cost. For example, the encoder 10 and the decoder 20 and corresponding modules 44, 58 can each support different prediction modes, such as intra-coded and inter-coded modes. The granularity at which the encoder and decoder switch between these prediction mode types may correspond to the subdivision of the pictures 12 and 12′, respectively, into coded segments or coded blocks. At these coded segment levels, for example, the picture may be subdivided into intra-coded blocks and inter-coded blocks. The intra-coded blocks are predicted based on their spatial already coded / decoded neighbors, as outlined in more detail below. Several intra-coding modes, including directional intra-coding modes or angular intra-coding modes, may be selected for each intra-coded segment, and each segment is filled by extrapolating neighboring sample values ​​along a specific direction specific to the directional intra-coding mode. The intra-coding modes may also include one or more additional modes, such as a DC coding mode in which prediction of each intra-coded block assigns a DC value to all samples within the respective intra-coded segment, and / or a planar intra-coding mode in which prediction of each block is approximated or determined based on neighboring samples as a spatial distribution of sample values ​​described by a two-dimensional linear function over the sample positions of the respective intra-coded block with a planar driving slope and offset defined by the two-dimensional linear function. In contrast, inter-coded blocks may be predicted, for example, temporally.For inter-coded blocks, motion vectors may be signaled within the data stream, indicating the spatial displacement of portions of previously coded pictures of the video to which picture 12 belongs, and the previously coded / decoded pictures are sampled to obtain a prediction signal for the respective inter-coded blocks. This means that in addition to the residual signal coding included in data stream 14, such as entropy-coded transform coefficient levels representing the quantized spectral domain prediction residual signal 24″, data stream 14 may also encode optional further parameters, such as coding mode parameters for assigning coding modes to various blocks, motion parameters for inter-coded segments, and some prediction parameters for the blocks, as well as parameters for controlling and signaling the subdivision of pictures 12 and 12′ into their respective segments. Decoder 20 uses these parameters to subdivide the pictures in the same manner as did the encoder, assign the same prediction modes to the segments, and perform the same predictions, resulting in the same prediction signals.

[0146] 14 shows the relationship between, on the one hand, the reconstructed signal, i.e., the reconstructed picture 12′, and, on the other hand, the combination of a prediction residual signal 24″″ and a prediction signal 26 signaled in the data stream 14. As already mentioned above, the combination may be additive. In FIG. 14, the prediction signal 26 is shown as a subdivision of the picture area into intra-coded blocks, exemplarily shown with hatching, and inter-coded blocks, exemplarily shown without hatching. The subdivision may be any subdivision, such as a regular subdivision of the picture area into rows and columns of square or non-square blocks, or a multi-tree subdivision of the picture 12 from a tree root block into a number of leaf blocks of various sizes, such as a quadtree subdivision, a mixture of which is shown in FIG. 14, in which the picture area is first subdivided into rows and columns of tree root blocks, which are then further subdivided into one or more leaf blocks according to a recursive multi-tree subdivision.

[0147] Again, data stream 14 may have an intra-coding mode coded for intra-coded blocks 80, which assigns one of several supported intra-coding modes to each intra-coded block 80. For inter-coded blocks 82, data stream 14 may have one or more motion parameters coded therein. Generally speaking, inter-coded blocks 82 are not limited to being temporally coded. Alternatively, inter-coded blocks 82 may be any blocks predicted from a previously coded portion beyond current picture 12 itself, such as a previously coded picture of the video to which picture 12 belongs, or, if the encoder and decoder are scalable encoder and decoder, respectively, a picture of another view or a hierarchically lower layer.

[0148] The prediction residual signal 24"" in Figure 14 is also shown as a subdivision of the picture region into blocks 84. These blocks are sometimes called transform blocks to distinguish them from the coded blocks 80 and 82. In fact, Figure 14 shows that the encoder 10 and the decoder 20 may use two different subdivisions of the picture 12 and the picture 12' into blocks, respectively: one subdivision into coded blocks 80 and 82, and the other subdivision into transform blocks 84. While both subdivisions may be the same, i.e., each coded block 80 and 82 may simultaneously form a transform block 84, Figure 14 also shows the case where, for example, the subdivision into transform blocks 84 forms an extension of the subdivision into coded blocks 80, 82, such that any boundary between two of the blocks 80 and 82 covers the boundary between the two blocks 84, or each block 80, 82 coincides with one of the transform blocks 84 or with a cluster of transform blocks 84. However, these subdivisions may also be determined or selected independently of one another, such that the transformation blocks 84 may alternatively cross the block boundaries between the blocks 80, 82. Thus, as far as the subdivision into transformation blocks 84 is concerned, similar statements are true as those presented with regard to the subdivision into blocks 80, 82, i.e., the blocks 84 may be the result of a regular subdivision of the picture area into blocks (with or without arrangement into rows and columns), a recursive multi-tree subdivision of the picture area, or a combination thereof, or any other type of blocking. It should be noted that the blocks 80, 82, and 84 are not limited to being square, rectangular, or any other shape.

[0149] FIG. 14 further illustrates that the combination of the prediction signal 26 and the prediction residual signal 24"" directly results in the reconstructed signal 12'. However, it should be noted that, according to alternative embodiments, multiple prediction signals 26 can be combined with the prediction residual signals 24"" into the picture 12'.

[0150] In Figure 14, the transform blocks 84 have the following meaning: The transformer 28 and the inverse transformer 54 perform transforms in units of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow for skipping the transform for some of the transform blocks 84 so that the prediction residual signal is directly coded in the spatial domain. However, according to the embodiments described below, the encoder 10 and the decoder 20 are configured to support several transforms. For example, the transforms supported by the encoder 10 and the decoder 20 may include: DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform. DST-IV, where DST stands for discrete sine transform DCT-IV DST-VII Identity Transformation (IT) Of course, the transformer 28 supports all of the forward transform versions of these transforms, while the decoder 20 or inverse transformer 54 supports the corresponding backward or inverse transform versions.

[0151] Inverse DCT-II (or Inverse DCT-III) Reverse DST-IV ·Inverse DCT-IV Reverse DST-VII Identity Transformation (IT) The following description provides further details about which transforms may be supported by the encoder 10 and decoder 20. Note that in any case, the set of supported transforms may include only one transform, such as one spectral-to-spatial transform or one spatial-to-spectral transform.

[0152] As already outlined above, Figures 12, 13, and 14 are presented as examples in which the inventive concepts further described below can be implemented to form specific examples of encoders and decoders according to the present application. To that extent, the encoders and decoders of Figures 12 and 13 may represent possible implementations of the encoders and decoders described later in this specification. However, Figures 12 and 13 are merely examples. However, an encoder according to an embodiment of the present application may perform block-based encoding of picture 12 using concepts outlined in more detail below, and differs from the encoder of Figure 12 in, for example, being a still picture encoder rather than a video encoder, not supporting inter-prediction, or in that the subdivision into blocks 80 is performed in a different manner than illustrated in Figure 14. Similarly, a decoder according to an embodiment of the present application may perform block-based decoding of picture 12′ from data stream 14 using the coding concepts further outlined below, but may differ from decoder 20 of FIG. 13, for example, in that it is a still picture decoder rather than a video decoder, in that it does not support intra-prediction or in that it subdivides picture 12′ into blocks in a different manner than described with respect to FIG. 14, and / or in that it does not derive prediction residuals from data stream 14 in the transform domain but does, for example, in the spatial domain.

[0153] Different examples for coding the residual blocks representing spatial residual blocks in the transform domain and their transform blocks, respectively, are given below. Although a codec may support only one of them, the video data stream may also include an entropy coding mode indicator that indicates whether the prediction residual data of the residual blocks should be decoded from the video data stream using a context-adaptive variable length coding mode or a context-adaptive binary arithmetic coding mode, examples of which can be derived from the following description.

[0154] Context-based adaptive variable length coding (CAVLC) This is the method used to code residual zigzag-ordered 4x4 (and 2x2) blocks of transform coefficients. CAVLC is designed to take advantage of several properties of quantized 4x4 blocks. 1. After prediction, transformation, and quantization, blocks are usually sparse (contain mostly zeros). CAVLC uses run-level coding to represent zero strings compactly. 2. The highest non-zero coefficient after zigzag scanning is often a sequence of + / - 1. CAVLC compactly signals the number of frequent + / - 1 coefficients ("trailing 1" or "T1"). 3. The number of non-zero coefficients in adjacent blocks is correlated. The number of coefficients is coded using a look-up table. The choice of look-up table depends on the number of non-zero coefficients in the adjacent blocks. 4. The levels (magnitudes) of non-zero coefficients tend to be high at the beginning of the sorted sequence (near the DC coefficient) and decrease towards higher frequencies. CAVLC exploits this by adapting the VLC lookup table selection of the "level" parameter depending on the magnitude of the most recently coded level.

[0155] CAVLC coding of a block of transform coefficients proceeds as follows. 1. Encode the number of coefficients and trailing ones (coeff_token). The first VLC, coeff_token, encodes both the total number of non-zero coefficients (TotalCoeffs) and the number of trailing + / - 1 values ​​(T1). TotalCoeffs is 0 (no coefficients in a 4x4 block). 1T1 can be anything from 0 to 16 (16 non-zero coefficients). T1 can be anything from 0 to 3. If there are more than three trailing + / - 1's, only the last three are treated as a "special case" and the others are coded as normal coefficients. Note: The coded_block_pattern (described above) indicates which 8x8 blocks within a macroblock contain non-zero coefficients. However, within a coded 8x8 block, there may be 4x4 sub-blocks that do not contain any coefficients, and therefore TotalCoeff can be 0 in any 4x4 sub-block. In practice, this value of TotalCoeff occurs most frequently and is assigned the shortest VLC. There are four choices of lookup tables to use to encode the coeff_token, described as Num-VLC0, Num-VLC1, Num-VLC2, and Num-FLC (three variable length code tables and a fixed length code). The choice of table depends on the number of non-zero coefficients in the above and left blocks Nu and NL that were previously coded. The parameter N is calculated as follows:

[0156] If blocks U and L are available (i.e., in the same coding slice), then N=(Nu+NL) / 2; if only block U is available, then N=NU; if only block L is available, then N=NL; and if neither is available, then N=0. N selects a lookup table (Table 34), and in this way the VLC selection adapts depending on the number of coded coefficients in the neighboring block (context adaptation). Num-VLC0 is "biased" towards small numbers of coefficients. Low values ​​of TotalCoeff (0 and 1) are assigned particularly short codes, and high values ​​of TotalCoeff are assigned particularly long codes. Num-VLC1 is biased towards medium numbers of coefficients (relatively short codes are assigned to TotalCoeff values ​​around 2-4), Num-VLC2 is biased towards higher numbers of coefficients, and FLC assigns a fixed 6-bit code to all values ​​of TotalCoeff. Table 34. Coeff_token lookup table selection JPEG2026016324000017.jpg29112

[0157] 2. Encode the code of each T1. For each T1 (trailing + / -1) signaled by the coeff_token, a single bit encodes the sign (0=+, 1=-). These are coded in reverse order, starting with the most frequent T1. 3. Code the levels of the remaining non-zero coefficients. The level (sign and magnitude) of each remaining non-zero coefficient of the block is coded in reverse order, starting with the most frequent and working towards the DC coefficient. The selection of the VLC table to code each level adapts according to the magnitude of each successive coding level (context adaptation). There are seven VLC tables to choose from: Level_VLC0 through Level_VLC6. Level_VLC0 is biased towards lower magnitudes. Level_VLC1 is biased towards slightly higher magnitudes, etc. The table selection is adapted in the following way:

[0158] (a) Initialize the table to Level_VLC0 (start at Level_VLC1 unless there are more than 10 non-zero coefficients and the trailing 1 is less than 3). (b) Encode the most frequent non-zero coefficients. (c) If the magnitude of this coefficient is greater than a predetermined threshold, move to the next VLC table. In this way, the level selection corresponds to the magnitude of the most recently coded coefficient. The thresholds are listed in Table 35. The first threshold is 0, which means that the table is always incremented after a coefficient level of the first mode is coded. Table 35. Thresholds for determining whether to increment the level table number JPEG2026016324000018.jpg46126

[0159] 4. Code the total number of zeros before the last coefficient. TotalZeros is the sum of all zeros preceding the highest non-zero coefficient in the sorted array. This is coded in the VLC. The reason for sending a separate VLC t to indicate TotalZeros is that many blocks contain some non-zero coefficients at the beginning of the array, and (as we'll see below) this approach means that the run of zeros at the beginning of the array does not need to be coded.

[0160] 5. Code each run of zeros. The number of zeros preceding each non-zero coefficient (run_before) is coded in reverse order: the run_before parameter is coded for each non-zero coefficient, starting with the most frequent, with two exceptions: (a) If there are no more zeros left to encode (ie, Σ[run_before]=TotalZeros), then there is no need to encode any more run_before values. (b) There is no need to code the run_before of the last (least frequent) non-zero coefficient. The VLC for each run of zeros is selected depending on (a) the number of zeros left to encode (ZerosLeft) and (b) run_before. For example, if there are only two zeros left to encode, run_before can only take three values ​​(0, 1, or 2), and therefore the VLC need not be more than two bits long. If there are six zeros left to encode, run_before can take seven values ​​(0 through 6), and the VLC table must be increased accordingly.

[0161] CAVLC Example In all the following examples, we assume that the table Num-VLC0 is used to encode the coff_token. Example 1 4x4 Block: JPEG2026016324000019.jpg2459 Reordered blocks: 0,3,0,1,-1,-1,0,1,0... TotalCoeff=5 (indexed from highest frequency [4] to lowest frequency [0]) TotalZeros=3 T1s=3 (actually there are 4 trailing ones, but only 3 can be coded as a "special case") Encoding: JPEG2026016324000020.jpg78165 The transmitted bitstream for this block is 000010001110010111101101.

[0162] Decryption: The output array is "built" from the decoded values ​​as shown below, where the values ​​added to the output array at each stage are shown in brackets. JPEG2026016324000021.jpg77170 The decoder inserted two zeros, however TotalZeros equals 3, so another zero is inserted before the lowest coefficient to form the final output array.

[0163] [0],3,0,1,-1,-1,0,1 Example 2 4x4 Block: JPEG2026016324000022.jpg2353 Reordered blocks: -2,4,3,-3,0,0,-1,... TotalCoeff=5 (indexed from highest frequency [4] to lowest frequency [0]) TotalZeros=2 T1s=1 Encoding: JPEG2026016324000023.jpg62168 The transmitted bitstream for this block is 000000011010001001000010111001100.

[0164] NOTE 1: Level (3) has a value of -3 and is coded as a special case. If there are less than three T1s, the first non-T1 level does not have a value of + / -1 (it would otherwise have been coded as T1). To save bits, this level is incremented if negative (decremented if positive), + / -2 is mapped to + / -1, + / -3 is mapped to + / -2, etc. In this way, a shorter VLC is used. Note 2: After encoding level (3), the level_VLC table is incremented because the magnitude of this level is greater than the first threshold (0). After encoding level (1), with a magnitude of 4, the table number is incremented again because level (1) is greater than the second threshold (which is 3). Note that the final level (-2) uses a different code than the first encoding level (also -2).

[0165] Decryption: JPEG2026016324000024.jpg78169 Now all zeros have been decoded, so the output array is: -2,4,3,-3,0,0,-1 (This example shows how bits are saved by encoding Total Zeros; even though there are five non-zero coefficients, only a single run needs to be coded.)

[0166] CABAC In CABAC, encoding and decoding can be done as follows: 1) The position of the first non-zero transform coefficient encountered when traversing the transform coefficients along a predetermined scan order is coded / decoded. 2) Following the scan order, the transform coefficients, including the non-zero transform coefficients at the coding / decoding positions, are coded / decoded from the data stream. In CABAC, encoding and decoding can alternatively be done as follows: 1) coding / decoding a significance map indicating the positions of non-zero transform coefficients by using significance flags and last significance flags; in a forward scan traversing the positions of the transform coefficients, coding / decoding significance flags indicating whether a non-zero transform coefficient is located at each position; if so, and if the position is not the end of the forward scan, coding / decoding a last significance flag indicating whether the non-zero transform coefficient located at each position is the last non-zero transform coefficient in the forward scan order; 2) Non-zero transform coefficient values ​​are coded / decoded sequentially backwards while reversing the forward scan order.

[0167] It should be noted that any of the embodiments described with reference to Figures 1 to 11 may be combined with any of the embodiments described with reference to Figures 12 to 14. In other words, the implementation of the video codec used by the encoder or decoder may be independent from the implementation of the reliability check of the video data stream / rendering of the reliability-checkable video data stream. While the description of Figures 1-14 is of an apparatus, the block diagrams of Figures 1-14 can alternatively be considered as flow diagrams of respective methods, with each block representing a step of the respective method. Accordingly, further disclosed in the above description are: A method 16 for checking a video data stream 14 in which video is encoded for authenticity, the method including: applying (31) a predetermined portion 13 of the video data stream, or data 62 derived therefrom, to a hash function 31 to obtain a hash value 33; obtaining (51) a unique identifier 45 (e.g., from the video data stream or a reference, e.g., using a URI) that uniquely identifies the media asset to which the predetermined portion 13 belongs; obtaining (51) a digital signature 43 based on the video data stream (e.g., from the video data stream, e.g., from a twcv_signature, or from a metadata file or manifest file, e.g., a C2PA file); and checking (41) whether a combination of the hash value 33 and the unique identifier 45 (e.g., a combination of multiple information including the hash value and the unique identifier 45) matches the digital signature 43 to determine whether the video data stream is authentic.

[0168] A method (20) for decoding a video data stream in which video is encoded, the method comprising: decoding (21) a syntax structure from the video data stream, the syntax structure containing information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream, the predetermined portion being subjected to a hash function (31) or used to derive data to be subjected to the hash function (31), the hash function (31) being for deriving a hash value (33) that serves to check the authenticity of the video data stream; from the video data stream (13), decoding (21) a unique identifier 45 or a reference pointing to the unique identifier 45, where the unique identifier 45 uniquely identifies the media asset to which the given portion 13 belongs; and decoding (21) an indication of a digital signature 43 from the video data stream (e.g., from the video data stream, e.g., from a twcv_signature, or from a metadata file or manifest file, e.g., a C2PA file), where the digital signature 43 is based on a combination of the hash value 33 and the unique identifier 45 (e.g., a combination of multiple information including the hash value and the unique identifier 45).

[0169] A method 15 for rendering a video data stream 14 in which the video is encoded in a manner that can be checked for authenticity, the method including: applying (31) a predetermined portion 13 of the video data stream 14, or data 62 from which the video data stream 14 is derived, to a hash function 31 to obtain a hash value 33; assigning to the predetermined portion 13 a unique identifier 45 that uniquely identifies the media asset to which the predetermined portion 13 belongs; and signing (71) a combination of the hash value 33 and the unique identifier 45 (e.g., a combination of multiple information including the hash value 33 and the unique identifier 45) to obtain a digital signature 43.

[0170] A method 16 for checking a video data stream 14 in which video is encoded for authenticity, the method determining whether the video data stream is authentic by applying (31) a predetermined portion 13 of the video data stream, or data 62 derived therefrom, to a hash function 31 to obtain a hash value 33, decrypting (46) a digital signature 43 using a public key 57 of an asymmetric decryption scheme to obtain a check value 47, and checking (49) whether the hash value 33 matches a digital signature 43 (e.g., derived from the video data stream and derived from criteria indicated in the video data stream), and if the video data stream includes an indication of an external resource 280 (e.g., a metadata structure, e.g., a manifest file) that includes an editor's track 231 of the video data stream (e.g., a track or record of edits or modifications from the generation of the video data stream to the current version, and / or the identity of the corresponding editor), querying (or retrieving) the editor's track for a certificate 233 of a content provider who is the last editor of the video data stream and deriving a public key based on the content provider's certificate.

[0171] A method 17 for transcoding a video data stream in which video is encoded, comprising receiving an input video data stream (14'), checking (15') the authenticity of the input video data stream (14'), transcoding (12) the input video data stream (14') to generate an output data stream (14), subjecting (31) a predetermined portion 13 of the output video data stream 14, or data 62 from which the output video data stream is derived, to a hash function 31 to obtain a hash value 33, signing (71) the hash value using a private key 58 of an asymmetric cryptography scheme to obtain a digital signature 43, and a certificate 233 of the content provider (e.g. identifying the device) to an external resource 280 (e.g. a track or record of edits or modifications to the current version, and / or the identity of the corresponding editor), the editor's track being provided in a metadata structure, e.g. a manifest file, at the external resource, the certificate including or pointing to a public key 57 of an asymmetric cryptography scheme, and providing a digital signature 43 to the output video data stream 14 (e.g. in an SEI message) or the external resource 280 (e.g. inserting the digital signature in the metadata structure or a further metadata structure and providing the same to the external resource) or the further external resource.

[0172] 1. A method 15 for rendering a video data stream 14 in which the video is encoded in a way that can be checked for authenticity, the method comprising: applying a predetermined portion 13 of the video data stream 14, or of a predetermined portion 13 of data 62 from which the video data stream 14 is derived, to a hash function 31 to obtain a hash value 33; signing the hash value using a private key 58 of an asymmetric cryptography scheme to obtain a digital signature 43; providing an editor's track 231 of the video data stream (e.g. a track or record of edits or modifications of the video data stream from its generation to its current version, and / or the identity of the corresponding editor) to an external resource 280; the editor's track (e.g. a metadata structure, e.g. a manifest file) at the external resource; a certificate 233 of the content provider (e.g. identifying the device), the certificate including or pointing to the public key of the asymmetric cryptography scheme; and providing the digital signature 43 to the output video data stream (e.g. in an SEI message) or to the external resource 280 (e.g. inserting the digital signature in a metadata structure or a further metadata structure and providing the same to the external resource) or the further external resource.

[0173] 1. A method for checking a video data stream 14 in which video is encoded for authenticity, the method comprising: applying a predetermined portion 13 of the video data stream, or data derived therefrom, to a hash function 31 to obtain a hash value 33; and checking whether the hash value 33 matches a digital signature 43 (e.g., derived from the video data stream and derived from a reference indicated in the video data stream) to determine whether the video data stream is authentic; the method further comprising: determining whether the video data stream is authentic by determining whether the video data stream is authentic; a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; a plurality of layer identifiers (e.g., layer_id in HEVC / VVC, dependency_id and / or quality_id in AVC) that identify the layer of the video data stream to which the respective picture belongs (e.g., the video data stream is a layered video data stream including multiple layers, such as, for example, a base layer and one or more enhancement layers, that represent video at different resolutions or from different viewpoints), and determining the predetermined portion 13 based on one or more layer identifiers, a combination of a temporal layer identifier and a layer identifier, a temporal frame identifier (e.g., picture order count, POC), a priority level identifier indicating a priority level of the picture (e.g., AVC priority_id), and a nal_ref_id in AVC.

[0174] A method 16 for checking a video data stream 14 in which video is encoded for authenticity, the method comprising: applying a predetermined portion 13 of the video data stream, or of data derived therefrom, to a hash function 31 to obtain a hash value 33; checking whether the hash value 33 matches a digital signature 43 (e.g., derived from the video data stream and derived from criteria indicated in the video data stream) to determine whether the video data stream is authentic; and deriving an indication from the video data stream, the indication indicating how the predetermined portion 13 is determined.

[0175] A method 20 for decoding a video data stream 14 in which video is encoded, the method including deriving from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion 13 of the video data stream (e.g., which predetermined portion 13 should be subjected to a hash function or used to derive data to be subjected to a hash function to derive a hash value 33 that functions to check the authenticity of the video data stream), the syntax structure including instructions indicating how to determine the predetermined portion 13 of the video data stream.

[0176] 1. A method 15 for rendering a video data stream 14 in which the video has been encoded in a way that can be checked for authenticity, the method comprising: subjecting a predetermined portion 13 of the video data stream, or data 62 from which a further portion of the video data stream is derived, to a hash function 31 to obtain a hash value 33; and signing the hash value 33 (e.g., by using a private key in an asymmetric cryptography scheme) to obtain a digital signature 43; the method comprising: a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; one or more layer identifiers associated with pictures of the video data stream (e.g., temporal_id); determining the predetermined portion 13 based on one or more of: a layer_id in HEVC / VVC, a dependency_id and / or a quality_id in AVC), identifying the layer of the video data stream to which the respective picture belongs (e.g., the video data stream is a layered video data stream including multiple layers, such as a base layer and one or more enhancement layers, which represent video at different resolutions or from different viewpoints), one or more layer identifiers, a combination of a temporal layer identifier and a layer identifier, a temporal frame identifier (e.g., a picture order count, POC), a priority level identifier indicating a priority level of the picture (e.g., a priority_id in AVC), a nal_ref_id in AVC.

[0177] A method 15 for rendering a video data stream 14 in which the video is encoded in a way that can be checked for authenticity, the method comprising: subjecting a predetermined portion 13 of the video data stream, or data 62 from which a further portion of the video data stream is derived, to a hash function 31 to obtain a hash value 33; signing the hash value 33 (e.g., by using a private key in an asymmetric encryption scheme) to obtain a digital signature 43; and inserting instructions into the video data stream, the instructions indicating a manner in which the predetermined portion 13 is determined.

[0178] A method 16 for checking a video data stream 14 in which video is encoded for authenticity, the method comprising: applying a predetermined portion 13 of the video data stream (e.g., video data associated with an access unit, e.g., a time frame) or data derived therefrom to a hash function 31 to obtain a hash value 33; deriving a digital signature 43 associated with the predetermined portion 13 from an external resource (e.g., a server); and checking whether the hash value 33 matches the digital signature 43 to determine whether the video data stream is authentic.

[0179] A method 16 for checking a video data stream 14 in which video is encoded for authenticity, the method comprising: applying a predetermined portion 13 of the video data stream (e.g., video data associated with an access unit, e.g., a time frame) or data derived therefrom to a hash function 31 to obtain a hash value 33; deriving a check value associated with the predetermined portion 13 from an external resource (e.g., a server) (e.g., a signed check value signed with a private key of an asymmetric cryptography scheme); and checking whether the hash value 33 matches the check value to determine whether the video data stream is authentic.

[0180] A method 20 for decoding a video data stream 14 in which video is encoded, the method comprising deriving from the video data stream a syntax structure containing information for checking the authenticity of the video data stream based on a predetermined portion 13 of the video data stream (e.g., to be subjected to a hash function or used to derive data to be subjected to a hash function to derive a hash value 33 that functions to check the authenticity of the video data stream), the syntax structure including a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a digital signature 43 associated with the predetermined portion 13.

[0181] A method 20 for decoding a video data stream 14 in which video is encoded, the method being configured to derive from the video data stream a syntax structure containing information for checking the authenticity of the video data stream based on a predetermined portion 13 of the video data stream (e.g., to be subjected to a hash function or used to derive data to be subjected to a hash function to derive a hash value 33 that functions to check the authenticity of the video data stream), the syntax structure including a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a check value associated with the predetermined portion 13 (e.g., a check value signed with a digital signature).

[0182] A method 15 for rendering a video data stream 14 in which the video is encoded in a way that can be checked for authenticity, the method comprising: applying a predetermined portion 13 of the video data stream, or data 62 from which the predetermined portion 13 of the video data stream is derived, to a hash function 31 to obtain a hash value 33; signing the hash value 33 (e.g., by using a private key in an asymmetric cryptography scheme) to obtain a digital signature 43; providing the digital signature 43 to an external resource; and inserting an indication of the external resource (e.g., a reference to the digital signature of the external resource) into the video data stream (e.g., a URI of the external resource or the digital signature).

[0183] A method 15 for rendering a video data stream 14 in which the video is encoded in a way that can be checked for authenticity, the method comprising: subjecting a predetermined portion 13 of the video data stream, or data 62 from which a further portion of the video data stream is derived, to a hash function 31 to obtain a hash value 33; signing the hash value 33 (e.g., by using a private key in an asymmetric cryptography scheme) to obtain a digital signature 43; providing the hash value 33 and the digital signature 43 to an external resource; and inserting an indication of the external resource (e.g., a reference to the external resource's digital signature) into the video data stream (e.g., a URI of the external resource or the digital signature).

[0184] Further embodiments It should be noted that the text in brackets is not intended to be a required part of the embodiment, but rather provides explanations, examples, or optional features that may optionally be integrated into the embodiment. The different aspects may be combined, i.e. any feature defined with respect to any of the aspects may be combined with any of the further aspects.

[0185] Embodiments of the First Aspect 1. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, the device comprising: applying a predetermined portion (13) of the video data stream or of data derived therefrom (62) to a hash function (31) to obtain a hash value (33); obtaining a unique identifier (45) [e.g., from a video data stream, or a reference, e.g., using a URI] that uniquely identifies the media asset to which the given portion (13) belongs; Obtaining a digital signature (43) based on the video data stream (e.g., from the video data stream, e.g., from twcv_signature, or from a metadata file or manifest file, e.g., C2PA file); An apparatus configured to check (41) whether a combination of a hash value (33) and a unique identifier (45) (e.g., a combination of multiple pieces of information including the hash value and the unique identifier (45)) matches a digital signature (43) to determine whether the video data stream is authentic.

[0186] 2. When checking whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43), decrypting the digital signature (43) (e.g., by using an asymmetric decryption method with a public key) to obtain a check value; 2. The apparatus of embodiment 1, configured to check whether a combination of the hash value (33) and the unique identifier (45) matches a check value. 3. When checking whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43), forming a verification string based on the hash value (33) and the unique identifier (45); 3. The apparatus of any of embodiments 1 or 2, configured to compare the verification string with the digital signature (43) using the public key (e.g., using a verification algorithm).

[0187] 4. An apparatus as described in any one of embodiments 1 to 3, configured to derive an indication of an external resource for obtaining a public key from the video data stream and to derive the public key from the external resource (e.g., a URI). 5. The apparatus of embodiment 4, configured to derive the unique identifier (45) from an external resource. 6. An apparatus as described in any one of embodiments 1 to 5, configured to check whether the unique identifier (45) matches a unique identifier (45) associated with one or more further media components (e.g., audio or subtitles) (e.g., additional media components signaled in a data stream constituting a video data stream).

[0188] 7. An apparatus as described in any one of embodiments 1 to 6, configured to derive a digital signature (43) from a video data stream (e.g., from payload packets interspersed in the video data stream between video payload packets carrying encoded video data) [e.g., from a supplemental enhancement information (SEI) message of the video data stream such as trustworthy_content_verification]. 8. Derive an indication of an external resource [e.g., a URI] from the video data stream [e.g., from a video data stream SEI message such as trustworthy_content_verification]; 8. An apparatus as recited in any preceding embodiment, configured to derive a digital signature (43) from an external resource.

[0189] 9. The apparatus of embodiment 8, wherein the indication of the external resource is a uniform resource identifier (URI) that points to a manifest file stored on a server. 10. An apparatus according to any one of embodiments 1 to 9, configured to derive a unique identifier (45) from a video data stream. 11. An apparatus according to any one of embodiments 1 to 10, configured to derive a unique identifier (45) from a payload packet signaled in a video data stream. 12. Payload packets are Instructions for hash functions (31), an indication of some portions of the video data stream for which digital signatures (43) are available for verifying the authenticity of the video data stream; and instructions indicating how to obtain a public key to check whether a combination of the hash value (33) and the unique identifier (45) matches the digital signature (43).

[0190] 13. When applying a predetermined portion (13) of a video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33), reconstructing the video with respect to the predetermined portion (13) to obtain a reconstructed portion of the video; 13. An apparatus as described in any of embodiments 1 to 12, configured to subject the reconstructed portion to a hash function (31). 14. The device Decode video from the video data stream; 14. An apparatus as in any preceding embodiment, wherein the apparatus is a decoder configured to decode an indication of a digital signature (43) from a video data stream.

[0191] 15. The apparatus of embodiment 14, configured to decode a digital signature (43) from a supplemental enhancement information message of a video data stream. 16. Performing a check of the authenticity of the video data stream sequentially on a plurality of portions of the video data stream, and obtaining a hash value (33) by applying a hash function (31) to a predetermined portion (13) of the video data stream or data derived therefrom; obtaining a further hash value by subjecting a further portion of the video data stream, or data derived therefrom, to a hash function (31) (e.g., the further portion is a portion preceding the given portion (13)); 16. The apparatus of any one of embodiments 1 to 15, configured to check whether a combination of the hash value (33), the further hash value, and the unique identifier (45) matches the digital signature (43) (e.g., the combination of the multiple pieces of information includes the hash value, the further hash value, and the unique identifier (45)).

[0192] 17. An apparatus according to any one of embodiments 1 to 16, configured to check whether a combination of a plurality of pieces of information including a hash value (33), a unique identifier (45), and an indication of a hash function (31) matches a digital signature (43) to determine whether a video data stream is trustworthy. 18. An apparatus (20) for decoding a video data stream in which video is encoded, the apparatus comprising: decoding a syntax structure (55) from the video data stream, the syntax structure (55) containing information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream, the predetermined portion being subjected to a hash function (31) or used to derive data to be subjected to the hash function (31), deriving a hash value (33) that serves to check the authenticity of the video data stream; decoding a unique identifier (45) or a reference pointing to the unique identifier (45) from the video data stream, the unique identifier (45) uniquely identifying the media asset to which the given portion (13) belongs; An apparatus configured to decode instructions for a digital signature (43) from a video data stream (e.g., from a video data stream, e.g., twcv_signature, or from a metadata file or manifest file, e.g., a C2PA file), the digital signature (43) being based on a combination of a hash value (33) and a unique identifier (45) (e.g., a combination of multiple pieces of information including the hash value and the unique identifier (45)).

[0193] 19. The apparatus of embodiment 18, wherein the unique identifier (45) is signaled in a syntax structure. 20. A device (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, the device comprising: applying a predetermined portion (13) of the video data stream (14) or of the data (62) from which the video data stream (14) is derived to a hash function (31) to obtain a hash value (33); assigning a unique identifier (45) to the given portion (13) that uniquely identifies the media asset to which the given portion (13) belongs; The device is configured to sign a combination of the hash value (33) and the unique identifier (45) (e.g., a combination of multiple pieces of information including the hash value (33) and the unique identifier (45)) to obtain a digital signature (43).

[0194] 21. Form a verification string based on the hash value (33) and the unique identifier (45); 21. The apparatus of embodiment 20, configured to sign the verification string using a private key (e.g., using a signature algorithm) to obtain a digital signature (43). 22. An apparatus as described in either embodiment 20 or 21, configured to provide, in the video data stream, an indication of an external resource that holds or indicates a private key. 23. The device of embodiment 22, configured to provide a unique identifier (45) in an external resource. 24. An apparatus as described in any of embodiments 20 to 23, configured to provide a digital signature (43) in a video data stream [e.g., from payload packets interspersed in the video data stream between video payload packets carrying encoded video data] [e.g., from a supplemental enhancement information (SEI) message of the video data stream such as trustworthy_content_verification].

[0195] 25. An apparatus as described in any of embodiments 20 to 23, configured to provide an indication of an external resource (e.g., a URI) in the video data stream (e.g., from an SEI message of the video data stream such as trustworthy_content_verification) and to provide a digital signature (43) at the external resource. 26. The apparatus of embodiment 25, wherein the indication of the external resource is a uniform resource identifier (URI) that points to a manifest file stored on a server. 27. An apparatus according to any one of embodiments 20 to 26, configured to insert a unique identifier (45) into a video data stream. 28. An apparatus according to any one of embodiments 20 to 27, configured to insert a unique identifier (45) into a payload packet signaled in a video data stream, for example, an SEI message.

[0196] 29. SEI message is, Instructions for hash functions (31), an indication of some portions of the video data stream for which digital signatures (43) are available for verifying the authenticity of the video data stream; and instructions indicating how to obtain a public key to check whether a combination of the hash value (33) and the unique identifier (45) matches the digital signature (43). 30. The device is Encoding the video into a video data stream; 30. An apparatus as described in any one of embodiments 20 to 29, which is an encoder configured to encode an indication of the digital signature (43) into a video data stream.

[0197] 31. The apparatus of embodiment 30, configured to encode a digital signature (43) into a supplemental enhancement information message of a video data stream. 32. Performing a rendering of the video data stream that allows the video data stream to be checked for authenticity sequentially for multiple portions of the video data stream; and obtaining a hash value (33) by applying a hash function (31) to a predetermined portion (13) of the video data stream, or data derived therefrom; obtaining a further hash value by subjecting a further portion of the video data stream, or data derived therefrom, to a hash function (31) (e.g., the further portion is a portion preceding the given portion (13)); An apparatus as described in any of embodiments 20 to 31, configured to obtain a digital signature (43) by signing a combination of the hash value (33), the further hash value, and the unique identifier (45) (e.g., a combination of multiple information including the hash value (33), the further hash value, and the unique identifier (45)). 33. The apparatus of any of embodiments 20 to 32, configured to sign a combination of a plurality of pieces of information including a hash value (33), a unique identifier (45), and instructions for a hash function (31) to obtain a digital signature (43).

[0198] Embodiments of the Second Aspect 34. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, the device comprising: applying a predetermined portion (13) of the video data stream, or data derived therefrom (62), to a hash function (31) to obtain a hash value (33); determining whether the video data stream is authentic by checking (41) whether the hash value (33) matches a digital signature (43) [e.g., derived from the video data stream and derived from criteria indicated in the video data stream]; Decrypting (46) the digital signature (43) using the public key (57) of the asymmetric decryption method to obtain a check value (47); It is configured to check (49) whether the hash value (33) matches the check value (47), The device is configured to check whether the video data stream includes an indication (55) of an external resource (280) [e.g., a metadata structure, e.g., a manifest file] that includes an editor's track (231) of the video data stream [e.g., a record of edits or modifications from the generation of the video data stream to its current version, and / or the identity of the corresponding editor], and if the video data stream includes an indication of an external resource that includes an editor's track of the video data stream, the device queries [or retrieves] the editor's track for a certificate (233) of a content provider that is the last editor of the video data stream and derives a public key based on the content provider's certificate.

[0199] 35. When checking whether the video data stream includes an indication of an external resource, including an editor track of the video data stream, deriving a syntax element from the video data stream, which is (1) The video data stream is a URI pointing directly to the content provider's certificate, or whether it contains a URI pointing to a register of certificates of the content provider and an index to a register pointing to the content provider of the video data stream, or (2) The device of embodiment 34, configured to indicate whether the video data stream includes an indication of an external resource including an editor's track. 36. The device of embodiment 35, wherein the syntax elements and, if present, the indication of external resources including the editor's track are transmitted in an SEI message of the video data stream.

[0200] 37. If present, an indication of the external resource, including the editor's track, is sent in an SEI message in the video data stream [e.g., SEI message], and the device: deriving a further digital signature (43) from the external resource, if there is an indication of the external resource containing the editor's track; 37. The apparatus of any of embodiments 34 to 36, configured to check whether the payload of the SEI message, or a predetermined portion thereof (13), matches a further digital signature (43).

[0201] 38. When checking whether the payload or a predetermined part thereof (13) matches a further digital signature (43), subjecting the payload or a predetermined portion thereof (13) to a further hash function to obtain a further hash value; 38. The apparatus of embodiment 37, configured to check whether the further hash value matches the further digital signature (43).

[0202] 39. The device of embodiment 37 or 38, wherein the predetermined portion (13) excludes indications of external resources, including editor tracks. 40. An apparatus according to any one of embodiments 37 to 39, wherein a predetermined portion (13) of the payload of the SEI message [e.g., a payload portion specific to a video data stream] includes a unique identifier. 41. The apparatus of any of embodiments 37 to 40, wherein the syntax structure further includes a media component identifier, and the apparatus is configured to use the media component identifier to select a further digital signature (43) from a set of one or more digital signatures included in the external resource (e.g., each of the one or more digital signatures is associated with a media component such as audio, video, subtitles, etc.).

[0203] 42. Syntax structure is Instructions for hash functions (31), An apparatus as described in any of embodiments 37 to 41, further comprising one or more of: an indication of some portions of the video data stream for which a digital signature is available for verifying the authenticity of the video data stream. 43. A device (17) for transcoding a video data stream in which the video is encoded, comprising: receiving an input video data stream (14') and checking (15') the authenticity of the input video data stream (14'); Transcoding (12) an input video data stream (14') to generate an output data stream (14); applying a predetermined portion (13) of the output video data stream (14) or of the data (62) from which the output video data stream is derived to a hash function (31) to obtain a hash value (33); Signing (71) the hash value using the private key (58) of the asymmetric cryptosystem to obtain a digital signature (43); providing a certificate (233) of the content provider (e.g., identifying the device) in an editor's track (231) of the output video data stream (e.g., a record of edits or modifications from the generation of the video data stream to the current version, and / or the identity of the corresponding editor), the editor's track being provided in an external resource (280) (e.g., a metadata structure in the external resource, e.g., a manifest file), the certificate including or pointing to a public key (57) of an asymmetric cryptography scheme; The apparatus is configured to provide a digital signature (43) to an output video data stream (14) [e.g., in an SEI message] or to an external resource (280) [e.g., by inserting the digital signature into a metadata structure or a further metadata structure and providing the same to the external resource] or to a further external resource.

[0204] 44. The device of embodiment 43, configured to insert into the output video data stream an indication of an external resource that includes an editor track for the video data stream, and to provide in the output video data stream (e.g., an SEI message) a syntax element indicating that the output video data stream (e.g., an SEI message) includes an indication of an external resource that includes an editor track. 45. The device of embodiment 44, wherein syntax elements and, if present, indications of external resources including editor tracks are transmitted in a syntax structure (e.g., an SEI message) of the output video data stream.

[0205] 46. ​​The apparatus is configured to insert, in a syntax structure (e.g., an SEI message) of an output video data stream, an indication of an external resource, including an editor track, applying a further hash function to the payload of the syntax structure, or a predetermined portion thereof (13), to obtain a further hash value; 46. ​​An apparatus as described in any of embodiments 43 to 45, configured to store the further hash value in an external resource (e.g., in an editor's track).

[0206] 47. The apparatus is configured to insert an indication of an external resource, including an editor's track, in a syntax structure (e.g., an SEI message) of an output video data stream, applying a further hash function to the payload of the syntax structure, or a predetermined portion thereof (13), to obtain a further hash value; obtaining a further digital signature using the private key to sign the syntax structure, or the payload of the syntax structure, or a predetermined portion thereof (13); 47. The apparatus of embodiment 46, configured to store a further digital signature in an external resource (e.g., in an editor's track).

[0207] 48. The device of embodiment 46 or 47, wherein the predetermined portion (13) excludes indications of external resources, including editor tracks. 49. An apparatus according to any one of embodiments 46 to 48, wherein a predetermined portion (13) of the payload of the SEI message [e.g., a payload portion specific to the output video data stream] includes a unique identifier. 50. An apparatus described in any of embodiments 46 to 49, configured to insert a media component identifier into a syntax structure, and the apparatus configured to store an association between the further digital signature or further hash value and the media component identifier in an external resource, for example, an editor track.

[0208] 51. Instructions for hash function (31) An apparatus as described in any of embodiments 46 to 50, configured to insert into a syntax structure one or more indications of some portions of the output video data stream for which a digital signature is available for verifying the authenticity of the output video data stream. 52. A device (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, the device comprising: applying a predetermined portion (13) of the video data stream (14) or of the data (62) from which the video data stream (14) is derived to a hash function (31) to obtain a hash value (33); Signing the hash value (33) using the private key (58) of the asymmetric cryptosystem to obtain a digital signature (43); providing a certificate (233) of the content provider (e.g., identifying the device) to an editor's track (231) of the video data stream (e.g., a record of edits or modifications from the generation of the video data stream to the current version, and / or the identity of the corresponding editor), the editor's track being provided in an external resource (280) (e.g., a metadata structure, e.g., a manifest file, in the external resource), the certificate including or pointing to a public key of an asymmetric cryptography scheme; The apparatus is configured to provide a digital signature (43) to an output video data stream (e.g., in an SEI message) or to an external resource (280) (e.g., by inserting the digital signature into a metadata structure or a further metadata structure and providing the same to the external resource) or a further external resource.

[0209] 53. The device of embodiment 52, configured to insert into the video data stream [e.g., in an SEI message] an indication of an external resource that includes an editor track for the video data stream, and to provide into the video data stream [e.g., in an SEI message] a syntax element indicating that the video data stream includes an indication of an external resource that includes an editor track. 54. The device of embodiment 53, wherein syntax elements and, if present, indications of external resources including editor tracks are transmitted in a syntax structure of the video data stream (e.g., an SEI message). 55. The apparatus is configured to insert, in a syntax structure (e.g., an SEI message) of a video data stream, an indication of an external resource, including an editor's track, applying a further hash function to the payload of the syntax structure, or a predetermined portion thereof (13), to obtain a further hash value; 55. An apparatus as described in any of embodiments 52 to 54, configured to store a further hash value in an external resource (e.g., in an editor's track).

[0210] 56. The apparatus is configured to insert an indication of an external resource, including an editor's track, in a syntax structure (e.g., an SEI message) of a video data stream, applying a further hash function to the payload of the syntax structure, or a predetermined portion thereof (13), to obtain a further hash value; obtaining a further digital signature using the private key to sign the syntax structure, or the payload of the syntax structure, or a predetermined portion thereof (13); 56. The apparatus of embodiment 55, configured to store a further digital signature in an external resource (e.g., in an editor's track). 57. The device of embodiment 55 or 56, wherein the predetermined portion (13) excludes indications of external resources, including editor tracks.

[0211] 58. An apparatus according to any one of embodiments 55 to 57, wherein a predetermined portion (13) of the payload of the SEI message [e.g., a payload portion specific to a video data stream] includes a unique identifier. 59. An apparatus described in any of embodiments 55 to 58, configured to insert a media component identifier into a syntax structure, and the apparatus configured to store an association between the further digital signature or further hash value and the media component identifier in an external resource, for example, an editor track. 60. An apparatus as described in any of embodiments 55 to 59, configured to insert into the syntax structure one or more of an indication of a hash function (31), an indication of some portions of the video data stream for which a digital signature is available to verify the authenticity of the video data stream.

[0212] Embodiments of the Third Aspect 61. An apparatus for checking a video data stream (14) in which video is encoded for authenticity, the apparatus being configured to apply a predetermined portion (13) of data of or derived from the video data stream to a hash function (31) to obtain a hash value (33), and to check whether the hash value (33) matches a digital signature (43) [e.g., derived from the video data stream and derived from criteria indicated in the video data stream] to determine whether the video data stream is authentic, the apparatus further comprising: a temporal layer [e.g., temporal_id] identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier [e.g., AVC priority_id] indicating the priority level of the picture, An apparatus configured to determine a predetermined portion (13) based on one or more of the AVC nal_ref_ids. 62. The apparatus of embodiment 61, configured to derive from the video data stream an indication of how to determine the predetermined portion (13).

[0213] 63. The instructions indicating the manner in which the predetermined portion (13) is determined are an indication associated with a time frame (e.g., an access unit) of the video data stream, indicating whether the time frame belongs to a given portion (13); a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], An apparatus as described in embodiment 62, which distinguishes one or more of AVC nal_ref_id. 64. The apparatus of embodiment 62 or 63, configured to derive an indication indicating a manner in which the predetermined portion (13) is determined from a syntax structure, for example, an SEI message.

[0214] 65. Syntax structure is Instructions for hash functions (31), The device of embodiment 64, further comprising one or more of: indications of several portions of the video data stream, for which respective digital signatures are provided for checking whether the video data stream is trustworthy (e.g., in the video data stream or criteria indicated in the video data stream). 66. The device is configured to determine a predetermined portion (13) based on a temporal layer identifier, a layer identifier, or a temporal frame identifier, and the device is configured to derive a range of values ​​from the video data stream, the range of values ​​indicating the value of each identifier, which value is associated with the predetermined portion (13) [e.g., such that pictures whose respective identifiers take values ​​within the range belong to the predetermined portion], as described in embodiment 62.

[0215] 67. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, the device comprising: applying a predetermined portion (13) of the video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); checking whether the hash value (33) matches the digital signature (43) (e.g., derived from the video data stream and derived from criteria indicated in the video data stream) to determine whether the video data stream is authentic; An apparatus configured to derive instructions from the video data stream, the instructions indicating how the predetermined portion (13) is determined. 68. The device of embodiment 67, wherein the instruction is a syntax element that distinguishes between multiple modes for determining the predetermined portion.

[0216] 69. The plurality of modes includes a first mode, and the device In a first mode, the device described in embodiment 68 is configured to determine the predetermined portion by determining whether to include a predetermined picture of the video data stream in the predetermined portion [or whether to assign the predetermined picture to the predetermined portion] depending on which layer of multiple layers of the video data stream the predetermined picture [e.g., predetermined from the currently considered perspective, such as a picture or an access unit that does not include a content selection SEI message] belongs to [e.g., depending on the value of the layer identifier of the layer to which the predetermined picture belongs].

[0217] 70. The plurality of modes further comprises a second mode, and the device: In a second mode, the device described in embodiment 69 is configured to determine the predetermined portion by including (e.g., by default) a predetermined picture of the video data stream (e.g., a picture or access unit that does not include a content selection SEI message, which is predetermined from the currently considered perspective) (e.g., the predetermined portion is a verification substream whose substream ID is equal to 0).

[0218] 71. The device of embodiment 70, wherein the plurality of modes comprises a first mode and a second mode [e.g., the indication is a syntax element [such as a flag] that distinguishes between the first mode and the second mode]. 72. An apparatus as described in embodiment 67, configured to check the authenticity of a video data stream in units of one or more portions (e.g., verification substreams), where one or more portions include predetermined portions, and the apparatus is configured to determine the one or more portions in a manner indicated by instructions.

[0219] 73. The plurality of modes includes a first mode, and the device An apparatus as described in embodiment 72, configured in a first mode to assign a given picture to one of one or more parts (e.g., including the given picture in an assigned part) depending on which layer of a plurality of layers of a video data stream the given picture belongs to (e.g., depending on the value of the layer identifier associated with the given picture (e.g., depending on the value of the layer identifier of the layer the given picture belongs to)). 74. The plurality of modes further includes a second mode, and the device: 74. The apparatus of embodiment 73, configured in the second mode to assign a predetermined picture to a predetermined one of the one or more portions (e.g., to a predetermined portion). 75. The plurality of modes further includes a second mode, and the device: The apparatus of embodiment 73, configured in a second mode to assign a given picture to a given one or more portions (e.g., to a predetermined portion) in the absence of dedicated signaling in the video data stream that signals which of the one or more portions the given picture should be assigned to.

[0220] 76. The instructions indicating the manner in which the predetermined portion (13) is determined are an indication associated with a time frame (e.g., an access unit) of the video data stream, indicating whether the time frame belongs to a given portion (13); a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], An apparatus as described in any of embodiments 67 to 75, which distinguishes one or more of AVC nal_ref_id.

[0221] 77. The apparatus of any of embodiments 67-76, configured to derive an indication from a syntax structure, e.g., an SEI message, indicating a manner in which the predetermined portion (13) is determined. 78. The syntax structure is Instructions for hash functions (31), The device of embodiment 77, further comprising one or more of: indications of several portions of the video data stream, for which respective digital signatures are provided for checking whether the video data stream is trustworthy (e.g., in the video data stream or criteria indicated in the video data stream). 79. An apparatus as described in any of embodiments 67 to 78, wherein the apparatus is configured to determine a predetermined portion (13) based on a temporal layer identifier, a layer identifier, or a temporal frame identifier, and the apparatus is configured to derive a range of values ​​from the video data stream, the range of values ​​indicating the value of each identifier, which value is associated with the predetermined portion (13) [e.g., a picture whose respective identifier takes a value within the range belongs to the predetermined portion].

[0222] 80. An apparatus (20) for decoding a video data stream (14) in which the video is encoded, comprising: configured to derive from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream (e.g., which predetermined portion (13) should be subjected to a hash function or used to derive data to be subjected to a hash function for deriving a hash value (33) that serves to check the authenticity of the video data stream); The syntax structure includes instructions indicating how to determine a predetermined portion (13) of the video data stream.

[0223] 81. The instructions indicating the manner in which the predetermined portion (13) is determined are a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream, the temporal layer identifier identifying the subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], AVC's nal_ref_id, An apparatus as described in embodiment 80, which distinguishes one or more of AVC nal_ref_id.

[0224] 82. The apparatus of embodiment 81, configured to decode, from a syntax structure, e.g., an SEI message, an indication indicating a manner in which the predetermined portion (13) is determined. 83. The syntax structure is Instructions for hash functions (31), The device of embodiment 82, further comprising one or more of: indications of several portions of the video data stream, for which respective digital signatures are provided for checking whether the video data stream is trustworthy (e.g., in the video data stream or criteria indicated in the video data stream).

[0225] 84. The device is configured to derive a range of values ​​from the video data stream, the range of values ​​indicating the value of a temporal layer identifier, layer identifier, or temporal frame identifier, and the value is associated with a predetermined portion (13) [e.g., such that pictures whose respective identifiers take values ​​within the range belong to the predetermined portion (13)], the device described in embodiment 81. 85. A device (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, the device comprising: subjecting a predetermined portion (13) of the video data stream, or of the data (62) from which a further portion of the video data stream is derived, to a hash function (31) to obtain a hash value (33); configured to sign the hash value (33) to obtain a digital signature (43) [e.g., by using a private key in an asymmetric cryptography scheme], The device is a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream, the temporal layer identifier identifying the subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], An apparatus configured to determine a predetermined portion (13) based on one or more of the AVC nal_ref_ids. 86. The apparatus of embodiment 85, configured to insert into the video data stream an indication indicating how the predetermined portion (13) is determined.

[0226] 87. The instructions indicating the manner in which the predetermined portion (13) is determined are an indication associated with a time frame (e.g., an access unit) of the video data stream, indicating whether the time frame belongs to a given portion (13); a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], An apparatus as described in embodiment 86, which distinguishes one or more of AVC nal_ref_id.

[0227] 88. The apparatus of embodiment 86 or 87, configured to insert an indication in a syntax structure, for example, an SEI message, indicating a manner in which the predetermined portion (13) is determined. 89. Instructions for hash functions (31) The device of embodiment 88, further configured to insert into the syntax structure one or more of the following indications of portions of the video data stream, for which respective digital signatures are provided for checking whether the video data stream is trustworthy (e.g., in the video data stream or criteria indicated in the video data stream). 90. The device is configured to determine a predetermined portion (13) based on a time layer identifier, a layer identifier, or a time frame identifier, and the device is configured to insert a range of values ​​into the video data stream, the range of values ​​indicating the value of each identifier, which value is associated with the predetermined portion (13) [e.g., such that pictures whose respective identifiers take values ​​within the range belong to the predetermined portion (13)], as described in embodiment 86.

[0228] 91. A device (15) for rendering a video data stream (14) in which the video is encoded in a way that is checkable for authenticity, the device comprising: subjecting a predetermined portion (13) of the video data stream, or of the data (62) from which a further portion of the video data stream is derived, to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43) [e.g., by using the private key in an asymmetric cryptography scheme], An apparatus configured to insert instructions into a video data stream, the instructions indicating a manner in which the predetermined portion (13) is determined.

[0229] 92. The device of embodiment 91, wherein the instruction is a syntax element that distinguishes between multiple modes for determining the predetermined portion. 93. The plurality of modes includes a first mode, and the device In a first mode, the device described in embodiment 92 is configured to determine the predetermined portion by determining whether to include a predetermined picture of the video data stream in the predetermined portion [or whether to assign the predetermined picture to the predetermined portion] depending on which layer of multiple layers of the video data stream the predetermined picture [e.g., predetermined from the currently considered perspective, such as a picture or access unit that does not include a content selection SEI message] belongs to [e.g., depending on the value of the layer identifier of the layer to which the predetermined picture belongs].

[0230] 94. The plurality of modes further comprises a second mode, and the device: In a second mode, the device described in embodiment 93 is configured to determine the predetermined portion by including (e.g., by default) a predetermined picture of the video data stream (e.g., a picture or access unit that does not include a content selection SEI message, which is predetermined from the currently considered perspective) (e.g., the predetermined portion is a verification substream whose substream ID is equal to 0). 95. The device of embodiment 94, wherein the plurality of modes comprises a first mode and a second mode [e.g., the indication is a syntax element [such as a flag] that distinguishes between the first mode and the second mode].

[0231] 96. An apparatus as described in embodiment 91, wherein the video data stream is configured to be rendered so that its authenticity can be checked in units of one or more parts [e.g., verification substreams], the one or more parts including predetermined parts, and the apparatus is configured to select a method for determining the one or more parts and indicate the selected method for determining the one or more parts of the data stream by instructions. 97. The plurality of modes includes a first mode, and the device: An apparatus as described in embodiment 96, configured in a first mode to assign a given picture to one of one or more parts (e.g., including the given picture in an assigned part) depending on which layer of a plurality of layers of a video data stream the given picture belongs to (e.g., depending on the value of the layer identifier associated with the given picture (e.g., depending on the value of the layer identifier of the layer the given picture belongs to)).

[0232] 98. The plurality of modes further includes a second mode, and the device: 98. The apparatus of embodiment 97, configured in the second mode to assign a predetermined picture to a predetermined one of the one or more portions (e.g., to a predetermined portion). 99. The plurality of modes further includes a second mode, and the device: The apparatus of embodiment 97, configured in a second mode to assign a predetermined picture to a predetermined one of one or more portions (e.g., to a predetermined portion) or to provide dedicated signaling that signals which of one or more portions of the video data stream a predetermined picture is assigned to.

[0233] 100. The instructions indicating the manner in which the predetermined portion (13) is determined are an indication associated with a time frame (e.g., an access unit) of the video data stream, indicating whether the time frame belongs to a given portion (13); a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], An apparatus as described in any of embodiments 91 to 99, which distinguishes one or more of AVC nal_ref_id.

[0234] 101. The apparatus of any of embodiments 91 to 100, configured to insert an indication in a syntax structure, for example, an SEI message, indicating a manner in which the predetermined portion (13) is determined. 102. Instructions for hash functions (31), The device of embodiment 101, further configured to insert into the syntax structure one or more of the following indications of portions of the video data stream, for which respective digital signatures are provided for checking whether the video data stream is trustworthy (e.g., in the video data stream or criteria indicated in the video data stream). 103. An apparatus described in any of embodiments 91 to 102, wherein the apparatus is configured to determine a predetermined portion (13) based on a time layer identifier, a layer identifier, or a time frame identifier, and the apparatus is configured to insert a range of values ​​into the video data stream, the range of values ​​indicating the value of each identifier, which value is associated with the predetermined portion (13) [e.g., such that a picture whose respective identifier takes a value within the range belongs to the predetermined portion (13)].

[0235] Embodiments of the Fourth Aspect 104. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, the device comprising: applying a predetermined portion (13) of the video data stream or of data derived therefrom (e.g., video data associated with an access unit, e.g., a time frame) to a hash function (31) to obtain a hash value (33); Deriving a digital signature (43) associated with the predetermined portion (13) from an external resource (e.g., a server); The apparatus is configured to check whether the hash value (33) matches the digital signature (43) to determine whether the video data stream is authentic. 105. The apparatus of embodiment 104, configured to derive a reference to an external resource from a video data stream.

[0236] 106. Decrypt the digital signature (43) using the public key in an asymmetric decryption scheme to obtain a check value; 106. The apparatus of embodiment 104 or 105, configured to check whether the hash value (33) matches the check value. 107. Derive a portion identifier [e.g., identifying a temporal portion of the video data stream, e.g., a time frame, e.g., an access unit] [e.g., a hash identifier or hash index, e.g., twcs_associated_hash_idx] from the video data stream, the portion identifier being associated with a given portion (13); Use the part identifier to identify the part of the check value, An apparatus as described in embodiment 106, configured to check whether the hash value (33) matches a portion of the check value when checking whether the hash value (33) matches the check value.

[0237] 108. Derive a media component identifier [e.g., twcs_associated_hash_group_id] from the video data stream that indicates the media type [e.g., video, audio, subtitle] of a given portion (13); 108. The apparatus of embodiment 107, configured to use a media component identifier to identify portions of the check value. 109. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, the device comprising: applying a predetermined portion (13) of the video data stream or of data derived therefrom (e.g., video data associated with an access unit, e.g., a time frame) to a hash function (31) to obtain a hash value (33); Derive a check value (e.g., a signed check value signed with a private key of an asymmetric encryption scheme) associated with the predetermined portion (13) from an external resource (e.g., a server); The apparatus is configured to check whether the hash value (33) matches a check value to determine whether the video data stream is authentic.

[0238] 110. The apparatus of embodiment 109, configured to derive a reference to an external resource from a video data stream. 111. Derive a digital signature (43) from an external resource; The device of embodiment 109 or 110, configured to verify the check value (e.g., the origin of the check value or the identity of the provider of the check value) using a digital signature (43) (e.g., using a public key in an asymmetric cryptography scheme). 112. Derive a portion identifier [e.g., identifying a temporal portion of the video data stream, e.g., a time frame, e.g., an access unit] [e.g., a hash identifier or hash index, e.g., twcs_associated_hash_idx] from the video data stream, the portion identifier being associated with a given portion (13); 112. An apparatus according to any one of embodiments 109 to 111, configured to obtain a check value from an external resource using the partial identifier.

[0239] 113. Performing a video data stream authenticity check sequentially on multiple portions of the video data stream; and applying a further portion of the video data stream, or data derived therefrom, to a hash function (31) to obtain a further hash value (e.g., the further portion is a previous portion relative to the given portion (13)); deriving a further check value (e.g., a signed check value signed with the private key of an asymmetric cryptography scheme) associated with the further part from an external resource (e.g., a server); Derive a digital signature (43) from an external resource; An apparatus described in any of embodiments 109 to 112, configured to verify the check value and a further check value (e.g., the origin of the check value or the identity of the provider of the check value) using a digital signature (43) (e.g., using a public key in an asymmetric encryption scheme).

[0240] 114. An apparatus (20) for decoding a video data stream (14) in which the video is encoded, comprising: configured to derive from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream (e.g., which predetermined portion (13) should be subjected to a hash function or used to derive data to be subjected to a hash function for deriving a hash value (33) that serves to check the authenticity of the video data stream); The syntax structure includes a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a digital signature (43) associated with a given portion (13).

[0241] 115. The apparatus of embodiment 114, configured to derive a portion identifier [e.g., identifying a time portion, e.g., a time frame, e.g., an access unit, of the video data stream] [e.g., from a syntax structure or from a further syntax structure] from the video data stream, the portion identifier being associated with a predetermined portion (13), and the portion identifier being associated with one or more digital signatures contained in an external resource.

[0242] 116. The device of embodiment 114 or 115, configured to derive a media component identifier [e.g., twcs_associated_hash_group_id] from the video data stream [e.g., from a syntax structure or from a further syntax structure], the media component identifier indicating the media type [e.g., video, audio, subtitle] of a given portion (13). 117. An apparatus (20) for decoding a video data stream (14) in which the video is encoded, comprising: configured to derive from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream (e.g., which predetermined portion (13) should be subjected to a hash function or used to derive data to be subjected to a hash function for deriving a hash value (33) that serves to check the authenticity of the video data stream); The syntax structure includes a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a check value (e.g., a check value signed with a digital signature) associated with a given portion (13).

[0243] 118. The device of embodiment 117, configured to derive a portion identifier [e.g., identifying a time portion, e.g., a time frame, e.g., an access unit, of the video data stream] [e.g., from a syntax structure or from a further syntax structure] from the video data stream, the portion identifier being associated with a predetermined portion (13), and the portion identifier being associated with one or more check values ​​contained in an external resource. 119. The device of embodiment 117 or 118, configured to derive a media component identifier [e.g., twcs_associated_hash_group_id] from the video data stream [e.g., from a syntax structure or from a further syntax structure], the media component identifier indicating the media type [e.g., video, audio, subtitle] of a given portion (13).

[0244] 120. A device (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, the device comprising: applying a predetermined portion (13) of the video data stream or of the data (62) from which the predetermined portion (13) of the video data stream is derived to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43) [e.g., by using the private key in an asymmetric cryptography scheme] and providing the digital signature (43) to the external resource; An apparatus configured to insert an indication of an external resource (e.g., a reference to a digital signature of the external resource) into a video data stream (e.g., a URI of the external resource or the digital signature).

[0245] 121. The apparatus of embodiment 120, configured to sign the hash value (33) using a private key of an asymmetric decryption scheme to obtain a digital signature (43). 122. Inserting into the video data stream a portion identifier [e.g., identifying a temporal portion of the video data stream, e.g., a time frame, e.g., an access unit] [e.g., a hash identifier or hash index, e.g., twcs_associated_hash_idx], the portion identifier being associated with a given portion (13); 122. The apparatus of embodiment 120 or 121, configured to create an association between the partial identifier and the digital signature (43) at an external resource. 123. Insert a media component identifier [e.g., twcs_associated_hash_group_id] into the video data stream; a media component identifier indicating the media type (e.g., video, audio, subtitle) of a given part (13); 123. An apparatus according to any one of embodiments 120 to 122, configured to create an association between a media component identifier and a digital signature (43) in an external resource.

[0246] 124. A device (15) for rendering a video data stream (14) in which the video is encoded in a way that is checkable with respect to authenticity, the device comprising: subjecting a predetermined portion (13) of the video data stream, or of the data (62) from which a further portion of the video data stream is derived, to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43) [e.g., by using the private key in an asymmetric cryptography scheme], and providing the hash value (33) and the digital signature (43) to an external resource; An apparatus configured to insert an indication of an external resource (e.g., a reference to a digital signature of the external resource) into a video data stream (e.g., a URI of the external resource or the digital signature).

[0247] 125. Inserting into the video data stream a portion identifier [e.g., identifying a temporal portion of the video data stream, e.g., a time frame, e.g., an access unit] [e.g., a hash identifier or hash index, e.g., twcs_associated_hash_idx], the portion identifier being associated with a given portion (13); 125. The apparatus of embodiment 124, configured to provide, at an external resource, an association between the partial identifier and the digital signature (43). 126. Rendering the video data stream sequentially for a plurality of portions of the video data stream to enable authenticity checks, and further subjecting further portions of the video data stream, or data (62) from which further portions of the video data stream are derived, to a hash function (31) to obtain further hash values ​​(e.g., the further portions are portions preceding the given portion (13)); signing the hash value (33) and the further hash value together to obtain a digital signature (43) [e.g., by using a private key in an asymmetric cryptography scheme], and providing the hash value (33), the further hash value, and the digital signature to an external resource; 126. The apparatus of embodiment 124 or 125, configured to insert an indication of an external resource [e.g., a reference to a digital signature of the external resource] [e.g., a URI of the external resource or the digital signature] into the video data stream.

[0248] Embodiments of all aspects 127. An apparatus according to any preceding embodiment, wherein the hash value (33) depends on all bits of a predetermined portion (13) of the video data stream. 128. An apparatus according to any of the preceding embodiments, wherein the hash value (33) is based on all bits of a predetermined portion (13) of the video data stream in an encoded domain (e.g., in a domain in which at least a portion of the video data stream is entropy coded). 129. The predetermined portion (13) of the video data stream extends across multiple access units (or time frames) of the video data stream such that the hash value (33) depends on bits from multiple access units; or 10. The apparatus of any of the preceding embodiments, wherein the predetermined portion (13) comprises video data of only one access unit (or time frame). 130. The apparatus of any one of the preceding embodiments, wherein the apparatus is a decoder for decoding a video data stream [e.g., a decoder compliant with H.264 / AVC, or H.265 / HEVC, or H.266 / VVC] [e.g., the decoder is configured to decode video from the video data stream by block-based predictive decoding and transform-based residual decoding].

[0249] 131. The device is The video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding prediction residual data of a residual block from the video data stream; a first syntax element indicating the total number of non-zero transform coefficients of a transform block representing a residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements that indicate the signs of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element that indicates the total number of zero-valued transform coefficient levels in the transform block, starting from the first encountered non-zero transform coefficient in scan order onward; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or decoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by decoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, decoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and 131. The apparatus of embodiment 130, using context-adaptive binary arithmetic coding by reversing the forward scan order and decoding the values ​​of non-zero transform coefficients sequentially in reverse scan order.

[0250] 132. The apparatus of any of embodiments 20 to 33, 52 to 60, 85 to 90, 120 to 123, 91 to 103, and 124 to 126, wherein the apparatus is an encoder for encoding a video data stream (e.g., an encoder for encoding a video data stream to be compliant with H.264 / AVC, H.265 / HEVC, or H.266 / VVC). 133. The device is encoding the video into a video data stream by block-based predictive coding; 1. An encoder configured for transform-based residual coding by encoding prediction residual data of a residual block into a video data stream, comprising: a first syntax element indicating the total number of non-zero transform coefficients of a transform block representing a residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements that indicate the signs of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element that indicates the total number of zero-valued transform coefficient levels in the transform block, starting from the first encountered non-zero transform coefficient in scan order onward; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and 127. The apparatus of any of embodiments 20 to 33, 52 to 60, 85 to 90, 120 to 123, 91 to 103, and 124 to 126, by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of non-zero transform coefficients in sequential reverse scan order.

[0251] 134. A method (16) for checking a video data stream (14) in which video is encoded for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream or of data derived therefrom (62) to a hash function (31) to obtain a hash value (33); obtaining (51) a unique identifier (45) [e.g., from a video data stream, or a reference, e.g., using a URI] that uniquely identifies the media asset to which the given portion (13) belongs; Obtaining (51) a digital signature (43) based on the video data stream [e.g., from the video data stream, e.g., twcv_signature, or from a metadata file or manifest file, e.g., C2PA file]; 1. A method comprising: checking (41) whether a combination of a hash value (33) and a unique identifier (45) (e.g., a combination of multiple pieces of information including the hash value and the unique identifier (45)) matches a digital signature (43) to determine whether the video data stream is authentic.

[0252] 135. A method (20) for decoding a video data stream in which video is encoded, the method comprising: decoding (21) a syntax structure from the video data stream, which contains information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream, the predetermined portion being subjected to a hash function (31) or used to derive data to be subjected to the hash function (31) to derive a hash value (33) which serves to check the authenticity of the video data stream; decoding (21) a unique identifier (45) or a reference pointing to the unique identifier (45) from the video data stream, the unique identifier (45) uniquely identifying the media asset to which the given portion (13) belongs; A method comprising: decoding (21) an indication of a digital signature (43) from a video data stream (e.g., from a video data stream, e.g., twcv_signature, or from a metadata file or manifest file, e.g., a C2PA file), wherein the digital signature (43) is based on a combination of a hash value (33) and a unique identifier (45) (e.g., a combination of multiple pieces of information including the hash value and the unique identifier (45)).

[0253] 136. A method (15) for rendering a video data stream (14) in which the video is encoded in a way that is checkable for authenticity, the method comprising: applying a hash function (31) to a predetermined portion (13) of the video data stream (14) or of the data (62) from which the video data stream (14) is derived to obtain a hash value (33); assigning a unique identifier (45) to the given portion (13) that uniquely identifies the media asset to which the given portion (13) belongs; A method comprising signing (71) a combination of a hash value (33) and a unique identifier (45) (e.g., a combination of multiple pieces of information including the hash value (33) and the unique identifier (45)) to obtain a digital signature (43).

[0254] 137. A method (16) for checking a video data stream (14) in which video is encoded for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream, or data derived therefrom (62), to a hash function (31) to obtain a hash value (33); determining whether the video data stream is authentic by checking (41) whether the hash value (33) matches a digital signature (43) [e.g., derived from the video data stream and derived from criteria indicated in the video data stream]; Decrypting (46) the digital signature (43) using the public key (57) of the asymmetric decryption method to obtain a check value (47); checking (49) whether the hash value (33) matches a check value (47); The method includes checking whether the video data stream includes an indication (55) of an external resource (280) [e.g., a metadata structure, e.g., a manifest file] that includes an editor's track (231) of the video data stream [e.g., a record of edits or modifications from the generation of the video data stream to the current version, and / or the identity of the corresponding editor], and if the video data stream includes an indication of an external resource that includes an editor's track of the video data stream, querying [or retrieving] the editor's track for a certificate (233) of a content provider that is the last editor of the video data stream, and deriving a public key based on the content provider's certificate.

[0255] 138. A method (17) for transcoding a video data stream in which video is encoded, the method comprising: receiving an input video data stream (14') and checking (15') the authenticity of the input video data stream (14'); Transcoding (12) an input video data stream (14') to generate an output data stream (14); applying a predetermined portion (13) of the output video data stream (14) or of the data (62) from which the output video data stream is derived to a hash function (31) to obtain a hash value (33); Signing (71) the hash value using the private key (58) of the asymmetric cryptosystem to obtain a digital signature (43); providing a certificate (233) of the content provider (e.g., identifying the device) in an editor's track (231) of the output video data stream (e.g., a record of edits or modifications from the generation of the video data stream to the current version, and / or the identity of the corresponding editor), the editor's track being provided in an external resource (280) (e.g., a metadata structure in the external resource, e.g., a manifest file), the certificate including or pointing to a public key (57) of an asymmetric cryptography scheme; providing a digital signature (43) to the output video data stream (14) [e.g., in an SEI message] or to an external resource (280) [e.g., inserting the digital signature in a metadata structure or a further metadata structure and providing the same to the external resource] or to a further external resource.

[0256] 139. A method (15) for rendering a video data stream (14) in which the video is encoded in a way that is checkable for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream (14) or of the data (62) from which the video data stream (14) is derived to a hash function (31) to obtain a hash value (33); Signing the hash value (33) using the private key (58) of the asymmetric cryptosystem to obtain a digital signature (43); providing a certificate (233) of the content provider (e.g., identifying the device) to an editor's track (231) of the video data stream (e.g., a record of edits or modifications from the generation of the video data stream to the current version, and / or the identity of the corresponding editor), the editor's track being provided in an external resource (280) (e.g., a metadata structure, e.g., a manifest file, in the external resource), the certificate including or pointing to a public key of an asymmetric cryptography scheme; providing a digital signature (43) to the output video data stream (e.g., in an SEI message) or to an external resource (280) (e.g., inserting the digital signature into a metadata structure or a further metadata structure and providing the same to the external resource) or to a further external resource.

[0257] 140. A method for checking a video data stream (14) in which video is encoded for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream or of data derived therefrom to a hash function (31) to obtain a hash value (33); and checking whether the hash value (33) matches a digital signature (43) [e.g., derived from the video data stream and derived from criteria indicated in the video data stream] to determine whether the video data stream is authentic; the method further comprising: a temporal layer [e.g., temporal_id] identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which each picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], determining a predetermined portion (13) based on one or more of the AVC nal_ref_ids.

[0258] 141. A method (16) for checking a video data stream (14) in which video is encoded for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); checking whether the hash value (33) matches the digital signature (43) (e.g., derived from the video data stream and derived from criteria indicated in the video data stream) to determine whether the video data stream is authentic; The method includes deriving instructions from the video data stream, the instructions indicating how the predetermined portion (13) is determined.

[0259] 142. A method (20) for decoding a video-encoded video data stream (14), the method comprising: deriving a syntax structure from the video data stream that includes information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream (e.g., which predetermined portion (13) should be subjected to a hash function or used to derive data that should be subjected to a hash function to derive a hash value (33) that functions to check the authenticity of the video data stream); The method, wherein the syntax structure includes instructions indicating how to determine a predetermined portion (13) of the video data stream.

[0260] 143. A method (15) for rendering a video data stream (14) in which the video is encoded in a way that is checkable for authenticity, the method comprising: subjecting a predetermined portion (13) of the video data stream, or of the data (62) from which a further portion of the video data stream is derived, to a hash function (31) to obtain a hash value (33); signing the hash value (33) to obtain a digital signature (43) [e.g., by using a private key in an asymmetric cryptography scheme]; The method is: a temporal layer (e.g., temporal_id) identifier associated with pictures of the video data stream, the temporal layer identifier identifying the subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream [e.g., layer_id for HEVC / VVC, dependency_id and / or quality_id for AVC] that identify the layer of the video data stream to which the respective picture belongs [e.g., the video data stream is a layered video data stream that includes multiple layers, e.g., a base layer and one or more enhancement layers, representing video at different resolutions or from different viewpoints], A combination of a time layer identifier and a layer identifier; a time frame identifier [e.g., Picture Order Count, POC], A priority level identifier indicating the priority level of the picture [e.g., priority_id in AVC], determining a predetermined portion (13) based on one or more of the AVC nal_ref_ids.

[0261] 144. A method (15) for rendering a video data stream (14) in which the video is encoded in a way that is checkable for authenticity, the method comprising: subjecting a predetermined portion (13) of the video data stream, or of the data (62) from which a further portion of the video data stream is derived, to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43) [e.g., by using the private key in an asymmetric cryptography scheme], A method comprising inserting instructions into a video data stream, the instructions indicating a manner in which the predetermined portion (13) is determined.

[0262] 145. A method (16) for checking a video data stream (14) in which video is encoded for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream or of data derived therefrom (e.g., video data associated with an access unit, e.g., a time frame) to a hash function (31) to obtain a hash value (33); Deriving a digital signature (43) associated with the predetermined portion (13) from an external resource (e.g., a server); checking whether the hash value (33) matches the digital signature (43) to determine whether the video data stream is authentic.

[0263] 146. A method (16) for checking a video data stream (14) in which video is encoded for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream or of data derived therefrom (e.g., video data associated with an access unit, e.g., a time frame) to a hash function (31) to obtain a hash value (33); Derive a check value (e.g., a signed check value signed with a private key of an asymmetric encryption scheme) associated with the predetermined portion (13) from an external resource (e.g., a server); checking whether the hash value (33) matches a check value to determine whether the video data stream is authentic.

[0264] 147. A method (20) for decoding a video-encoded video data stream (14), comprising: deriving a syntax structure from the video data stream that includes information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream (e.g., which predetermined portion (13) should be subjected to a hash function or used to derive data that should be subjected to a hash function to derive a hash value (33) that functions to check the authenticity of the video data stream); The syntax structure includes a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a digital signature (43) associated with a given portion (13).

[0265] 148. A method (20) for decoding a video-encoded video data stream (14), comprising: configured to derive from the video data stream a syntax structure including information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream (e.g., which predetermined portion (13) should be subjected to a hash function or used to derive data to be subjected to a hash function for deriving a hash value (33) that serves to check the authenticity of the video data stream); The syntax structure includes a reference to an external resource (e.g., a metadata structure or a manifest file) for obtaining a check value (e.g., a check value signed with a digital signature) associated with a given portion (13).

[0266] 149. A method (15) for rendering a video data stream (14) in which the video is encoded in a way that is checkable for authenticity, the method comprising: applying a predetermined portion (13) of the video data stream or of the data (62) from which the predetermined portion (13) of the video data stream is derived to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43) [e.g., by using the private key in an asymmetric cryptography scheme] and providing the digital signature (43) to the external resource; A method comprising inserting an indication of an external resource into a video data stream [e.g., a reference to a digital signature of the external resource] [e.g., a URI of the external resource or the digital signature].

[0267] 150. A method (15) for rendering a video data stream (14) in which the video is encoded in a manner that allows checking for authenticity, the method comprising: subjecting a predetermined portion (13) of the video data stream, or of the data (62) from which a further portion of the video data stream is derived, to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43) [e.g., by using the private key in an asymmetric cryptography scheme], and providing the hash value (33) and the digital signature (43) to an external resource; A method comprising inserting an indication of an external resource into a video data stream [e.g., a reference to a digital signature of the external resource] [e.g., a URI of the external resource or the digital signature].

[0268] 151. A method for storing video, comprising: A method comprising storing a data stream on a digital storage medium, the data stream being generated by a method according to any one of embodiments 136, 139, 143, 144, 149, or 150. 152. A method for transmitting a data stream generated by the method of any one of embodiments 136, 139, 143, 144, 149, or 150. 153. A computer program [or computer program product, for example a computer program stored on a non-transitory digital storage medium] for performing the method according to any of embodiments 134 to 152 when the computer program is run on a computer or signal processor. 154. A video data stream (e.g., a non-transitory digital storage medium containing a video data stream) generated by a method according to any one of embodiments 136, 139, 143, 149, or 144, 150.

[0269] Implementation alternatives While some aspects are described as features in the context of an apparatus, it will be apparent that such description may also be considered a description of corresponding features of a method. Although some aspects are described as features in the context of a method, it will be apparent that such description may also be considered a description of corresponding features with respect to the functionality of the apparatus. In particular, block diagrams illustrating apparatuses may also be considered illustrations of respective methods that include the steps described by the blocks in the block diagram. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0270] The coded image signal of the present invention can be stored on a digital storage medium or transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet. In other words, a further embodiment provides a video bitstream product, e.g., a digital storage medium having a video bitstream stored thereon, comprising a video bitstream according to any of the embodiments described herein. Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware, or at least partially in software. Implementations can be performed using digital storage media, such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memories, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage media can be computer-readable.

[0271] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein. Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer. The program code can, for example, be stored on a machine-readable carrier.

[0272] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or sequence of signals can for example be adapted to be transferred via a data communication connection, for example via the Internet.

[0273] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. Further embodiments according to the invention comprise an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0274] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0275] In the foregoing Detailed Description, it can be seen that various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than are expressly recited in each claim. Rather, as the following claims reflect, subject matter may lie in fewer than all features of a single disclosed example. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate example. While each claim may stand on its own as a separate example, and a dependent claim may refer to a specific combination with one or more other claims in the claim, it should be noted that other examples may also include a combination of a dependent claim with the subject matter of each other dependent claim, or a combination of each feature with other dependent or independent claims. Such combinations are suggested herein unless it is stated that a specific combination is not intended. Furthermore, it is also intended to include features of any other independent claim, even if that claim is not directly dependent on that independent claim.

[0276] The above-described embodiments are merely illustrative of the principles of the present disclosure. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented as descriptions and explanations of the embodiments herein.

Claims

1. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, said device comprising: applying (31) a predetermined portion (13) of the video data stream or of data derived therefrom to a hash function (31) to obtain a hash value (33); obtaining a unique identifier (45) that uniquely identifies the media asset to which the predetermined portion (13) belongs; obtaining a digital signature (43) based on said video data stream; configured to check (41) whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43) to determine whether the video data stream is authentic; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

2. When checking whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43), decrypting said digital signature (43) to obtain a check value; 2. The apparatus of claim 1, configured to check whether the combination of the hash value (33) and the unique identifier (45) matches the check value.

3. When checking whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43), forming a verification string based on the hash value (33) and the unique identifier (45); 3. The apparatus of claim 1, configured to compare the verification string with the digital signature (43) using a public key.

4. 4. Apparatus according to claim 1, configured to derive instructions for an external resource to obtain the public key from the video data stream, and to derive the public key from the external resource.

5. The apparatus of claim 4 , configured to derive the unique identifier (45) from the external resource.

6. 6. The apparatus of claim 1, configured to check whether the unique identifier (45) matches a unique identifier (45) associated with one or more further media components.

7. 7. Apparatus according to any of the preceding claims, configured to derive the digital signature (43) from the video data stream.

8. deriving an indication of an external resource from the video data stream; 8. The apparatus of claim 1, configured to derive the digital signature (43) from the external resource.

9. The apparatus of claim 8 , wherein the indication of the external resource is a uniform resource identifier (URI) that points to a manifest file stored on a server.

10. 10. Apparatus according to any preceding claim, configured to derive the unique identifier (45) from the video data stream.

11. 11. Apparatus according to any of the preceding claims, configured to derive the unique identifier (45) from payload packets signalled in the video data stream.

12. The payload packet comprises: an indication of the hash function (31); an indication of some portions of the video data stream for which digital signatures (43) are available for verifying the authenticity of the video data stream; and instructions indicating how to obtain a public key to check whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43).

13. When the predetermined portion (13) of the video data stream or data derived therefrom is subjected to a hash function (31) to obtain the hash value (33), reconstructing the video with respect to the predetermined portion (13) to obtain a reconstructed portion of the video; 13. Apparatus according to any of the preceding claims, configured to subject the reconstructed portion to the hash function (31).

14. The device comprises: Decoding the video from the video data stream; 14. Apparatus according to any preceding claim, which is a decoder configured to decode the indication of the digital signature (43) from the video data stream.

15. 15. The apparatus of claim 14, configured to decode the digital signature (43) from a supplemental enhancement information message of the video data stream.

16. performing said check of authenticity of said video data stream sequentially on a plurality of portions of said video data stream; and applying a predetermined portion (13) of the video data stream, or data derived therefrom, to a hash function (31) to obtain a hash value (33); applying a further portion of the video data stream, or data derived therefrom, to the hash function (31) to obtain a further hash value; 16. The apparatus of claim 1, configured to check whether a combination of the hash value (33), the further hash value, and the unique identifier (45) matches the digital signature (43).

17. 17. The apparatus of claim 1, configured to check whether a combination of a plurality of pieces of information including the hash value (33), the unique identifier (45), and an indication of the hash function (31) matches the digital signature (43) to determine whether the video data stream is trustworthy.

18. An apparatus (20) for decoding a video data stream in which video is encoded, said apparatus comprising: decoding a syntax structure (55) from the video data stream, the syntax structure (55) containing information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream, the predetermined portion being subjected to a hash function (31) or used to derive data to be subjected to a hash function (31), deriving a hash value (33) that serves to check the authenticity of the video data stream; decoding from the video data stream a unique identifier (45) or a reference pointing to the unique identifier (45), the unique identifier (45) uniquely identifying the media asset to which the predetermined portion (13) belongs; decoding an indication of a digital signature (43) from the video data stream, the digital signature (43) being configured based on a combination of the hash value (33) and the unique identifier (45); The device comprises: configured to decode the video from the video data stream by block-based prediction and transform-based residual decoding by decoding the predictive residual data of the residual blocks from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number followed by a one indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients along a scan order; one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and one or more fifth syntax elements indicating a position of the non-zero transform coefficients along the scan order by indicating a number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or decoding a significance map indicating the location of non-zero transform coefficients of the transform block representing the residual block by decoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, decoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic decoding by reversing the forward scan order and decoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

19. 19. The apparatus of claim 18, wherein the unique identifier (45) is signaled in the syntax structure.

20. An apparatus (15) for rendering a video data stream (14) in which the video is coded in a way that allows it to be checked for authenticity, said apparatus comprising: applying a predetermined portion (13) of said video data stream (14) or of the data (62) from which said video data stream (14) is derived to a hash function (31) to obtain a hash value (33); assigning a unique identifier (45) to the predetermined portion (13) that uniquely identifies the media asset to which the predetermined portion (13) belongs; configured to sign a combination of said hash value (33) and said unique identifier (45) to obtain a digital signature (43); The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

21. forming a verification string based on said hash value (33) and said unique identifier (45); 21. The apparatus of claim 20, configured to sign the verification string using a private key to obtain the digital signature (43).

22. 22. Apparatus according to claim 20 or 21, arranged to provide, in the video data stream, an indication of an external resource which holds or indicates the private key.

23. 23. The apparatus of claim 22, configured to provide the unique identifier (45) at the external resource.

24. 24. Apparatus according to any of claims 20 to 23, arranged to provide said digital signature (43) in said video data stream.

25. 24. Apparatus according to any of claims 20 to 23, arranged to provide an indication of an external resource in the video data stream and to provide the digital signature (43) at the external resource.

26. 26. The apparatus of claim 25, wherein the indication of the external resource is a uniform resource identifier (URI) that points to a manifest file stored on a server.

27. 27. Apparatus according to any of claims 20 to 26, configured to insert the unique identifier (45) into the video data stream.

28. 28. Apparatus according to any of claims 20 to 27, configured to insert the unique identifier (45) into payload packets signalled in the video data stream.

29. In the payload packet: an indication of the hash function (31); an indication of some portions of the video data stream for which digital signatures (43) are available for verifying the authenticity of the video data stream; and instructions indicating how to obtain a public key to check whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43).

30. The device, encoding said video into said video data stream; 30. Apparatus according to any of claims 20 to 29, being an encoder configured to encode an indication of said digital signature (43) into said video data stream.

31. 31. The apparatus of claim 30, configured to encode the digital signature (43) into a supplemental enhancement information message of the video data stream.

32. performing a rendering of the video data stream that allows sequential authenticity checks on portions of the video data stream; and obtaining a hash value (33) by subjecting a predetermined portion (13) of the video data stream, or data derived therefrom, to a hash function (31); obtaining a further hash value by subjecting a further portion of the video data stream, or data derived therefrom, to the hash function (31); 32. The apparatus of any of claims 20 to 31, configured to obtain a digital signature (43) by signing a combination of the hash value (33), the further hash value, and the unique identifier (45).

33. 33. The apparatus of any of claims 20 to 32, configured to sign a combination of information including the hash value (33), the unique identifier (45), and an instruction to the hash function (31) to obtain the digital signature (43).

34. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, said device comprising: applying a predetermined portion (13) of said video data stream, or data derived therefrom (62), to a hash function (31) to obtain a hash value (33); checking (41) whether said hash value (33) matches a digital signature (43) to determine whether said video data stream is authentic; decrypting (46) said digital signature (43) using a public key (57) in an asymmetric decryption scheme to obtain a check value (47); configured to check (49) whether said hash value (33) matches said check value (47); The device is configured to check whether the video data stream contains an indication (55) of an external resource (280) that contains an editor's track (231) of the video data stream, and if the video data stream contains an indication of an external resource that contains an editor's track of the video data stream, query the editor's track for a certificate (233) of a content provider that is the last editor of the video data stream, and derive the public key based on the certificate of the content provider, The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

35. When checking whether the video data stream includes an indication of an external resource including an editor's track of the video data stream, deriving a syntax element from the video data stream, which includes: (1) the video data stream is a URI that points directly to the certificate of the content provider; or whether it contains a URI pointing to a register of certificates of content providers and an index into said register pointing to said content provider of said video data stream, or 35. The apparatus of claim 34, further configured to: (2) indicate whether the video data stream includes the indication of the external resource that includes the editor's track.

36. 36. The apparatus of claim 35, wherein the syntax elements and, if present, the indication of the external resources including the editor's track are transmitted in an SEI message of the video data stream.

37. If present, the indication of the external resource including the editor's track is sent in an SEI message of the video data stream, and the device: deriving a further digital signature (43) from the external resource, if said indication of said external resource containing said editor's track is present; 37. An apparatus according to any of claims 34 to 36, configured to check whether the payload of the SEI message, or a predetermined part thereof (13), matches the further digital signature (43).

38. When checking whether the payload or the predetermined part thereof (13) matches the further digital signature (43), subjecting the payload or the predetermined portion thereof (13) to a further hash function to obtain a further hash value; 38. Apparatus according to claim 37, configured to check whether the further hash value matches the further digital signature (43).

39. 39. Apparatus according to claim 37 or 38, wherein the predetermined portion (13) excludes the indication of the external resource including the editor's track.

40. 40. The apparatus of any of claims 37 to 39, wherein the predetermined portion (13) of the payload of the SEI message includes a unique identifier.

41. 41. The apparatus of claim 37, wherein the syntax structure further comprises a media component identifier, and wherein the apparatus is configured to use the media component identifier to select the further digital signature (43) from a set of one or more digital signatures included in the external resource.

42. The syntax structure is: an indication of the hash function (31); 42. An apparatus according to any of claims 37 to 41, further comprising one or more of: an indication of portions of the video data stream for which digital signatures are available to verify the authenticity of the video data stream.

43. An apparatus (17) for transcoding a video data stream in which video is encoded, comprising: receiving an input video data stream (14') and checking (15') the authenticity of said input video data stream (14'); transcoding (12) the input video data stream (14') to generate an output data stream (14); applying (31) a predetermined portion (13) of said output video data stream (14) or of the data (62) from which said output video data stream is derived to a hash function (31) to obtain a hash value (33); Signing (71) the hash value using a private key (58) of an asymmetric cryptosystem to obtain a digital signature (43); providing an editor's track (231) of the output video data stream with a content provider's certificate (233) for the editor's track provided to an external resource (280), the certificate including or pointing to a public key (57) of the asymmetric encryption scheme; configured to provide said digital signature (43) to said output video data stream (14) or to said external resource (280) or to a further external resource; The video data stream may include the video encoding said video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

44. 44. The apparatus of claim 43, configured to insert into the output video data stream an indication of an external resource that includes an editor's track for the video data stream, and to provide in the output video data stream a syntax element indicating that the output video data stream includes the indication of the external resource that includes the editor's track.

45. 45. The apparatus of claim 44, wherein the syntax elements and, if present, the indication of the external resources including the editor's track are transmitted in a syntax structure of the output video data stream.

46. and configured to insert the indication of the external resource, including the editor's track, in a syntax structure of the output video data stream, the device comprising: subjecting the payload of the syntax structure, or a predetermined portion thereof (13), to a further hash function to obtain a further hash value; 46. ​​Apparatus according to any of claims 43 to 45, configured to store the further hash value in the external resource.

47. and configured to insert the indication of the external resource, including the editor's track, in a syntax structure of the output video data stream, the device comprising: subjecting the payload of the syntax structure, or a predetermined portion thereof (13), to a further hash function to obtain a further hash value; obtaining a further digital signature using said private key to sign said syntax structure, or said payload of said syntax structure, or said predetermined portion thereof (13); 47. The apparatus of claim 46, configured to store the further digital signature in the external resource.

48. 48. Apparatus according to claim 46 or 47, wherein the predetermined portion (13) excludes the indication of the external resource including the editor's track.

49. 49. The apparatus of any of claims 46 to 48, wherein the predetermined portion (13) of the payload of the SEI message includes a unique identifier.

50. 50. An apparatus according to any one of claims 46 to 49, configured to insert a media component identifier into the syntax structure, the apparatus being configured to store an association between the further digital signature or the further hash value and the media component identifier in the external resource.

51. an indication of the hash function (31); 51. An apparatus according to any one of claims 46 to 50, configured to insert into the syntax structure one or more indications of portions of the output video data stream for which digital signatures are available for verifying the authenticity of the output video data stream.

52. An apparatus (15) for rendering a video data stream (14) in which the video is coded in a way that allows it to be checked for authenticity, said apparatus comprising: applying a predetermined portion (13) of the video data stream (14) or a predetermined portion (13) of the data (62) from which the video data stream (14) is derived to a hash function (31) to obtain a hash value (33); signing said hash value (33) using a private key (58) of an asymmetric cryptosystem to obtain a digital signature (43); providing a content provider certificate (233) to an editor's track (231) of the output video data stream, the content provider certificate (233) including or pointing to a public key of the asymmetric encryption scheme, the content provider certificate (233) being provided to an external resource (280) in the editor's track (231), the content provider certificate including or pointing to a public key of the asymmetric encryption scheme; configured to provide said digital signature (43) to said output video data stream or to said external resource (280) or to a further external resource; The video data stream may include the video encoding said video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

53. 53. The apparatus of claim 52, configured to insert into the video data stream an indication of an external resource that includes an editor's track for the video data stream, and to provide in the video data stream a syntax element indicating that the video data stream includes the indication of the external resource that includes the editor's track.

54. 54. The apparatus of claim 53, wherein the syntax elements and, if present, the indication of the external resources including the editor's track are transmitted in a syntax structure of the video data stream.

55. and configured to insert the indication of the external resource, including the editor's track, in a syntax structure of the video data stream, the device comprising: subjecting the payload of the syntax structure, or a predetermined portion thereof (13), to a further hash function to obtain a further hash value; 55. Apparatus according to any of claims 52 to 54, configured to store the further hash value in the external resource.

56. and configured to insert the indication of the external resource, including the editor's track, in a syntax structure of the video data stream, the device comprising: subjecting the payload of the syntax structure, or a predetermined portion thereof (13), to a further hash function to obtain a further hash value; obtaining a further digital signature using said private key to sign said syntax structure, or said payload of said syntax structure, or said predetermined portion thereof (13); 56. The apparatus of claim 55, configured to store the further digital signature in the external resource.

57. 57. Apparatus according to claim 55 or 56, wherein the predetermined portion (13) excludes the indication of the external resource including the editor's track.

58. 58. The apparatus of any of claims 55 to 57, wherein the predetermined portion (13) of the payload of the SEI message includes a unique identifier.

59. 59. An apparatus according to any one of claims 55 to 58, configured to insert a media component identifier into the syntax structure, the apparatus being configured to store an association between the further digital signature or the further hash value and the media component identifier in the external resource.

60. 60. An apparatus as claimed in any one of claims 55 to 59, configured to insert into the syntax structure one or more of an indication of the hash function (31), an indication of several parts of the video data stream for which a digital signature is available to verify the authenticity of the video data stream.

61. 1. An apparatus for checking a video data stream (14) in which video is encoded with respect to authenticity, said apparatus comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); configured to check whether the hash value (33) matches a digital signature (43) to determine whether the video data stream is authentic; the apparatus comprising: a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; a combination of the temporal layer identifier and the layer identifier; a time frame identifier, a priority level identifier indicating the priority level of the picture; configured to determine the predetermined portion (13) based on one or more of an AVC nal_ref_id; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

62. 62. Apparatus according to claim 61, configured to derive from the video data stream indications of how to determine the predetermined portion (13).

63. The instructions indicating the manner in which the predetermined portion (13) is determined are: an indication associated with a time frame of said video data stream, indicating whether said time frame belongs to said predetermined portion (13); the temporal layer identifiers associated with pictures of the video data stream, the temporal layer identifiers identifying a subset of temporal frames of the video data stream to which the respective pictures belong; the one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; the combination of the time layer identifier and the layer identifier; the time frame identifier; the priority level identifier indicating the priority level of the picture; 63. The apparatus of claim 62, wherein one or more of the AVC nal_ref_ids are differentiated.

64. 64. Apparatus according to claim 62 or 63, configured to derive the indication indicating the manner in which the predetermined portion (13) is determined from a syntax structure.

65. The syntax structure is: an indication of the hash function (31); 65. The apparatus of claim 64, further comprising one or more of: an indication of portions of the video data stream, for which portions are provided respective digital signatures for checking whether the video data stream can be trusted.

66. 63. The apparatus of claim 62, wherein the apparatus is configured to determine the predetermined portion (13) based on the temporal layer identifier, the layer identifier, or the time frame identifier, and the apparatus is configured to derive a range of values ​​from the video data stream, the range of values ​​indicating the value of the respective identifier, which value is associated with the predetermined portion (13).

67. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, said device comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); checking whether said hash value (33) matches a digital signature (43) to determine whether said video data stream is authentic; deriving instructions from said video data stream, said instructions being configured to indicate how said predetermined portion (13) is determined; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

68. 68. The apparatus of claim 67, wherein the indication is a syntax element that distinguishes between a plurality of modes for determining the predetermined portion.

69. The plurality of modes includes a first mode, and the device:

69. The device of claim 68, configured to, in the first mode, determine the predetermined portion by determining whether to include a given picture of the video data stream in the predetermined portion depending on which layer of a plurality of layers of the video data stream the given picture belongs to.

70. The plurality of modes further comprises a second mode, and the apparatus 70. The apparatus of claim 69, configured to, in the second mode, determine the predetermined portion by including the predetermined picture of the video data stream in the predetermined portion.

71. 71. The apparatus of claim 70, wherein the plurality of modes comprises the first mode and the second mode, and wherein the first mode and the second mode are distinguished from each other.

72. 68. The device of claim 67, configured to check the video data stream for authenticity in units of one or more portions, the one or more portions including the predetermined portion, the device configured to determine the one or more portions in a manner indicated by the instruction.

73. The plurality of modes includes a first mode, and the device:

73. The apparatus of claim 72, configured, in the first mode, to allocate the given picture to one of the one or more portions depending on which layer of a plurality of layers of the video data stream the given picture belongs to.

74. The plurality of modes further includes a second mode, and the device:

74. The apparatus of claim 73, configured to, in the second mode, assign the given picture to a given one of the one or more portions.

75. The plurality of modes further includes a second mode, and the device:

74. The apparatus of claim 73, configured, in the second mode, to assign the given picture to a given one of the one or more portions in the absence of dedicated signaling in the video data stream that signals to which of the one or more portions the given picture should be assigned.

76. The instructions indicating the manner in which the predetermined portion (13) is determined are: an indication associated with a time frame of said video data stream, indicating whether said time frame belongs to said predetermined portion (13); a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; a combination of the temporal layer identifier and the layer identifier; a time frame identifier, a priority level identifier indicating the priority level of the picture; 76. The apparatus of claim 67, wherein one or more of the AVC nal_ref_ids are differentiated.

77. 77. Apparatus according to any of claims 67 to 76, configured to derive the indication indicating the manner in which the predetermined portion (13) is determined from a syntax structure.

78. The syntax structure is: an indication of the hash function (31); 78. The apparatus of claim 77, further comprising one or more of: an indication of portions of the video data stream, for which portions are provided respective digital signatures for checking whether the video data stream can be trusted.

79. 79. An apparatus as described in any of claims 67 to 78, wherein the apparatus is configured to determine the predetermined portion (13) based on the temporal layer identifier, the layer identifier, or the time frame identifier, and the apparatus is configured to derive a range of values ​​from the video data stream, the range of values ​​indicating the value of the respective identifier, which value is to be associated with the predetermined portion (13).

80. An apparatus (20) for decoding a video data stream (14) in which video is decoded, comprising: configured to derive from said video data stream a syntax structure containing information for checking authenticity of said video data stream based on a predetermined portion (13) of said video data stream, the syntax structure includes an indication of how to determine the predetermined portion (13) of the video data stream; The device comprises: the video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding the prediction residual data of the residual block from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or decoding a significance map indicating the location of non-zero transform coefficients of the transform block representing the residual block by decoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, decoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and decoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

81. the indication of the manner in which the predetermined portion (13) is determined is a temporal layer identifier associated with a picture of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream, the layer identifiers identifying the layer of the video data stream to which the respective picture belongs; a combination of the temporal layer identifier and the layer identifier; a time frame identifier, A priority level identifier indicating the priority level of the picture, nal_ref_id of AVC, the nal_ref_id of AVC, 81. The apparatus of claim 80, wherein the apparatus distinguishes between one or more of:

82. 82. The apparatus of claim 81, configured to decode the indication indicating the manner in which the predetermined portion (13) is determined from a syntax structure.

83. The syntax structure is: an indication of the hash function (31); 83. The apparatus of claim 82, further comprising one or more of: an indication of portions of the video data stream for which respective digital signatures are provided for checking whether the video data stream can be trusted.

84. 82. The apparatus of claim 81, wherein the apparatus is configured to derive a range of values ​​from the video data stream, the range of values ​​indicating values ​​of the temporal layer identifier, the layer identifier, or the temporal frame identifier associated with the predetermined portion (13).

85. An apparatus (15) for rendering a video data stream (14) in which the video is encoded with respect to reliability, said apparatus comprising: subjecting a predetermined portion (13) of said video data stream or of the data (62) from which said further portion of said video data stream is derived to a hash function (31) to obtain a hash value (33); configured to sign the hash value (33) to obtain a digital signature (43); the apparatus comprising: a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; a combination of the temporal layer identifier and the layer identifier; a time frame identifier, a priority level identifier indicating the priority level of the picture; configured to determine the predetermined portion (13) based on one or more of an AVC nal_ref_id; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

86. 86. Apparatus according to claim 85, configured to insert into said video data stream an indication of how said predetermined portion (13) is determined.

87. The instructions indicating the manner in which the predetermined portion (13) is determined are: said indications associated with time frames of said video data stream, indicating whether said time frames belong to said predetermined portion (13); the temporal layer identifiers associated with pictures of the video data stream, the temporal layer identifiers identifying a subset of temporal frames of the video data stream to which the respective pictures belong; the one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; the combination of the time layer identifier and the layer identifier; the time frame identifier; the priority level identifier indicating the priority level of the picture; 87. The apparatus of claim 86, wherein one or more of the AVC nal_ref_ids are differentiated.

88. 88. Apparatus according to claim 86 or 87, configured to insert the indication indicating the manner in which the predetermined portion (13) is determined into a syntax structure.

89. an indication of the hash function (31); 90. The apparatus of claim 88, further configured to insert into the syntax structure one or more indications of portions of the video data stream for which respective digital signatures are provided for checking whether the video data stream is authentic.

90. 87. The apparatus of claim 86, wherein the apparatus is configured to determine the predetermined portion (13) based on the time layer identifier, the layer identifier, or the time frame identifier, and the apparatus is configured to insert a range of values ​​into the video data stream, the range of values ​​indicating the value of the respective identifier, which value is associated with the predetermined portion (13).

91. An apparatus (15) for rendering a video data stream (14) in which the video is coded in a way that allows it to be checked for authenticity, said apparatus comprising: subjecting a predetermined portion (13) of said video data stream, or of the data (62) from which said further portion of said video data stream is derived, to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43); inserting instructions into said video data stream, said instructions being configured to indicate how said predetermined portion (13) is determined; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

92. 92. The apparatus of claim 91, wherein the indication is a syntax element that distinguishes between a plurality of modes for determining the predetermined portion.

93. The plurality of modes includes a first mode, and the device:

93. The device of claim 92, configured to, in the first mode, determine the predetermined portion by determining whether to include a given picture of the video data stream in the predetermined portion depending on which layer of a plurality of layers of the video data stream the given picture belongs to.

94. The plurality of modes further comprises a second mode, and the apparatus 94. The apparatus of claim 93, configured to, in the second mode, determine the predetermined portion by including the predetermined picture of the video data stream in the predetermined portion.

95. 95. The apparatus of claim 94, wherein the plurality of modes comprises the first mode and the second mode, and wherein the first mode and the second mode are distinguished from each other.

96. 92. The device of claim 91, configured to render the video data stream capable of checking reliability in units of one or more portions, the one or more portions including the predetermined portion, the device configured to select a manner of determining the one or more portions, and to indicate the selected manner of determining the one or more portions of the data stream by the instruction.

97. The plurality of modes includes a first mode, and the device:

97. The apparatus of claim 96, configured, in the first mode, to allocate the given picture to one of the one or more portions depending on which layer of a plurality of layers of the video data stream the given picture belongs to.

98. The plurality of modes further includes a second mode, and the device:

98. The apparatus of claim 97, configured to, in the second mode, assign the given picture to a given one of the one or more portions.

99. The plurality of modes further includes a second mode, and the device:

98. The apparatus of claim 97, configured in the second mode to assign the predetermined picture to a predetermined one of the one or more portions or to provide dedicated signaling that signals which of the one or more portions of the video data stream the predetermined picture is assigned to.

100. The instructions indicating the manner in which the predetermined portion (13) is determined are: said indications associated with time frames of said video data stream, indicating whether said time frames belong to said predetermined portion (13); the temporal layer identifiers associated with pictures of the video data stream, the temporal layer identifiers identifying a subset of temporal frames of the video data stream to which the respective pictures belong; the one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; the combination of the time layer identifier and the layer identifier; the time frame identifier; the priority level identifier indicating the priority level of the picture; 100. The apparatus of claim 91, wherein one or more of the AVC nal_ref_ids are differentiated.

101. 101. Apparatus according to any of claims 91 to 100, configured to insert the indication indicating the manner in which the predetermined portion (13) is determined into a syntax structure.

102. an indication of the hash function (31); 102. The apparatus of claim 101, further configured to insert into the syntax structure one or more indications of portions of the video data stream for which respective digital signatures are provided for checking whether the video data stream is authentic.

103. 103. An apparatus as described in any of claims 91 to 102, wherein the apparatus is configured to determine the predetermined portion (13) based on the time layer identifier, the layer identifier, or the time frame identifier, and the apparatus is configured to insert a range of values ​​into the video data stream, the range of values ​​indicating the value of the respective identifier, which value is associated with the predetermined portion (13).

104. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, said device comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); deriving a digital signature (43) associated with said predetermined portion (13) from an external resource; configured to check whether the hash value (33) matches the digital signature (43) to determine whether the video data stream is authentic; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

105. 105. The apparatus of claim 104, configured to derive a reference to the external resource from the video data stream.

106. decrypting said digital signature (43) using the public key in an asymmetric decryption scheme to obtain a check value; 106. Apparatus according to claim 104 or 105, configured to check whether the hash value (33) matches the check value.

107. deriving a portion identifier from said video data stream, said portion identifier being associated with said predetermined portion (13); using the portion identifier to identify a portion of the check value; 107. The apparatus of claim 106, wherein when checking whether the hash value (33) matches the check value, the apparatus is configured to check whether the hash value (33) matches the portion of the check value.

108. deriving from said video data stream a media component identifier indicating a media type of said predetermined portion (13); 108. The apparatus of claim 107, configured to use the media component identifier to identify the portion of the check value.

109. A device (16) for checking a video data stream (14) in which video is encoded for authenticity, said device comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); deriving a check value associated with said predetermined portion (13) from an external resource; configured to check whether the hash value (33) matches the check value to determine whether the video data stream is authentic; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

110. 110. An apparatus according to claim 109, configured to derive a reference to the external resource from the video data stream.

111. Deriving a digital signature (43) from the external resource; 111. Apparatus according to claim 109 or 110, configured to verify the check value using the digital signature (43).

112. deriving a portion identifier from said video data stream, said portion identifier being associated with said predetermined portion (13); 112. An apparatus according to any of claims 109 to 111, configured to obtain the check value from the external resource using the partial identifier.

113. performing said check of authenticity of said video data stream sequentially on a plurality of portions of said video data stream; and applying a further portion of the video data stream, or data derived therefrom, to the hash function (31) to obtain a further hash value; deriving a further check value associated with the further portion from the external resource; Deriving a digital signature (43) from the external resource; 113. Apparatus according to any of claims 109 to 112, configured to verify the check value and the further check value using the digital signature (43).

114. An apparatus (20) for decoding a video data stream (14) in which video is encoded, comprising: configured to derive from said video data stream a syntax structure containing information for checking authenticity of said video data stream based on a predetermined portion (13) of said video data stream, the syntax structure includes a reference to an external resource for obtaining a digital signature (43) associated with the predetermined portion (13); The device comprises: the video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding the prediction residual data of the residual block from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

115. 115. The apparatus of claim 114, further configured to derive a portion identifier from the video data stream, the portion identifier being associated with the predetermined portion (13), the portion identifier being configured to associate the predetermined portion (13) with one or more digital signatures included in the external resource.

116. 116. Apparatus according to claim 114 or 115, configured to derive a media component identifier from the video data stream, the media component identifier indicating a media type of the predetermined portion (13).

117. An apparatus (20) for decoding a video data stream (14) in which video is encoded, comprising: configured to derive from said video data stream a syntax structure containing information for checking authenticity of said video data stream based on a predetermined portion (13) of said video data stream, the syntax structure includes a reference to an external resource for obtaining a check value associated with the predetermined portion (13); The device comprises: the video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding the prediction residual data of the residual block from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length decoding to decode by using one or more fifth syntax elements to indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic decoding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

118. 118. The apparatus of claim 117, configured to derive a portion identifier from the video data stream, the portion identifier being associated with the predetermined portion (13), the portion identifier associating the predetermined portion (13) with one or more check values ​​contained in the external resource.

119. 119. Apparatus according to claim 117 or 118, configured to derive a media component identifier from the video data stream, the media component identifier indicating a media type of the predetermined portion (13).

120. An apparatus (15) for rendering a video data stream (14) in which the video is coded in a way that allows it to be checked for authenticity, said apparatus comprising: subjecting the predetermined portion (13) of the video data stream, or of the data (62) from which the predetermined portion (13) of the video data stream is derived, to a hash function (31) to obtain a hash value (33); signing the hash value (33) to obtain a digital signature (43), and providing the digital signature (43) to an external resource; configured to insert an indication of the external resource into the video data stream; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

121. 121. The apparatus of claim 120, configured to sign the hash value (33) using a private key of an asymmetric decryption scheme to obtain the digital signature (43).

122. inserting a portion identifier into said video data stream, said portion identifier being associated with said predetermined portion (13); 122. Apparatus according to claim 120 or 121, configured to provide, at the external resource, an association between the partial identifier and the digital signature (43).

123. inserting a media component identifier into the video data stream; the media component identifier indicates a media type of the predetermined portion (13); 123. An apparatus according to any of claims 120 to 122, configured to provide, at the external resource, an association between the media component identifier and the digital signature (43).

124. An apparatus (15) for rendering a video data stream (14) in which the video is coded in a way that allows it to be checked for authenticity, said apparatus comprising: subjecting a predetermined portion (13) of said video data stream, or of the data (62) from which said further portion of said video data stream is derived, to a hash function (31) to obtain a hash value (33); signing the hash value (33) to obtain a digital signature (43), and providing the hash value (33) and the digital signature (43) to an external resource; configured to insert an indication of the external resource into the video data stream; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

125. inserting a portion identifier into said video data stream, said portion identifier being associated with said predetermined portion (13); 125. The apparatus of claim 124, configured to provide, at the external resource, an association between the partial identifier and the digital signature (43).

126. performing a rendering of the video data stream in a manner that allows for an authenticity check of the video data stream for a plurality of portions of the video data stream in turn, and subjecting further portions of the video data stream or further portions of data (62) from which the further portions of the video data stream are derived to the hash function (31) to obtain further hash values; signing the hash value (33) and the further hash value together to obtain the digital signature (43), and providing the hash value (33), the further hash value, and the digital signature to an external resource; 126. Apparatus according to claim 124 or 125, configured to insert an indication of the external resource into the video data stream.

127. 127. Apparatus according to any of claims 124 to 126, wherein the hash value (33) depends on all bits of the predetermined portion (13) of the video data stream.

128. 128. Apparatus according to any of claims 124 to 127, wherein the hash value (33) depends on all bits of the predetermined portion (13) of the video data stream in the coded domain.

129. the predetermined portion (13) of the video data stream extends across multiple access units of the video data stream such that the hash value (33) depends on bits of the multiple access units; or 129. Apparatus according to any of claims 124 to 128, wherein the predetermined portion (13) comprises video data of only one access unit.

130. 130. Apparatus according to any one of claims 124 to 129, wherein the apparatus is a decoder for decoding the video data stream.

131. The device comprises: the video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding the prediction residual data of the residual block from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or decoding a significance map indicating the location of non-zero transform coefficients of the transform block representing the residual block by decoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, decoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and 131. The apparatus of claim 130, by using context-adaptive binary arithmetic coding by reversing the forward scan order and decoding values ​​of the non-zero transform coefficients in sequential reverse scan order.

132. 127. The apparatus of any of claims 20 to 33, 52 to 60, 85 to 90, 120 to 123, 91 to 103, and 124 to 126, wherein the apparatus is an encoder for encoding the video data stream.

133. The device comprises: encoding the video into the video data stream using block-based predictive coding; an encoder configured to perform transform-based residual coding by encoding the prediction residual data of the residual block into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and 127. The apparatus of any of claims 20-33, 52-60, 85-90, 120-123, 91-103, and 124-126, by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in a sequential reverse scan order.

134. A method (16) for checking a video data stream (14) in which video is encoded with respect to authenticity, said method comprising: applying (31) a predetermined portion (13) of the video data stream or of data derived therefrom to a hash function (31) to obtain a hash value (33); obtaining (51) a unique identifier (45) that uniquely identifies the media asset to which the predetermined portion (13) belongs; obtaining (51) a digital signature (43) based on said video data stream; checking (41) whether the combination of the hash value (33) and the unique identifier (45) matches the digital signature (43) to determine whether the video data stream is authentic; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

135. A method (20) for decoding a video data stream in which video is encoded, said method comprising: decoding (21) a syntax structure from the video data stream, which contains information for checking the authenticity of the video data stream based on a predetermined portion (13) of the video data stream, which predetermined portion is subjected to a hash function (31) or is used to derive data to be subjected to the hash function (31), deriving a hash value (33) which serves to check the authenticity of the video data stream; decoding (21) from the video data stream a unique identifier (45) or a reference pointing to said unique identifier (45), said unique identifier (45) uniquely identifying the media asset to which said predetermined portion (13) belongs; decrypting (21) an indication of a digital signature (43) from the video data stream, the digital signature (43) being configured based on a combination of the hash value (33) and the unique identifier (45); The device comprises: configured to decode the video from the video data stream by block-based prediction and transform-based residual decoding by decoding the predictive residual data of the residual blocks from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number followed by a one indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients along a scan order; one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and one or more fifth syntax elements indicating a position of the non-zero transform coefficients along the scan order by indicating a number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or decoding a significance map indicating the location of non-zero transform coefficients of the transform block representing the residual block by decoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, decoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic decoding by reversing the forward scan order and decoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

136. A method (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, said method comprising: applying (31) a predetermined portion (13) of said video data stream (14) or of the data (62) from which said video data stream (14) is derived to a hash function (31) to obtain a hash value (33); assigning a unique identifier (45) to the predetermined portion (13) that uniquely identifies the media asset to which the predetermined portion (13) belongs; signing (71) the combination of said hash value (33) and said unique identifier (45) to obtain a digital signature (43); The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

137. A method (16) for checking a video data stream (14) in which video is encoded with respect to authenticity, said method comprising: applying (31) a predetermined portion (13) of said video data stream, or data derived therefrom (62), to a hash function (31) to obtain a hash value (33); checking (41) whether said hash value (33) matches a digital signature (43) to determine whether said video data stream is authentic; decrypting (46) said digital signature (43) using a public key (57) in an asymmetric decryption scheme to obtain a check value (47); checking (49) whether said hash value (33) matches said check value (47); The method includes checking whether the video data stream includes an indication (55) of an external resource (280) that includes an editor's track (231) of the video data stream, and if the video data stream includes an indication of an external resource that includes an editor's track of the video data stream, querying the editor's track for a certificate (233) of a content provider that is the last editor of the video data stream, and deriving the public key based on the certificate of the content provider; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

138. A method (17) for transcoding a video data stream in which video is encoded, said method comprising: receiving an input video data stream (14') and checking (15') the authenticity of said input video data stream (14'); transcoding (12) the input video data stream (14') to generate an output data stream (14); applying (31) a predetermined portion (13) of said output video data stream (14) or of the data (62) from which said output video data stream is derived to a hash function (31) to obtain a hash value (33); Signing (71) the hash value using a private key (58) of an asymmetric cryptosystem to obtain a digital signature (43); providing an editor's track (231) of the output video data stream with a content provider's certificate (233) for the editor's track provided to an external resource (280), the certificate including or pointing to a public key (57) of the asymmetric encryption scheme; providing said digital signature (43) to said output video data stream (14) or to said external resource (280) or to a further external resource; The video data stream may include the video encoding said video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

139. A method (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, said method comprising: applying a predetermined portion (13) of the video data stream (14) or data (62) derived from the video data stream (14) to a hash function (31) to obtain a hash value (33); signing said hash value (33) using a private key (58) of an asymmetric cryptosystem to obtain a digital signature (43); providing a content provider certificate (233) to an editor's track (231) of the output video data stream, the content provider certificate (233) including or pointing to a public key of the asymmetric encryption scheme, the content provider certificate (233) being provided to an external resource (280) in the editor's track (231), the content provider certificate including or pointing to a public key of the asymmetric encryption scheme; providing said digital signature (43) to said output video data stream or to said external resource (280) or to a further external resource; The video data stream may include the video encoding said video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

140. A method for checking a video data stream (14) in which video is encoded with respect to authenticity, said method comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); checking whether the hash value (33) matches a digital signature (43) to determine whether the video data stream is authentic; The method includes the steps of: a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; a combination of the temporal layer identifier and the layer identifier; a time frame identifier, a priority level identifier indicating the priority level of the picture; determining the predetermined portion (13) based on one or more of AVC nal_ref_id; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

141. A method (16) for checking a video data stream (14) in which video is encoded with respect to authenticity, said method comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); checking whether said hash value (33) matches a digital signature (43) to determine whether said video data stream is authentic; deriving instructions from said video data stream, said instructions indicating how to determine said predetermined portion (13); The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

142. A method (20) for decoding a video-encoded video data stream (14), said method comprising: deriving from said video data stream a syntax structure containing information for checking authenticity of said video data stream based on a predetermined portion (13) of said video data stream, the syntax structure includes an indication of how to determine the predetermined portion (13) of the video data stream; The device comprises: the video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding the prediction residual data of the residual block from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

143. A method (15) for rendering a video data stream (14) in which the video is encoded with respect to reliability, said method comprising: subjecting a predetermined portion (13) of said video data stream or of the data (62) from which said further portion of said video data stream is derived to a hash function (31) to obtain a hash value (33); signing the hash value (33) to obtain a digital signature (43); The method includes the steps of: a temporal layer identifier associated with pictures of the video data stream, the temporal layer identifier identifying a subset of temporal frames of the video data stream to which the respective picture belongs; one or more layer identifiers associated with pictures of the video data stream, the one or more layer identifiers identifying the layer of the video data stream to which the respective picture belongs; a combination of the temporal layer identifier and the layer identifier; a time frame identifier, a priority level identifier indicating the priority level of the picture; determining the predetermined portion (13) based on one or more of AVC nal_ref_id; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

144. A method (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, said method comprising: subjecting a predetermined portion (13) of said video data stream, or of the data (62) from which said further portion of said video data stream is derived, to a hash function (31) to obtain a hash value (33); Signing the hash value (33) to obtain a digital signature (43); inserting instructions into said video data stream, said instructions indicating how said predetermined portion (13) is determined; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

145. A method (16) for checking a video data stream (14) in which video is encoded with respect to authenticity, said method comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); deriving a digital signature (43) associated with said predetermined portion (13) from an external resource; checking whether the hash value (33) matches the digital signature (43) to determine whether the video data stream is authentic; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

146. A method (16) for checking a video data stream (14) in which video is encoded with respect to authenticity, said method comprising: applying a predetermined portion (13) of said video data stream or data derived therefrom to a hash function (31) to obtain a hash value (33); deriving a check value associated with said predetermined portion (13) from an external resource; checking whether the hash value (33) matches the check value to determine whether the video data stream is authentic; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

147. A method (20) for decoding a video-encoded video data stream (14), comprising: The method comprises deriving from the video data stream a syntax structure comprising information for checking authenticity of the video data stream based on a predetermined portion (13) of the video data stream, the syntax structure includes a reference to an external resource for obtaining a digital signature (43) associated with the predetermined portion (13); The device comprises: the video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding the prediction residual data of the residual block from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

148. A method (20) for decoding a video-encoded video data stream (14), comprising: The method comprises: deriving from said video data stream a syntax structure containing information for checking authenticity of said video data stream based on a predetermined portion (13) of said video data stream, the syntax structure includes a reference to an external resource for obtaining a check value associated with the predetermined portion (13); The device comprises: the video is decoded from the video data stream by block-based prediction; configured to perform transform-based residual decoding by decoding the prediction residual data of the residual block from the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

149. A method (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, said method comprising: subjecting the predetermined portion (13) of the video data stream, or of the data (62) from which the predetermined portion (13) of the video data stream is derived, to a hash function (31) to obtain a hash value (33); signing the hash value (33) to obtain a digital signature (43), and providing the digital signature (43) to an external resource; inserting an indication of the external resource into the video data stream; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

150. A method (15) for rendering a video data stream (14) in which the video is coded in a way that is checkable with respect to authenticity, said method comprising: subjecting a predetermined portion (13) of said video data stream, or of the data (62) from which said further portion of said video data stream is derived, to a hash function (31) to obtain a hash value (33); signing the hash value (33) to obtain the digital signature (43), and providing the hash value (33) and the digital signature (43) to an external resource; inserting an indication of the external resource into the video data stream; The video data stream encoding the video into the video data stream using block-based predictive coding; the video is encoded using transform-based residual coding by encoding the prediction residual data of the residual blocks into the video data stream; a first syntax element indicating a total number of non-zero transform coefficients of a transform block representing the residual block; and a number of trailing ones indicating the number of non-zero transform coefficients having an absolute value of one when traversing the coefficients in scan order. one or more second syntax elements indicating the signs of the non-zero transform coefficients having an absolute value of one when traversing the coefficients along the scan order; one or more third syntax elements indicating values ​​of the non-zero transform coefficients, excluding the number of non-zero transform coefficients that have an absolute value of one when traversing the coefficients along the scan order; a fourth syntax element indicating a total number of zero-valued transform coefficient levels for the transform block, beginning with a first encountered non-zero transform coefficient in the scan order or later; and by using context-adaptive variable length coding by using one or more fifth syntax elements that indicate the position of the non-zero transform coefficients along the scan order by indicating the number of consecutive zero-valued transform coefficients in the scan order between consecutively encountered non-zero transform coefficients; or encoding a significance map indicating the locations of non-zero transform coefficients of the transform block representing the residual block by encoding a significance flag indicating whether a non-zero transform coefficient is at a current position in a forward scan traversing the transform coefficients of the transform block, and if so and the current position is not the end of the forward scan, encoding a last significance flag indicating whether the non-zero transform coefficient at the current position is the last non-zero transform coefficient in the forward scan order; and by using context-adaptive binary arithmetic coding by reversing the forward scan order and encoding the values ​​of the non-zero transform coefficients in sequential reverse scan order.

151. 1. A method for storing video, said method comprising:

151. A method comprising storing a data stream on a digital storage medium, the data stream being generated by the method of any of claims 136, 139, 143, 144, 149, or 150.

152. 151. A method for transmitting a data stream generated by the method of any of claims 136, 139, 143, 144, 149, or 150.

153. 153. A computer program for performing the method according to any of claims 134 to 152 when the computer program is run on a computer or signal processor.

154. A video data stream produced by the method of any of claims 136, 139, 143, 149, or 144, 150.