Editable video data signed on uncompressed data
Patent Information
- Application Number
- JP2023205832
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-06
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-12-06
AI Technical Summary
Existing methods for editing signed video bitstreams require re-encoding and re-signing the entire video sequence, which is computationally intensive and inefficient, especially when only a small portion of the video needs to be edited.
A method for editing signed video bitstreams that minimizes the need for re-signing by maintaining a one-to-one relationship between data units and macroblocks, allowing for localized editing and re-encoding only the affected areas, while preserving the digital signature through the use of archive objects and fingerprints.
This approach reduces computational overhead and maintains data security by limiting the need for re-signing, ensuring efficient and secure editing of video bitstreams with minimal disruption to the original signature.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to the field of security devices for protecting video data from unauthorized activity, particularly in connection with data storage and transmission. The present disclosure proposes methods and devices for editing signed video bitstreams and verifying signed video bitstreams that may result from such editing. [Background technology]
[0002] Digital signatures can verify the authenticity or integrity of a message and ensure non-repudiation by providing a layer of verification and security to digital messages transmitted over insecure channels. With particular reference to video coding, there are secure and highly efficient methods described in the prior art for digitally signing predictively coded video sequences. See, for example, our previous published patent applications EP 4164173 and EP 4164230. See also US 20140010366, which proposes a cryptographic video verification technique specifically adapted for predictively coded video data with a group of pictures structure.
[0003] A video sequence may need to be edited after it has been signed. In addition to visual improvements, the edits can aim to ensure privacy protection by cropping, masking, blurring, or similar image processing that makes visual features less recognizable. In most available methods, this requires re-encoding and re-signing the entire edited frame. The re-encoding and re-signing is preferably extended to several neighboring frames as well, so as not to disturb predictive coding dependencies (inter- / intra-frame references) that may exist, even if the neighboring frames are not directly affected by the edits. These steps can consume significant computational resources and introduce annoying delays for the user.
[0004] US7437007 discloses a method for performing a region of interest edit of a video stream in the compressed domain. The compressed video stream comprises compressed video stream frames representing video stream frames having an unwanted portion and a region of interest portion. According to the method, the compressed video stream frames are edited to modify said unwanted portion while maintaining the original structure of the video stream and to obtain compressed video stream frames including said region of interest portion. To achieve this, the edit comprises skipping macroblocks located above, below and to the right of said region of interest portion of predictively coded (P) frames and bidirectionally predictively coded (B) frames. The video stream under consideration in US7437007 is not a signed video stream. Summary of the Invention
[0005] One object of the present disclosure is to make available a method for editing a signed video bitstream obtained by predictive coding of a video sequence, which largely avoids the need to re-sign the bitstream outside the parts affected by the edits, as is the case in some available methods. A particular object is to make available such a video editing method, which preserves the signatures of all macroblocks except those that are edited. A further object is to enable video editing without significantly compromising the data security of the original signed video bitstream. A further object is to provide a method for verifying a signed video bitstream obtained by predictive coding of a video sequence. It is a still further object to provide a device and a computer program for these purposes.
[0006] At least some of these objects are achieved by the invention as defined by the independent claims. The dependent claims relate to advantageous embodiments of the invention.
[0007] In a first aspect of the present disclosure, a method for editing a signed video bitstream obtained by predictive coding of a video sequence is provided. The signed video bitstream shall include data units and associated signature units. Each data unit represents (i.e. encodes) at most one macroblock in a video frame of the predictively coded video sequence. Each signature unit includes a digital signature of a bit string derived from a number of fingerprints, each fingerprint calculated from a macroblock reconstructed from one data unit associated with the signature unit. The signature unit may optionally include the bit string to which the digital signature pertains ("document approach"). For a signed video bitstream having these characteristics, the method includes receiving a request to replace a region of at least one video frame; reconstructing a first set of macroblocks in which the region is contained and a second set of macroblocks that directly or indirectly reference the macroblocks in the first set; adding an archival object to the signed video bitstream, the archival object including a fingerprint calculated from the reconstructed first set of macroblocks; editing the first set of macroblocks in accordance with the request to replace and encoding the edited first set of macroblocks as a first set of new data units; re-encoding the second set of macroblocks as a second set of new data units; and adding the first and second sets of new data units to the signed video bitstream.
[0008] Due to the one-to-one relationship between data units and macroblocks, a data unit does not represent multiple macroblocks (nor parts of multiple macroblocks), and therefore the direct effect of editing some macroblocks is limited to one or more data units (first set) of the edited macroblock. Furthermore, since each fingerprint is calculated from a macroblock reconstructed from one associated data unit, the need to re-sign data units after editing is limited. More precisely, the method according to the first aspect preserves any predictive coding dependencies connecting pairs or groups of macroblocks, i.e. by re-encoding a set (second set) of data units representing macroblocks that directly or indirectly reference the edited macroblock. If the re-encoding is limited to this second set of data units, the method makes efficient use of the available computational capacity. This allows the method to be executed with good performance on conventional processing devices.
[0009] The method according to the first aspect includes a further advantage on the receiver side. Due to the archive object from which the fingerprints associated with the first set of data units can be obtained, the receiver can verify all data units of the signed video bitstream that have not been affected by the editing. This allows a significant portion of the existing signatures to be preserved, and in this respect the video editing method according to the first aspect is said to be minimally destructive. Verification on the receiver side is described in detail within the second aspect of the present disclosure. Importantly, the fingerprints associated with the second set of macroblocks do not need to be archived, since re-encoding is an operation that should not change the visual appearance of these macroblocks.
[0010] In some embodiments, the data security of the edited signed video bitstream is improved by adding to the bitstream one or more signature units associated with the first set of edited macroblocks. This avoids scenarios of unauthorized modification of the first set of edited macroblocks. The one or more signature units may be new signature units added to the bitstream or edited (alternative) versions of signature units included in the signed video bitstream before editing. The integrity of the second set of macroblocks is protected by the digital signature already present on the signature unit in the bitstream. Indeed, the receiver can verify the integrity of the second set of macroblocks using the existing digital signature, since the fingerprint is calculated from the reconstructed macroblocks and the re-encoding operation should negligibly or not at all change their visual appearance.
[0011] In some embodiments, the archive object further includes the locations of the first set of macroblocks. The locations may refer to the locations of the macroblocks within the frame, for example in frame coordinates. If static macroblock partitioning is used, the locations of the macroblocks may be expressed as macroblock sequence numbers or another identifier. This provides one way of assisting a recipient of an edited video bitstream to determine whether a particular macroblock has been altered and therefore to select an appropriate method of obtaining a fingerprint of the data unit representing that macroblock.
[0012] In some embodiments, the second set of macroblocks is re-encoded losslessly, thereby obtaining new data units. The use of lossless encoding ensures that macroblocks subsequently reconstructed from these new data units do not differ from the second set of macroblocks. As a result, the new data units remain consistent with the signature units already present in the video bitstream before editing, since the fingerprint is calculated from the reconstructed pixel / plaintext data. In other embodiments, the second set of macroblocks is re-encoded using reduced data compression, and the fingerprint of the second set of macroblocks in the signed video bitstream comprises a robust hash. Robust hashing refers to a class of algorithms that have a tolerance such that the algorithm accepts a data set as authentic even if it is subject to small differences. The algorithm may be optimized for hashing image or video data. The tolerance of the algorithm may be configurable and may be set such that errors corresponding to tampering are detected, while normal errors expected from data compression are not detectable (they are small differences in the above sense).
[0013] In some embodiments, the second set of macroblocks is re-encoded non-predictively. Non-predictive encoding may correspond to using only I-frames. This is a simple and robust way of re-encoding the second set of macroblocks. If editing is assumed to be relatively infrequent, the additional memory or bitrate costs are unlikely to be significant. In other embodiments, the second set of macroblocks is re-encoded predictively with reference to the first set of edited macroblocks. For example, the same GoP structure can be preserved for continuity. However, these embodiments may not perform very well in terms of data compression, since they include an attempt to predict across discontinuities introduced by editing. If the discontinuities are large, the B or P frames may be larger than usual, or the quality may be locally degraded.
[0014] In some embodiments, the first set of edited macroblocks is losslessly and / or encoded using reduced data compression and / or non-predictively. Each of these data compressions limits or avoids further loss of image quality beyond that caused by the original image encoding operation. The additional computational and / or memory costs for this encoding are likely justified because the edited macroblocks are likely to be used or studied more carefully than the rest of the video sequence. Furthermore, in important use cases involving video surveillance, the size of the edited portion is typically negligible compared to the amount of data generated by continuous surveillance.
[0015] In some embodiments, each time a fingerprint is calculated in connection with encoding a macroblock, a fingerprinting operation is performed on image data (e.g., plaintext data, pixel data) in a reference buffer of the encoder. This can be implemented in an encoder device that prepares a signed video bitstream. It can also be implemented in an editing tool that performs the editing method according to the first aspect. Since a reference buffer is a necessary component of most encoder implementations (i.e., to ensure correct predictive encoding), the reconstructed video data used for fingerprinting can be obtained without additional computational cost.
[0016] In some embodiments, the second set of macroblocks is non-predictively re-encoded, while in other embodiments, the second set of macroblocks is predictively re-encoded with reference to the first set of edited macroblocks.
[0017] According to a generalization of the first aspect, there is provided a video editing method performed on a signed video sequence comprising data units and associated signature units. Each macroblock in a video frame of a predictively coded video sequence is coded by at most M data units. Edits made to a macroblock 110 directly affect at most M data units 120, where M is a small integer such as 1, 2, 3, 4, 5 or at most 10. Each signature unit comprises a digital signature of a bit string derived from a number of fingerprints, each fingerprint calculated from a macroblock reconstructed from each of at most M associated data units, which optionally also comprise the bit string. The video editing method includes receiving a request to replace a region of at least one video frame, reconstructing a first set of macroblocks in which the region is included and a second set of macroblocks that directly or indirectly reference the macroblocks in the first set, adding to a signed video bitstream an archival object including a fingerprint calculated from the reconstructed first set of macroblocks, editing the first set of macroblocks in accordance with the replacement request, re-encoding the second set of macroblocks, and adding the first and second sets of new data units thus obtained to the signed video bitstream.
[0018] According to a further generalization of the first aspect, there is provided a video editing method performed on a signed video sequence comprising data units and associated signature units. Each data unit represents (encodes) at most N macroblocks in a video frame of the predictively coded video sequence. This means that the effect of editing performed on a macroblock is not limited to the macroblock itself, but may require re-encoding and / or re-signing one or more data units that the edited macroblock shares with further macroblocks, where N is a small integer such as 1, 2, 3, 4, 5, or at most 10. Each signature unit comprises a digital signature of a bit string derived from a number of fingerprints, each associated with exactly one associated data unit, and optionally a bit string. The video editing method includes receiving a request to replace a region of at least one video frame, reconstructing a first set of macroblocks in which the region is included and a second set of macroblocks that directly or indirectly reference the macroblocks in the first set, adding to a signed video bitstream an archival object including a fingerprint calculated from the reconstructed first set of macroblocks, editing the first set of macroblocks in accordance with the replacement request, re-encoding the second set of macroblocks, and adding the first and second sets of new data units thus obtained to the signed video bitstream.
[0019] In a second aspect of the present disclosure, a method for verifying a signed video bitstream obtained by predictive coding of a video sequence is provided. It is understood that the signed video bitstream includes data units, signature units respectively associated with some of the data units, and an archival object. Each data unit represents one macroblock in a frame of the predictively coded video sequence. Each signature unit includes a bit string and optionally a digital signature of the bit string itself. It is expected that the bit string is at least partially verified by performing the verification method, that the bit string is derived from a plurality of fingerprints, and further, each fingerprint is calculated from a macroblock reconstructed from one data unit associated with the signature unit. Finally, the archival object includes at least one archived fingerprint. The method for verifying the signed video bitstream includes reconstructing macroblocks from data units associated with the signature unit, calculating respective fingerprints from at least some of the reconstructed macroblocks, obtaining the at least one archived fingerprint from the archival object, deriving a bit string from the calculated and obtained fingerprint, and verifying the data unit associated with the signature unit using the digital signature in the signature unit.
[0020] Verification of the data unit may involve verifying the derived bit string using the digital signature. Alternatively (the "document approach"), the verification step involves verifying the bit string in the signature unit using the digital signature and then comparing the derived bit string with the verified bit string.
[0021] Although the archival object may have been added by performing the editing method according to the first aspect, the method according to the second aspect can be implemented without reliable knowledge of such pre-processing. Thus, the method according to the second aspect achieves verification of the authenticity of the video sequence in that it verifies that the digital signature (and any bit sequence) carried in the signature unit actually matches the fingerprint for the associated data unit. Thus, the data unit may not have been modified, since this means that the corresponding reconstructed macroblock has been modified.
[0022] The method according to the second aspect includes two options for obtaining fingerprints: by direct calculation from reconstructed macroblocks or by recovery from the archive object. This supports a process that minimizes the destruction of existing fingerprints during the editing phase (first aspect). The fact that each data unit represents one macroblock tends to limit the number of fingerprints that need to be archived for a given replacement request, thus limiting the size of the archive object.
[0023] In some embodiments, the archive object further indicates the location of the macroblock represented by the data unit to which the archived fingerprint relates. In other words, the archived fingerprint is a fingerprint of a macroblock reconstructed from the data unit, the location of the macroblock being indicated in the archive object. During execution of the method according to the second aspect, to obtain a fingerprint for the data unit, it is determined whether to calculate the fingerprint based on the location indicated by the archive object or to obtain the fingerprint from the archive object.
[0024] As with the first aspect of the present disclosure, the second aspect can be generalized to the cases discussed above, i.e., numerical ratios of macroblocks to corresponding data units of 1:M or N:1.
[0025] The third aspect of the present disclosure relates to devices configured to perform the methods of the first and / or second aspects. These devices may be incorporated into systems with different main purposes (e.g., video recording, video content management, video playback), or may be dedicated to the aforementioned editing and verification, respectively. The devices within the third aspect of the present disclosure generally share the effects and advantages of the first and second aspects, which may be embodied with a comparable degree of technical variation.
[0026] In a fourth aspect, a signed video bitstream includes data units and associated signature units, each data unit representing at most one macroblock in a video frame of a predictively coded video sequence, and each signature unit includes a digital signature s(H1) of a bit string H1 derived from a plurality of fingerprints, each fingerprint h1, h2, ... being calculated from a macroblock reconstructed from one or more data units associated with the signature unit. In particular, a signed video bitstream includes data units and associated signature units, each data unit representing exactly one macroblock in a video frame of a predictively coded video sequence, and each signature unit includes a digital signature s(H1) of a bit string H1 derived from a plurality of fingerprints, each fingerprint h1, h2, ... being calculated from a macroblock reconstructed from exactly one data unit associated with the signature unit. As mentioned above, when a portion of a video bitstream is expected to be edited at a later point in time, a video bitstream having this format can provide certain advantages, including the ability to reuse existing signature units for purposes of verifying the unedited portion. Since the signature is applied to the decrypted image data, there is no need to handle dependencies between macroblocks. Yet another advantage is that fingerprinting (e.g., hashing) can be implemented at a very fine granularity, such as at the level of a single pixel.
[0027] The invention further relates to a computer program comprising instructions for making a computer execute the above method. The computer program can be stored or distributed on a data carrier. As used herein, a "data carrier" may be a transitory data carrier, such as a modulated electromagnetic or light wave, or a non-transient data carrier. Non-transient data carriers include volatile and non-volatile memories, such as permanent and non-permanent storage media of magnetic, optical, or solid-state type. Still within the scope of a "data carrier", such memories may be fixedly attached or mobile.
[0028] In a further aspect of the disclosure, a signed video bitstream is provided that includes data units and associated signature units. Each data unit represents at most N macroblocks in a video frame of a predictively coded video sequence, where N is a small integer such as 1, 2, 3, 4, 5, or at most 10. Each signature unit includes a digital signature of a bit string derived from multiple fingerprints, each associated with exactly one associated data unit, and optionally the bit string itself. The signed video bitstream is edit-friendly because the direct effect of editing one macroblock is limited to at most N data units of the edited macroblock, and each fingerprint is a fingerprint of exactly one associated data unit. This limits the propagation of edits to a limited number of data units, resulting in fewer data units needing to be re-signed after editing.
[0029] It should be noted that a "macroblock" as used in this disclosure may advantageously be a coded macroblock. However, the invention is also applicable to non-predictively coded video, and in a more generalized case, a macroblock can therefore be any contiguous group of pixels. Since the signature of the video is performed on the decoded frames, the grouping of pixels need not be limited to any coded group division.
[0030] In general, all terms used in the claims should be interpreted according to their ordinary meaning in the art, unless expressly specified otherwise herein. "All references to a / an / the element, apparatus, component, means, step, etc. should be interpreted without limitation as referring to at least one example of the element, apparatus, component, means, step, etc., unless expressly stated. The steps of any method disclosed herein do not have to be performed in the exact order described, unless expressly stated.
[0031] Aspects and embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief description of the drawings]
[0032] [Figure 1] 2 illustrates two example patterns of intraframe referencing within macroblocks in a video frame. [Diagram 2] 1 shows a sequence of frames representing one Group of Pictures (GoP) with an exemplary pattern of inter-frame referencing. [Diagram 3] 4 illustrates four exemplary correspondence patterns between data units and the macroblocks they represent. [Figure 4] It shows an editing operation that directly affects some macroblocks in the first frame (first column) and leads to consequent changes in further frames (second and third columns) until the end of the GoP. [Diagram 5] Shown from top to bottom are a signed video bitstream, macroblocks reconstructed from the bitstream, and the effects of edit operations performed on portions of the video bitstream. [Figure 6] 1 is a flowchart of a method for editing a signed video bitstream according to an embodiment of the present specification. [Figure 7]1 is a flowchart of a method for verifying a signed video bitstream according to an embodiment herein. [Figure 8] A device suitable for carrying out the methods shown in Figures 6 and 7 is shown. [Figure 9] Several such devices are shown connected via local area and / or wide area networks. [Figure 10A] (Document Approach) The signed video bitstream and some operations in the method shown in FIG. [Figure 10B] Illustrates a signed video bitstream and some operations within the method shown in FIG. [Figure 11] 6 and 7. The signal processing operations and data flows occurring during the execution of the methods shown in FIGS. 6 and 7 are shown, as well as suitable functional units for performing the signal processing operations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0033] Aspects of the present invention are described more fully below with reference to the accompanying drawings, in which specific embodiments of the invention are shown. These aspects may, however, be embodied in many different forms and should not be construed as limiting, but rather these embodiments are provided as examples so that this disclosure will be thorough and complete, and will fully convey the scope of all aspects of the present invention to those skilled in the art. Like numerals refer to like elements throughout the description.
[0034] In the terminology of this disclosure, a "video bitstream" includes any substantially linear data structure that may resemble a sequence of bit values. A video bitstream may be carried by a transitory medium (e.g., modulated electromagnetic or optical waves), as in some streaming use cases, or the video bitstream may be stored in a non-transitory medium, such as volatile or non-volatile memory.
[0035] A video bitstream represents a video sequence, which may be understood to be a sequence of video frames played sequentially at a nominal time interval. Each video frame may be divided into macroblocks. Further in this disclosure, a "macroblock" may be a block having a transform block or a predictive block, or both, within a frame of a video sequence. The use of "frame" and "macroblock" herein is intended to be consistent with the H.26x video coding standard or similar specifications. A "macroblock" may further be a coding block. As mentioned above, although the term macroblock is used, the grouping of pixels need not be limited to any partition used for coding, but rather, a macroblock may be any group of adjacent pixels. When applied in the case of prediction-based coding, it may be advantageous to use the same partition used for coding.
[0036] 1A and 1B show an exemplary partition of a video frame 100 into a 4×4 uniform arrangement of macroblocks 110. It should be noted that this is a simplification for illustrative purposes. In practice, a video frame 100 is typically divided into a much larger number of macroblocks, e.g., a macroblock may be 8×8 pixels, 16×16 pixels, or 64×64 pixels. Curly arrows are used consistently herein to indicate intraframe or interframe references used in predictive coding. Without departing from the scope of this disclosure, the partition into macroblocks seen in FIGS. 1A and 1B can be significantly modified to include non-square arrangements and / or arrangements of macroblocks 110 that are not homogenous to the video frame 100 and / or mixed arrangements of different macroblocks 110 having different sizes or shapes. It is understood that some video coding formats support dynamic macroblock partitioning, i.e., the partitions may be different for different video frames in a sequence. This is true, for example, for H.265.
[0037] In FIG. 1A, the pattern of intraframe references is limited to a single row of macroblocks 110. In fact, each macroblock 110 points to the macroblock immediately to its left (if present in frame 100) in the sense that it is represented by a data unit that predictively represents image data in the macroblock, i.e., image data in this macroblock is represented relative to image data in the macroblock to the left. Conceptually, with some simplification, the data unit represents the macroblock 110 in terms of change or movement relative to the left macroblock. Another possible understanding is that the data unit represents the correction of a given predictive operation that derives the macroblock 110 from the left macroblock. In alternative examples within the scope of this disclosure, the macroblock 110 may point to another macroblock 110 to its right, or above or below it.
[0038] In FIG. 1B, the pattern of intraframe references is denser. Here, each macroblock 110 points to its immediate left macroblock (if present in frame 100) and its immediate above macroblock (if present in frame 100). Thus, the macroblock 110 is represented by a data unit that predictively represents image data in the macroblock, e.g., in terms of differences or corrections of predictions of this image data, based on a predetermined interpolation operation (or some other predetermined combination) that operates on image data in the left and above macroblocks. Interpolation may include post-processing operations such as smoothing. A further alternative within the scope of the present disclosure is to use intraframe references with a direction opposite to that shown in FIG. 1B, i.e., starting from the bottom row.
[0039] Image / video formats with a predefined pattern of intraframe references can be associated with a specified scan order, which represents a feasible sequence for decoding macroblocks. In Fig. 1A, the macroblock scan order is not unique, i.e., because each row can be decoded independently. In the case of the pattern according to Fig. 1B, macroblocks can be decoded row-wise from the top or column-wise from the left, and video formats with this reference pattern can specify a column-wise or row-wise scan order, so that any reconstruction errors can be anticipated at the encoder side, which benefits the coding efficiency. Furthermore, video formats with arbitrary macroblock orders or so-called slicing exist.
[0040] FIG. 2 illustrates a video sequence V that includes a sequence of frames 100. In the video sequence V, there are independently decodable frames (I-frames) and predicted frames, including unidirectionally predicted frames (P-frames) and bidirectionally predicted frames (B-frames). The International Telecommunications Union's ITU-T Recommendation H.264 (06 / 2019) "Advanced video coding for generic audiovisual services" specifies a video coding standard in which both forward and bidirectionally predicted frames are used. As can be seen in FIG. 2, the independently decodable frames do not reference any other frames. The unidirectionally predicted frames (P-frames) of FIG. 2 are forward predicted in that they directly reference at least one other preceding or immediately preceding frame. The bidirectionally predicted frames (B-frames) may further directly reference a subsequent or immediately succeeding frame in the video sequence V. A first frame indirectly references a second frame if the video sequence includes a third frame (or a subsequence of frames) that the first frame directly references and directly references the second frame. In predictive video coding, a Group of Pictures (GoP) is defined as a subsequence of video frames that does not reference any video frames outside the subsequence, and that can be decoded without referencing any other I, P, or B frames. The video frames in Figure 2 form a GoP. The GoP in Figure 2 is the smallest because it cannot be subdivided into further GoPs.
[0041] In a simpler implementation, a video sequence V may consist of only independently decodable frames (I) and unidirectionally predicted frames (P). Such a video sequence may have the appearance IPPIPPPPIPPPIPPP, with each P frame pointing to the previous I or P frame. In this example, GoPs such as IPP, IPPPP, IPPP, IPPP, etc. may be identified.
[0042] There are several options for adjusting inter- and intra-frame predictive coding. For example, if static (fixed) macroblock partitions are used in all video frames, inter-frame references such as the illustrated ones may be defined at the level of one macroblock position at a time (e.g., the top left macroblock in Fig. 1A and Fig. 1B). Some video formats allow dynamic macroblock partitions, e.g., a macroblock can be predicted from a corresponding pixel in a preceding video frame or from a spatially shifted pixel in a preceding frame. Alternatively, inter-frame references are defined for the entire video frame. I-frames may consist of only I-blocks, and P-frames may consist of only P-blocks or a mix of I- and P-blocks.
[0043] With reference now to FIG. 3, attention is directed to a video bitstream encoding the video sequence under consideration. Sub-FIGS. 3A, 3B, 3C, and 3D show different correspondence patterns between the data units 120 and the macroblocks 110 that they represent (encode). For the purposes of this disclosure, the data units may have any suitable format and structure, and no assumption is made other than that the data units can be separated (or extracted) from the video bitstream, e.g., to enable processing, without the need to decode that data unit or any surrounding data units. In addition to the data units, the signed video bitstream further includes signature units that are separable from the signed video bitstream in the same or similar manner. More details regarding signature units are presented below with reference to FIG. 5.
[0044] Under one option, as shown in Figure 3A, data units 120 correspond one-to-one to macroblocks 110 (dashed lines). This correspondence pattern ensures that each macroblock 110 is always reconstructed from one data unit 120. Furthermore, when a macroblock 110 is edited (which may in fact require changes to signature units, metadata, etc.), no other data units 120 other than the corresponding data unit 120 need be modified.
[0045] Alternatively, as shown in Figure 3B, each macroblock 110 is coded by multiple data units 120, where each data unit 120 represents at most one macroblock 110. Thus, if each macroblock 110 is coded by at most M data units 120, then any macroblock 110 can be reconstructed from at most M data units 120. Edits made to a macroblock 110 directly affect at most M data units 120, where M is a small integer such as 1, 2, 3, 4, 5, or at most 10.
[0046] In a further option, as shown in FIG. 3C, each data unit 120 encodes multiple macroblocks 110. This means that the effect of editing performed on a macroblock 110 is not limited to the macroblock 110 itself, but may require re-encoding and / or re-signing data units 120 that the edited macroblock 110 shares with further macroblocks 110. This correspondence pattern, although quite feasible to implement, is not applied in the highest performance embodiment of this disclosure, since performing re-signing on an unnecessarily large data set may be computationally wasteful. If a data unit 120 is allowed to be shared by at most a predefined number N of macroblocks 110, the total amount of computation added may remain limited. Again, N may be specified to be a small integer, such as 1, 2, 3, 4, 5, or at most 10.
[0047] Combinations of the patterns found in Figures 3B and 3C are possible within the scope of this disclosure. As a result, the techniques proposed herein may be applied to video bitstreams in which the ratio of macroblocks 110 to data units 120 may be 2:1, 1:1, 2:2, 3:3, 4:2, 2:4, or 4:4.
[0048] In yet another alternative, as shown in Figure 3D, each data unit 120 can represent any number of macroblocks 110 in a video sequence, and each macroblock 110 can be encoded by any number of data units 120. This correspondence pattern, as suggested by the dashed line, can mean that in the worst case, even limited editing operations on macroblocks 110 require a complete re-encoding and re-signing of the video sequence. The techniques disclosed herein should not be implemented on video sequences having the structure shown in Figure 3D.
[0049] FIG. 4 illustrates the editing operations described below.
[0050] FIG. 5 illustrates a section of a video sequence V including a series of macroblocks 110 belonging to one or more video frames. For example, the macroblocks 110 may occupy fixed positions in successive video frames (e.g., the top left macroblock in FIG. 1A and FIG. 1B). The macroblocks 110 are assumed to be predictively coded according to the references shown by the curly arrows. It can be seen that the first five macroblocks 110 belong to one GoP and the following five macroblocks 110 belong to the next GoP. The video sequence V is coded as a signed video bitstream B including data units 120 and signature units 130. For illustrative purposes and not for limitation, FIG. 5 illustrates the data units such that the data units 120 code the video macroblocks 110 according to the corresponding pattern shown in FIG. 3A, i.e., the one-to-one relationship between the macroblocks 110 and the data units 120. In other words, each data unit 120.n has been created by applying the encoder Enc to the corresponding macroblock 110.n. The data units 120 may conform to a proprietary or standardized video encoding format, such as ITU-T H.264, H.265, or AV1. Bitstream B may further include additional types of units (e.g., dedicated metadata units) without departing from the scope of this disclosure.
[0051] Each of the signature units 130 may be associated with multiple data units 120. It is understood that in FIG. 5, the data units 120 between two consecutive signature units 130 are associated with the subsequent signature unit 130, which is not an essential feature of the present invention and other arrangements are possible without departing from the scope of the present disclosure. A signature unit 130 may be associated with a set of data units 120 that are all included in one GoP, although other association patterns are possible. Furthermore, the set of data units 120 associated with one signature unit 130 is preferably selected taking into account the applicable macroblock scanning order. For example, the set of data units 120 associated with a signature unit 130 may represent the number of macroblocks that should be sequentially scanned during decoding, thereby minimizing the number of macroblocks that need to be re-referenced if a signature unit 130 fails to verify.
[0052] The signature unit 130 includes at least one bit string (e.g., H1) and a digital signature of that bit string (e.g., s(H1)). As suggested by the use of dashed lines, the presence of the bit string is optional. If the signature unit 130 includes multiple bit strings, the signature unit 130 can have one digital signature for all of these bit strings, or multiple digital signatures, each for a single bit string or each for a subgroup of bit strings. The bit string on which the digital signature is formed may be a combination of fingerprints calculated on the basis of the macroblock 111 reconstructed from the data unit 120 associated with the signature unit 130, or the bit string may be a fingerprint of said combination of fingerprints. More precisely, the fingerprint is a fingerprint of the reconstructed macroblock, which may be obtained by reading a so-called reference buffer in the encoder or by performing an independent decoding operation. The combination of fingerprints (or "references") may be a list or other concatenation of string representations of fingerprints. In the ITU-T H.264 and H.265 formats, the signature unit may be included as a Supplemental Enhancement Information (SEI) message in the video bitstream. In the AV1 standard, the signature may be included in a Metadata Open Bitstream Unit (OBU).
[0053] Each fingerprint may be a hash or a salted hash. A salted hash may be a hash of a combination of a data unit (or a portion of a data unit) and a cryptographic salt, the presence of the salt may prevent a malicious party with access to multiple hashes from guessing which hash function is being used. Potentially useful cryptographic salts include the value of an active internal counter, a random number, and the time and location of the signature. The hash may be generated by a hash function (or one-way function) h, which is a cryptographic function that provides a level of security that is deemed appropriate given the confidentiality of the video data being signed and / or the value that would be compromised if the video data were manipulated by an unauthorized party. Three examples are SHA-256, SHA3-512 and RSA-1204. The hash function shall be predefined (e.g., reproducible) so that the fingerprint can be regenerated when the recipient is attempting to verify it. In the example of FIG. 5, the bit string is given by: H1=h([h1, h2, h3, h4, h5]) and H2 = h([h6, h7, h8, h9, h 10 ]), where h1, h2, ... are hashes of macroblocks 111.1, 111.2, ... and [·] represents concatenation. The concatenation operation may be linear (juxtaposition) or may provide a staggered arrangement of the data. The successive operations may further include arithmetic operations on the data, such as bitwise OR, XOR, multiplication, division, or modulo operations. An exemplary salted hash may be defined as follows: TIFF2024086623000002.tif5170 or In the first example, the hash function h has a parametric dependence on the second argument, which is assigned a salt σ.
[0054] In some embodiments, each of the fingerprints h1, h2... is calculated from macroblocks (e.g., pixel values or other plain data) reconstructed from data unit 120. The fingerprints are calculated as h1=h(Y 111.1 ) or h1=h([Y 111.1 , σ]) or h1 = h(Y 111.1 , σ), where Y 111.1 represents data from the first one of the reconstructed macroblocks 111, and σ is an optional cryptographic salt. In a third option, the hash function h has a parametric dependence on the second argument to which the salt σ is assigned. The fingerprints can be calculated from the entire macroblock or from a subset thereof extracted according to a pre-agreed rule. In a variant of these embodiments, the fingerprints h1, h2, ... are calculated not at the plaintext level but from intermediate reconstructed data derived from the data units. More precisely, when an encoder is used that includes a frequency domain transform (e.g. DCT, DST, DFT, wavelet transform) followed by a coding process (e.g. entropy, Huffman, Lempel-Ziv, run-length, binary or non-binary arithmetic coding, e.g. Context Adaptive Variable Length Coding, CAVLC, Context Adaptive Binary Arithmetic Coding, CABAC), the transform coefficients are usually available as intermediate reconstructed data at the decoder side. The transform coefficients can be restored from the coded representation. If the encoder further comprises a quantization process immediately downstream of the transform, the quantized transform coefficients become available at the decoder side. In more complex codecs, at a greater number of successive processing stages, there may be further types of intermediate reconstructed data, which can be used for the fingerprint calculation. It is particularly convenient to use types of intermediate reconstructed data that appear identically in the encoding process, as well as the quantized transform coefficients. Common to all the embodiments considered in this paragraph, the fingerprint relates to exactly one data unit 120, one of the data units 120 associated with the signature unit 130.
[0055] Optionally, the fingerprints can be linked together in sequence to detect unauthorized removal or insertion of data units. That is, each fingerprint depends on the next or previous fingerprint, e.g., the input to the hash includes the hash of the next or previous fingerprint. The linking can be done, for example, as follows: h1=h(Y 111.1 ), h2=h([h1, Y 111.2 ]), h3=h([h2, Y 111.3 ]), and Y 111.1 , Y 111.2 , Y 111.3 denotes data from the first, second, and third macroblocks of the reconstructed macroblock 111. Another way to link the fingerprints is to use h1 = h(Y 111.1 ), h 12 =h([Y 111.1 , Y 111.2 ]), h 13 =h([Y 111.2 , Y 111.3 ]) etc.
[0056] With further reference to signature unit 130 of Figure 5, a cryptographic element (not shown) with a pre-stored private key can be utilized to generate the digital signature s(H1). The recipient of the signed video bitstream can be expected to hold a public key belonging to the same key pair (see also Figure 10), allowing the recipient to verify that the signature generated by the cryptographic element is authentic but not generate a new signature. The public key can also be included in the signed video bitstream as metadata, in which case there is no need to store it at the recipient's end.
[0057] Next, a method 600 for editing a signed video bitstream B obtained by predictive coding of a video sequence V is described with reference to Fig. 6. It is assumed that the non-optional steps of the method 600 are performed after the original signing of the video bitstream. For example, if the signed video bitstream is originally generated on a recording device, the editing method 600 can be performed on a video management system (VMS). Another exemplary use case is when the signed video bitstream is generated on a device, stored in memory, and then re-referenced for editing using the same device. Editing may also be performed at a later time, for example, after the need to perform privacy masking is known.
[0058] As mentioned above, a device performing the editing method 600 may be a dedicated application or system, but may have the basic functional structure shown in FIG. 8. As shown, the device 800 includes a processing circuit 810, a memory 820, and an external interface 830. The memory 820 may be suitable for storing a computer program 821 having instructions for implementing the editing method 600. The external interface 830 may be a communication interface that allows the device 800 to communicate with a similar device (not shown) maintained by a recipient and / or a video content author (e.g., a recording device), or may enable read and write operations in an external memory 890 suitable for storing a video bitstream.
[0059] Figure 9 illustrates the case where the bitstream is transferred between multiple devices. Note that the device performing the editing method 600 may be connected to the recipient device via a local area network (connecting lines in the lower half of Figure 9) or a wide area network 990. An attack on the bitstream B can occur in either type of network, which justifies the signature.
[0060] Returning to Figure 6, one embodiment of method 600 begins with step 612 of receiving a request to replace a region of at least one video frame 100 in a video sequence V. The request may be received via a human machine interface or in an automated manner, for example in a message from a controlling application running on the same device 800 or remotely. The region to be replaced may be a set of replacement pixel values, such as a privacy mask, for replacing similarly located pixels in the video sequence V.
[0061] For the avoidance of doubt, it should be noted that a video sequence V to be edited is encoded by predictive coding as a signed video bitstream B comprising data units 120 and associated signature units 130, each data unit representing at most one macroblock 110 in a video frame 100 of the predictively coded video sequence V, and each signature unit comprising a digital signature of a bit string derived from a number of fingerprints, each associated with exactly one associated data unit. Such a bitstream format is illustrated with reference to FIG.
[0062] In the next step 614 of the method 600, a first set of macroblocks in which the aforementioned region is included and a second set of macroblocks that directly or indirectly reference the macroblocks in the first set are determined and reconstructed. In FIG. 5, the reconstruction corresponds to an arrow symbolizing a decoding operation Dec. Recalling that bidirectionally predicted frames (B-frames) may be defined in some video coding formats, it is understood that the second set of macroblocks may be located before or after the first set of macroblocks, or may occupy both of these positions. It is understood that the first and second sets are defined to be mutually disjoint. For example, it may be stipulated that a macroblock belongs to the second set only if it does not belong to the first set, i.e. if this macroblock is not needed to form the set of macroblocks that includes the region to be replaced. Thus, if the first set of macroblocks extends to the GoP boundary, the second set of macroblocks is usually empty. Furthermore, it is understood that the second set of macroblocks may contain macroblocks in more than one P-frame or more than two B-frames, since additional frames may be used as references for the region to be replaced, depending on the video encoder initially used. In step 614, as a result of inter-frame or intra-frame prediction references between macroblocks, it may be necessary to reconstruct more than just the first and second sets of macroblocks. More precisely, one or more macroblocks located earlier in the chain of prediction references leading to the first and second sets of macroblocks may have to be reconstructed first.
[0063] If the region to be replaced is limited to a single video frame, the first set of macroblocks can be determined with reference only to the macroblock partitions of the frame. More precisely, the first set are all macroblocks with which the region overlaps (i.e., macroblocks with which the region has a non-empty intersection in pixel space). If the region extends to multiple frames, this operation is repeated for each frame. In the special case where the region is repeated identically in all of the video frames and furthermore the macroblock partitions are constant across all such frames, the first set of macroblocks is a copy of that determined (by the overlap criterion) for the initial frame for each of the subsequent frames. The second set of macroblocks can be determined based on the first set and the pattern of intra-frame and inter-frame references in the signed predictive coded video sequence. Since such references by definition do not extend beyond GoP boundaries, the search for macroblocks to be included in the second set can be limited to the GoP or those GoPs to which the first set of macroblocks belongs.
[0064] A possible result of step 614 is illustrated in FIG. 4, where each column represents one video frame of the video sequence V and each row represents one macroblock at a particular position in the frame (e.g., the top left macroblock). In FIG. 4, furthermore, the references between macroblocks are shown as curly arrows, and the boundary between two consecutive GoPs, GoP1 and GoP2, is shown as a dashed vertical line. It should be noted that the inter-frame references are defined at the level of one macroblock position in FIG. 4. Furthermore, the diagonally hashed macroblocks are the macroblocks that are directly affected by the request to replace a region, which are all located in the first frame and form the first set of macroblocks 401. The macroblocks with dotted shading are all the macroblocks that directly (second frame) or indirectly (third frame) reference macroblocks in the first set, which are identified as the second set of macroblocks 402. In accordance with expectations, the second set of macroblocks does not extend beyond the GoP boundary.
[0065] It should be noted that the arrangement of the first and second sets of macroblocks seen in Figure 4 may, but need not necessarily, be changed when intra-frame referencing is introduced: for example, if a macroblock position corresponding to a first row references a macroblock position corresponding to a second row, the first and second sets of macroblocks 401, 402 remain unchanged.
[0066] In FIG. 5, the first set consists of macroblocks 111.2 and 111.3, and the second set consists of macroblocks 111.4 and 111.5.
[0067] In the next step 616, the archival object 140 is added to the signed video bitstream B. The archival object 140 includes the fingerprints h2, h3 calculated from the first set of reconstructed macroblocks. At the level of the signed video bitstream B, the archival object 140 can have a similar format to the data units 120 and the signature units 130, in that the archival object 140 can be separated from the video bitstream without decoding. The fingerprints h2, h3 are not necessarily calculated by the entity executing the method 600. Indeed, in the case of a bitstream B according to the "document approach" in which the signature units 130 include bit strings, these fingerprints are already available from one of the signature units 130. It should further be noted that in the case of a signature unit 130 in the bitstream B including such bit strings, the step 616 can be performed as soon as the first set of macroblocks is determined, i.e. before the completion of the step 614.
[0068] Optionally, each archival object 140 may include digital signatures of these fingerprints, or of a combination of these fingerprints in this archival object 140, or may include digital signatures of the fingerprints of the combination. Further optionally, the archival object 140 may also include the locations of the first and second sets of macroblocks whose signatures are archived. The locations may refer to the locations of the macroblocks in the frame, for example in frame coordinates, which correspond to the locations in the bit string. If static macroblock partitions are used, the locations of the macroblocks 111 may be represented as macroblock sequence numbers or another identifier. For example, the bit string may be formed by concatenating the fingerprints in the same order as the macroblock sequence in the frame.
[0069] A further step 618 of the method 600 is illustrated with reference to the lower left part of Fig. 5. In this step, an edit operation Edit transforms the first set of macroblocks 111.2, 111.3 into an edited first set of macroblocks 112.2, 112.3. The editing of the first set of macroblocks is in accordance with a request to replace an area of a video frame. Then, in a further step 618, the first set of edited macroblocks 112.2, 112.3 is coded as a first set of new data units 121.2, 121.3.
[0070] The first set of edited macroblocks 112.2, 112.3 may be encoded using normal encoder settings, normal GoP patterns, etc., i.e., in the same way as the video bitstream B was created. Optionally, the first set of edited macroblocks 112.2, 112.3 resulting after substitution are instead encoded as independently decodable data units. Each independently decodable data unit may correspond to an I-frame in the H.264 or H.265 encoding specification, a coded macroblock that does not reference another macroblock, or an equivalent data unit. This is consistent with the inventors' recognition that substitution introduces abrupt temporal changes in a video sequence. In particular, the edited macroblock 112.2 is likely to be significantly different from the immediately preceding unedited macroblock 111.1, which may degrade the performance of predictive coding. A further option is to encode the first set of edited macroblocks 112.2, 112.3 losslessly and / or using reduced data compression. This should be understood against the background that the video sequence V is encoded with a predefined normal level of data compression. More precisely, the first set of edited macroblocks 112.2, 112.3 are expected to be encoded with a reduced data compression level compared to the normal data compression.
[0071] Step 620 of method 600 is optional and will be described separately.
[0072] In the next non-optional step 622, the second set of macroblocks 112.4, 112.5 are re-encoded as a second set of new data units 121.4, 121.5. The re-encoding should preferably minimally change the visual appearance of the second set of macroblocks 112.4, 112.5. Ideally, the image data (e.g. pixel data or other plaintext data) obtained by decoding the new data units 121.4, 121.5 should be identical or visually inseparable from the second set of macroblocks 112.4, 112.5. However, to achieve this, the re-encoding operation in step 622 may modify the predictive coding settings and / or modify the coding process. In particular, the coding process may be modified in terms of the level of data compression, with lossy coding (when normal data compression occurs) being replaced by lossless coding (reduced data compression) or lossless coding. Lossless encoding may involve representing the second set of macroblocks 112.4, 112.5 as unencoded "raw" blocks, such as a list of the original values for each position in the macroblock in the appropriate color space. If some type of lossy encoding is used for the second set of macroblocks 112.4, 112.5, it may be advantageous to combine this with a robust hash, and in particular a robust hash verification. In this way, macroblocks reconstructed from the new data units 121.4, 121.5 are accepted as authentic with respect to the original signature unit 130, even if the image quality of these macroblocks is slightly degraded.
[0073] With regard to changing predictive coding settings, in important use cases (e.g. masking, blurring) it may be advantageous to use non-predictive coding. Again, the substitution introduces an abrupt time change to the video sequence in that the edited macroblock 112.3 is likely to be significantly different from the immediately following unedited macroblock 112.4, which may degrade coding performance if predictive coding is applied. If coding per performance is not a primary concern, or if the editing operations are of a less obtrusive nature (e.g. filtering, enhancement), predictive coding can be used. For predictive reference (curly arrows in FIG. 5), here the image data in the second set of macroblocks 112.4, 112.5 is expressed relative to the image data in the first set of edited macroblocks 112.2, 112.3 that would normally require updating. As a result, the content of the second set of new data units 121.4, 121.5 is different from the data units 120.4, 120.5.
[0074] Then, in step 624, the first and second sets of new data units 121.2, 121.3, 121.4, 121.5 are added to the signed video bitstream B. At the same time, the corresponding original data units 120.2, 120.3, 120.4, 120.5 may be removed from the video bitstream B. In some embodiments, step 624 is the last operation in the editing method 600.
[0075] In some embodiments, the method 600 uses fingerprints calculated from the first set of edited macroblocks 112.2, 112.3. The fingerprint further includes a step 620 of appending TIFF2024086623000004.tif5170 to the video bitstream B. The fingerprint is calculated by adding the signature unit 130 associated with the data units 120.2, 120.3 that encode the first set of macroblocks to at least the calculated fingerprint TIFF2024086623000005.tif5170 may be supplemented by replacing it with a substitute signature unit (not shown) that includes a digital signature of a bit string derived from the TIFF2024086623000005.tif5170. For example, the bit string may be the calculated fingerprint The fingerprint of the TIFF2024086623000006.tif5170 and one or more unedited macroblocks can be derived, so that the complete frame 100 or a portion can be conveniently verified using a single signature unit. Alternative signature units can be obtained by editing an existing signature unit, in particular by extending it with a further digital signature. Alternatively, the calculated fingerprint At least one new signature unit 131 containing a bit string derived from the digital signature of TIFF2024086623000007.tif5170 is added to the video bitstream B. The new signature unit 131 may have the same structure as the signature unit 130 described above.
[0076] Optionally (the "document approach"), the substitute signature unit and the new signature unit 131 may further comprise the digitally signed bit string itself.
[0077] Note that it is typically not necessary to compute and include fingerprints for the second set of macroblocks, since these remain susceptible to verification using appropriate ones of the existing signature units 130 in the video bitstream B. Optional step 620 may be performed at any point in the method 600 after the edited first set of macroblocks 112.2, 112.3 is available.
[0078] In yet other embodiments, method 600 further includes an initial step 610 of providing at least one signature unit 130. It is understood that in use cases believed to be of primary interest, step 610 is performed by a different entity than steps 612, 614, 616, 618, 620, 622, and 624 of method 600 and / or step 610 is performed at an earlier point in time. In either approach, step 610 is separated from subsequent steps 612, 614, 616, 618, 620, 622, and 624 by a relatively insecure data transfer and / or a retention period that justifies the signature to ensure a desired level of data security.
[0079] Optional step 610 may include substep 610.1 of reconstructing a number of macroblocks from the respective data units 120 associated with the signature unit, substep 610.2 of calculating a number of fingerprints from the respective reconstructed macroblocks, substep 610.3 of deriving a bit string from the calculated fingerprints, which is a combination of the number of fingerprints or a fingerprint of the combination, and substep 610.4 of obtaining a digital signature of the bit string. Suitable implementations of fingerprint calculation 610.2, bit string derivation 610.3, and digital signature 610.4 are described in detail above. In particular, the bit string to which the digital signature in the signature unit 130 relates may be a combination of the fingerprints of the associated data units 120, or may be a fingerprint of the combination of the fingerprints of the associated data units 120. The combination (or "document") may be a list or another concatenation of the respective string representations of the fingerprints.
[0080] Having now described the editing method 600, attention is now directed to the receiver side. More precisely, a method 700 for verifying a signed video bitstream B is described with reference to the flow chart of FIG. 7. It is again assumed that the signed video bitstream B has been obtained by predictive coding of a video sequence V and optionally by subsequent editing operations. It is not essential that the signed video bitstream B has been processed according to the editing method 600. It is further assumed that the signed video bitstream comprises data units 120, associated signature units 130 and an archive object 140, where each data unit 120 represents one macroblock 110 in a frame 100 of the predictively coded video sequence V, each signature unit 130 comprises a digital signature (e.g. s(H1), s(H2)) of a bit string (e.g. H1, H2) and optionally the bit string itself, and the archive object 140 comprises at least one fingerprint, which may be an archived fingerprint for a data unit not currently present in the bitstream B and / or which has undergone editing. Irrelevant to the verification method 700, it is typically not possible for the recipient to determine whether a particular signature unit 130 was added in connection with an edit (e.g., by the editing method 600) or whether it was part of the original, unedited bitstream B.
[0081] In a first step 710 of the method 700, a macroblock 113 is reconstructed from data units 120, 121 associated with a signature unit 130. The reconstruction involves a decoding process symbolized by the top row of downward arrows in Figures 10A and 10B.
[0082] The next step 712 of the method 700 is to extract from at least some of the reconstructed macroblocks 113 their respective fingerprints. TIFF2024086623000008.tif5170 is calculated. Because of the editing in method 600, no fingerprints are calculated from the second and third macroblocks 113.2, 113.3 (first set) which are different from the corresponding original macroblocks in the video sequence V. The fact that the second and third macroblocks 113.2, 113.3 belong to the first set may be indicated in metadata in the edited video bitstream B or may be evident from the encoder timestamps on the corresponding data units 121.1, 121.3. Yet another option may be to include the original edit request (a request to replace a region of a video frame) in the edited video bitstream B, from which a recipient can determine which macroblocks are changed.
[0083] In a third step 714, which in principle could be performed before or overlapping with step 712, at least one archived fingerprint is obtained from archive object 140. In the example shown in Figures 10A and 10B, fingerprints h2, h3 for the second and third macroblocks 113.2, 113.3 are obtained in this way. Note that only TIFF2024086623000009.tif5170 can fail the validation in the following fourth step 716, in which case failure would suggest that incorrect manipulation of bitstream B has taken place.
[0084] In step 716, after the fingerprints for all macroblocks have been obtained in steps 712 and 714, a bit string is generated from these fingerprints: TIFF2024086623000010.tif5170 is derived. This can be done according to pre-agreed rules, for example by a procedure similar to that described in step 610.3. It is recalled that the bit string may be a combination of the obtained fingerprints or a fingerprint of said combination of fingerprints.
[0085] In a next step 718, the data units 120, 121 associated with the signature unit 130 are verified using the digital signature s(H1) in the signature unit 130. For the avoidance of doubt, it should be noted that the verification in step 718 of the data units is indirect, without any necessary processing acting on the data units themselves.
[0086] In embodiments where the signature unit 130 does not include the bit string H1, step 718 generates a bit string derived using the digital signature s(H1). This can be done by examining TIFF2024086623000011.tif5170. For example, the derived bit sequence TIFF2024086623000012.tif5170 can be verified using a public key that belongs to the same key pair as the private key used to generate the digital signature s(H1). In FIG. 10B, this is the derived bit string This is demonstrated by supplying TIFF2024086623000013.tif5170 and the digital signature s(H1) to a public key-stored cryptographic entity 1001 which outputs a binary result W1 representing the outcome of the verification.
[0087] Alternatively, in embodiments where the signature unit 130 includes a bit string H1 ("document approach"), step 718 may be performed in two sub-steps. In a first sub-step, the bit string H1 is verified as described above, for example using a public key belonging to the same key pair as the private key used to generate the digital signature s(H1). This is illustrated in FIG. 10A by function block 1001 and binary result V2. Then, in a second sub-step, the verified bit string H1 is compared to the bit string derived in step 716. TIFF2024086623000014.tif5170. The comparison may be a bitwise equality check, as suggested by function block 1002 in Figure 10A, resulting in an output V2 of a true or false value. If both results V1, V2 are true, then it can be concluded that the signed video bitstream 100 is authentic as far as this signing unit 130 is concerned.
[0088] Execution of method 700 may then include repeating the relevant ones of steps 710, 712, 714, 716, 718 described above for any additional signature units 130 in signed video bitstream B. If the result is positive for all signature units 130, then signed video bitstream B is concluded to be valid and it may be consumed or processed further. If not, signed video bitstream B shall be considered as inauthentic and may be quarantined from further use or processing.
[0089] As previously noted, the steps of any method disclosed herein do not have to be performed in the exact order described unless explicitly stated. This is particularly illustrated by verification method 700, where it is clear that step 714 could be performed before, between, or after steps 710 and 712, as desired.
[0090] Note that verification of the data units in the first set is based on a different trust relationship than verification of the data units in the second set: the data units in the first set are verified by trusting the entity that created the digital signature s(H1), i.e., the owner of the private key if asymmetric key cryptography is used; the data units in the second set are verified by trusting the entity that compiled the signed bitstream B and created the archive object.
[0091] In some embodiments of the verification method 700, the derivation 716 of the bit string includes a decision 716.1 of whether to calculate the required fingerprint from the reconstructed macroblock 111 or to obtain it from the archival object 140. This decision can be guided by location information in the archival object 140, which indicates the location of the macroblock 110 represented by the data unit 120 to which the archived fingerprint pertains. Having access to these macroblock locations allows the receiver to perform a reliable integrity check, based on the assumption that any macroblock 110 in the video frame 100 that cannot be reconstructed from the data unit 120 in the signed video bitstream B is encoded by another data unit whose fingerprint can necessarily be obtained from the archival object 140. If the archival object 140 does not indicate the location of these macroblocks, the receiver can, for example, insert the missing fingerprints, i.e. fingerprints that cannot be calculated from the data units 120 in the signed video bitstream B, by a trial and error approach. A trial and error approach may include performing steps 714 and 716 for each possible way of inserting the archived fingerprint from archive object 140 (each such way of insertion may be inferred to be a permutation of the locations of the missing macroblocks), and concluding that signed video bitstream B is fraudulent only if all of these attempts fail.
[0092] As an overview, FIG. 11 shows the signal processing operations and data flows occurring during the execution of the editing method shown in FIG. 6 and the verification method of FIG. 7, as well as functional units suitable for performing the signal processing operations. The time evolution of the signal processing flows is generally oriented from left to right. Data and signals are depicted as simple frames, while functional units are represented as frames with double vertical lines. Each functional unit may correspond to a programmable processor that executes a segment of software code, or a network of such processors. It is also possible to implement each functional unit as a dedicated hardware circuit, such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). From another perspective, the functional units can be considered to represent respective portions of software code (e.g., modules, routines) that are executed on a common processor or processor network.
[0093] Assume initially that an input image 100 in plaintext form is fed to an encoder 1110 configured for a predictive video coding format. The encoder 1110 outputs an encoded image 110A. The encoded image 110A may be formatted like the video bitstream described with reference to FIG. 5 to include, among other things, data units corresponding to macroblocks linked by inter-frame or intra-frame prediction references. To ensure accurate predictive coding, the encoder 1110 comprises a reference decoder 1111 configured to reconstruct image data from the data units as they are generated, for example quasi-sequentially, by the encoding process in the encoder 1110. The reconstructed image data, which may be considered to form a decoded reference image 100C, is temporarily stored in a reference buffer 1112 associated with the encoder 1110. More precisely, to represent image data in a first macroblock relative to image data in a second macroblock, the encoder 1110 calculates an increment from the second macroblock in the decoded reference image 100C, rather than from the original second macroblock (i.e., as provided in the input image 100). The hash function 1113 computes fingerprints h1, h2, h3 from the macroblocks of the decoded reference image 100C and collects (e.g., concatenates) them into a hash list H1. The hash list H1 may be a string of bits. As described above, the hash list H1 may be digitally signed and the resulting digital signature s(H1) may be carried in a signature unit in the video bitstream.
[0094] Suppose a request to mask an area of an image is received from a user. To execute the user's request, the encoded image 100A is input to an editing tool 1120 comprising a decoder 1121, an optional verifier 1122, and a masker 1123. The decoder 1121 is configured to reconstruct image data, in particular macroblocks, from the encoded image 100A. Nominally, the image reconstructed from the encoded image 100A is identical to the decoded reference image 100C. The presence of the verifier 1122 in the editing tool 1120 can be justified in particular if the encoded image 100A is transferred over an untrusted connection, in which case a verification can be performed before the masking is performed. The expected result of the verification is that the encoded image 100A is authentic, in which case it makes sense to perform the masking. The masker 1123 is configured to replace pre-specified colors or patterns in areas of the image. The substitution produces a masked decoded image 100B, which is transferred (via a trusted connection) to the second encoder 1130. The masked decoded image 100B has at least the edited parts (edited macroblocks) in plaintext form. It may be advantageous to refrain from decoding the non-edited parts as far as feasible. The second encoder 1130 outputs a further encoded image 100F, which may be made available to a recipient. Similar to the first encoder 1110, the second encoder 1130 comprises a reference decoder 1131 configured to reconstruct image data from data units created by the encoding process in the second encoder 1130. The reconstructed image data, i.e. the decoded reference image 100D, is temporarily stored in a reference buffer 1132 associated with the second encoder 1130.
[0095] To ensure that the edited and then encoded image can be verified, a hash list is generated from the masked decoded image 100B using a hash function 1124. TIFF2024086623000015.tif5170 is calculated. (This corresponds to optional step 620 of the editing method 600 described above.) Note that the edited macroblock (here The fingerprint for the original image (h1, h2) has likely changed as a result of the masking, but the remaining fingerprints (h1, h3 here) may match the fingerprints calculated from the original decoded reference image 100C. The new hash list includes the digital signature TIFF2024086623000017.tif5170 may be added. Alternatively, a hash list TIFF2024086623000018.tif5170 is calculated from the decoded reference image 100D in the reference buffer 1132 associated with the encoder 1130.
[0096] It should be noted that the encoder 1110 may be controlled by a different entity (e.g., an author) than the editing tool 1120 and the second decoder 1130. Because the masked decoded image 100B may be tampered with before reaching the encoder 1130, the editing tool 1120 and the second decoder 1130 are preferably co-located or linked by a reliable data connection.
[0097] At the receiving end, a decoder 1140 is provided which computes a decoding process adapted to reconstruct the decoded reference image 100E from the further encoded image 100F. From the decoded reference image 100E, a hash list 1142 is generated using a further hash function 1142. The verifier 1143 associated with the decoder 1140 can then use this hash list TIFF2024086623000020.tif5170 is a hash list calculated from the masked decoded image 100B or the decoded reference image 100D. TIFF2024086623000021.tif5170. As explained above, this can be done directly (the "document approach") as in Figure 10A, or indirectly as in Figure 10B. If the verifier 1143 concludes that the two hash lists do in fact match, then the further encoded image 100F (or equivalently the decoded image 100E) may be released for reproduction or further processing.
[0098] It should be noted that the verification of the further encoded image 100F is based on a trust relationship between the further encoder 1130 and the decoder 1140 or between persons or entities controlling these devices. Additionally or alternatively, if it is desired to verify the further encoded image 100F based on a trust relationship between the first encoder 1110 and the decoder 1140 or between persons or entities controlling these devices (e.g., between the author and the consumer), the further encoded image 100F may be given an archived object containing fingerprints of the edited macroblocks. It should be noted that due to the editing, the entire further encoded image 100F, only its unedited parts, cannot be verified based on this latter trust relationship.
[0099] It will be understood that the hash functions 1113, 1124, and 1142 appearing in FIG. 11 are equivalent in the sense that they provide equal outputs for equal inputs.
[0100] Aspects of the present disclosure have been described above primarily with reference to some embodiments. However, as will be readily understood by those skilled in the art, other embodiments than those disclosed above are equally possible within the scope of the inventive concept, as defined by the appended claims. In particular, it is noted that the above description of various embodiments focuses on predictively coded videos. This is because aspects of the present disclosure are expected to have particular advantages in predictive-based coding, where both the encoder and the decoder have access to the same frame, in the form of a reference frame in the encoder and decoder reference buffers. Thus, in predictive-based coding, no additional steps are required to obtain the reconstructed or decoded frame to be signed and verified. However, the same approach can be used for any video, not just predictively coded videos, as long as the entity that signs the video and the entity that verifies the video have access to the reconstructed or decoded frame of the video. In the case of predictive-based coding, it proves practically convenient to utilize fingerprints of groups of pixels that are also used as macroblocks in the coding. However, in general, the fingerprints may be calculated from groups of pixels grouped in other ways. Once the frame is decoded, it does not matter how the pixels were divided for coding. For example, depending on the type of editing expected, it may be useful to divide the decoded image into smaller or larger groups of pixels than those used for encoding. For example, if the masking is always performed in a rectangular form, a coarser division of pixels than that used for encoding may be sufficient for the signature process. On the other hand, if it is assumed that the masking can be performed more closely following the contours of the object to be masked, a finer division may be useful for the signature process.
Claims
1. 1. A method for editing a signed video bitstream obtained by predictive coding of a video sequence, comprising: the signed video bitstream includes a plurality of data units and a plurality of associated signature units; each data unit representing one macroblock in a video frame of the video sequence that has been predictively coded, the data unit being obtained by applying an encoder to a corresponding macroblock; each signature unit comprising a digital signature of a bit string derived from a plurality of fingerprints; each fingerprint is calculated from macroblocks reconstructed from one data unit associated with said signature unit; The method comprises: receiving a request to replace a region of at least one video frame; reconstructing a first set of macroblocks in which the region is included and a second set of macroblocks that directly or indirectly reference macroblocks in the first set, each reconstructed macroblock being obtained by applying a decoder to a corresponding data unit; adding an archival object to the signed video bitstream, the archival object including a fingerprint calculated from the first set of reconstructed macroblocks and indicating the location of the macroblock represented by the data unit to which the included fingerprint relates; editing the first set of macroblocks in accordance with the request to replace the region of the at least one video frame and encoding the edited first set of macroblocks as a first set of new data units; re-encoding the second set of macroblocks as a second set of new data units; adding the first and second sets of new data units to the signed video bitstream; A method comprising:
2. adding a fingerprint calculated from the first set of edited macroblocks to the signed video bitstream; a signature unit associated with a data unit encoding the first set of macroblocks is replaced in the signed video bitstream with a substitute signature unit comprising a digital signature of a bit string derived from at least the calculated fingerprint; or The method of claim 1 , wherein at least one new signature unit is added to the signed video bitstream, the new signature unit comprising a bit string derived from a digital signature of the calculated fingerprint.
3. The method of claim 1 , wherein the archive object further includes the locations of the first set of macroblocks.
4. The method of claim 1 , wherein the second set of macroblocks is losslessly re-encoded.
5. the second set of macroblocks are re-encoded using reduced data compression; The method of claim 1 , wherein the fingerprints of the second set of the macroblocks in the signed video bitstream include robust hashes.
6. The method of claim 1 , wherein the second set of macroblocks are non-predictively re-encoded.
7. The method of claim 1 , wherein the second set of macroblocks are predictively re-encoded with reference to the edited first set of macroblocks.
8. editing and encoding the first set of macroblocks; encoding the first set of edited macroblocks losslessly and / or using reduced data compression and / or non-predictively; The method of claim 1 , comprising:
9. first providing a signature unit, reconstructing a plurality of macroblocks from respective data units associated with the signature unit; calculating a plurality of fingerprints from each of the reconstructed macroblocks; deriving a bit string from the calculated fingerprint; obtaining a digital signature of said bit string; First provide a signature unit, including The method of claim 1 further comprising:
10. The method of claim 9 , wherein the bit string is a combination of the fingerprints or a fingerprint of the combination.
11. The method of claim 1 , wherein calculating the fingerprint comprises obtaining the reconstructed macroblock from a reference decoder buffer.
12. 1. A method for verifying a signed media bitstream obtained by predictive coding of a video sequence, comprising: the signed video bitstream includes a plurality of data units, a plurality of associated signature units, and an archive object; each data unit representing one macroblock in a video frame of the video sequence that has been predictively coded, the data unit being obtained by applying an encoder to a corresponding macroblock; each signature unit comprising a digital signature of a bit string derived from a plurality of fingerprints; each fingerprint is calculated from macroblocks reconstructed from one data unit associated with said signature unit; the archive object includes at least one archived fingerprint and indicates the location of a macroblock represented by a data unit to which the archived fingerprint relates; The method comprises: reconstructing the macroblock by applying a decoder to the data units associated with the signature unit; calculating respective fingerprints from at least some of said reconstructed macroblocks; obtaining at least one archived fingerprint from the archive object; deriving a bit string from the calculated fingerprint and the obtained fingerprint; verifying the data unit associated with the signature unit using the digital signature in the signature unit; A method comprising:
13. Deriving the bit string includes determining, for each data unit associated with the signature unit, whether to calculate a fingerprint from a reconstructed macroblock or to obtain a corresponding archived fingerprint from the archive object based on the position according to the archive object. The method of claim 12.
14. A device comprising processing circuitry arranged to carry out the method of any one of claims 1 to 13.
15. A computer program comprising instructions which, when executed by a computer, cause the computer to perform a method according to any one of claims 1 to 13.