Editable signed video data

The method addresses the inefficiencies in editing signed video bitstreams by selectively re-encoding only necessary macroblocks, preserving signatures and computational efficiency, thus enabling efficient and secure video editing.

JP7682982B2Active Publication Date: 2025-05-26AXIS

Patent Information

Application Number
JP2023207480
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-12-15
Filing Date
2023-12-08
Publication Date
2025-05-26
Estimated Expiration
2043-12-08

AI Technical Summary

Technical Problem

Existing methods for editing signed video bitstreams require re-encoding and re-signing the entire frame, which consumes significant computational resources and introduces delays, especially when preserving prediction coding dependencies.

Method used

A method for editing a signed video bitstream that selectively re-encodes only the macroblocks that directly or indirectly reference the edited macroblocks, while preserving the signatures of unaffected macroblocks, thereby minimizing computational overhead.

Benefits of technology

The method efficiently edits signed video bitstreams with minimal disruption to data security and computational resources, allowing for real-time editing with preserved predictive coding dependencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682982000010
    Figure 0007682982000010
  • Figure 0007682982000011
    Figure 0007682982000011
  • Figure 0007682982000012
    Figure 0007682982000012
Patent Text Reader

Abstract

To provide a method and a device for verifying a video bitstream with signature.SOLUTION: A method for editing a video bitstream with signature obtained by predictive coding of a video sequence includes: receiving a request to replace a region of the bitstream; determining a first set of macro-blocks containing the region and a second set of macro-blocks that references the macro-blocks in the first set; adding an archive object containing fingerprints of first and second sets of data units representing the first and second sets of macro-blocks; editing the first set of data units according to the request to replace; and re-coding the second set of data units.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of security devices for protecting video data from unauthorized activities, particularly in relation to the storage and transmission of data. The present disclosure proposes a method and device for editing a signed video bitstream and verifying the signed video bitstream that may result from such editing.

Background Art

[0002] Digital signatures can verify the authenticity or integrity of a message and guarantee non-repudiation by means of a digital signature that provides a layer of verification and security to a digital message transmitted over an insecure channel. Particularly with respect to video encoding, there are secure and highly efficient methods for digitally signing a predicted video sequence as described in the prior art. See, for example, European Patent Application Publication No. 212013601 (European Patent Application Publication No. 4164173) and European Patent Application Publication No. 212013627 (European Patent Application Publication No. 4164230), which are previous patent applications by the inventors. Also see U.S. Patent Application Publication No. 20140010366, which proposes an encrypted video verification technique particularly adapted to predicted video data having a picture group structure.

[0003] A video sequence may need to be edited after being signed. In addition to visual improvement, editing can be aimed at ensuring privacy protection by trimming, masking, blurring, or similar image processing that makes visual features difficult to recognize. In most available methods, this requires re-encoding and re-signing the entire edited frame. Re-encoding and re-signing are preferably also extended to several adjacent frames so as not to disrupt the existing prediction coding dependencies (inter-frame / intra-frame references), even if the adjacent frames are not directly affected by the editing. These steps consume significant computational resources and can introduce annoying delays for the user.

[0004] U.S. Patent No. 7,437,007 discloses a method for performing region-of-interest editing of a video stream in a compressed domain. The compressed video stream includes compressed video stream frames representing video stream frames having unwanted portions and region-of-interest portions. According to the method, the compressed video stream frames are edited to obtain compressed video stream frames that correct the aforementioned unwanted portions and include the aforementioned region-of-interest portions while maintaining the original structure of the video stream. To achieve this, the editing includes skipping macroblocks located above, below, and to the right of the aforementioned region-of-interest portions of the predicted (P) frames and the bi-directionally predicted (B) frames. The video stream under consideration in U.S. Patent No. 7,437,007 is not a signed video stream. SUMMARY OF THE INVENTION

[0005] One object of the present disclosure is to enable a method for editing a signed video bitstream obtained by predictive encoding of a video sequence that significantly avoids the need to re-sign the bitstream outside of the portions affected by the edit, as in some available methods. A particular object is to enable a video editing method that preserves the signatures of all macroblocks except those edited and any further macroblocks that reference them, whether directly or indirectly. A further object is to enable video editing without significantly compromising the data security of the original signed video bitstream. A further object is to provide a method for verifying a signed video bitstream obtained by predictive encoding of a video sequence. Providing a device and computer program for these purposes is yet a further object.

[0006] At least some of these objects are achieved by the present invention as defined by the independent claims. The dependent claims relate to advantageous embodiments of the present invention.

[0007] In a first aspect of the present disclosure, a method for editing a signed video bitstream obtained by predictive coding of a video sequence is provided. The signed video bitstream is assumed to include data units and associated signature units. Each data unit represents (i.e., encodes) at most one macroblock within a video frame of the predictive-coded video sequence. Each signature unit includes a digital signature of a bit sequence derived from a plurality of fingerprints of exactly one associated data unit. Optionally, the signature unit can include the bit sequence to which the digital signature pertains. For a signed video bitstream having these characteristics, the method includes receiving a request to replace an area of at least one video frame (e.g., replacing a privacy mask for pixel values of the area within one or more video frames), determining a first set of macroblocks that include the aforementioned area and a second set of macroblocks that directly or indirectly reference the macroblocks within the first set, adding an archive object to the signed video bitstream, the archive object including fingerprints of a first set and a second set of data units that respectively represent the determined first set and second set of macroblocks, editing the first set of data units according to the replacement request, and re-encoding the second set of data units. Optionally, the signature unit can include the derived bit sequence to which the digital signature pertains ("document approach").

[0008] Since each data unit represents at most one macroblock, as a result, there is a one-to-one relationship between the data unit and the macroblock, or some macroblocks are represented by two or more data units and can be reconstructed from two or more data units. A data unit does not represent a plurality of macroblocks (or a portion of a plurality of macroblocks), and thus, the direct effect of editing some macroblocks is limited to one or more data units (a first set) of the edited macroblock. Further, since each fingerprint is exactly the fingerprint of one associated data unit, the need to re-sign the data unit after editing is limited. More precisely, the method according to the first aspect preserves any predictive coding dependencies that connect pairs or groups of macroblocks, i.e., by re-encoding a set (a second set) of data units that represent macroblocks that directly or indirectly reference the edited macroblock(s). When the re-encoding is limited to this second set of data units, the method efficiently utilizes the available computational resources. Thereby, the method can be executed with good performance of a normal processing device.

[0009] The method according to the first aspect includes further advantages on the receiver side. For an archive object capable of obtaining the fingerprints of the first set and the second set of data units, the receiver can verify all data units of the signed video bitstream that have not been affected by the editing. Thereby, a significant portion of the existing signature can be preserved, and in this regard, it can be said that the video editing method according to the first aspect has minimal disruption. Verification on the receiver side is described in detail within the second aspect of the present disclosure.

[0010] In some embodiments, the archive object further includes the positions of the first set and the second set of macroblocks. The positions can refer to the positions of the macroblocks within the frame, for example, within the frame coordinates. When static macroblock partitioning is used, the positions of the macroblocks can be represented as macroblock sequence numbers or another identifier. This provides one way of assisting a recipient of an edited video bitstream in determining whether a particular macroblock has been changed and thus in selecting an appropriate method for obtaining the fingerprint of the data unit representing the macroblock.

[0011] Some embodiments provide an advantageous procedure for editing a first set of data units. More precisely, such editing may include decoding the data units into reconstructed macroblocks, providing the edited macroblocks by performing the required substitutions on the reconstructed macroblocks, and providing the edited data units by encoding the edited macroblocks. If the area to be replaced encompasses the entire reconstructed macroblock, the edited data unit may be provided, in particular, by encoding a subset of the area, for example, the intersection of the macroblock and the area.

[0012] In such embodiments, it is optional to encode the macroblocks resulting after the substitution as intra-coded macroblocks (I-blocks), thereby obtaining independently decodable data units. This explains the fact that the substitution can introduce sudden temporal changes into the video sequence. This tends to reduce the temporal continuity of the video sequence, and thus most known predictive coding techniques do not function well.

[0013] Furthermore, within the above outlined general procedure for editing a first set of data units, it is possible to decode data units within a second set using the reconstructed macroblocks. Since the second set of data units represents a second set of macroblocks that directly or indirectly reference the first set of macroblocks, the decoding is predictive. The output of decoding the aforementioned data units within the second set is used in the re-encoding step.

[0014] In different embodiments, re-encoding of the second set of data units can be performed by predictive encoding with reference to the first set of edited data units or by non-predictive encoding. The choice of one of these two options may correspond to achieving a desired balance between quality and bitrate. Additionally, the second set of data units may be re-encoded using reduced data compression. In non-reversible video encoding formats, data compression is achieved by discarding some of the information within the video sequence, such as by quantization of pixel values. The degree of quantization can correspond, for example, to the value of a quantization parameter (QP) that represents the fineness of the quantized pixel values. The degree of quantization may further depend on the entries within the definition of the scaling matrix. Assuming that the video sequence is encoded at a given normal data compression level, it is expected that the second set of macroblocks is encoded at a reduced data compression level relative to normal data compression. The use of a reduced data compression level means, for example, that by mapping the pixel values of the macroblocks to a relatively fine set of quantized pixel values, less information within the macroblocks is discarded during the encoding step. In other words, the macroblocks within the second set are encoded at a relatively higher bitrate and can be reproduced with fewer residual errors than when a normal data compression level is used. The additional memory or bandwidth cost is likely to be acceptable, even more so in use cases where editing occurs infrequently and / or in isolated portions of the video sequence.

[0015] In some embodiments, the data security of an edited signed video bitstream is improved by providing to the bitstream one or more signature units associated with a first set of edited data units and / or a second set of re-encoded data units. This avoids scenarios of unauthorized modification of the first set of edited data units and the second set of re-encoded data units. The one or more signature units may be new signature units added to the bitstream or edited versions of signature units included in the signed video bitstream prior to editing.

[0016] According to a generalization of the first aspect, a video editing method is provided that is performed on a signed video sequence including data units and associated signature units. Each data unit represents at most N macroblocks within a video frame of a prediction-encoded video sequence. Here, N is a small integer such as 1, 2, 3, 4, 5, or at most 10. Each signature unit includes a digital signature of a bit sequence derived from multiple fingerprints of exactly one associated data unit, and optionally includes the bit sequence. The video editing method includes receiving a request to replace an area of at least one video frame, determining a first set of macroblocks included in the area and a second set of macroblocks that directly or indirectly reference the macroblocks within the first set, adding to the signed video bitstream an archive object that includes fingerprints of the first set and the second set of data units that respectively represent the determined first set and second set of macroblocks, editing the first set of data units according to the request to replace, and re-encoding the second set of data units.

[0017] According to this generalization, since each data unit represents at most N macroblocks, the direct effect of editing one macroblock is limited to at most N data units of the edited macroblock. Further, since each fingerprint is exactly the fingerprint of one associated data unit, the need to re-sign the data units after editing is limited. In particular, the archive object holds at most N times the fingerprint of the first and second sets of macroblocks, and as a result, the video editing method has a realizable complexity.

[0018] In a second aspect of the present disclosure, a method for verifying a signed video bitstream obtained by predictive coding of a video sequence is provided. The signed video bitstream is understood to include data units, signature units respectively associated with some of the data units, and an archive object. Each data unit represents at most one macroblock within a frame of the predicted video sequence. Each signature unit includes a bit string and optionally a digital signature of the bit string itself. Finally, the archive object includes fingerprints of at least one archived data unit. The method for verifying the signed video bitstream includes obtaining the fingerprint of each data unit associated with the signature unit by calculating the fingerprint of the data unit or by obtaining the archived fingerprint from the archive object, deriving a bit string from the obtained fingerprint, and verifying the data unit associated with the signature unit using the digital signature within the signature unit. The final verification step can include verifying the bit string derived using the digital signature. Alternatively (the "document approach"), the verification step includes verifying the bit string within the signature unit using the digital signature and comparing the derived bit string with the verified bit string.

[0019] Archive objects may be added by implementing an editing method according to a first aspect, but the method according to the second aspect can be implemented without such reliable knowledge of preprocessing. Thus, the method according to the second aspect achieves verification of the authenticity of a video sequence in terms of verifying that the digital signature (and any bit string) carried by the signature unit actually matches the fingerprint of the data unit it relates to. Thus, there is also a possibility that the data unit has not been modified.

[0020] The method according to the second aspect includes two options for obtaining the fingerprint of a data unit, either by direct calculation or by obtaining it from an archive object. This supports a process that minimizes the destruction of existing fingerprints during the editing phase (first aspect). The fact that each data unit represents at most one macroblock tends to limit the number of fingerprints that need to be archived for a given replacement request, and thus limits the size of the archive object.

[0021] In some embodiments, the archive object further indicates the position of the macroblock represented by the data unit to which the archived fingerprint relates. In other words, the archived fingerprint is the fingerprint of the data unit, and the data unit represents (encodes) the macroblock whose position is indicated within the archive object. During the execution of the method according to the second aspect, it is determined whether to calculate the fingerprint or obtain the fingerprint from the archive object based on the position indicated by the archive object in order to obtain the fingerprint of the data unit.

[0022] A third aspect of the present disclosure relates to a device configured to execute the methods of the first aspect and / or the second aspect. These devices may be incorporated into systems having different primary purposes (e.g., video recording, video content management, video playback), or may be dedicated to the aforementioned editing and verification respectively. Devices within the third aspect of the present disclosure generally share the effects and advantages of the first and second aspects, and they can be embodied with an equivalent degree of technical variations.

[0023] The present invention further relates to a computer program comprising instructions for causing a computer to execute the above method. The computer program can be stored or distributed on a data carrier. As used herein, a "data carrier" may be a transient data carrier such as a modulated electromagnetic wave or light wave, or a non-transient data carrier. The non-transient data carrier includes volatile and non-volatile memories such as magnetic, optical, or solid-state type permanent and non-permanent storage media. Still within the scope of "data carrier", such memories may be fixedly attached or removable.

[0024] In a further aspect of the present disclosure, a signed video bitstream is provided that includes data units and associated signature units. Each data unit represents at most N macroblocks within a video frame of a prediction-encoded video sequence. Here, N is a small integer such as 1, 2, 3, 4, 5, or at most 10. Each signature unit includes a digital signature of a bit sequence derived from a plurality of fingerprints of exactly one associated data unit, and optionally the bit sequence itself. The signed video bitstream is suitable for editing because the direct effect of editing one macroblock is limited to at most N data units of the edited macroblock, and each fingerprint is the fingerprint of exactly one associated data unit. This limits the propagation of the edit to a limited number of data units, resulting in fewer data units that need to be re-signed after editing.

[0025] Note that "macroblock" as used in the present disclosure can advantageously be an encoded macroblock. However, the present invention is also applicable to non-predictively encoded video, and in a more generalized case, a macroblock can thus be any group of adjacent pixels. Since the signature of the video is performed on the decoded frame, the grouping of pixels need not be limited to the splitting of any coding group.

[0026] In general, all terms used in the claims should be construed according to their ordinary meaning in the technical field, unless otherwise explicitly defined herein. All references to "an / a / the element, apparatus, component, means, step, etc." should be construed non-limitingly as referring to at least one example of the element, apparatus, component, means, step, etc., unless otherwise specified. The steps of any method disclosed herein need not be performed in the exact order described, unless explicitly stated.

[0027] Hereinafter, embodiments and implementations will be described by way of example with reference to the accompanying drawings.

Brief Description of the Drawings

[0028]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

DETAILED DESCRIPTION OF THE INVENTION

[0029] Aspects of the present invention are more fully described below with reference to the accompanying drawings in which specific embodiments of the invention are shown. However, these aspects may be embodied in many different forms and should not be construed as limiting. Rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete and will fully convey the scope of all aspects of the invention to those skilled in the art. Throughout the description, like reference numerals refer to like elements.

[0030] In the terms of the present disclosure, a "video bitstream" includes any substantially linear data structure that may be similar to a sequence of bit values. The video bitstream can be carried by a transient medium (e.g., a modulated electromagnetic or optical wave) as in some streaming use cases, or the video bitstream can be stored in a non-transient medium such as volatile or non-volatile memory.

[0031] The video bitstream represents a video sequence, which can be understood as a sequence of video frames that are sequentially played back at a nominal time interval. Each video frame can be divided into macroblocks. In the present disclosure, further, a "macroblock" can be a transform block or a prediction block within a frame of the video sequence, or a block having both of these uses. The use of "frame" and "macroblock" herein is intended to be consistent with the H.26x video coding standard or similar specifications. A "macroblock" can further be an encoded block. As described above, although the term macroblock is used, the grouping of pixels need not be limited to any particular division used for encoding; rather, a macroblock can be any group of adjacent pixels. In the case of prediction-based encoding, it may be advantageous to use the same division as that used for encoding.

[0032] Figures 1A and 1B show an exemplary partition of video frame 100 into a 4×4 uniform arrangement of macroblocks 101. Note that this is a simplification for illustrative purposes. In reality, video frame 100 is generally divided into a much larger number of macroblocks. For example, the macroblocks can be 8×8 pixels, 16×16 pixels, or 64×64 pixels. In this specification, curly arrows are consistently used to indicate intra-frame references or inter-frame references used in predictive coding. Without departing from the scope of the present disclosure, the partition into macroblocks seen in Figures 1A and 1B can be significantly changed to include non-square arrangements and / or arrangements of macroblocks 101 that are not of the same type as video frame 100 and / or mixed arrangements of different macroblocks 101 having different sizes or shapes. It is understood that some video coding formats support dynamic macroblock partitioning, i.e., the partition can be different for different video frames within a sequence. This applies, for example, to H.265.

[0033] In Figure 1A, the pattern of intra-frame references is limited to a single row of macroblocks 101. In fact, each macroblock 101 refers to the macroblock immediately to its left (if it exists within frame 100) in the sense that it is represented by a data unit that predictively represents the image data within the macroblock, i.e., the image data within this macroblock is represented with respect to the image data within the left macroblock. Conceptually, although somewhat simplified, the data unit represents macroblock 101 with respect to changes or movements relative to the left macroblock. Another possible understanding is that the data unit represents the correction of a predetermined prediction operation that derives macroblock 101 from the left macroblock. In alternative examples within the scope of the present disclosure, macroblock 101 may refer to another macroblock 101 to its right, or above or below it.

[0034] In FIG. 1B, the pattern of intra-frame references is denser. Here, each macroblock 101 refers to the macroblock immediately to its left (if it exists in frame 100) and the macroblock immediately above it (if it exists in frame 100). Thus, the macroblock 101 is represented by a data unit that predictively represents the image data within the macroblock, for example, with respect to differences or with respect to the correction of the prediction of this image data, based on a predetermined interpolation operation (or any other predetermined combination) acting on the image data in the left and upper macroblocks. The interpolation can include post-processing operations such as smoothing. A further alternative within the scope of the present disclosure is to use an intra-frame reference having a direction opposite to that shown in FIG. 1B, i.e., a direction starting from the bottom row.

[0035] An image / video format having a predetermined pattern of intra-frame references can be associated with a specified scanning order representing an executable sequence for decoding macroblocks. In FIG. 1A, the macroblock scanning order is not unique, i.e., because each row can be decoded independently. In the case of the pattern according to FIG. 1B, the macroblocks can be decoded in a row direction from top to bottom or in a column direction from left to right, and a video format having this reference pattern can specify a scanning order for each column or each row such that any reconstruction error can be expected on the encoder side, which results in a benefit in coding efficiency. Furthermore, there are video formats having any macroblock order or so-called slicing.

[0036] Figure 2 shows a video sequence V that includes a sequence of frames 100. The video sequence V includes independently decodable frames (I-frames) and predicted frames, including unidirectional prediction frames (P-frames) and bidirectional prediction frames (B-frames). The ITU-T Recommendation H.264 (06 / 2019) "Advanced video coding for generic audiovisual services" of the International Telecommunication Union specifies a video coding standard in which both forward prediction frames and bidirectional prediction frames are used. As can be seen in Figure 2, an independently decodable frame does not reference any other frames. The unidirectional prediction frame (P-frame) in Figure 2 is forward predicted in that it directly references at least one other preceding or immediately preceding frame. A bidirectional prediction frame (B-frame) can further directly reference subsequent or immediately following frames within the video sequence V. If the video sequence includes a third frame (or subsequence of frames) that is directly referenced by a first frame and a second frame is directly referenced, the first frame indirectly references the second frame. In predictive video coding, a group of pictures (GoP) is defined as a subsequence of video frames that does not reference any video frames outside the subsequence and that can be decoded without referencing any other I, P, or B frames. The video frames in Figure 2 form a GoP. The GoP in Figure 2 is the smallest because it cannot be further subdivided into additional GoPs.

[0037] In a simpler implementation, the video sequence V can consist only of independently decodable frames (I) and unidirectional prediction frames (P). Such a video sequence can have the appearance of IPPIPPPPIPPPIPPP, where each P-frame points to the immediately preceding I-frame or P-frame. In this example, GoPs such as IPP, IPPPP, IPPP, IPPP, etc. can be identified.

[0038] There are several options for adjusting inter-frame and intra-frame predictive coding. For example, if a static (fixed) macroblock partition is used for all video frames, an inter-frame reference such as that illustrated can be defined at the level of one macroblock position at a time (e.g., the top-left macroblock in FIGS. 1A and 1B). Some video formats allow for dynamic macroblock partitioning, e.g., a macroblock can be predicted from corresponding pixels in a preceding video frame or from spatially shifted pixels in a preceding frame. Alternatively, the inter-frame reference is defined for the entire video frame. An I-frame consists of only I-blocks, and a P-frame can consist of only P-blocks or a mixture of I-blocks and P-blocks.

[0039] Referring now to FIG. 3, attention is directed to a video bitstream that encodes the video sequence under consideration. Subfigures 3A, 3B, 3C, and 3D show different corresponding patterns between the data unit 102 they represent (encode) and the macroblock 101. For purposes of the present disclosure, the data unit can have any suitable format and structure, and it is assumed only that the data unit can be separated (or extracted) from the video bitstream, e.g., to enable processing, without the need to decode that data unit or any surrounding data units. A signed video bitstream further includes, in addition to the data unit, a signature unit separable from the signed video bitstream in the same or a similar manner. Details regarding the signature unit are presented below with reference to FIG. 5A.

[0040] Under one option, as shown in Figure 3A, data unit 102 corresponds one-to-one to macroblock 101 (dashed line). With this correspondence pattern, each macroblock 101 is always reconstructed from one data unit 102. Further, when a macroblock 101 is edited (certainly, there may be cases where changes such as a signature unit and metadata are required), other data units 102 other than the corresponding data unit 102 do not need to be modified.

[0041] Alternatively, as shown in Figure 3B, each macroblock 101 is encoded by a plurality of data units 102. Note that each data unit 102 represents at most one macroblock 101. Thus, if each macroblock 101 is encoded by at most M data units 102, as a result, any macroblock 101 can be reconstructed from at most M data units 102. Edits made to a macroblock 101 directly affect at most M data units 102. Here, M is a small integer such as 1, 2, 3, 4, 5, or at most 10.

[0042] In a further option, as shown in Figure 3C, each data unit 102 encodes a plurality of macroblocks 101. This means that the effect of an edit made to a macroblock 101 is not limited to the macroblock 101 itself, and it may be necessary to re-encode and / or re-sign the data units 102 shared by the edited macroblock 101 and further macroblocks 101. Since performing re-signing on an unnecessarily large data set can be computationally wasteful, this correspondence pattern is quite possible to implement but is not applied in the highest-performance embodiments in the present disclosure. If a data unit 102 can be shared by at most a predetermined number N of macroblocks 101, the additional total computational amount can remain limited. Here too, N can be specified to be a small integer such as 1, 2, 3, 4, 5, or at most 10.

[0043] The combination of patterns found in FIGS. 3B and 3C is possible within the scope of the present disclosure. As a result, the techniques proposed herein may be applied to video bitstreams where the ratio of macroblock 101 to data unit 102 can be 2:1, 1:1, 2:2, 3:3, 4:2, 2:4, or 4:4.

[0044] In yet another option, as shown in FIG. 3D, each data unit 102 can represent any number of macroblocks 101 within a video sequence, and each macroblock 101 can be encoded by any number of data units 102. This corresponding pattern, as suggested by the dashed lines, may mean that in the worst case, even limited editing operations on macroblock 101 require a complete re-encoding and re-signing of the video sequence. The techniques disclosed herein should not be implemented on video sequences having the structure shown in FIG. 3D.

[0045] FIG. 4 shows the editing operations described below.

[0046] FIG. 5A shows a section of a video sequence V including a series of macroblocks 101 belonging to one or more video frames. For example, the macroblock 101 can occupy a fixed position (e.g., the upper left macroblock in FIGS. 1A and 1B) within consecutive video frames. Assume that the macroblock 101 is prediction-coded according to the reference indicated by the winding arrow. The video sequence V is encoded as a signed video bitstream B including a data unit 102 and a signature unit 103. For purposes of illustration and not limitation, FIG. 5A can be described as a hybrid of the patterns shown in FIGS. 3A and 3B, with a regular or irregular alternation therebetween (e.g., based on the macroblock size, larger macroblocks 101 correspond to multiple data units 102, and smaller or more compressed macroblocks 101 correspond to a single data unit 102), showing a data unit 102 that encodes the video macroblock 101 according to a corresponding pattern that allows this. Nevertheless, each data unit 102 represents at most one macroblock 101. The data unit 102 can conform to a proprietary or standardized video coding format such as ITU-T H.264, H.265, or AV1. The bitstream B can further include additional types of units (e.g., dedicated metadata units) without departing from the scope of the present disclosure.

[0047] Each of the signature units 103 can be associated with a plurality of data units 102. In FIG. 5A, it can be understood that the data unit 102 between two consecutive signature units 103 is associated with the subsequent signature unit 103, which is not an essential feature of the present invention and other arrangements are possible without departing from the scope of the present disclosure. The signature units 103 can all be associated with a set of data units 102 that are all included in one GoP, but other association patterns as seen in FIG. 5A are also possible. Further, the set of data units 102 associated with one signature unit 103 is preferably selected in consideration of the applicable macroblock scanning order. For example, the set of data units 102 associated with the signature unit 103 can represent the number of macroblocks to be sequentially scanned during decoding, thereby minimizing the number of macroblocks that need to be re-referenced if the signature unit 103 fails verification.

[0048] The signature unit 103 includes at least one bit string (e.g., H 1 ) and a digital signature of that bit string (e.g., s(H 1)) and the like. As suggested by the use of the dashed line, the presence of the bit string is optional. If the signature unit 103 includes a plurality of bit strings, the signature unit 103 can have one digital signature for all of these bit strings, or each can have a plurality of digital signatures for a single bit string or each for a subgroup of bit strings. The bit string for which the digital signature is formed may be a combination of the fingerprints of the associated data unit 102, or may be the fingerprint of the combination of the fingerprints of the associated data unit 102. The combination of fingerprints (or "documents") may be a list of the literal representations of the fingerprints or other concatenation. In the ITU-T H.264 and H.265 formats, the signature unit may be included as a Supplemental Enhancement Information (SEI) message within the video bitstream. In the AV1 standard, the signature may be included in a Metadata Open Bitstream Unit (OBU).

[0049] Fingerprints may each be a hash or a salted hash. A salted hash may be a hash of a combination of a data unit (or a portion of a data unit) and a cryptographic salt, and the presence of the salt can prevent an unauthorized person with access to multiple hashes from inferring which hash function was used. Potentially useful cryptographic salts include the value of an active internal counter, a random number, and the time and location of a signature. The hash can be generated by a hash function (or one-way function) h, which is a cryptographic function that provides a security level considered appropriate in view of the confidentiality of the video data being signed and / or in view of values that would be put at risk if the video data were manipulated by an unauthorized person. Three examples are SHA-256, SHA3-512, and RSA-1024. The hash function is to be predefined (e.g., reproducible) so that the recipient can regenerate the fingerprint when attempting to verify it. In the example of FIG. 5A, the bit string is given by the following. H 1 =h([h 1 、h 2 、h 3 、h 4 、h 5 、h 6 ) and H 2 =h([h 7 、h 8 、h 9 、h 10 )、 where h 1 、h 2 、... are hashes of data units and [·] represents concatenation. An exemplary salted hash can be defined as follows. TIFF0007682982000001.tif5170 or In TIFF0007682982000002.tif5170, σ is an encryption salt. In the first example, the hash function h has a parametric dependency on a second argument to which the salt σ is assigned.

[0050] In some embodiments, the fingerprint h 1 , h 2 , etc. are each calculated directly from the data unit 102, for example, from the encoded transform coefficients or other video data therein. The fingerprint can be calculated from the entire data unit or from a subset thereof extracted according to a pre-agreed rule. In some embodiments, the fingerprint h 1 , h 2 ... are calculated from the reconstructed macroblocks obtained by decrypting the data unit 102, for example, pixel values or other plaintext data. In yet other embodiments, the fingerprint h 1 , h 2,... are calculated from intermediate reconstructed data derived from data units, rather than at the plaintext level or the bitstream level. More precisely, when an encoder including a frequency domain transform (e.g., DCT, DST, DFT, wavelet transform) followed by an encoding process (e.g., entropy, Huffman, Lempel-Ziv, run-length, binary or non-binary arithmetic coding, e.g., context adaptive variable length coding, CAVLC, context adaptive binary arithmetic coding, CABAC) is used, the transform coefficients are usually available as intermediate reconstructed data on the decoder side. The transform coefficients can be restored from the encoded representation. When the encoder further includes a quantization process immediately downstream of the transform, the quantized transform coefficients become available on the decoder side. In more complex codecs, there may be additional types of intermediate reconstructed data in a larger number of consecutive processing stages, and these can be used for fingerprint calculation. Similar to the quantized transform coefficients, it is particularly convenient to use the types of intermediate reconstructed data that appear identically in the encoding process. Common to all embodiments considered in this paragraph, the fingerprint relates to exactly one data unit 102 associated with the signature unit 103.

[0051] Optionally, the fingerprints can be sequentially linked to each other to detect unauthorized removal or insertion of data units. That is, each fingerprint depends on the next or previous fingerprint. For example, the input to a hash includes the hash of the next or previous fingerprint. The link can be realized, for example, as h 1 = h(X 102.1 ), h 2 = h([h 1 , X 102.2 ), h 3 = h([h 2 , X 102.3 ), etc., where X 102.1 , X 102.2 , X 102.3 represent data from the first, second, and third units of the data units 102.

[0052] Referring further to the signature unit 103 of FIG. 5A, to generate the digital signature s(H 1 ), an encryption element (not shown) having a pre-stored private key can be utilized. The recipient of the signed video bitstream can be expected to hold the public key belonging to the same key pair (see also FIG. 10), whereby the recipient can verify that the signature generated by the encryption element is authentic and does not generate a new signature. The public key can also be included as metadata in the signed video bitstream, in which case there is no need for the recipient to store it.

[0053] Next, referring to FIG. 6, a method 600 for editing a signed video bitstream B obtained by predictive encoding of a video sequence V will be described. The non-optional steps of method 600 are assumed to be performed after the original signature of the video bitstream. For example, if the signed video bitstream is originally generated by a recording device, the editing method 600 can be performed in a video management system (VMS). Another exemplary use case is when the signed video bitstream is generated by a device, stored in memory, and then re-referenced for editing using the same device. The editing may be performed at a later point in time, for example, after the need to perform privacy masking is known.

[0054] As described above, the device that executes the editing method 600 may be a specific-purpose application or system, or may have the basic function structure shown in FIG. 8. As shown in the figure, the device 800 includes a processing circuit 810, a memory 820, and an external interface 830. The memory 820 may be suitable for storing a computer program 821 having instructions for implementing the editing method 600. The external interface 830 may be a communication interface that enables the device 800 to communicate with a similar device (not shown) held by a recipient and / or video content author (e.g., a recording device), or may enable read and write operations in an external memory 890 suitable for storing a video bitstream.

[0055] FIG. 9 shows the case where a bitstream is transferred between multiple devices. It should be noted that the device that executes the editing method 600 may be connected to the recipient device via a local area network (the connection line in the lower half of FIG. 9) or a wide area network 990. An attack on the bitstream B may occur in any type of network, which justifies the signature.

[0056] Returning to FIG. 6, one embodiment of the method 600 begins with step 612 of receiving a request to replace an area of at least one video frame 100 within the video sequence V. The request may be received via a human-machine interface or in an automated manner, for example, as a message from a control application running on the same device 800 or remotely. The area to be replaced may be a set of replacement pixel values such as a privacy mask for replacing pixels located similarly within the video sequence V.

[0057] To avoid misunderstandings, the video sequence V to be edited is encoded by predictive coding as a signed video bitstream B that includes data units 102 and associated signature units 103, where each data unit represents at most one macroblock 101 within a video frame 100 of the predicted video sequence V, and each signature unit includes a digital signature of a bit sequence derived from a plurality of fingerprints of exactly one associated data unit. Note that such a bitstream format is illustrated with reference to FIG. 5A.

[0058] In the next step 614 of method 600, a first set of macroblocks that includes the aforementioned region, and a second set of macroblocks that directly or indirectly reference the macroblocks within the first set are determined. Recall that bidirectional prediction frames (B-frames) can be defined in some video coding formats, and it is understood that the second set of macroblocks can be located before or after the first set of macroblocks, or can occupy both of these positions. It is understood that the first set and the second set are defined to be disjoint. For example, it may be stipulated that a macroblock belongs to the second set only if the macroblock does not belong to the first set, i.e., if this macroblock is not needed to form the set of macroblocks that includes the region to be replaced. Thus, if the first set of macroblocks extends to the boundary of the GoP, the second set of macroblocks is usually empty. Further, it is understood that the second set of macroblocks can include macroblocks in two or more P-frames or three or more B-frames, since the second set of macroblocks can be used with reference to the region where additional frames are replaced, depending on the video encoder used first.

[0059] If the area to be replaced is limited to a single video frame, the first set of macroblocks can be determined by referring only to the macroblock partitioning of the frame. More precisely, the first set is all macroblocks where the area overlaps (i.e., macroblocks that have a non-empty intersection within the pixel space). If the area extends over multiple frames, this operation is repeated for each frame. In the special case where the area is identically repeated throughout all video frames and the macroblock partitioning is constant over all those frames, the first set of macroblocks is a copy of what was determined (by the overlap criterion) for the initial frame for each of the subsequent frames. The second set of macroblocks can be determined based on the first set and the in-frame and inter-frame reference patterns within the signed predictive coded video sequence. Since such references by definition do not extend beyond GOP boundaries, the search for macroblocks to be included in the second set can be limited to the GOP or those GOPs to which the first set of macroblocks belongs.

[0060] The possible results of step 614 are shown in FIG. 4, where each column represents one video frame of video sequence V, and each row represents one macroblock at a specific position within the frame (e.g., the top left macroblock). In FIG. 4, further, the references between macroblocks are shown as curling arrows, and the boundary between two consecutive GoPs, GoP1 and GoP2, is shown as a dashed vertical line. Note that the inter-frame reference is defined at the level of one macroblock position in FIG. 4. Further, the macroblocks that are diagonally hashed are the macroblocks that are directly affected by the request to replace the region, and they are all located within the first frame and form the first set 401 of macroblocks. The macroblocks with dotted shading are all macroblocks that directly (second frame) or indirectly (third frame) reference the macroblocks within the first set, and these are identified as the second set 402 of macroblocks. As expected, the second set of macroblocks does not extend beyond the GoP boundary.

[0061] Note that the configuration of the first and second sets of macroblocks seen in FIG. 4 may be changed when intra-frame reference is introduced, but it is not necessarily so. For example, when the macroblock position corresponding to the first row references the macroblock position corresponding to the second row, the first and second sets of macroblocks 401, 402 remain unchanged.

[0062] The next step 616 of method 600 is shown with respect to FIG. 5B showing the video sequence V and the signed video bitstream B after the region replacement has been performed. In FIG. 5B, diagonal hashing is used for the first set of macroblocks 101.2, 101.3, and dotted shading is used for the macroblocks 101.4, 101.5 of the second set of macroblocks.

[0063] In step 616, the archive object 104 is added to the signed video bitstream B. The archive object 104 includes the first and second sets of data units, i.e., the fingerprints of the data units representing the first and second sets of macroblocks determined in step 614. The fingerprints are not strictly required, but preferably each is an individual fingerprint belonging to exactly one data unit. In an implementation where the fingerprints are calculated from macroblocks reconstructed from a data unit that is one of the above-described options and where the re-encoding is expected to faithfully preserve these macroblocks, the fingerprints of the second set of data units need not be included in the archive object 104. At the level of the signed video bitstream B, the archive object 104 can have a format similar to that of the data unit 102 and the signature unit 103 in that the archive object 104 can be separated from the video bitstream without decoding.

[0064] In FIG. 5B, the first set of data units corresponds to the third, fourth, and fifth data units 102, and the second set of data units corresponds to the sixth and seventh data units 102. Thus, the (two) archive objects 104 added to the signed video bitstream B in step 616 together have fingerprints h 3 、h 4 、h 5 、h 6 、h 7Store it. Optionally, each archive object 104 can include a digital signature of these fingerprints, or a digital signature of a combination of these fingerprints within this archive object 104, or can include a digital signature of the combined fingerprints. Further optionally, the archive object 104 may also include the positions of the first and second sets of macroblocks whose signatures are archived. The position can refer to, for example, the position of the macroblock within the frame in frame coordinates, which corresponds to the position within the bit string. When a static macroblock partition is used, the position of the macroblock 101 can be represented as a macroblock sequence number or another identifier. For example, a bit string may be formed by concatenating the fingerprints in the same order as the macroblock sequence within the frame.

[0065] Next, the execution flow of the editing method 600 proceeds to step 618, where the first set of data units is edited according to the request to replace the region. In other words, step 618 leaves a signed video bitstream B in which some data units are replaced or modified.

[0066] In some embodiments, step 618 may include decrypting the data unit into the reconstructed macroblock 618.1, providing the edited macroblock by performing the requested replacement on the reconstructed macroblock 618.2, and providing the edited data unit by encoding the edited macroblock 618.3.

[0067] Optionally, the macroblocks that occur after substitution 618.2 (e.g., macroblocks 101.2 and 101.3 in FIG. 5B) may be encoded as independently decodable data units. An independently decodable data unit may correspond to an I-frame in the H.264 or H.265 encoding specification, an encoded macroblock that does not reference another macroblock, or a data unit equivalent thereto. This is consistent with the recognition that the substitution may introduce a sudden temporal change in the video sequence and may degrade the performance of predictive encoding.

[0068] In the next step 620 of method 600, a second set of data units is re-encoded. A generally desirable goal is to decode the second set of re-encoded data units into macroblocks that are as close as possible to the second set of (original) macroblocks. Referring to the use of the decoder, the goal is that the second set of re-encoded data units generates substantially the same reference buffer content and / or substantially the same decoder state. However, while the macroblocks within the second set (e.g., macroblocks 101.4 and 101.5 in FIG. 5B) are assumed not to change, they include references to the macroblocks within the first set, which generally become unusable when the first set of macroblocks changes. Due to the sudden temporal change introduced by the editing, the first set of macroblocks ceases to provide a promising basis for predicting the second set of macroblocks. Re-encoding 620 can be facilitated by an optional sub-step 618.1 that outputs a macroblock reconstructed based thereon, i.e., by decoding the second set of data units.

[0069] The second set of macroblocks can advantageously be re-encoded 620 using reduced data compression. This is to be understood against the background that the video sequence V is encoded with a given normal level of data compression. More precisely, the second set of macroblocks is expected to be encoded at a reduced data compression level as compared to normal data compression.

[0070] In step 620, in order to encode the second set of macroblocks, the main options are non-predictive encoding and predictive encoding. If non-predictive encoding is used, the second set of macroblocks is encoded in a form that can be decoded independently as an intra-encoded block, also called an I-block. Thus, the second set of encoded macroblocks is represented by independently decodable data units. Under this option, it is further possible to use reversible encoding for the second set of macroblocks. For example, the second set of macroblocks can be represented by an "unencoded" "raw" block, such as a plane list of the original values for each position within the macroblock in an appropriate color space.

[0071] In the case of predictive coding, which is the second option, step 620 may be performed by re-encoding a second set of data units using predictive coding with reference to the first set of edited data units. More precisely, applying step 620 to a data unit can include: 620.1 obtaining a macroblock reconstructed from a further data unit directly referenced by the data unit; 620.2 decoding the data unit using the reconstructed macroblock; 620.3 obtaining a reconstructed edited version of the macroblock; and 620.4 providing an edited data unit by re-encoding the aforementioned data unit by predictive coding with reference to the reconstructed edited version of the macroblock. Sub-step 620.1 can include decoding a further data unit (see step 618.1), and the further data unit can belong to the first or second set of data units. In sub-step 620.3, the reconstructed edited version of the macroblock can correspond to the image data generated by performing the required replacement on the reconstructed macroblock in sub-step 618.2 above. Sub-step 620.4 may include representing the macroblock in the second set (i.e., the macroblock being processed) with respect to the difference or correction to the reconstructed edited version of the macroblock (derived from the aforementioned further data unit). As described above, this option is mainly useful when the editing performed on the first set of macroblocks is relatively limited or when predictive coding is not fully executed.

[0072] In an optional final step 622 of method 600, one or more signature units 105 associated with the first set of edited data units and the second set of re-encoded data units are provided to the signed video bitstream B. The signature unit 105 provided in step 622 can have the same structure as the signature unit 103 described above. Thus, the signature unit 105 may include a bit sequence derived from the fingerprint of one edited data unit and a digital signature of that bit sequence, or the signature unit 105 may include only the digital signature. The signature unit 105 provided in step 622 may be a newly generated signature unit, as proposed in FIG. 5B. Alternatively, the signature unit 105 is provided by editing an existing signature unit, in particular by extending it with a further digital signature.

[0073] As already mentioned, the steps of any method disclosed herein need not be performed in the exact order described, unless explicitly stated otherwise. This is particularly shown by the editing method 600, and it is clearly possible to perform step 616 before, between, or after the partial sequence of steps 618 and 620, as desired.

[0074] In some embodiments, method 600 further includes an initial step 610 of providing at least one signature unit 103. In the use case that is mainly considered of interest, it is understood that step 610 is performed by a different entity than steps 612, 614, 616, 618, 620, and 622 of method 600, and / or step 610 is performed at an earlier time. In any case, step 610 is separated from the subsequent steps 612, 614, 616, 618, 620, and 622 by a storage period that justifies the signature for relatively insecure data transfer and / or to ensure a desired level of data security.

[0075] Optional step 610 can include sub-step 610.1 of calculating a plurality of fingerprints from each data unit associated with the signature unit, deriving a bit sequence from the plurality of fingerprints 610.2, and obtaining a digital signature of the bit sequence 610.3, where the bit sequence is a combination of the aforementioned plurality of fingerprints, or a fingerprint of the combination of fingerprints, obtaining a digital signature of the bit sequence 610.3. Appropriate implementations of fingerprint calculation 610.1, bit sequence derivation 610.2, and digital signature 610.3 are described in detail above. In particular, the bit sequence related to the digital signature within the signature unit 103 may be a combination of fingerprints of the associated data unit 102, or a fingerprint of the combination of fingerprints of the associated data unit 102. The combination (or "document") may be a list of the literal representations of each fingerprint or another concatenation.

[0076] Having finished the description of the editing method 600 above, next, attention is paid to the receiver side. More precisely, with reference to the flowchart of FIG. 7, a method 700 for verifying the signed video bitstream B will be described. Here too, it is assumed that the signed video bitstream B is obtained by predictive coding of the video sequence V and optionally by subsequent editing operations. It is not essential that the signed video bitstream B has been processed according to the editing method 600. Furthermore, the signed video bitstream is assumed to include a data unit 102, an associated signature unit 103, and an archive object 104. Here, each data unit 102 represents at most one macroblock 101 within a frame 100 of the predictively coded video sequence V, and each signature unit 103 is a digital signature of a bit sequence (e.g., H 1 、H 2 ) (e.g., s(H 1 )、s(H 2)(and optionally including the bit string itself, the archive object 104 may include at least one fingerprint that is an archived fingerprint of a data unit that does not currently exist in the bitstream B and / or has been edited. It is irrelevant to the verification method 700 and, usually, on the receiver side, it is not possible to determine whether a particular signature unit 103 was added in relation to an edit (e.g., by the editing method 600) or was part of the original unedited bitstream B.)

[0077] In an optional first step 710 of method 600 that is only executed in some embodiments (the "document approach"), the bit string H within one signature unit 103 1 is verified using the digital signature s(H 1 ) to verify that the fingerprint contained therein is authentic in a manner known per se. As shown in FIG. 10A, the verification can be performed using a cryptographic element 1001 located in the receiver device 800 and storing the public key. This can be described as an asymmetric signature setup, where signing and verification are separate cryptographic operations corresponding to a private key / public key. Without departing from the scope of the present disclosure, other combinations of symmetric and / or asymmetric verification operations are possible. If the result V1 of the bit string verification is negative (rejected), the execution of method 700 ends. Instead, if the result V1 is positive (approved), the execution of method 700 proceeds to the second step 712.)

[0078] In the second step 712, the fingerprint h 1 of each data unit 102 associated with the signature unit 103, h 2,... is obtained. An independent determination 712.1 regarding how to obtain the fingerprint can be made for each data unit. More precisely, the fingerprint is calculated 712.2 from the data unit or 712.3 the fingerprint is obtained from the archive object 104 within the bitstream. As described above, the fingerprint can be calculated from the reconstructed macroblocks obtained by decoding the data unit 102 directly from this data unit 102 (a subset thereof), for example, from the transform coefficients or other video data therein. In FIG. 10, the fingerprint TIFF0007682982000003.tif5170 is calculated from the data unit 102, and the remaining fingerprints h 3 , h 4 , h 5 , h 6 are found to be obtained from the archive object 104. Therefore, only the fingerprint TIFF0007682982000004.tif5170 may fail the verification in the next fourth step 716, in which case the failure suggests that an improper operation has been performed on the bitstream B.

[0079] In the third step 714, the bit sequence TIFF0007682982000005.tif5170 is derived from the fingerprint thus obtained. This can be done according to a pre-agreed rule, for example, by a procedure similar to that described within step 610.2. It is recalled that the bit sequence may be a combination of the obtained fingerprints or may be the fingerprint of the combination of the fingerprints.

[0080] Finally, in a fourth step 716, the data unit associated with the signature unit 103 is verified using the digital signature within the signature unit 103. To avoid misunderstanding, it should be noted that the verification in step 716 of the data unit is indirect without any processing acting on the data unit itself.

[0081] In embodiments where the signature unit 103 does not include the bit string H 1 , step 716 is performed by verifying the bit string 1 TIFF0007682982000006.tif5170 derived using the digital signature s(H . For example, the derived bit string TIFF0007682982000007.tif5170 can be verified using the public key belonging to the same key pair as the private key used to generate the digital signature s(H 1 ). In Figure 10B, this is shown by supplying the derived bit string TIFF0007682982000008.tif5170 and the digital signature s(H 1 ) to the cryptographic entity 1001 storing the public key that outputs a binary result W1 representing the result of the verification.

[0082] Alternatively, in embodiments where the signature unit 103 includes the bit string H 1 (the "document approach"), the bit string H 1 is first verified in step 710 and then the verified bit string H 1 in step 716 is compared with the derived bit string TIFF0007682982000009.tif5170. The comparison may be a bit-by-bit equivalence check, as suggested by the function block 1002 in Figure 10A, resulting in a true or false output V2. If the result V2 of the comparison is true, it can be concluded that the signed video bitstream 100 is authentic as far as this signature unit 103 is concerned.

[0083] Next, the execution of method 700 may include repeating the relevant ones of steps 710, 712, 714 described above for any further signature units 103 within the signed video bitstream 100. If the result is affirmative for all signature units 103, the signed video bitstream 100 is concluded to be valid and it may be further consumed or processed. Otherwise, the signed video bitstream 100 is considered to be non-genuine and may be isolated from further use or processing.

[0084] Note that the verification of data units within the first set is based on a different trust relationship than the verification of data units within the second set. The data units within the first set are verified by trusting the entity that created the digital signature s(H 1 ), i.e., the owner of the private key when asymmetric key cryptography is used. The data units within the second set are verified by trusting the entity that edited the signed bitstream B and created the archival object.

[0085] In some embodiments of verification method 700, the determination of sub-step 712.1 is guided by the positions indicated in archive object 104. These positions are the positions of macroblocks 101 represented by data units 102 to which the archived fingerprints pertain. By having access to these macroblock positions, a recipient can perform a trustable integrity check based on the assumption that any macroblock 101 within video frame 100 that cannot be reconstructed from data unit 102 within signed video bitstream B is encoded by another data unit that can always obtain its fingerprint from archive object 104. If archive object 104 does not indicate the positions of these macroblocks, the recipient can insert, for example, a missing fingerprint, i.e., a fingerprint that cannot be computed from data unit 102 within signed video bitstream B, by a trial-and-error approach. The trial-and-error approach includes performing steps 714 and 716 for each possible way of inserting the archived fingerprint from archive object 104 (each such way of insertion can be presumed to be a permutation of the positions of the missing macroblocks), and concluding that the signed video bitstream B is corrupt only if all of these executions fail.

[0086] Aspects of the present disclosure have been mainly described above with reference to several embodiments. However, as will be readily understood by those skilled in the art, other embodiments other than those disclosed above, as defined by the appended claims, are equally possible within the scope of the inventive concept. In particular, note that the above description of the various embodiments focuses on the predicted-encoded video. However, as long as the entity that signs the video and the entity that verifies the video can access the reconstructed or decoded frames of the video, the same technique can be used not only for the predicted-encoded video but also for any video. In the case of prediction-based encoding, it can be seen that it is practically convenient to utilize the fingerprint of a group of pixels that is also used as a macroblock in encoding. However, in general, the fingerprint may be calculated from a group of pixels grouped in other ways. When the frame is decoded, it does not matter how the pixels were divided for encoding. For example, depending on the type of editing expected, it may be useful to divide the decoded image into groups of pixels that are smaller or larger than those used for encoding. For example, if masking is always performed in a rectangular form, a coarser pixel division than that used for encoding may be sufficient for the signing process. On the other hand, if it is assumed that masking can be performed more closely following the contour of the object to be masked, a finer division may be useful for the signing process.

Claims

1. A method for editing a signed video bitstream obtained by predictive coding of a video sequence, comprising: The signed video bitstream includes data units and associated signature units, each data unit representing at most one macroblock within a video frame of the predicted-coded video sequence, and each signature unit including a digital signature of a bit sequence derived from a plurality of fingerprints of exactly one associated data unit respectively; The method includes: Receiving a request to replace an area of at least one video frame; Determining a first set of macroblocks included in the area and a second set of macroblocks that directly or indirectly reference macroblocks within the first set; Adding an archive object to the signed video bitstream, the archive object including fingerprints of a first set and a second set of data units respectively representing the determined first set and second set of macroblocks, and the archive object including positions of the first set and second set of macroblocks; Editing the first set of data units according to the request to replace the area of the at least one video frame, including: Decoding the data units into reconstructed macroblocks; Providing edited macroblocks by performing the requested replacement on the reconstructed macroblocks; and Providing edited data units by encoding the edited macroblocks; Editing the first set of data units; Re-encoding the second set of data units; A method.

2. The method according to claim 1, wherein after the replacement, the macroblocks are encoded as independently decodable data units.

3. The method according to claim 1, further comprising decoding data units within the second set using the reconstructed macroblocks to facilitate the re-encoding.

4. The method according to claim 1, wherein a second set of the data units is re-encoded using reduced data compression. **Claim 5** The method according to claim 1, wherein a second set of the data units is re-encoded as independently decodable data units. **Claim 6** Providing, in the signed video bitstream, one or more signature units associated with the first set of the edited data units and the second set of the re-encoded data units The method according to claim 1, further comprising. **Claim 7** Calculating a plurality of fingerprints of each data unit associated with the signature unit; Deriving a bit string from the plurality of fingerprints; Obtaining a digital signature of the bit string; The method further comprising initially providing the signature unit, wherein The method according to claim 1, wherein the bit string is a combination of the plurality of fingerprints or a fingerprint of the combination. **Claim 8** Each fingerprint is a) from the data unit, or b) from a macroblock reconstructed from the data unit, or c) from intermediate reconstructed data derived from the data unit, The method according to claim 7, wherein the fingerprint is calculated. **Claim 9** The method according to claim 1, wherein the received request is to replace regions within a plurality of video frames. **Claim 10** A method for verifying a signed video bitstream obtained by predictive encoding of a video sequence, wherein The signed video bitstream includes data units, associated signature units, and an archive object, each data unit representing at most one macroblock within a frame of the predictively encoded video sequence, each signature unit including a digital signature of a bit string, the archive object including at least one archived fingerprint of a data unit, and the archived fingerprint indicating a position of a macroblock represented by the related data unit, The method includes Calculating a fingerprint of the data unit, or Obtaining an archived fingerprint from the archive object obtaining a fingerprint of each data unit associated with the signature unit; deriving a bit sequence from the obtained fingerprint; verifying the data unit associated with the signature unit using the digital signature within the signature unit, verifying the bit sequence derived using the digital signature, or if the signature unit includes a bit sequence, verifying the bit sequence within the signature unit using the digital signature and comparing the derived bit sequence with the verified bit sequence including verifying the data unit; including a method. **Claim 11**: The obtaining of the fingerprint includes determining whether to calculate the fingerprint of the data unit based on the position indicated by the archive object or obtain the fingerprint from the archive object. The method according to claim 10. **Claim 12** A device comprising a processing circuit arranged to execute the method according to any one of claims 1 to 11. **Claim 13** A computer program comprising instructions that, when executed by a computer, cause the computer to execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Device and method for signing video segment including at least one picture group

    JP2022174726A

  • Signed video data with linked hash

    JP2023056491A

  • Signed video data using salted hash

    JP2023056492A

  • System and method for providing cryptographic video verification

    US20140010366A1

  • Verifying provenance of digital content

    US20200275166A1

Cited By

  • Editable video data signed on uncompressed data

    JP2024086623A