Method and system for applying digital signature to audio and video data

The method and system digitally sign video and audio sequences by generating and inserting digests and signatures to ensure authenticity and simultaneous capture, addressing the challenge of tampering and synchronization in surveillance data.

JP2025106190AActive Publication Date: 2025-07-15AXIS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024204121
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-22
Publication Date
2025-07-15
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

Existing methods for authenticating video and audio data captured by surveillance systems are insufficient to ensure that the data has not been tampered with and were captured simultaneously from the same scene, as clock synchronization can be manipulated to forge evidence.

Method used

A method and system for digitally signing video and audio sequences by generating digests and signatures for each sequence portion, inserting signatures into corresponding portions of the other sequence, ensuring that both sequences are linked and authentic.

Benefits of technology

Ensures that video and audio data are authenticated as originating from the same scene and captured simultaneously, robust against tampering and transmission errors, even when data is partially transmitted or stored.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106190000001_ABST
    Figure 2025106190000001_ABST
Patent Text Reader

Abstract

To provide a method, a digital signature system, and a storage medium which guarantee that audio data and video data are not altered after being taken in.SOLUTION: A method generates a first video digest and a first audio digest by applying digest algorithm to a first video section of a video sequence and to a first audio section of an audio sequence S1 and S2; generates a first video signature and a first audio signature by applying a digital signature to each digest S3 and S4; inserts each digest to a first target audio section and to a first target video section S5 and S6; generates a second video digest and a second audio digest from the first target video section and the first target audio section S7 and S8; and generates a second video signature and a second audio signature by applying a digital signature to each of the second video digest and the second audio digest S9 and S10.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of authentication of audio data and video data, and more particularly to protection against the fact that audio data and video data have not been tampered with after being captured.

Background Art

[0002] Cameras for capturing video are widely used for surveillance. Such surveillance or observation is often carried out for security interests, but may also be carried out for, for example, safety or industrial process control. In some scenarios, it is useful to capture not only images but also audio. A stand-alone microphone may be used to capture audio. However, as a practical solution, many surveillance cameras incorporate one or more microphones. In this way, the same device can be used to observe a scene.

[0003] If an event such as a crime occurs in the observed scene, the video data and audio data captured by the surveillance camera may be useful for investigating and proving what happened. In such a situation, it is important to ensure that neither the video data nor the audio data has been tampered with after being captured by the camera's image sensor and microphone. One way to prove that video data or audio data has not been tampered with is to attach a digital signature to the data. By signing the video data and audio data in the camera, it becomes possible to authenticate the data later and ensure that no one has tampered with the data after it has been transmitted from the camera. For example, in an evidence management system where data is transmitted from a camera, there are also systems that attach digital signatures. In such a system, the digital signature can be used to prove that the data has not been tampered with after being input into the system.

[0004] Digital signature methods for video data and audio data are known. However, when video and audio from the same observation or captured scene are used as evidence, it may not be sufficient to simply show that the video data and the audio data are each authentic on their own. It may also be necessary to show that the video data and the audio data were actually captured at the same location at the same time. Some systems address this by relying on timestamps in the video data and the audio data. If the clocks of the cameras and other devices within a surveillance or observation system are synchronized and the timestamp of a video frame is equal to the timestamp of an audio frame, it can be assumed that the video frame and the audio frame were captured at the same time. However, in general, the user of a camera or surveillance system can adjust the clocks of the cameras and other devices within the surveillance system. In this way, it may be possible to forge surveillance data by replacing the original audio sequence captured at the same time as the video sequence with another audio sequence captured at a different time but after illegally resetting the clock. Therefore, there is still a need for a convenient and secure way to digitally sign video data and audio data in a way that enables the video data and the audio data to be authentic, to have originated from the same captured scene, and to have been captured at the same time. SUMMARY OF THE INVENTION

[0005] An object of the present invention is to provide a method capable of authenticating a video sequence and an audio sequence captured simultaneously so as to represent the same captured scene. Another object is to provide a method for digitally signing a video sequence and an audio sequence that can be applied to a data stream from a camera without the need to encapsulate the video data and the audio data in a file or a container. An object of the present invention is also to provide a system that enables the authentication of a video sequence and an audio sequence that simultaneously capture an observed scene.

[0006] According to a first aspect, these and other objects are achieved, fully or at least in part, by the method according to claim 1. Thus, this is achieved by a method of digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously so as to represent the same scene in which they were captured, the video sequence including successive video portions and the audio sequence including successive audio portions. The method includes generating a first video digest by applying a digest algorithm to a first video portion of the video sequence, generating a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence, generating a first video signature by digitally signing the first video digest, generating a first audio signature by digitally signing the first audio digest, inserting the first video signature into a first target audio portion of the audio sequence, inserting the first audio signature into a first target video portion of the video sequence, generating a second video digest by applying a digest algorithm to the first target video portion including the first audio signature, generating a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature, generating a second video signature by digitally signing the second video digest, and generating a second audio signature by digitally signing the second audio digest. In this way, the video sequence and the audio sequence are linked. To authenticate the video sequence, the audio sequence is required, and vice versa. For example, if the audio sequence is replaced after capture in order to give the impression that the people in the video sequence are saying something other than what is actually being said, an attempt to verify the audio signature included in the video sequence will fail.Such authentication failures can be due to intentional tampering of the video or audio, but can also be caused by transmission errors. Thus, when authentication is successful, it indicates that the video sequence and the audio sequence are genuine, but when authentication fails, it only indicates that something has occurred after the video sequence and the audio sequence have been captured, and does not indicate whether the video and / or audio has been deliberately tampered with or whether there have been accidental changes to the video and / or audio.

[0007] Note that the video signature and the audio signature do not necessarily need to be generated synchronously here. The video signature can be generated, for example, for each group of pictures (abbreviated as GOP). The duration of the GOP depends on the frame rate of the image sensor and the GOP length used to encode the video data. As an example, at a frame rate of 30 frames per second and a GOP length of 62, the video signature of the GOP is generated approximately once every 2 seconds. The audio sequence may be signed more frequently or less frequently than the video sequence. As an example, if audio frames or packets are created every 10 ms and an audio signature is generated for every 100 groups of audio frames, the audio signature is generated once per second. When the video signature is generated, it is inserted into the target audio portion of the audio sequence. This target audio portion may be the currently captured and encoded audio portion when the video signature is generated, or it may be the next audio portion for which capture and encoding have not yet started when the video signature is generated. Similarly, when the audio signature is generated, it is inserted into the target video portion, which may be the currently captured and encoded video portion, or the next video portion. When the video signature and the audio signature are generated at the same rate, that is, when the video portion and the audio portion for which the signature is generated are captured during equal time lengths, each target video portion contains one audio signature after insertion, and each target audio portion contains one video signature. On the other hand, when the audio signature and the video signature are generated at different intervals, there is no one-to-one relationship of signatures between the video sequence and the audio sequence. When the audio signature is generated more frequently than the video signature, each target video portion can contain two or more audio signatures, but each target audio portion contains one video signature, and there are audio portions that do not contain a video signature between the target audio portions. As will be understood by those skilled in the art, the situation is reversed when the video signature is generated more frequently than the audio signature.In this regard, it should also be noted that even in a scenario where video signatures and audio signatures are generated at the same rate, they do not necessarily need to be generated simultaneously. When video signatures and audio signatures are generated by the same signature circuit or processor, they compete for computing resources and must be generated one after another.

[0008] As used herein, a "target video portion" is a video portion into which an audio signature is inserted. As can be understood from the foregoing description, when an audio signature is generated less than once per video portion, there may be other video portions between the target video portions. Similarly, a "target audio portion" is an audio portion into which a video signature is inserted. When a video signature is generated less frequently than once per audio portion, there may be other audio portions between the target audio portions.

[0009] It should be noted that an audio sequence that captures the same observation or scene as the video sequence can capture very well the sound generated outside the field of view of the camera that captures the video sequence. A microphone placed inside or near the camera can capture, for example, the voices of people within and just outside the field of view of the camera, or the sound of a window breaking in front of, above, or behind the camera. All such sounds are considered to belong to the same observation or captured scene as those within the field of view of the camera.

[0010] Variations of this method are defined in the dependent claims.

[0011] In some variations, the method further includes inserting a second video signature into a second target audio portion of the audio sequence and inserting a second audio signature into a second target video portion of the video sequence. Thereby, the second target audio portion includes the second video signature that is affected by the first audio signature. As a result, verification of the second video signature confirms a link to the audio sequence, more precisely to the first audio portion. Similarly, the second target video portion includes the second audio signature that is affected by the first video signature. Thereby, verification of the second audio signature confirms a link to the video sequence, more precisely to the first video portion.

[0012] The method can further include inserting the first video signature into the video portion of the video sequence and inserting the first audio signature into the audio portion of the audio sequence. Thereby, it becomes possible to authenticate the first video portion even when the first target audio portion is lost, and it becomes possible to authenticate the first audio portion even when the first target video portion is lost.

[0013] Similarly, the method can include inserting a second video signature into the video portion of the video sequence and inserting a second audio signature into the audio portion of the audio sequence, thereby enabling authentication of the second video portion when there is no second target audio portion and authentication of the second audio portion when there is no second target video portion. This makes the authentication procedure more robust against packet loss during transmission or pruning of data storage. For example, when the transmission bandwidth or storage capacity is insufficient, it may be determined to transmit only the audio sequence, which generally consumes fewer bits, and omit the transmission of the video sequence. As will be understood by those skilled in the art, even in a situation where only video data or only audio data is transmitted, the presence of an audio signature within the video sequence or a video signature within the audio sequence still enables indication that the video sequence was linked to the audio sequence at the time of capture, or that the audio sequence was linked to the video sequence at the time of capture.

[0014] The digest algorithm may be a hash function. Hash data is a well-known practical method for creating a digest.

[0015] Each video portion may be a group of pictures, also referred to as a GOP.

[0016] Each audio portion may be a group of audio packets or audio frames.

[0017] In a video sequence, each audio signature may be inserted into each SEI message or open bitstream unit of the video sequence. SEI messages are available in the h.26X coding standard and are also referred to as SEI frames or SEI NAL units. A particular SEI NAL unit of unregistered type of user data may be used for signature insertion. Open bitstream units are available in the AV1 coding standard and may also be referred to as OBUs or metadata OBUs.

[0018] In an audio sequence, each video signature is inserted into each data stream element or header of the audio sequence.

[0019] In some variations, the video sequence and the audio sequence are captured by a single device comprising an image sensor and a microphone. This is a practical and efficient way to capture video data and audio data in the observed scene. When the image sensor and the microphone are incorporated into the same device, the fact that the video sequence and the audio sequence are captured by the same device can also make it possible to indicate that they captured the same scene.

[0020] According to a second aspect, the above object is achieved, wholly or at least in part, by the digital signature system according to claim 11. Thus, this is achieved by a digital signature system for digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously so as to represent the same scene in which they were captured, the system being arranged to receive a video sequence comprising successive video portions, a video input function arranged to receive an audio sequence comprising successive audio portions, a video digest function arranged to apply a digest algorithm to the video portions of the video sequence, thereby generating a video digest, an audio digest function arranged to apply a digest algorithm to the audio portions of the audio sequence, thereby generating an audio digest, a signature function arranged to generate a video signature by digitally signing the video digest and an audio signature by digitally signing the audio digest, a video signature insertion function arranged to insert the video signature into the audio sequence, and an audio signature insertion function arranged to insert the audio signature into the video sequence, the system comprising a circuit configured to perform the functions, the video digest function being arranged to generate a first video digest by applying a digest algorithm to a first video portion of the video sequence, the audio digest function being arranged to generate a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence, the signature function being arranged to generate a first video signature by digitally signing the first video digest and a first audio signature by digitally signing the first audio digest, the video signature insertion function being arranged to insert the first video signature into a first target audio portion of the audio sequence, and the audio signature insertion function being arranged to insert the first audio signature into a first target video portion of the video sequence.The video digest function is arranged to generate a second video digest by applying a digest algorithm to a first target video portion including a first audio signature, and the audio digest function is arranged to generate a second audio digest by applying a digest algorithm to a first target audio portion including a first video signature. The signature function is arranged to generate a second video signature by digitally signing the second video digest and to generate a second audio signature by digitally signing the second audio digest. Using such a system, it is possible to determine whether a video sequence and an audio sequence capturing the same scene are authentic. The ability to insert an audio signature into the video sequence and an audio signature into the audio sequence makes it possible to verify that the video sequence and the audio sequence were actually captured at the same scene at the same time.

[0021] Embodiments of the digital signature system are defined in the dependent claims.

[0022] The digital signature system of the second aspect can generally be implemented in the same way as the method of the first aspect, with the attendant advantages.

[0023] According to a third aspect, the above object is achieved by a camera comprising an image sensor, a microphone, and a digital signature system according to the second aspect.

[0024] According to a fourth aspect, the above object is achieved by a computer-readable storage medium including instructions that, when executed by a device having processing capabilities, cause the device to execute the method of the first aspect.

[0025] The further scope of application of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that various changes and modifications within the scope of the present invention will be apparent to those skilled in the art from this detailed description, and that the detailed description and specific examples, while indicating preferred embodiments of the present invention, are given by way of illustration only.

[0026] Accordingly, it should be understood that the present invention is not limited to the specific components of the described apparatus or the steps of the described method, and that such apparatus and methods may vary. It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein and in the appended claims, the articles "a," "an," "the," and "said" are intended to mean that there is one or more of the elements unless the context clearly indicates otherwise. Thus, for example, reference to "an object" or "the object" may include several objects, and the like. Further, the word "comprising" does not exclude other elements or steps.

[0027] Here, with reference to the accompanying schematic drawings, the present invention will be described in more detail by way of example.

Brief Description of the Drawings

[0028]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

DETAILED DESCRIPTION OF THE INVENTION

[0029] FIG. 1 is a view of a scene 1 observed using a camera 2 disposed on a lighting pole. As will be further described below with reference to FIG. 7, the camera 2 includes an image sensor (not shown in FIG. 1) and a microphone (not shown in FIG. 1). In the example shown in FIG. 1, the camera 2 is directed at the door of a building 3. The image sensor of the camera 2 can capture objects in the space between the camera 2 and the building. Also, the microphone within the camera 2 can capture sounds around the camera 2. The microphone can generally simply have a range wider than the field of view of the camera 2. For example, the microphone can capture sounds originating from behind the camera or on the opposite side of the building 3.

[0030] The video sequence and the audio sequence captured by camera 2 may be useful as evidence in the event that a crime was committed in the observed scene 1. In some cases, for example, the video may be sufficient to confirm what happened in order to identify who entered the building through the window. In other examples, it may also be necessary to listen to the audio from the scene of the crime, such as when a person was assaulted and the alleged perpetrator claims to have acted in self-defense because the victim of the assault was verbally threatening. In order for the video sequence and the audio sequence to be usefully combined as evidence, it may be necessary to prove that they were actually captured simultaneously in the same scene, or even by the same device. Therefore, it may not be sufficient to simply authenticate the video sequence and the audio sequence separately to show that they have not been tampered with after they were captured. It may also be necessary to authenticate that they are actually related.

[0031] According to an embodiment of the present invention, a video signature may be generated for a video sequence, and an audio signature may be generated for an audio sequence in a manner similar to that performed when digitally signing video data and audio data separately. The video signature may be included in the audio sequence and the audio signature may be included in the video sequence so that it is also possible to authenticate that the video sequence and the audio sequence are complete. Further, when a video signature is generated, it should include data from a preceding audio signature, and when an audio signature is generated, it should include data from a preceding video signature. In this way, authentication of the video signature fails not only when the video sequence has been tampered with, but also when the audio sequence has been tampered with, and vice versa. In order to successfully authenticate the video signature, both the video data and the audio data must be intact. Similarly, for authentication of the audio signature, both the audio data and the video data must be intact. If the video sequence and the audio sequence are actually genuine, but were combined some time after capture and thus were not captured simultaneously in the same scene, authentication will fail.

[0032] FIG. 2 shows a video sequence 10 and an audio sequence 20. The video sequence consists of a sequence of encoded image frames IF1, IF2...IF n which are arranged in groups of pictures, or GOPs for short. Each GOP starts with an intra-coded image frame, also called an I-frame, followed by several inter-coded image frames. The inter-coded image frames may be forward prediction frames, also called P-frames, or bidirectional prediction frames, also called B-frames.

[0033] The audio sequence 20 consists of encoded audio frames or packets AF1, AF2,...AF nIt is composed of a sequence. The audio sequence does not have the same type of hierarchical structure of frames as a video sequence having key frames (intra-coded frames) and delta frames (inter-coded frames), and thus does not have the same type of specific grouping. However, for the purpose of the digital signature method of the present invention, the audio sequence is also divided into groups of audio frames. It may be worth mentioning that the video part and the audio part for signature can be arbitrarily small, such as a single image frame or a single audio sample. However, for bitrate efficiency and computational efficiency, generally, it is preferable to sign several image frames or audio frames at a time.

[0034] Digital signatures are generated to enable authentication of video sequences and audio sequences, as well as the connection between video sequences and audio sequences. Each digital signature is generated for a part of each sequence. In this example, for the video sequence, a video signature is generated for the video part in the form of a GOP, and an audio signature is generated for the audio part of the audio sequence in the form of a group of audio frames.

[0035] Next, with reference to the flowchart of FIG. 4, an example of a method for digitally signing video sequences and audio sequences will be described.

[0036] As a basis for the first video signature VS1, as shown in FIG. 3, a first video digest is generated by applying a digest algorithm to the first video portion 11. In this example, the digest is a hash. Therefore, hereinafter, the video digest will be referred to as a video hash. Thus, the first video portion 11 is hashed to generate a first video hash (step S1). The first video signature VS1 is generated by digitally signing the first video hash (step S3). This is exactly the first video signature of the video sequence, and if it is generated before any audio signature is inserted into the video sequence, this first video signature cannot directly authenticate any link to the audio sequence, but note that it is still useful to indicate that the first video portion itself has not been tampered with. Once the first video signature VS1 is generated, it is inserted into the first target audio portion of the audio sequence (S5). This insertion will be described in more detail below with reference to FIG. 3.

[0037] Similarly, as the basis for the first audio signature AS1, the first audio digest is generated by applying a digest algorithm to the first audio portion 21 of the audio sequence 20. For the video sequence, in this example, the digest is a hash. Thus, the first audio portion 21 is hashed to generate the first audio hash (step S2). The first audio signature AS1 is generated by digitally signing the first audio hash (step S4). Similar to the video sequence, this is exactly the first audio signature of the audio sequence, and if any video signature is generated before being inserted into the audio sequence, this first audio signature cannot authenticate any link to the video sequence, but note that it is still useful to indicate that the first audio portion itself has not been tampered with. Once the first audio signature AS1 is generated, it is inserted into the first target video portion of the audio sequence (S5). This insertion will be described in more detail below with reference to FIG. 3.

[0038] The second video hash is generated by hashing the first target video portion (step S7). The first target video portion is the video portion into which the first audio signature AS1, which was captured some time after the first video portion 11, is inserted. Thus, the second video hash is the hash of the combination of the video data and the first audio signature AS1, thereby linking the video sequence 10 and the audio sequence 21 to each other. The second video signature VS2 is generated by digitally signing the second video hash (step S9).

[0039] Further, a second audio hash is generated by hashing the first target audio portion (step S8). The first target audio portion is an audio portion into which the first video signature VS1, which is captured some time after the first audio portion 21, is inserted. Accordingly, the second audio hash becomes a hash of the combination of the audio data and the first video signature VS1, thereby linking the video sequence 10 and the audio sequence 21 to each other. The second audio signature AS2 is generated by digitally signing the second audio hash (step S10).

[0040] Referring briefly to the flowchart of FIG. 5, before returning to FIG. 3, when the second video signature VS2 is generated, it may be inserted into the second target audio portion (step S11). Similarly, the second audio signature AS2 may be inserted into the second target video portion (step S12). The digital signatures of the consecutive video portions of the video sequence 10 and the consecutive audio portions of the audio sequence 20 can continue in the same way until the end of the sequence. Such "cross-fertilization" of the video sequence and the audio sequence by each signature makes it possible not only to authenticate the video sequence and the audio sequence separately, but also to authenticate their common origin.

[0041] Next, referring to FIG. 3, the insertion of the video signature into the audio sequence and the insertion of the audio signature into the video sequence will be described. Again, the video sequence 10 and the audio sequence 20 are shown. As described above, the image frames are grouped into GOPs, and the audio frames or audio packets are similarly grouped into groups of audio frames or audio packets. Note that the diagram in FIG. 3 is very schematic, and the groups generally actually consist of many more frames.

[0042] As described above, a first video signature VS1 is generated for the first video portion 11. This first video signature VS1 is inserted into the first target audio portion 41 of the audio sequence 20, and in this case, it coincides accidentally with the second audio portion 22. In the audio sequence 20, the position where the first video signature VS1 is inserted is indicated by hatching. Depending on the frequency at which the video signature is generated compared to the length of the audio portion, the first target audio portion 41 can be an earlier or later audio portion than the second audio portion 22. Thus, the first target audio portion can be the first audio portion, the second audio portion, or a later audio portion. A second video signature VS2 is generated for the second video portion 12, and this second video signature VS2 is inserted into the second target audio portion 42. In the illustrated example, the second target audio portion 42 coincides with the third audio portion 23, but as explained above, this depends on the relationship between the frequency at which the video signature is generated and the length of the audio portion. Similarly, a third video signature VS3 is generated for the third video portion 13. The third video signature VS3 is inserted into the third target audio portion 43, which is the fourth audio portion 24 in the illustrated example.

[0043] Looking at the audio sequence 20, it follows the same principle. For the first audio portion 21, a first audio signature AS1 is generated. The first audio signature AS1 is inserted into the first target video portion 31. In the illustrated example, the first target video portion is the second video portion, but as explained for the video signature, this depends on the frequency at which the audio signature is generated compared to the length of the video portion. For the second audio portion, a second audio signature AS2 is generated and inserted into the second target video portion 32, which is the third video portion 13 in this example. For the third audio portion 23, a third audio signature AS3 is generated. The third audio signature AS3 is inserted into the third target video portion 33. In the illustrated example, the third target video portion 33 happens to be the same as the second target video portion 32, i.e., the third video portion 13. Thus, two audio signatures AS2 and AS3 are inserted into the third video portion. Continuing with the fourth audio portion, which is not shown in full length in FIG. 3, a fourth audio signature AS4 is generated and inserted into the fourth target video portion. However, this fourth target video portion is not shown in FIG. 3.

[0044] In the video sequence 10, the audio signature AS n may be inserted into an already available element of the video compression format. For example, if the video sequence is encoded using the h.264 or another h.26X coding format, the audio signature can be inserted into a specified type of SEI message, sometimes called an SEI frame or SEI NAL unit. Other video compression formats provide corresponding elements. As an example, if the video sequence is encoded using AV1, the audio signature can be inserted into a specified type of open bitstream unit, also called an OBU or metadata OBU.

[0045] Similarly, in the audio sequence 20, the video signature VS nIt may be inserted into appropriate existing elements in the audio coding format. For example, when an audio sequence is encoded using AAC, the video signature may be inserted into the data stream element, or abbreviated as DSE. In the PCM audio coding format, the video signature may be inserted into the header of the WAV container in which the audio data is encapsulated. Other audio coding formats offer similar possibilities.

[0046] The video signature may be generated, for example, by cryptographic operations in accordance with the method described in the applicant's European Patent Application No. 4164230. The audio signature may be inserted into the SEI frame or the corresponding element of the video sequence as described in the applicant's European Patent Application No. 4192018. The audio signature can be generated and inserted using the same principle as the video signature.

[0047] Figure 6 is a simplified block diagram of a digital signature system 60 that can be used to execute the digital signature method described above. The digital signature system 60 includes a video input function 61 arranged to receive a video sequence 10 and an audio input function 62 configured to receive an audio sequence 20. The digital signature system 60 further includes a video digest function 63 arranged to apply a digest algorithm to the video portion of the video sequence, thereby generating a video digest. Similar to the above description, in this example, the digest algorithm is a hash function, and thus the video digest function can be called the video hash function 63. The digital signature system 60 also includes an audio digest function 64 arranged to apply a digest algorithm to the audio portion of the audio sequence, thereby generating an audio digest. Regarding the video digest, the audio digest is a hash in this example, and thus the audio digest function can be called the audio hash function 64.

[0048] Furthermore, the digital signature system 60 is arranged with a signature function 65 that generates a video signature by digitally signing a video digest or hash and an audio signature by digitally signing an audio digest or hash, a video signature insertion function 66 arranged to insert the video signature into the audio sequence, and an audio signature insertion function 67 arranged to insert the audio signature into the video sequence.

[0049] In accordance with what has been described above in connection with FIGS. 3-5, the video hash digest function 63 is arranged to generate a first video hash by hashing a first video portion 11 of the video sequence 10, and the audio hash function 64 is arranged to generate a first audio hash by hashing a first audio portion 21 of the audio sequence 20. The signature function 65 is arranged to generate a first video signature VS1 by digitally signing the first video hash and a first audio signature AS1 by digitally signing the first audio hash. The video signature insertion function 66 is arranged to insert the first video signature VS1 into a first target audio portion 41 of the audio sequence 20. The audio signature insertion function 67 is arranged to insert the first audio signature AS1 into a first target video portion 31 of the video sequence 10. The video hash function 63 is further arranged to generate a second video hash by hashing the first target video portion 31 including the first audio signature AS1, and the audio hash function is arranged to generate a second audio hash by hashing the first target audio portion 41 including the first video signature VS1. The signature function 65 is arranged to generate a second video signature VS2 by digitally signing the second video hash and a second audio signature AS1 by digitally signing the second audio hash.

[0050] In the same manner as described in connection with FIG. 5, the video signature insertion function 66 may be further arranged to insert a second video signature VS2 into the second target audio portion 42, and the audio signature insertion function 67 may be further arranged to insert a second audio signature AS2 into the second target video portion 32.

[0051] The functions 61-67 of the digital signature system 60 may be implemented using a general-purpose processor, a dedicated processor, an integrated circuit, an ASIC ("application-specific integrated circuit"), a conventional circuit, and / or a circuit or processing circuit including a combination thereof configured or programmed to implement the disclosed functionality. In the present disclosure, a circuit is hardware programmed or configured to execute or implement the recited functionality. The hardware may be any hardware disclosed herein or other known hardware programmed or configured to execute the recited functionality.

[0052] In a pure hardware implementation, each function can be dedicated and have a corresponding circuit specially designed to implement the function. The circuit may be in the form of one or more application-specific integrated circuits or one or more integrated circuits such as one or more field programmable gate arrays.

[0053] In embodiments that also include software, the circuit can include a processor. Since the processor includes transistors and other circuits therein, it is regarded as a processing circuit or a circuit. In this case, the circuit can be regarded as a combination of hardware and software, and the software is used to configure the hardware and / or the processor.

[0054] It should also be understood that it is possible to have a combination of hardware and software implementations, which means that some of the functions are implemented by dedicated circuits and others are implemented in the form of computer code executed by software, i.e., by a processor.

[0055] Figure 7 shows a simplified block diagram of camera 2. Camera 2 has a lens 71 and an image sensor 72 for capturing an image of the observed scene. Camera 2 further includes a video encoder 73 for encoding the video data captured by the image sensor 72. To capture the audio within the scene, camera 2 is equipped with a microphone 74. Further, camera 2 includes an audio encoder 75 for encoding the audio data captured by the microphone 74. Also, camera 2 includes a digital signature system 60 as shown in FIG. 6. The network interface 76 enables the transmission of the encoded video data and audio data including video signature and audio signature. As will be understood by those skilled in the art, camera 2 can have additional components, but here only the components relevant to the description of the present invention are shown and described.

[0056] To make the above-described methods and systems more resilient to transmission losses, video signatures can also be inserted into the video sequence, and audio signatures can also be inserted into the audio sequence. When a video signature is generated, in this scenario, after the video signature is generated, it may be inserted into the subsequent video portion or the nearest video portion in other ways. Similarly, when an audio signature is generated, it may be inserted into the subsequent audio portion or the nearest audio portion in other ways after the audio signature is generated. In situations where the transmission bandwidth from the camera is low, it may be determined whether to transmit only the video or only the audio. In such cases, it is impossible to authenticate the transmitted sequence unless the corresponding signature is also inserted into the sequence from which it originated. Therefore, if only the video is transmitted, the video signature is not transmitted unless it is also included in the video sequence. Similarly, if only the audio is transmitted, the audio signature is not transmitted unless it is also included in the audio sequence. As a further adaptation, if it is determined to return to "video only" transmission during a temporary bandwidth limitation, it is possible to temporarily suspend the insertion of the audio signature into the video sequence during the "video only" transmission period, thereby enabling the video sequence to be authenticated independently. Of course, this can be applied to "audio only" transmission with the necessary changes. During such periods of "video only" or "audio only" transmission, when the network state improves, the unsent sequence can be locally stored in the camera for later transmission. In such cases, it should be noted that it is impossible to securely authenticate the link between the video sequence and the audio sequence if the insertion of the signature into the "companion sequence" is temporarily suspended during "video only" or "audio only" transmission. It may be useful to insert metadata indicating that the transmission of other sequences has been temporarily suspended into the transmitted sequence.

[0057] Those skilled in the art will understand that the above embodiments can be modified in many ways and the advantages of the present invention shown in the above embodiments can be further used. As an example, video signatures and audio signatures are described as being generated for groups of images and groups of audio frames or audio packets. However, it is also possible to generate signatures more or less frequently. For example, a video signature can be generated for each image frame, and an audio signature can be generated for each audio frame or audio packet. Still, generally, generating signatures for multiple frames or packets is more bitrate-efficient. Also, it is more computationally resource-efficient not to generate signatures too frequently.

[0058] In the above, video digests and audio digests were described as being the hashes of video data and audio data respectively. Other types of digests, such as a list of hashes of individual frames of a group of signed frames or packets, may also be useful.

[0059] The signature function may use the same secure element to generate video signatures and audio signatures. It is also possible to use separate secure elements to sign two different sequences. In such a case, the link between them can be ensured by linking the two secure elements within the same device to each other.

[0060] Although the above image frames and audio frames or audio packets are described as being encoded, generally, the same concept of digital signatures can be used for unencoded data or raw data as long as the signed data is not subsequently encoded. Encoding inevitably involves some information loss that can cause authentication to fail, even if there is no tampering or packet loss.

[0061] As described above, the video sequence can be encoded using h.264, h.265 or another h.26x encoding standard, or the AV1 encoding standard. In a variant, other video encoding standards such as LCEVC or MJPEG may be used.

[0062] As described above, the audio sequence can be encoded using the AAC encoding standard or the PCM encoding standard. PCM-encoded audio data can be packaged in a WAV container as described above. In a variant, generally any encoding standard capable of processing metadata can be used.

[0063] Therefore, the present invention should not be limited to the disclosed embodiments and should be defined only by the appended claims.

Claims

1. A method for digitally signing a video sequence and an audio sequence, wherein the audio sequence and the video sequence are simultaneously captured so as to represent the same scene in which the audio sequence and the video sequence are captured, the video sequence includes consecutive video portions, the audio sequence includes consecutive audio portions, and the method includes: generating a first video digest by applying a digest algorithm to a first video portion of the video sequence; generating a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence; generating a first video signature by digitally signing the first video digest; generating a first audio signature by digitally signing the first audio digest; inserting the first video signature into a first target audio portion of the audio sequence; inserting the first audio signature into a first target video portion of the video sequence; generating a second video digest by applying a digest algorithm to the first target video portion including the first audio signature; generating a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature; generating a second video signature by digitally signing the second video digest; generating a second audio signature by digitally signing the second audio digest; A method including the above steps.

2. inserting the second video signature into a second target audio portion of the audio sequence; inserting the second audio signature into a second target video portion of the video sequence; The method according to claim 1, further including the above steps.

3. inserting the first video signature into the video portions of the video sequence; inserting the first audio signature into the audio portions of the audio sequence; The method according to claim 1, further including the above steps.

4. inserting the second video signature into the video portions of the video sequence; Inserting the second audio signature also into the audio portion of the audio sequence; The method according to claim 2, further comprising.

5. The method according to claim 1, wherein the digest algorithm is a hash function.

6. The method according to claim 1, wherein each video portion is a group of pictures.

7. The method according to claim 1, wherein each audio portion is an audio frame or an audio packet.

8. The method according to claim 1, wherein each audio signature is inserted into a respective SEI message or open bitstream unit of the video sequence.

9. The method according to claim 1, wherein each video signature is inserted into a respective data stream element or header of the audio sequence.

10. The method according to claim 1, wherein the video sequence and the audio sequence are captured by a single device comprising an image sensor and a microphone.

11. A digital signature system for digitally signing a video sequence and an audio sequence, wherein the audio sequence and the video sequence are captured simultaneously to represent the same captured scene, and the system comprises: A video input function arranged to receive the video sequence comprising consecutive video portions; An audio input function arranged to receive the audio sequence comprising consecutive audio portions; A video digest function arranged to apply a digest algorithm to the video portions of the video sequence, thereby generating a video digest; An audio digest function arranged to apply a digest algorithm to the audio portions of the audio sequence, thereby generating an audio digest; A signature function arranged to generate a video signature by digitally signing the video digest and to generate an audio signature by digitally signing the audio digest; A video signature insertion function arranged to insert the video signature into the audio sequence; An audio signature insertion function arranged to insert the audio signature into the video sequence; Comprising a circuit configured to perform. The video digest function is arranged to generate a first video digest by applying a digest algorithm to a first video portion of the video sequence. The audio digest function is configured to generate a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence. The signature function is arranged to generate a first video signature by digitally signing the first video digest and generate a first audio signature by digitally signing the first audio digest. The video signature insertion function is arranged to insert the first video signature into a first target audio portion of the audio sequence. The audio signature insertion function is arranged to insert the first audio signature into a first target video portion of the video sequence. The video digest function is arranged to generate a second video digest by applying a digest algorithm to the first target video portion including the first audio signature. The audio digest function is arranged to generate a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature. The signature function is arranged to generate a second video signature by digitally signing the second video digest and generate a second audio signature by digitally signing the second audio digest. Digital signature system.

12. The video signature insertion function is arranged to insert the second video signature into a second target audio portion of the audio sequence. The audio signature insertion function is arranged to insert the second audio signature into a second target video portion of the video sequence. The digital signature system according to claim 11.

13. A video camera comprising an image sensor, a microphone, and the digital signature system according to claim 11.

14. A computer-readable storage medium including instructions that, when executed by a device having processing capabilities, cause the device to execute the method according to claim 1.

Citation Information

Patent Citations

  • Data processing detecting system, additional information embedding device, additional information detector, digital contents, music contents processor, additional data embedding method and contents processing detecting method, storage medium and program transmitter

    JP2002091465A

  • Video audio synchronization method and system thereof

    JP2003259314A

  • Systems, methods, and devices for media content tamper protection and detection

    US20220132178A1