Method and system for digitally signing audio and video data
The method and system digitally sign video and audio sequences by inserting signatures into each other's frames, addressing the challenge of authenticating simultaneous capture and origin, enhancing security and reliability of surveillance data.
Patent Information
- Application Number
- JP2024204121
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-22
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing surveillance systems lack a secure method to authenticate video and audio data captured simultaneously from the same scene, as clock synchronization can be manipulated to falsify evidence, and existing digital signatures do not adequately link audio and video data to ensure they were captured at the same time and location.
A method and system for digitally signing video and audio sequences by generating and inserting signatures into each other's frames or packets, ensuring that video and audio data are linked through a digest algorithm, allowing authentication of their simultaneous capture and origin from the same scene.
Ensures the authenticity of video and audio data by linking them through cross-referencing signatures, preventing tampering and ensuring they were captured simultaneously, even in the presence of transmission errors or data loss.
Smart Images

Figure 0007785153000001 
Figure 0007785153000002 
Figure 0007785153000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of authentication of audio and video data, and more particularly to protecting audio and video data from being tampered with after it has been captured. [Background technology]
[0002] Cameras that capture video are widely used for surveillance. Such surveillance or observation is often performed for security reasons, but may also be performed for safety or industrial process control, for example. In some scenarios, it is useful to capture not only images but also audio. A standalone microphone may be used to capture audio. However, as a practical solution, many surveillance cameras incorporate one or more microphones. In this way, the same device can be used for observing a scene.
[0003] When a crime or other incident occurs at a scene being observed, video and audio data captured by a surveillance camera can be useful in investigating and verifying what occurred. In such situations, it is important to ensure that the video and audio data have not been tampered with after being captured by the camera's image sensor and microphone. One way to prove that the video or audio data has not been tampered with is to digitally sign the data. Signing the video and audio data in the camera allows the data to be authenticated later, ensuring that no one has tampered with the data after it was transmitted from the camera. For example, some evidence management systems use digital signatures to transmit data from cameras. In such systems, the digital signature can be used to prove that the data has not been tampered with after it entered the system.
[0004] Digital signature schemes for video and audio data are known. Nevertheless, when video and audio from the same observed or captured scene are used as evidence, it may not be sufficient to demonstrate that the video and audio data are authentic by themselves. It may also be necessary to demonstrate that the video and audio data were actually captured at the same time and in the same location. Some systems address this by relying on timestamps in the video and audio data. If the clocks of cameras and other devices in a surveillance or observation system are synchronized and the timestamps of video and audio frames are equal, it can be assumed that the video and audio frames were captured simultaneously. However, users of cameras or surveillance systems are typically able to adjust the clocks of cameras and other devices in the surveillance system. In this way, it may be possible to falsify surveillance data by replacing an original audio sequence captured simultaneously with the video sequence with another audio sequence captured at a different time but after fraudulently resetting the clock. Therefore, there remains a need for a convenient and secure method of digitally signing video and audio data in a manner that enables the video and audio data to be demonstrated as authentic, originating from the same captured scene, and captured simultaneously. Summary of the Invention
[0005] It is an object of the present invention to provide a method by which video and audio sequences captured simultaneously can be authenticated as representing the same captured scene. Another object is to provide a method for digitally signing video and audio sequences that can be attached to the data stream from a camera without the need to encapsulate the video and audio data in a file or container. It is also an object of the present invention to provide a system that allows authentication of video and audio sequences that simultaneously capture an observed scene.
[0006] According to a first aspect, these and other objects are achieved in whole or at least in part by a method as set forth in claim 1. This is therefore achieved by a method for digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously so as to represent the same scene from which they were captured, the video sequence comprising successive video portions and the audio sequence comprising successive audio portions. The method includes generating a first video digest by applying a digest algorithm to a first video portion of a video sequence, generating a first audio digest by applying a digest algorithm to a first audio portion of an audio sequence, generating a first video signature by digitally signing the first video digest, generating a first audio signature by digitally signing the first audio digest, inserting the first video signature into a first target audio portion of the audio sequence, inserting the first audio signature into the first target video portion of the video sequence, generating a second video digest by applying a digest algorithm to the first target video portion including the first audio signature, generating a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature, digitally signing the second video digest, and digitally signing the second audio digest, thereby linking the video sequence and the audio sequence. An audio sequence is required to authenticate a video sequence, and vice versa: for example, if the audio sequence has been altered after capture to give the impression that people in the video sequence are saying something other than they actually are, attempts to verify the audio signature included in the video sequence will fail.Such authentication failures may be due to intentional tampering with the video or audio, but may also be caused by transmission errors. Thus, successful authentication indicates that the video and audio sequences are authentic, while unsuccessful authentication only indicates that something occurred after the video and audio sequences were captured, but does not indicate whether the video and / or audio were intentionally tampered with or whether there was accidental modification of the video and / or audio.
[0007] Note that the video signature and audio signature do not need to be generated synchronously. A video signature can be generated for each group of pictures (GOP, for short). The duration of a GOP depends on the frame rate of the image sensor and the GOP length used to encode the video data. As an example, with a frame rate of 30 frames per second and a GOP length of 62, a video signature for a GOP is generated approximately once every two seconds. Audio sequences may be signed more frequently or less frequently than video sequences. As an example, if an audio frame or packet is created every 10 ms and an audio signature is generated for a group of 100 audio frames, the audio signature is generated once per second. Once a video signature is generated, it is inserted into a target audio portion of the audio sequence. This target audio portion may be the audio portion currently being captured and encoded when the video signature is generated, or it may be the next audio portion whose capture and encoding has not yet begun when the video signature is generated. Similarly, once an audio signature is generated, it is inserted into a target video portion, which may be the video portion currently being captured and encoded or the next video portion. If the video and audio signatures are generated at the same rate, i.e., if the video and audio portions for which the signatures are generated are captured over equal time periods, each target video portion will contain one audio signature after insertion, and each target audio portion will contain one video signature. On the other hand, if the audio and video signatures are generated at different intervals, there will not be a one-to-one relationship between the signatures of the video and audio sequences. If the audio signatures are generated more frequently than the video signatures, each target video portion may contain more than one audio signature, but each target audio portion will contain one video signature, and there will be audio portions between the target audio portions that do not contain a video signature. As will be appreciated by those skilled in the art, the situation is reversed if the video signatures are generated more frequently than the audio signatures.In this regard, it should also be noted that even in scenarios where the video and audio signatures are generated at the same rate, they do not necessarily have to be generated simultaneously: if the video and audio signatures are generated by the same signature circuit or processor, they will compete for computing resources and must be generated one after the other.
[0008] As used herein, a "target video portion" is a video portion into which an audio signature is being inserted. As can be understood from the foregoing description, if the audio signature is generated less frequently than once per video portion, there may be other video portions between the target video portions. Similarly, a "target audio portion" is an audio portion into which a video signature is being inserted. If the video signature is generated less frequently than once per audio portion, there may be other audio portions between the target audio portions.
[0009] It should be noted that an audio sequence capturing the same observed or captured scene as a video sequence can very well capture sounds originating from outside the field of view of the camera capturing the video sequence. A microphone placed in or near the camera can, for example, capture the voices of people in the field of view of the camera as well as just outside, or the sound of a window breaking in front of, above, or behind the camera. All such sounds are considered to belong to the same observed or captured scene as that in the field of view of the camera.
[0010] Variants of the method are defined in the dependent claims.
[0011] In some variations, the method further includes inserting a second video signature into a second target audio portion of the audio sequence and inserting the second audio signature into a second target video portion of the video sequence, whereby the second target audio portion includes the second video signature influenced by the first audio signature. As a result, verification of the second video signature confirms a link to the audio sequence, more precisely to the first audio portion. Similarly, the second target video portion includes the second audio signature influenced by the first video signature. As a result, verification of the second audio signature confirms a link to the video sequence, more precisely to the first video portion.
[0012] The method may further include inserting a first video signature into a video portion of the video sequence and a first audio signature into an audio portion of the audio sequence, thereby enabling authentication of the first video portion even if the first target audio portion is lost, and enabling authentication of the first audio portion even if the first target video portion is lost.
[0013] Similarly, the method may include inserting a second video signature into the video portion of the video sequence and a second audio signature into the audio portion of the audio sequence, thereby enabling authentication of the second video portion in the absence of the second target audio portion and authentication of the second audio portion in the absence of the second target video portion. This makes the authentication procedure more robust against packet loss in transmission or pruning of data storage. For example, when transmission bandwidth or storage capacity is insufficient, it may be decided to transmit only the audio sequence, which generally consumes fewer bits, and omit transmission of the video sequence. As will be understood by those skilled in the art, even in situations where only video data or only audio data is transmitted, the presence of an audio signature in the video sequence or a video signature in the audio sequence can still indicate that the video sequence was linked to the audio sequence at the time of capture, or that the audio sequence was linked to the video sequence at the time of capture.
[0014] The digest algorithm may be a hash function, which is a well-known and practical way to create a digest.
[0015] Each video portion may be a group of pictures, also called a GOP.
[0016] Each audio portion may be a group of audio packets or audio frames.
[0017] In a video sequence, each audio signature can be inserted into a respective SEI message or open bitstream unit of the video sequence. SEI messages are available in the h.26X coding standard and are also called SEI frames or SEI NAL units. Specific SEI NAL units of unregistered types of user data may be used for signature insertion. Open bitstream units are available in the AV1 coding standard and may also be called OBUs or metadata OBUs.
[0018] In an audio sequence, each video signature is inserted into a respective data stream element or header of the audio sequence.
[0019] In some variations, the video and audio sequences are captured by a single device that includes an image sensor and a microphone. This is a practical and efficient way to capture video and audio data of an observed scene. When the image sensor and microphone are incorporated into the same device, being able to indicate that the video and audio sequences were captured by the same device also makes it possible to indicate that they captured the same scene.
[0020] According to a second aspect, the above object is achieved, in whole or at least in part, by a digital signature system as set forth in claim 11. This is therefore achieved by a digital signature system for digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously such that they represent the same scene in which they were captured, the system comprising a video input function arranged to receive a video sequence comprising successive video portions, an audio input function arranged to receive an audio sequence comprising successive audio portions, a video digest function arranged to apply a digest algorithm to the video portions of the video sequence, thereby generating a video digest, an audio digest function arranged to apply a digest algorithm to the audio portions of the audio sequence, thereby generating an audio digest, a signing function arranged to generate a video signature by digitally signing the video digest and to generate an audio signature by digitally signing the audio digest, and a circuit configured to implement a video signature insertion function arranged to insert a signature and an audio signature insertion function arranged to insert an audio signature into the video sequence, wherein the video digest function is arranged to generate a first video digest by applying a digest algorithm to a first video portion of the video sequence, the audio digest function is arranged to generate a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence, the signing function is arranged to generate a first video signature by digitally signing the first video digest and generate a first audio signature by digitally signing the first audio digest, the video signature insertion function is arranged to insert the first video signature into a first target audio portion of the audio sequence, and the audio signature insertion function is arranged to insert the first audio signature into the first target video portion of the video sequence;The video digest function is arranged to generate a second video digest by applying a digest algorithm to the first target video portion including the first audio signature, the audio digest function is arranged to generate a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature, and the signing function is arranged to generate a second video signature by digitally signing the second video digest and generate a second audio signature by digitally signing the second audio digest. Such a system can be used to determine whether video and audio sequences capturing the same scene are authentic. The ability to insert audio signatures into video sequences and audio signatures into audio sequences allows for verification that the video and audio sequences were actually captured at the same time and of the same scene.
[0021] Embodiments of the digital signature system are defined in the dependent claims.
[0022] The digital signature system of the second aspect may generally be implemented in the same manner as the method of the first aspect, with attendant advantages.
[0023] According to a third aspect, the above object is achieved by a camera comprising an image sensor, a microphone and a digital signature system according to the second aspect.
[0024] According to a fourth aspect, the above object is achieved by a computer-readable storage medium comprising instructions which, when executed by an apparatus having processing capability, cause the apparatus to perform the method of the first aspect.
[0025] Further scope of applicability of the present invention will become apparent from the detailed description given hereinafter. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the scope of the invention will become apparent to those skilled in the art from this detailed description.
[0026] Thus, it is to be understood that the present invention is not limited to the particular components of the described apparatus or steps of the described methods, and that such apparatus and methods may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It should be noted that, as used in this specification and the appended claims, the articles "a," "an," "the," and "said" are intended to mean that there are one or more of an element, unless the context clearly dictates otherwise. Thus, for example, reference to an "object" or "the object" may include several objects, etc. Furthermore, the word "comprising" does not exclude other elements or steps.
[0027] The invention will now be described in more detail, by way of example, with reference to the accompanying schematic drawings, in which: [Brief explanation of the drawings]
[0028] [Figure 1] FIG. 1 is a perspective view of a scene observed by a camera incorporating a microphone. [Figure 2] 2A and 2B are diagrams of video and audio sequences capturing the scene of FIG. 1; [Figure 3] 1 is a diagram of how a video signature is inserted into an audio sequence and how an audio signature is inserted into a video sequence. [Figure 4] 3 is a flowchart of a method for digitally signing video and audio sequences such as those shown in FIG. 2; [Figure 5] 5 is a flow chart illustrating steps that can be followed after the steps shown in FIG. 4. [Figure 6] 1 is a box diagram of a system for digitally signing video and audio sequences; [Figure 7] FIG. 1 is a box diagram of a camera incorporating a digital signature system. DETAILED DESCRIPTION OF THE INVENTION
[0029] FIG. 1 is a diagram of a scene 1 observed using a camera 2 positioned on a light pole. As described further below with reference to FIG. 7, camera 2 includes an image sensor (not shown in FIG. 1) and a microphone (not shown in FIG. 1). In the example shown in FIG. 1, camera 2 is pointed at the door of building 3. The image sensor of camera 2 allows it to capture objects in the space between camera 2 and the building. Additionally, the microphone in camera 2 allows it to capture audio around camera 2. The microphone can generally have a range wider than simply the field of view of camera 2. For example, the microphone can capture audio originating behind the camera or on the other side of building 3.
[0030] The video and audio sequences captured by camera 2 may be useful as evidence if a crime is committed at the observed scene 1. In some cases, video may be sufficient to confirm what happened, for example, to identify who entered a building through a window. In other instances, it may be necessary to also hear audio from the crime scene, such as when a person is assaulted and the accused perpetrator claims to have acted in self-defense because the assault victim was verbally threatening. For the video and audio sequences to be useful in combination as evidence, it may be necessary to prove that they were actually captured at the same time, at the same scene, or even by the same device. Therefore, it may not be sufficient to authenticate the video and audio sequences alone to show that they were not tampered with after they were captured; it may also be necessary to authenticate that they are actually associated.
[0031] According to an embodiment of the present invention, a video signature may be generated for a video sequence, and an audio signature may be generated for an audio sequence in a manner similar to that performed when digitally signing video data and audio data separately. The video signature is included in the audio sequence, and vice versa, so that the video and audio sequences can be authenticated as a complete set. Furthermore, when a video signature is generated, it should include data from the preceding audio signature, and when an audio signature is generated, it should include data from the preceding video signature. In this way, authentication of a video signature fails not only if the video sequence has been tampered with, but also if the audio sequence has been tampered with, and vice versa. For successful authentication of a video signature, both the video and audio data must be intact. Similarly, for authentication of an audio signature, both the audio and video data must be intact. If the video and audio sequences are in fact authentic, but were combined some time after capture, and therefore were not captured simultaneously in the same scene, authentication fails.
[0032] 2 shows a video sequence 10 and an audio sequence 20. The video sequence is made up of coded image frames IF1, IF2, . . . IF n Image frames are arranged in groups of pictures, or GOPs for short. Each GOP begins with an intra-coded image frame, also called an I-frame, followed by several inter-coded image frames. The inter-coded image frames may be forward-predicted frames, also called P-frames, or bidirectionally predicted frames, also called B-frames.
[0033] The audio sequence 20 is made up of coded audio frames or packets AF1, AF2, ... AF nAudio sequences consist of a series of frames. Audio sequences do not have the same kind of hierarchical structure of frames as video sequences, which have key frames (intra-coded frames) and delta frames (inter-coded frames), and therefore do not have the same kind of inherent grouping. However, for the purposes of the digital signature scheme of the present invention, audio sequences are also divided into groups of audio frames. It may be worth noting that the video and audio portions for signature can be arbitrarily small, such as a single image frame or a single audio sample. However, for bitrate and computational efficiency, it is generally preferable to sign several image frames or audio frames at a time.
[0034] To enable authentication of the video and audio sequences and the connections between them, digital signatures are generated, each for a portion of the respective sequence. In this example, for a video sequence, a video signature is generated for the video portion in the form of a GOP, and an audio signature is generated for the audio portion of the audio sequence in the form of a group of audio frames.
[0035] An example of a method for digitally signing video and audio sequences will now be described with reference to the flow chart of FIG.
[0036] As shown in FIG. 3, a first video digest is generated by applying a digest algorithm to the first video portion 11 as the basis for the first video signature VS1. In this example, the digest is a hash. Therefore, the video digest will be referred to as a video hash hereinafter. The first video portion 11 is hashed to generate a first video hash (step S1). The first video signature VS1 is generated by digitally signing the first video hash (step S3). Note that if this is the very first video signature of the video sequence and is generated before any audio signatures are inserted into the video sequence, this first video signature cannot directly authenticate any link to the audio sequence, but it is still useful for demonstrating that the first video portion itself has not been tampered with. Once the first video signature VS1 is generated, it is inserted into the first target audio portion of the audio sequence (step S5). This insertion will be described in more detail below with reference to FIG. 3.
[0037] Similarly, as the basis for the first audio signature AS1, a first audio digest is generated by applying a digest algorithm to the first audio portion 21 of the audio sequence 20. For video sequences, in this example, the digest is a hash. The first audio portion 21 is hashed to generate a first audio hash (step S2). The first audio signature AS1 is generated by digitally signing the first audio hash (step S4). Note that, similar to video sequences, if this is the very first audio signature for the audio sequence and is generated before any video signature is inserted into the audio sequence, this first audio signature cannot authenticate any link to the video sequence, but it is still useful for demonstrating that the first audio portion itself has not been tampered with. Once the first audio signature AS1 is generated, it is inserted into the first target video portion of the audio sequence (step S5). This insertion is described in more detail below with reference to FIG. 3.
[0038] A second video hash is generated by hashing the first target video portion (step S7), which is a video portion captured some time after the first video portion 11 and into which the first audio signature AS1 has been inserted. The second video hash is therefore a hash of the combination of the video data and the first audio signature AS1, thereby linking the video sequence 10 and the audio sequence 21 together. A second video signature VS2 is generated by digitally signing the second video hash (step S9).
[0039] Furthermore, a second audio hash is generated by hashing the first target audio portion (step S8). The first target audio portion is the audio portion into which the first video signature VS1 was inserted, captured some time after the first audio portion 21. The second audio hash is therefore a hash of the combination of the audio data and the first video signature VS1, thereby linking the video sequence 10 and the audio sequence 21 together. A second audio signature AS2 is generated by digitally signing the second audio hash (step S10).
[0040] Referring briefly to the flowchart of Figure 5, and before returning to Figure 3, once the second video signature VS2 is generated, it may be inserted into the second target audio portion (step S11). Similarly, a second audio signature AS2 may be inserted into the second target video portion (step S12). The digital signing of successive video portions of video sequence 10 and successive audio portions of audio sequence 20 may continue in the same manner until the end of the sequence. This "cross-fertilization" of the video and audio sequences with their respective signatures makes it possible not only to authenticate the video and audio sequences separately, but also to authenticate their joint origin.
[0041] The insertion of a video signature into an audio sequence and the insertion of an audio signature into a video sequence will now be described with reference to Figure 3. Again, a video sequence 10 and an audio sequence 20 are shown. As mentioned above, image frames are grouped into GOPs, and audio frames or packets are similarly grouped into groups of audio frames or packets. It should be noted that the illustration in Figure 3 is very schematic, and that in practice groups will typically consist of many more frames.
[0042] As described above, a first video signature VS1 is generated for the first video portion 11. This first video signature VS1 is inserted into a first target audio portion 41 of the audio sequence 20, which in this case happens to coincide with the second audio portion 22. The location in the audio sequence 20 where the first video signature VS1 is inserted is indicated by hatching. Depending on the frequency with which the video signature is generated compared to the length of the audio portion, the first target audio portion 41 may be an earlier or later audio portion than the second audio portion 22. Thus, the first target audio portion may be the first audio portion, the second audio portion, or a later audio portion. A second video signature VS2 is generated for the second video portion 12, and this second video signature VS2 is inserted into the second target audio portion 42. In the illustrated example, the second target audio portion 42 coincides with the third audio portion 23, but as explained above, this depends on the relationship between the frequency with which the video signature is generated and the length of the audio portion. Similarly, a third video signature VS3 is generated for the third video portion 13. The third video signature VS3 is inserted into the third target audio portion 43, which in the illustrated example is the fourth audio portion 24.
[0043] Looking at audio sequence 20, the same principle is followed: a first audio signature AS1 is generated for first audio portion 21. The first audio signature AS1 is inserted into first target video portion 31. In the illustrated example, the first target video portion is the second video portion, but as explained for video signatures, this depends on the frequency with which audio signatures are generated compared to the length of the video portions. A second audio signature AS2 is generated for the second audio portion and inserted into second target video portion 32, which in this example is third video portion 13. A third audio signature AS3 is generated for third audio portion 23. The third audio signature AS3 is inserted into third target video portion 33. In the illustrated example, third target video portion 33 happens to be the same as second target video portion 32, i.e., third video portion 13. Therefore, two audio signatures, AS2 and AS3, are inserted into the third video portion. Continuing with a fourth audio portion, the full length of which is not shown in Figure 3, a fourth audio signature AS4 is generated and inserted into a fourth target video portion, which is not shown in Figure 3.
[0044] In video sequence 10, the audio signature AS n The audio signature may be inserted into an already available element of the video compression format. For example, if the video sequence is encoded using h.264 or another h.26X encoding format, the audio signature may be inserted into a specified type of SEI message, also called an SEI frame or SEI NAL unit. Other video compression formats provide corresponding elements. As an example, if the video sequence is encoded using AV1, the audio signature may be inserted into a specified type of open bitstream unit, also called an OBU or metadata OBU.
[0045] Similarly, in audio sequence 20, the video signature VS nThe video signature may be inserted into an appropriate existing element in the audio coding format. For example, if the audio sequence is coded using AAC, the video signature may be inserted into the data stream element, or DSE for short, and in the PCM audio coding format, the video signature may be inserted into the header of the WAV container in which the audio data is encapsulated. Other audio coding formats offer similar possibilities.
[0046] The video signature may be generated by a cryptographic operation, for example along the lines of the method described in applicant's European Patent Application No. 4164230, and the audio signature may be inserted into the SEI frame or corresponding element of the video sequence, as described in applicant's European Patent Application No. 4192018. The audio signature can be generated and inserted using the same principles as the video signature.
[0047] FIG. 6 is a simplified box diagram of a digital signature system 60 that can be used to implement the digital signature method described above. Digital signature system 60 has a video input function 61 arranged to receive video sequence 10 and an audio input function 62 configured to receive audio sequence 20. Digital signature system 60 further includes a video digest function 63 arranged to apply a digest algorithm to the video portion of the video sequence, thereby generating a video digest. Similar to the description above, in this example, the digest algorithm is a hash function, and therefore, the video digest function can be referred to as a video hash function 63. Digital signature system 60 also includes an audio digest function 64 arranged to apply a digest algorithm to the audio portion of the audio sequence, thereby generating an audio digest. With respect to the video digest, the audio digest is a hash in this example, and therefore, the audio digest function can be referred to as an audio hash function 64.
[0048] Furthermore, the digital signature system 60 includes a signing function 65 arranged to generate a video signature by digitally signing a video digest or hash and to generate an audio signature by digitally signing an audio digest or hash, a video signature insertion function 66 arranged to insert the video signature into an audio sequence, and an audio signature insertion function 67 arranged to insert the audio signature into the video sequence.
[0049] 3-5, the video hash digest function 63 is arranged to generate a first video hash by hashing a first video portion 11 of the video sequence 10, and the audio hash function 64 is arranged to generate a first audio hash by hashing a first audio portion 21 of the audio sequence 20. The signing function 65 is arranged to generate a first video signature VS1 by digitally signing the first video hash and to generate a first audio signature AS1 by digitally signing the first audio hash. The video signature insertion function 66 is arranged to insert the first video signature VS1 into a first target audio portion 41 of the audio sequence 20. The audio signature insertion function 67 is arranged to insert the first audio signature AS1 into a first target video portion 31 of the video sequence 10. The video hash function 63 is further arranged to generate a second video hash by hashing the first target video portion 31 including the first audio signature AS1, and the audio hash function is arranged to generate a second audio hash by hashing the first target audio portion 41 including the first video signature VS1. The signing function 65 is arranged to generate a second video signature VS2 by digitally signing the second video hash and to generate a second audio signature AS1 by digitally signing the second audio hash.
[0050] In the same manner as described in relation to Figure 5, the video signature insertion function 66 may be further arranged to insert a second video signature VS2 into the second target audio portion 42, and the audio signature insertion function 67 may be further arranged to insert a second audio signature AS2 into the second target video portion 32.
[0051] Functions 61-67 of digital signature system 60 may be implemented using circuitry or processing circuitry, including general-purpose processors, special-purpose processors, integrated circuits, ASICs ("application-specific integrated circuits"), conventional circuitry, and / or combinations thereof configured or programmed to perform the disclosed functionality. In this disclosure, circuitry is hardware that performs or is programmed to perform the recited functionality. Hardware may be any hardware disclosed herein or otherwise known that is programmed or configured to perform the recited functionality.
[0052] In a pure hardware implementation, each function is dedicated and may have a corresponding circuit specifically designed to implement the function, which may be in the form of one or more integrated circuits, such as one or more application specific integrated circuits or one or more field programmable gate arrays.
[0053] In embodiments that also include software, the circuitry may include a processor. A processor is considered a processing circuit or circuit because it includes transistors and other circuitry within it. In this case, the circuitry may be viewed as a combination of hardware and software, with the software being used to configure the hardware and / or processor.
[0054] It should be understood that it is also possible to have a combination of hardware and software implementation, meaning that some of the functions are implemented by dedicated circuitry and others are implemented in software, i.e. in the form of computer code executed by a processor.
[0055] 7 shows a simplified box diagram of camera 2. Camera 2 has a lens 71 and an image sensor 72 for capturing images of the observed scene. Camera 2 further includes a video encoder 73 for encoding video data captured by image sensor 72. Camera 2 includes a microphone 74 for capturing audio within the scene. Camera 2 also includes an audio encoder 75 for encoding audio data captured by microphone 74. Camera 2 also includes a digital signature system 60 as shown in FIG. 6. Network interface 76 enables transmission of encoded video and audio data, including video and audio signatures. As will be appreciated by those skilled in the art, camera 2 may have additional components; however, only those components relevant to the description of the present invention have been shown and described herein.
[0056] To make the above-described method and system more resilient to transmission losses, a video signature can also be inserted into a video sequence, and an audio signature can also be inserted into an audio sequence. When a video signature is generated, in this scenario, it may be inserted into a subsequent video portion, or the nearest video portion, after the video signature is generated. Similarly, when an audio signature is generated, it may be inserted into a subsequent audio portion, or the nearest audio portion, after the audio signature is generated. In situations where the transmission bandwidth from the camera is low, a decision may be made to transmit only video or only audio. In such cases, it is impossible to authenticate the transmitted sequence unless the corresponding signature is also inserted into the sequence from which it originated. Thus, if only video is transmitted, the video signature is not transmitted unless it is also included in the video sequence. Similarly, if only audio is transmitted, the audio signature is not transmitted unless it is also included in the audio sequence. As a further adaptation, if it is decided to return to "video-only" transmission during temporary bandwidth limitations, it is possible to suspend the insertion of audio signatures into video sequences during the "video-only" transmission, thereby allowing the video sequence to be independently authenticated. Of course, this can be applied mutatis mutandis to "audio-only" transmissions. During such periods of "video-only" or "audio-only" transmission, untransmitted sequences can be stored locally in the camera for later transmission when network conditions improve. Note that in such cases, if the insertion of signatures into the "companion sequence" is paused during a "video-only" or "audio-only" transmission, it will not be possible to securely authenticate the link between the video and audio sequences. It may be useful to insert metadata into the transmitted sequence indicating that the transmission of the other sequence has been paused.
[0057] Those skilled in the art will appreciate that the above embodiments can be modified in many ways and still utilize the advantages of the present invention as illustrated in the above embodiments. As an example, video and audio signatures are described as being generated for a group of images and a group of audio frames or packets. However, it is also possible to generate signatures more or less frequently. For example, a video signature can be generated for each image frame, and an audio signature can be generated for each audio frame or packet. Nevertheless, generating signatures for multiple frames or packets is generally more bit-rate efficient. It is also more computationally resource efficient to not have to generate signatures too frequently.
[0058] While video and audio digests have been described above as being hashes of the video and audio data, respectively, other types of digests may also be useful, such as a list of hashes of individual frames in a group of signed frames or packets.
[0059] The signing function may use the same secure element to generate the video signature and the audio signature. It is also possible to use separate secure elements to sign the two different sequences. In such cases, the link between them can be ensured by linking the two secure elements together in the same device.
[0060] Although the image and audio frames or packets above are described as being encoded, the same digital signature concept can generally be used for unencoded or raw data, as long as the signed data is not subsequently encoded. Encoding more or less inevitably involves information loss that will cause authentication to fail, even in the absence of tampering or packet loss.
[0061] As mentioned above, the video sequence may be encoded using the h.264, h.265, or another h.26x encoding standard, or the AV1 encoding standard. Alternatively, other video encoding standards such as LCEVC or MJPEG may be used.
[0062] As mentioned above, the audio sequence can be encoded using the AAC or PCM encoding standards. The PCM-encoded audio data can be packaged in a WAV container, as mentioned above. In general, any encoding standard capable of handling metadata can be used.
[0063] Accordingly, the present invention should not be limited to the illustrated embodiments, but should be defined only by the appended claims.
Claims
1. 1. A method of digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously to represent the same captured scene, the video sequence comprising successive video portions and the audio sequence comprising successive audio portions, the method comprising: generating a first video digest by applying a digest algorithm to a first video portion of the video sequence; generating a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence; generating a first video signature by digitally signing the first video digest; generating a first audio signature by digitally signing the first audio digest; inserting the first video signature into a first target audio portion of the audio sequence; inserting the first audio signature into a first target video portion of the video sequence; generating a second video digest by applying a digest algorithm to the first target video portion including the first audio signature; generating a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature; generating a second video signature by digitally signing the second video digest; generating a second audio signature by digitally signing the second audio digest; A method comprising:
2. inserting the second video signature into a second target audio portion of the audio sequence; inserting the second audio signature into a second target video portion of the video sequence; The method of claim 1 further comprising:
3. inserting the first video signature into a video portion of the video sequence; inserting the first audio signature into an audio portion of the audio sequence; The method of claim 1 further comprising:
4. inserting the second video signature into a video portion of the video sequence; inserting the second audio signature also into an audio portion of the audio sequence; The method of claim 2 further comprising:
5. The method of claim 1 , wherein the digest algorithm is a hash function.
6. The method of claim 1 , wherein each video portion is a group of pictures.
7. The method of claim 1 , wherein each audio portion is an audio frame or an audio packet.
8. The method of claim 1 , wherein each audio signature is inserted into a respective SEI message or open bitstream unit of the video sequence.
9. The method of claim 1 , wherein each video signature is inserted into a respective data stream element or header of the audio sequence.
10. The method of claim 1 , wherein the video sequence and the audio sequence are captured by a single device comprising an image sensor and a microphone.
11. 1. A digital signature system for digitally signing video and audio sequences, the audio and video sequences being captured simultaneously to represent the same captured scene, the system comprising: a video input function arranged to receive a video sequence comprising successive video portions; an audio input function arranged to receive an audio sequence comprising successive audio portions; a video digest function arranged to apply a digest algorithm to video portions of the video sequence, thereby generating a video digest; an audio digest function arranged to apply a digest algorithm to the audio portions of the audio sequence, thereby generating an audio digest; a signing function arranged to generate a video signature by digitally signing the video digest and to generate an audio signature by digitally signing the audio digest; a video signature insertion function arranged to insert a video signature into said audio sequence; an audio signature insertion function arranged to insert an audio signature into said video sequence; a circuit configured to perform the the video digest function is arranged to generate a first video digest by applying a digest algorithm to a first video portion of the video sequence; the audio digest function is arranged to generate a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence; the signing function is arranged to digitally sign the first video digest to generate a first video signature and to digitally sign the first audio digest to generate a first audio signature; the video signature insertion function is configured to insert a first video signature into a first target audio portion of the audio sequence; the audio signature insertion function is configured to insert the first audio signature into a first target video portion of the video sequence; the video digest function is arranged to generate a second video digest by applying a digest algorithm to the first target video portion including the first audio signature; the audio digest function is arranged to generate a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature; the signing function is arranged to generate a second video signature by digitally signing the second video digest and to generate a second audio signature by digitally signing the second audio digest. Digital signature system.
12. the video signature insertion function is configured to insert the second video signature into a second target audio portion of the audio sequence; the audio signature insertion function is configured to insert the second audio signature into a second target video portion of the video sequence. The digital signature system of claim 11.
13. A video camera comprising an image sensor, a microphone and a digital signature system according to claim 11 or 12.
14. A computer readable storage medium comprising instructions that, when executed by a device having processing capability, cause said device to perform the method of any one of claims 1 to 10.
Citation Information
Patent Citations
Data processing detecting system, additional information embedding device, additional information detector, digital contents, music contents processor, additional data embedding method and contents processing detecting method, storage medium and program transmitter
JP2002091465A
Video audio synchronization method and system thereof
JP2003259314A
Systems, methods, and devices for media content tamper protection and detection
US20220132178A1