Method and system for digitally signing audio and video data
By digitizing the video and audio data and inserting the signature into the corresponding data stream, the problem of whether the video and audio data have not been tampered with after capture is solved, and the authenticity of the data being captured simultaneously in the same scenario is ensured.
Patent Information
- Application Number
- CN202411582363.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-07
AI Technical Summary
The prior art is difficult to effectively verify whether video data and audio data have not been tampered with after being captured, and it is difficult to ensure that video data and audio data are captured in the same scenario at the same time.
By digitizing the video sequence and the audio sequence, and inserting the video signature into the audio sequence, the audio signature is inserted into the video sequence, so that the integrity and relevance of the video and audio data need to be verified at the same time during verification.
The valid digital signature of video data and audio data is realized, ensuring that the data is not tampered with during transmission, and the authenticity of the data being captured simultaneously in the same scenario can be verified.
Smart Images

Figure CN120074850A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of verification of audio data and video data, and more particularly, to protecting audio data and image data from being tampered with after they are captured. Background Art
[0002] Cameras used to capture video are widely used for surveillance. Such surveillance or monitoring is usually carried out for security reasons, but it may also be carried out for safety or industrial process control. In some scenarios, it is not only useful to capture images, but also helpful to capture sound. Stand-alone microphones can be used to capture sound. However, as a practical solution, one or more microphones are incorporated into many surveillance cameras. In this way, it is possible to use one and the same device to monitor a scene.
[0003] If an event such as a crime occurs in the monitored scene, the video data and audio data captured by the surveillance camera may help in the investigation and prove what has happened. In such a situation, it is important to ensure that the video data and audio data have not been tampered with after they are captured by the image sensor and microphone of the camera. One way to prove that video data or audio data has not been tampered with is to apply a digital signature to the data. By signing the video data and audio data in the camera, the data can be verified later, thus ensuring that no one has tampered with the data after it is transmitted from the camera. There are also systems in which digital signatures are applied, such as evidence management systems, to which data is transmitted from the camera. In such a system, the digital signature can be used to prove that the data has not been tampered with after it enters the system.
[0004] Digital signature schemes for video data and audio data are known. Nevertheless, when video and audio from one and the same monitored or captured scene are used as evidence, it may not be sufficient to merely show that the video data and the audio data are each independently authentic. It may also be necessary to show that the video data and the audio data were actually captured at one and the same location at the same time. Some systems address this problem by relying on timestamps in the video data and the audio data. If the clocks of the cameras and other devices in a monitoring or surveillance system are synchronized and the timestamps of the video frames and the audio frames are equal, it can be expected that the video frames and the audio frames were captured at the same time. However, the user of a camera or a monitoring system can typically adjust the clocks of the cameras and other devices in the monitoring system. In this way, it is possible to forge surveillance data by replacing the original audio sequence captured simultaneously with the video sequence with another audio sequence captured at another time, but this is after fraudulently resetting the clock. Accordingly, there is still a need for a convenient and secure way to digitally sign video data and audio data in such a way that it is possible to show that the video data and the video data are authentic and originated from one and the same captured scene and were captured at the same time. Summary of the Invention
[0005] An object of the present invention is to provide a method by which it is possible to verify video sequences and audio sequences captured at the same time such that they represent one and the same captured scene. Another object is to provide a method for digitally signing video sequences and audio sequences that can be applied to data streams from cameras without having to encapsulate the video data and the audio data in a file or a container. Another object of the present invention is to provide a system that can verify video sequences and audio sequences of a monitored scene captured at the same time.
[0006] According to a first aspect, these and other objects are achieved in whole or at least in part by the method according to claim 1. Thus, this is achieved by a method of digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously such that they represent one and the same captured scene, the video sequence including successive video portions, and the audio sequence including successive audio portions. The method includes: generating a first video digest by applying a digest algorithm to a first video portion of the video sequence; generating a first audio digest by applying a digest algorithm to a first audio portion of the audio sequence; generating a first video signature by digitally signing the first video digest; generating a first audio signature by digitally signing the first audio digest; inserting the first video signature into a first target audio portion of the audio sequence; inserting the first audio signature into a first target video portion of the video sequence; generating a second video digest by applying the digest algorithm to the first target video portion including the first audio signature; generating a second audio digest by applying the digest algorithm to the first target audio portion including the first video signature; generating a second video signature by digitally signing the second video digest; and generating a second audio signature by digitally signing the second audio digest. In this way, the video sequence and the audio sequence are linked together. For the video sequence to be verified, the audio sequence is necessary and vice versa. For example, if the audio sequence is replaced after capture to give the impression that the people in the video sequence are not saying what they are actually saying, then an attempt to verify the audio signature contained in the video sequence will fail. Such a verification failure may be due to deliberate tampering with the video or audio, but may also be due to a transmission error. Thus, a successful verification will indicate that the video sequence and the audio sequence are authentic, but a failed verification will only indicate that something has happened after the video sequence and the video sequence were captured, it will not indicate whether the video and / or audio have been deliberately tampered with, nor will it indicate whether the video and / or video have undergone accidental changes.
[0007] It should be noted here that video signatures and audio signatures do not need to be generated synchronously. For example, video signatures can be generated for each group of pictures (abbreviated as GOP). The duration of the GOP will depend on the frame rate of the image sensor and the GOP length used when encoding the video data. As an example, at a frame rate of 30 frames per second and a GOP length of 62, the video signature of the GOP will be generated approximately once every two seconds. The signature frequency of the audio sequence can be higher or lower than that of the video sequence. As an example, if an audio frame or audio data packet is generated every 10 ms and the audio signature is generated for a group of 100 audio frames, the audio signature will be generated once per second. When the video signature is generated, it is inserted into the target audio part of the audio sequence. The target audio part can be the audio part that is currently being captured and encoded when the video signature has been generated, or it can be the upcoming audio part whose capture and encoding have not yet started when the video signature is generated. In the same way, when the audio signature is generated, it is inserted into the target video part, which can be the video part that is currently being captured and encoded, or it can be the upcoming video part. If the video signature and the audio signature are generated at the same rate, that is, if the video part and the audio part for which the signature is generated are captured within equal lengths of time, then each target video part will contain one audio signature and each target audio part will contain one video signature after insertion. On the other hand, if the audio signature and the video signature are generated at different intervals, there will be no one-to-one correspondence between the signatures in the video sequence and the audio sequence. If the audio signature is generated more frequently than the video signature, each target video part can contain more than one audio signature, while each target audio part will contain one video signature, and there will be audio parts between the target audio parts that do not contain any video signatures. As those skilled in the art will appreciate, if the video signature is generated more frequently than the audio signature, the situation will be the opposite. In this context, it can also be noted that even in the scenario where the video signature and the audio signature are generated at the same rate, they do not need to be generated simultaneously. If the video signature and the audio signature are generated by one and the same signature circuit or processor, they will compete for computing resources and will have to be generated one after another.
[0008] As used herein, a "target video part" is the video part into which the audio signature is inserted. From the foregoing discussion, it can be understood that if the audio signature is generated less frequently than once per video part, there can be other video parts between the target video parts. Similarly, a "target audio part" is the audio part into which the video signature is inserted. If the video signature is generated less frequently than once per audio part, there can be other audio parts between the target audio parts.
[0009] It should be noted that an audio sequence that captures the same surveillance scene or capture scene as the video sequence can well capture sounds originating outside the field of view of the camera from which the video sequence is captured. For example, a microphone arranged inside or near the camera can capture the voices of people both inside and outside the camera's field of view, or the sound of a window being smashed in front of, above, or behind the camera. These sounds will all be considered to belong to the same surveillance scene or capture scene as the sounds within the camera's field of view.
[0010] Variations of the method are defined in the dependent claims.
[0011] In some variations, the method further includes inserting a second video signature into a second target audio portion of the audio sequence and inserting a second audio signature into a second target video portion of the video sequence. Thus, the second target audio portion will contain the second video signature, which in turn will be affected by the first audio signature. Therefore, verification of the second video signature will confirm the link to the audio sequence and, more precisely, the link to the first audio portion. Similarly, the second target video portion will contain the second audio signature, which in turn is affected by the first video signature. Therefore, verification of the second audio signature will confirm the link to the video sequence and, more precisely, the link to the first video portion.
[0012] The method may further include inserting the first video signature also into the video portion of the video sequence and inserting the first audio signature also into the audio portion of the audio sequence. This will enable verification of the first video portion even if the first target audio portion has been lost and verification of the first audio portion even if the first target video portion has been lost.
[0013] Similarly, the method may include inserting the second video signature also into the video portion of the video sequence and inserting the second audio signature also into the audio portion of the audio sequence, so that the second video portion can be verified in the absence of the second target audio portion and the second audio portion can be verified in the absence of the second target video portion. This will make the verification process more robust to packet loss during transmission or pruning of data storage. For example, if the transmission bandwidth or storage capacity is insufficient, it may be decided to transmit only the audio sequence, which typically consumes fewer bits, and omit transmitting the video sequence. As those skilled in the art will appreciate, even in the case of transmitting only video data or only audio data, the presence of an audio signature in the video sequence or a video signature in the audio sequence will still enable it to indicate that the video sequence was linked to the audio sequence at the time of capture, or the audio sequence was linked to the video sequence at the time of capture.
[0014] The digest algorithm can be a hash function. Hashing data is a well-known and practical method for creating a digest.
[0015] Each video portion can be a group of pictures, also known as a GOP.
[0016] Each audio portion can be a grouping of audio data packets or audio frames.
[0017] In a video sequence, each audio signature can be inserted into the corresponding SEI message or open bitstream unit of the video sequence. The SEI message is available in the h.26X coding standard and can also be referred to as an SEI frame or SEI NAL unit. A specific SEI NAL unit of an unregistered user data type can be used to insert the signature. The open bitstream unit is available in the AV1 coding standard and can also be referred to as an OBU or metadata OBU.
[0018] In an audio sequence, each video signature is inserted into the corresponding data stream element or header of the audio sequence.
[0019] In some variations, the video sequence and the audio sequence are captured by a single device including an image sensor and a microphone. This is a practical and efficient method for capturing video data and audio data in a surveillance scenario. If the image sensor and the microphone are incorporated into one and the same device, it can be indicated that the video sequence and the audio sequence are captured by the same device, and it can also be indicated that they capture the same scene.
[0020] According to a second aspect, the above-mentioned objectives are achieved in whole or at least in part by the digital signature system according to claim 11. Thus, this is achieved by a digital signature system for digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously such that they represent one and the same captured scene, the digital signature system being configured to perform: a video input function, arranged to receive a video sequence, the video sequence comprising successive video portions; an audio input function, arranged to receive an audio sequence, the audio sequence comprising successive audio portions; a video digest function, arranged to apply a digest algorithm to the video portions of the video sequence so as to generate a video digest; an audio digest function, arranged to apply a digest algorithm to the audio portions of the audio sequence so as to generate an audio digest; a signature function, arranged to generate a video signature by digitally signing the video digest and to generate an audio signature by digitally signing the audio digest; a video signature insertion function, arranged to insert the video signature into the audio sequence; and an audio signature insertion function, arranged to insert the audio signature into the video sequence, wherein the video digest function is arranged to generate a first video digest by applying the digest algorithm to a first video portion of the video sequence; the audio digest function is arranged to generate a first audio digest by applying the digest algorithm to a first audio portion of the audio sequence; the signature function is arranged to generate a first video signature by digitally signing the first video digest and to generate a first audio signature by digitally signing the first audio digest; the video signature insertion function is arranged to insert the first video signature into a first target audio portion of the audio sequence; the audio signature insertion function is arranged to insert the first audio signature into a first target video portion of the video sequence; the video digest function is arranged to generate a second video digest by applying the digest algorithm to the first target video portion comprising the first audio signature; the audio digest function is arranged to generate a second audio digest by applying the digest algorithm to the first target audio portion comprising the first video signature; the signature function is arranged to generate a second video signature by digitally signing the second video digest and to generate a second audio signature by digitally signing the second audio digest. Using such a system, it is possible to determine whether the video sequence and the audio sequence capturing the same scene are authentic. The ability to insert an audio signature into the video sequence and a video signature into the audio sequence makes it possible to verify that the video sequence and the audio sequence were indeed captured simultaneously in the same scene.
[0021] Embodiments of the digital signature system are defined in the dependent claims.
[0022] The digital signature system of the second aspect can generally be implemented in the same manner as the method of the first aspect and has corresponding advantages.
[0023] According to a third aspect, the above-mentioned object is achieved by a camera comprising an image sensor, a microphone, and a digital signature system according to the second aspect.
[0024] According to a fourth aspect, the above-mentioned object is achieved by a computer-readable storage medium comprising instructions which, when executed by a device having processing capabilities, cause the device to perform the method of the first aspect.
[0025] The further scope of application of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the scope of the invention will become apparent to those skilled in the art from this detailed description.
[0026] Therefore, it should be understood that the present invention is not limited to the specific parts of the devices described or the steps of the methods described, as such devices and methods can be varied. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It must be noted that, as used in this specification and the appended claims, the words "a", "an", "the" and "said" are all intended to mean that there is one or more elements, unless the context clearly dictates otherwise. Thus, for example, a reference to "an object" or "the object" may include several objects, etc. Further, the term "comprising" does not exclude other elements or steps. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The present invention will now be described in more detail, by way of example, and with reference to the accompanying schematic drawings, in which:
[0028] Figure 1 is a perspective view of a scene monitored by a camera comprising a microphone,
[0029] Figure 2 is an illustration of Figure 1 capturing a video sequence and an audio sequence of the scene in
[0030] Figure 3 is an illustration of how to insert a video signature into an audio sequence and how to insert an audio signature into a video sequence,
[0031] Figure 4 is a flowchart of a method for digitally signing a video sequence and an audio sequence (such as the video sequence and the audio sequence shown in Figure 2 )
[0032] Figure 5 is a diagram showing what can follow after Figure 4Flowchart of the steps after the steps shown in
[0033] Figure 6 is a block diagram of a system for digitally signing a video sequence and an audio sequence, and
[0034] Figure 7 is a block diagram of a camera incorporating the digital signature system. Detailed implementation
[0035] Figure 1 is an illustration of Scenario 1, where Scenario 1 is monitored using Camera 2 arranged on a lamppost. As will be further discussed with reference to Figure 7 Camera 2 includes an image sensor ( Figure 1 not shown in Figure 1 ) and a microphone ( Figure 1 not shown in
[0036] In the example shown in
[0037] According to an embodiment of the present invention, a video signature can be generated for a video sequence and an audio signature can be generated for an audio sequence in a manner similar to what is done when digitally signing video data and audio data separately. To also be able to verify that the video sequence and the audio sequence belong together, the video signature can be included in the audio sequence and the audio signature can be included in the video sequence. In addition, when the video signature is generated, it should include data from the previous audio signature, and when the audio signature is generated, it should include data from the previous video signature. In this way, the verification of the video signature will fail not only in the case where the video sequence has been tampered with, but also in the case where the audio sequence has been tampered with, and vice versa. To successfully verify the video signature, both the video data and the audio data must remain intact. In the same way, to verify the audio signature, both the audio data and the video data must remain intact. If the video sequence and the audio sequence are each actually authentic, but they are combined at some point after capture such that the video sequence and the video sequence were not captured in the same scene at the same time, the verification will fail.
[0038] Figure 2 FIG. illustrates video sequence 10 and audio sequence 20. The video sequence consists of a series of encoded image frames IF 1 、IF 2 …IF n These image frames are arranged in groups of pictures (or simply referred to as GOPs). Each GOP starts with an intra-coded image frame (also known as an I-frame), followed by a plurality of inter-coded image frames. The inter-coded image frames can be forward-predicted frames (also known as P-frames) or bi-directionally predicted frames (also known as B-frames).
[0039] The audio sequence 20 consists of a series of encoded audio frames or encoded audio data packets AF 1 、AF 2 、…AF n The audio sequence does not have the same kind of frame hierarchy as the video sequence (including key frames (intra-coded frames) and delta frames (inter-coded frames)) and thus does not have the same kind of inherent grouping. However, for the digital signature scheme of the present invention, the audio sequence is also divided into groups of audio frames. It is worth mentioning that the video part and the audio part for signature can be arbitrarily small, such as a single image frame or a single audio sample. However, for considerations of bitrate efficiency and computational efficiency, it is generally preferred to sign several image frames or audio frames at a time.
[0040] Digital signatures are generated to be able to verify the video sequence and the audio sequence and the connection between the video sequence and the audio sequence. Each digital signature is generated for a portion of the corresponding sequence. In this example, for the video sequence, a video signature is generated for the video portion in the form of a GOP, and an audio signature is generated for the audio portion in the audio sequence in the form of a group of audio frames.
[0041] Reference Figure 4 to the flowchart in
[0042] As the basis for the first video signature VS 1 a first video digest is generated by applying a digest algorithm to the first video portion 11, such as Figure 3 shown in 1 . In this example, the digest is a hash. Therefore, hereinafter, the video digest will be referred to as the video hash. Thus, the first video hash is generated by hashing the first video portion 11 (step S1). The first video signature VS 1 is generated by digitally signing the first video hash (step S3). It can be noted that if this is the first video signature for the video sequence and is generated before any audio signature is inserted into the video sequence, then this first video signature will not be able to directly verify any link to the audio sequence, but it will still help to show that the first video portion itself has not been tampered with. When the first video signature VS Figure 3 has been generated, it is inserted into the first target audio portion of the audio sequence (S5). The insertion will be discussed in further detail below with reference to
[0043] In a similar manner, as the basis for the first audio signature AS 1 a first audio digest is generated by applying a digest algorithm to the first audio portion 21 of the audio sequence 20. For the video sequence, in this example, the digest is a hash. Therefore, the first audio hash is generated by hashing the first audio portion 21 (step S2). The first audio signature AS 1 is generated by digitally signing the first audio hash (step S4). In the same way as for the video sequence, it can be noted that if this is the first audio signature for the audio sequence and is generated before any video signature is inserted into the audio sequence, then this first audio signature will not be able to verify any link to the video sequence, but it will still help to show that the first audio portion itself has not been tampered with. When the first audio signature AS 1 has been generated, it is inserted into the first target video portion of the audio sequence (S5). This will also be discussed below with reference to Figure 3Let's discuss the insertion in more detail.
[0044] A second video hash is generated by hashing a first target video portion (step S7). The first target video portion is captured at some time after the first video portion 11 and is the video portion in which the first audio signature AS 1 is inserted. Thus, the second video hash will be the hash of the combination of the video data and the first audio signature AS 1 , thereby linking the video sequence 10 and the audio sequence 21 to each other. A second video signature VS 2 is generated by digitally signing the second video hash (step S9).
[0045] In addition, a second audio hash is generated by hashing a first target audio portion (step S8). The first target audio portion is captured at some time after the first audio portion 21 and is the audio portion in which the first video signature VS 1 is inserted. Thus, the second audio hash will be the hash of the combination of the audio data and the first video signature VS 1 , thereby linking the video sequence 10 and the audio sequence 21 to each other. A second audio signature AS 2 is generated by digitally signing the second audio hash (step S10).
[0046] Briefly referring to Figure 5 the flowchart in, before returning Figure 3 , once the second video signature VS 2 is generated, it can be inserted into the second target audio portion (step S11). Similarly, the second audio signature AS 2 can be inserted into the second target video portion (step S12). Subsequently, the digital signatures of the consecutive video portions of the video sequence 10 and the consecutive audio portions of the audio sequence 20 can continue in the same manner until the sequence ends. By this "cross-fusion" of the video sequence and the audio sequence using the corresponding signatures, not only can the video sequence and the video sequence be verified separately, but also their common origin can be verified.
[0047] Now we will refer to Figure 3 to explain the insertion of the video signature in the audio sequence and the insertion of the audio signature in the video sequence. The video sequence 10 and the audio sequence 20 are shown again. As discussed above, the image frames are grouped into GOPs, and the audio frames or audio data packets are similarly grouped into groups of audio frames or groups of audio data packets. It should be noted that Figure 3 the illustration in
[0048] As described above, a first video signature VS is generated for the first video portion 11 1 . This first video signature VS 1 is inserted into the first target audio portion 41 of the audio sequence 20, in which case the first target audio portion 41 exactly coincides with the second audio portion 22. The position where the first video signature VS is inserted in the audio sequence 20 is indicated by hatching 1 . Depending on the comparison of the generation frequency of the video signature with the length of the audio portion, the first target audio portion 41 can be an audio portion earlier or later than the second audio portion 22. Thus, the first target audio portion can be the first audio portion, the second audio portion, or a later audio portion. A second video signature VS is generated for the second video portion 12 2 , and this second video signature VS 2 is inserted into the second target audio portion 42. In the example shown, the second target audio portion 42 coincides with the third audio portion 23, but as just explained, this will depend on the relationship between the generation frequency of the video signature and the length of the audio portion. In the same way, a third video signature VS is generated for the third video portion 13 3 . This third video signature VS 3 is inserted into the third target audio portion 43, which is the fourth audio portion 24 in the example shown
[0049] Observing the audio sequence 20, the same principle is followed. A first audio signature AS is generated for the first audio portion 21 1 . This first audio signature AS 1 is inserted into the first target video portion 31. In the example shown, the first target video portion is the second video portion, but as discussed for the video signature, this will depend on the comparison of the generation frequency of the audio signature with the length of the video portion. A second audio signature AS is generated for the second audio portion 2 , and it is inserted into the second target video portion 32, which is the third video portion 13 in this example. A third audio signature AS is generated for the third audio portion 23 3 . This third audio signature AS 3 is inserted into the third target video portion 33. In the example shown, the third target video portion 33 exactly coincides with the second target video portion 32, i.e., the third video portion 13. Thus, the third video portion will receive two audio signatures AS 2 、AS 3 inserted therein. Continuing with the fourth audio portion ( Figure 3 the full length of which is not shown), a fourth audio signature AS is generated 4and insert it into the fourth target video portion. However, in Figure 3 the fourth target video portion is not shown.
[0050] In video sequence 10, the audio signature AS n can be inserted into an already available element in a video compression format. For example, if the video sequence is encoded using h.264 or another h.26X encoding format, the audio signature can be inserted into a specified type of SEI message, which is sometimes referred to as an SEI frame or an SEI NAL unit. Other video compression formats provide corresponding elements. As an example, if the video sequence is encoded using AV1, the audio signature can be inserted into a specified type of open bitstream unit (also known as an OBU or a metadata OBU).
[0051] Similarly, in audio sequence 20, the video signature VS n can be inserted into an appropriate existing element in an audio encoding format. For example, if the audio sequence is encoded using AAC, the video signature can be inserted into a data stream element (or simply referred to as a DSE), and in the PCM audio encoding format, the video signature can be inserted into the header of a WAV container that encapsulates the audio data. Other audio encoding formats provide similar possibilities.
[0052] The video signature can be generated by cryptographic operations (e.g., according to the method described in the applicant's European patent application No. 4164230), and the audio signature can be inserted into the SEI frame or the corresponding element of the video sequence, as described in the applicant's European application No. 4192018. The same principle as for the video signature can be used to generate and insert the audio signature.
[0053] Figure 6 is a simplified block diagram of a digital signature system 60 that can be used to perform the digital signature methods described above. The digital signature system 60 has: a video input function 61, arranged to receive the video sequence 10; and an audio input function 62, arranged to receive the audio sequence 20. The digital signature system 60 further includes: a video digest function 63, arranged to apply a digest algorithm to the video portion of the video sequence to generate a video digest. As in the above description, in this example, the digest algorithm is a hash function, and thus, the video digest function can be referred to as a video hash function 63. The digital signature system 60 also includes: an audio digest function 64, arranged to apply the digest algorithm to the audio portion of the audio sequence to generate an audio digest. As for the video digest, in this example, the audio digest is a hash, and thus, the audio digest function can be referred to as an audio hash function 64.
[0054] In addition, the digital signature system 60 includes: a signature function 65, arranged to generate a video signature by digitally signing a video digest or a video hash; a video signature insertion function 66, arranged to insert the video signature into an audio sequence; and an audio signature insertion function 67, arranged to insert an audio signature into a video sequence.
[0055] According to the above combination Figures 3 to 5 As described above, the video hash digest function 63 is arranged to generate a first video hash by hashing a first video portion 11 of the video sequence 10, and the audio hash function 64 is arranged to hash a first audio portion 21 of the audio sequence 20 to generate a first audio hash. The signature function 65 is arranged to generate a first video signature VS 1 by digitally signing the first video hash, and to generate a first audio signature AS 1 by digitally signing the first audio hash. The video signature insertion function 66 is arranged to insert the first video signature VS 1 into a first target audio portion 41 of the audio sequence 20. The audio signature insertion function 67 is arranged to insert the first audio signature AS 1 into a first target video portion 31 of the video sequence 10. The video hash function 63 is further arranged to generate a second video hash by hashing the first target video portion 31 including the first audio signature AS 1 , and the audio hash function is arranged to hash the first target audio portion 41 including the first video signature VS 1 to generate a second audio hash. The signature function 65 is arranged to generate a second video signature VS 2 by digitally signing the second video hash, and to generate a second audio signature AS1 by digitally signing the second audio hash.
[0056] In the same manner as discussed with respect to Figure 5 above, the video signature insertion function 66 can alternatively be arranged to insert the second video signature VS 2 into a second target audio portion 42, and the audio signature insertion function 67 can alternatively be arranged to insert the second audio signature AS 2 into a second target video portion 32.
[0057] The functions 61 to 67 of the digital signature system 60 can be implemented using circuitry or processing circuitry including a general-purpose processor, a special-purpose processor, an integrated circuit, an ASIC (“application-specific integrated circuit”), conventional circuitry, and / or combinations thereof that are configured or programmed to perform the disclosed functions. In the present disclosure, such circuitry is hardware that executes or is programmed to execute the functions. Such hardware can be any hardware that is programmed or configured to perform the functions disclosed herein or otherwise known.
[0058] In a pure hardware implementation, each of the functions can have corresponding circuitry that is specifically designed for that function. Such circuitry can be in the form of one or more integrated circuits (e.g., one or more application-specific integrated circuits or one or more field-programmable gate arrays).
[0059] In an implementation that also includes software, such circuitry can include a processor. Processors are considered processing circuitry or circuitry because they include transistors and other circuitry therein. In such cases, the circuitry can be considered a combination of hardware and software that is used to configure the hardware and / or the processor.
[0060] It should be understood that a combination of hardware implementation and software implementation can also be had, meaning that some functions are implemented by dedicated circuitry and some functions are implemented in software, i.e., in the form of computer code executed by a processor.
[0061] In Figure 7 is shown a simplified block diagram of a camera 2. The camera 2 has a lens 71 and an image sensor 72 for capturing an image of a monitored scene. The camera 2 further includes a video encoder 73 for encoding video data captured by the image sensor 72. To capture sound in the scene, the camera 2 includes a microphone 74. Additionally, the camera 2 includes an audio encoder 75 for encoding audio data captured by the microphone 74. Additionally, the camera 2 includes a digital signature system 60, such as Figure 6 the digital signature system shown in. A network interface 76 is capable of transmitting the encoded video data and audio data, including video signatures and audio signatures. As will be appreciated by those skilled in the art, the camera 2 can have additional components, but only those components relevant to the description of the present invention are shown and discussed here.
[0062] To make the methods and systems described above more resilient to transmission losses, video signatures can be inserted into the video sequence and audio signatures can be inserted into the audio sequence. When generating a video signature, in this scenario, it can be inserted into the video portion following the generation of the video signature or otherwise the most recent video portion. In the same way, when generating an audio signature, it can be inserted into the audio portion following the generation of the audio signature or otherwise the most recent audio portion. In cases where the transmission bandwidth from the camera is low, it may be decided to transmit only the video or only the audio. In such cases, the transmitted sequence cannot be verified unless the corresponding signature is also inserted into the sequence from which it originated. Thus, if only the video is transmitted, the video signatures will not be transmitted unless they are also included in the video sequence. Similarly, if only the audio is transmitted, the audio signatures will not be transmitted unless they are also included in the audio sequence. As a further adjustment, if it is decided to revert to "video-only" transmission during a temporary bandwidth limitation, the insertion of audio signatures into the video sequence can be paused during the period of "video-only" transmission, so that the video sequence can be verified independently. Of course, this can be applied mutatis mutandis to "audio-only" transmission. During such periods of "video-only" or "audio-only" transmission, the sequences not being transmitted can be locally stored in the camera for subsequent transmission when network conditions improve. It should be noted that in this case, if the insertion of signatures into the "accompanying sequence" is paused during "video-only" or "audio-only" transmission, the link between the video sequence and the audio sequence cannot be safely verified. It may be useful to insert metadata into the transmitted sequence to indicate that the transmission of other sequences has been paused.
[0063] It should be appreciated that those skilled in the art can modify the embodiments described above in various ways and still use the advantages of the invention shown in the above embodiments. For example, video signatures and audio signatures are described as being generated for groups of pictures and for groupings of audio frames or audio data packets. However, it is also possible to generate signatures more or less frequently. For example, video signatures can be generated for each image frame and audio signatures can be generated for each audio frame or audio data packet. Nevertheless, generating signatures for multiple frames or packets is generally more bitrate-efficient. Not generating signatures too frequently will also be more computationally resource-saving.
[0064] Above, video digests and audio digests are described as hashes of video data and audio data respectively. Other types of digests may also be useful, such as a list of hashes of individual frames in signed groups of frames or packets.
[0065] The signature function can use one and the same security element to generate video signatures and audio signatures. It is also possible to use separate security elements to sign two different sequences. In this case, the link between them can be ensured by linking the two security elements in one and the same device to each other.
[0066] Although the above image frames and audio frames or audio data packets are described as being encoded, the same digital signature concept can generally be used for unencoded data or raw data as long as the data to be signed is not subsequently encoded. Encoding more or less inevitably results in information loss, which can lead to verification failures even without tampering or data packet loss.
[0067] As mentioned above, the video sequence can be encoded using h.264, h.265 or another h.26x encoding standard, or using the AV1 encoding standard. In a variant, other video encoding standards such as LCEVC or MJPEG can be used.
[0068] As mentioned above, the audio sequence can be encoded using the AAC encoding standard or the PCM encoding standard. As mentioned above, PCM-encoded audio data can be packed in a WAV container. In a variant, any encoding standard capable of handling metadata can generally be used.
[0069] Therefore, the present invention should not be limited to the embodiments shown, but should be defined only by the appended claims.
Claims
1. A method for digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously so that they express one and the same captured scene, the video sequence comprising a continuous video portion and the audio sequence comprising a continuous audio portion, the method comprising: generating a first video summary by applying a summarization algorithm to a first video portion of the video sequence, generating a first audio summary by applying a summary algorithm to a first audio portion of the audio sequence, generating a first video signature by digitally signing the first video summary, generating a first audio signature by digitally signing the first audio summary, inserting the first video signature into a first target audio portion of the audio sequence, inserting the first audio signature into a first target video portion of the video sequence, generating a second video summary by applying a summary algorithm to the first target video portion including the first audio signature, generating a second audio digest by applying a digest algorithm to the first target audio portion including the first video signature, generating a second video signature by digitally signing the second video summary, and A second audio signature is generated by digitally signing the second audio summary.
2. The method according to claim 1, further comprising: inserting the second video signature into a second target audio portion of the audio sequence, and The second audio signature is inserted into a second target video portion of the video sequence.
3. The method according to claim 1, further comprising: inserting the first video signature also into the video portion of the video sequence, and The first audio signature is also inserted into the audio portion of the audio sequence.
4. The method according to claim 2, further comprising: inserting the second video signature also into the video portion of the video sequence, and The second audio signature is also inserted into the audio portion of the audio sequence.
5. The method according to claim 1, wherein: The digest algorithm is a hash function.
6. The method according to claim 1, wherein: Each video section is a group of pictures.
7. The method according to claim 1, wherein: Each audio portion is an audio frame or an audio packet.
8. The method according to claim 1, wherein: Each audio signature is inserted into a corresponding SEI message or open bitstream unit of the video sequence.
9. The method according to claim 1, wherein: Each video signature is inserted into a corresponding data stream element or header of the audio sequence.
10. The method according to claim 1, wherein: The video sequence and the audio sequence are captured by a single device comprising an image sensor and a microphone.
11. A digital signature system for digitally signing a video sequence and an audio sequence, the audio sequence and the video sequence being captured simultaneously such that they express one and the same captured scene, the system comprising circuitry configured to perform: a video input function arranged to receive a video sequence, the video sequence comprising consecutive video portions, an audio input function arranged to receive an audio sequence, the audio sequence comprising continuous audio portions, a video summarization function arranged to apply a summarization algorithm to the video portion of the video sequence, thereby generating a video summary, an audio summarization function arranged to apply a summarization algorithm to the audio portion of the audio sequence, thereby generating an audio summary, a signing function arranged to generate a video signature by digitally signing the video summary and to generate an audio signature by digitally signing the audio summary, a video signature insertion function arranged to insert a video signature into said audio sequence, and an audio signature insertion function arranged to insert an audio signature into said video sequence, in, The video summarization function is arranged to generate a first video summary by applying a summarization algorithm to a first video portion of the video sequence, the audio summarization function being arranged to generate a first audio summary by applying a summarization algorithm to a first audio portion of the audio sequence, the signing function being arranged to generate a first video signature by digitally signing the first video summary, and to generate a first audio signature by digitally signing the first audio summary, the video signature insertion functionality being arranged to insert the first video signature into a first target audio portion of the audio sequence, the audio signature insertion functionality being arranged to insert the first audio signature into a first target video portion of the video sequence, the video summarization function being arranged to generate a second video summary by applying a summarization algorithm to the first target video portion comprising the first audio signature, the audio summarization function being arranged to generate a second audio summary by applying a summarization algorithm to the first target audio portion comprising the first video signature, The signing function is arranged to generate a second video signature by digitally signing the second video summary, and to generate a second audio signature by digitally signing the second audio summary.
12. The digital signature system according to claim 11, wherein: The video signature insertion functionality is arranged to insert the second video signature into a second target audio portion of the audio sequence, and The audio signature insertion functionality is arranged to insert the second audio signature into a second target video portion of the video sequence.
13. A camera comprising an image sensor, a microphone, and the digital signature system according to claim 11.
14. A computer-readable storage medium comprising instructions which, when executed by a device having processing capabilities, cause the device to perform the method of claim 1.
Citation Information
Patent Citations
Method and device for signing an encoded video sequence
EP4192018A1
Method and apparatus for providing combined-summary in imaging apparatus
CN105516651A
Audio and video fingerprint identification method and tampering prevention system based on audio and video fingerprint streaming media
CN105550257A
Signed video data with linked hashes
EP4164230A1
Data processing detecting system, additional information embedding device, additional information detector, digital contents, music contents processor, additional data embedding method and contents processing detecting method, storage medium and program transmitter
JP2002091465A