Microphone
By converting audio signals into digital signals and dividing them into blocks within the microphone, generating signatures and embedding watermarks, the problem of difficult audio signal authentication in existing technologies is solved, achieving automated and efficient audio signal authentication.
Patent Information
- Application Number
- CN202480067151.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-21
- Filing Date
- 2024-09-18
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies are insufficient for effectively authenticating audio signals, especially verifying the source of the audio signal, which usually requires extensive manual research.
The audio signal is converted into a digital signal using the microphone's built-in analog-to-digital converter. The signal is then divided into multiple blocks by a block generator. Audio features are extracted and a signature is generated. The hash value or audio feature is signed using a private key, and a watermark is embedded to authenticate the audio signal.
It enables automated and robust authentication of audio signals, capable of identifying the authenticity and source of audio signals, reducing latency, and improving the efficiency and accuracy of authentication.
Smart Images

Figure CN122055720A_ABST
Abstract
Description
[0001] This invention relates to microphones and methods for authenticating audio signals.
[0002] Given the modern ability to create deepfake videos and audio, there is an increasing need to verify and authenticate audio signals such as those used in politicians' speeches.
[0003] Authenticating human audio signals has been extremely difficult until now. Typically, extensive manual research is required to verify or authenticate such audio signals and to verify the source of the audio signals.
[0004] Therefore, the object of the present invention is to provide a microphone, a method for authenticating audio signals, and an audio signal signature module for authenticating audio signals in a simple and robust manner.
[0005] This objective is achieved by the microphone according to claim 1, the method for generating an authenticable audio signal according to claim 6, and the audio signal signature module according to claim 13.
[0006] Therefore, a microphone is provided, including a microphone capsule for capturing audio signals. The microphone also includes an analog-to-digital converter configured to convert the captured audio signals into digital signals. A block generator is configured to divide the digital audio signals into multiple blocks. Audio features of the digital audio signals can be extracted, and a private key is provided to generate a signature based on at least one of the multiple audio blocks. Furthermore, the microphone includes an audio output configured to output the captured audio signals and the generated signature.
[0007] According to one aspect, the signature is generated based on audio blocks or the audio features of audio blocks.
[0008] On one hand, a hash generator is provided to generate a hash value based on at least one audio block or based on the audio characteristics of at least one audio block. A private key can be provided to generate a signature for the hash value. Using hash values is advantageous because it achieves efficient data reduction.
[0009] Therefore, it can be determined whether the transmitted audio signal was captured or signed by a specific microphone or a specific software entity.
[0010] According to one aspect of the invention, the microphone further includes a watermark generator to generate a watermark based on a signature of at least one audio block, a signature based on audio features of at least one audio block, a signature based on a hash value of at least one audio block, and / or a signature based on a hash value of audio features of at least one audio block, and to incorporate the watermark into at least one block such that an audio signal with the embedded watermark is output via an audio output. The watermark generator may also generate the watermark based on a signed value and / or a signature based on audio features of at least one audio block.
[0011] Therefore, at the microphone output, the detected audio signal and the generated signature are output. Thus, the detected audio signal and the generated signature can be transmitted together. In the example, the watermark generator generates a watermark based on the generated signature and incorporates the watermark into at least one audio block. Preferably, the watermark is incorporated into each audio block such that each audio block can be used to verify the authenticity of the audio signal based on the values embedded in the audio block.
[0012] Cryptographic signatures of hash or audio characteristics and verification of cryptographic signatures can be performed using private and public key pairs.
[0013] Preferably, the signed value (i.e., the signature) is transmitted together with multiple audio blocks. Alternatively, they can be transmitted via different methods or channels.
[0014] According to one aspect of the invention, a watermark containing a generated signature for an audio block can be transmitted in subsequent audio blocks. This is advantageous if the audio signal is to be transmitted in real time, allowing for the reduction of any latency. It is also possible if the signed value is not embedded in the watermark but is transmitted in a different manner (e.g., via real-time streaming based on IP, TCP, or UDP protocols).
[0015] According to one aspect of the invention, the watermark may include metadata of the captured audio signal. The metadata may include microphone identifier, microphone location / position information, recording time and data, and / or microphone model, etc.
[0016] Alternatively, metadata can be transmitted via different communication channels. Alternatively, a signature of the metadata can be generated based on the private key.
[0017] As illustrated in the example, the signed metadata value of a previous block can be embedded in or transmitted with the current audio block. Because audio blocks are arranged in a chain and optionally numbered sequentially, it is possible to detect if a block is missing or has been removed from the audio block chain. Even without sequential numbering, the removal of an audio block will be noticed when the signed metadata of one block is transmitted with the next block. The removal of a block causes a mismatch between the metadata and the audio block, thus indicating manipulation of the signal.
[0018] The metadata values of an audio block (e.g., audio characteristics, recording time, sequence number, recording location, etc.) are signed with a signature using the private key of the signing entity (such as a microphone or software), and can be authenticated by checking the signature with the public key.
[0019] On the receiver side, metadata (e.g., a signed audio feature or hash value of the received audio signal) can be authenticated by checking the signature of the value. Furthermore, the received audio feature or hash value is compared with the audio feature or hash value of the received audio signal. If the signature is verified and the value comparison is positive, then the received audio signal is authentic. A similarity threshold can also be defined to accept audio signals as authentic.
[0020] If hash values have been extracted from the audio features of the audio signal, they are authenticated at the receiver by checking the signatures of these hash values. Subsequently, the hash values of the audio features of the received audio signal are compared with the received hash values.
[0021] Without using hash values, the audio features of an audio signal or audio block can be signed (i.e., a signature can be generated) and embedded into the transmitted audio signal or transmitted on an alternative channel (e.g., stored on a cloud server). On the receiver side, the embedded audio features are extracted or received via different channels (e.g., downloaded from a cloud server) and compared with the audio features of the received signal to authenticate the received signal. The comparison algorithm can adapt to variations in, for example, the volume and / or equalization of the received signal. Therefore, the comparison can be more robust. Additionally, the signature of the signed metadata can be checked to authenticate the signer of the audio file.
[0022] The present invention also relates to a conferencing system including a microphone for detecting audio signals from participants. This microphone may correspond to the microphone mentioned above, such that it outputs blocks of audio data along with signed metadata.
[0023] The authenticity check of the received audio signal can be performed by a decoder, which can be implemented, for example, in a cloud service, computer, tablet computer, or smart device.
[0024] A block divider can be used to divide an audio signal into multiple audio blocks. The length of an audio block can be, for example, between 0.5 seconds and 20 seconds, or between 1 second and 10 seconds. A digital watermark is embedded in each audio block, or in each nth audio block. The digital watermark may include the hash value of the current or previous audio block, audio characteristics or hashed audio characteristics, and optional additional metadata such as the recorded location or time, which is signed with the microphone's private key. Preferably, this is performed before the audio signal is output via the microphone's audio output. In other words, according to one example, the generation of signed metadata and the embedding of the signed metadata or watermark into the audio signal can be performed within the microphone.
[0025] Alternatively, the generation and embedding of signed metadata can be performed externally to the microphone, such as in a smart device coupled to the microphone, or later in audio processing software.
[0026] The audio blocks of the audio signal can be arranged sequentially, allowing each audio block to have a sequence number. This is advantageous because later, during the authentication of the audio signal, the sequence number of the audio block can be checked to determine whether an audio block has been removed from the audio block chain. This can also detect the order in which changes were made to the original blocks.
[0027] Optionally, metadata about the audio block (such as the microphone user, date, time, GPS location, etc.) can also be embedded in the watermark. This increases the probability of reliable authentication results.
[0028] According to one aspect of the present invention, a method for generating an authenticated audio signal is provided. The method involves capturing an audio signal or receiving a captured audio signal. The captured or received audio signal is converted into a digital signal when needed. The digital audio signal is divided into multiple audio blocks. The audio blocks or their hashes are signed with a private key, or the audio characteristics of the audio blocks are signed with a private key. At least one audio block and the signed audio block or its signed audio characteristics are output.
[0029] This invention also relates to an audio signal signing module. The audio signal signing module is used to sign the metadata of an audio signal using a private key. This metadata contains audio characteristics of the audio signal or audio block. The metadata may contain additional information such as location, time, or sequence number. The signed audio signal or audio signal and signed audio characteristics can then be transmitted or stored. Afterwards, the signature of the received audio signal or received audio characteristics can be checked using a public key to determine whether the audio signal is authentic.
[0030] To reduce the data size of the signature unit, the hash value of the audio block of the audio signal or the hash value of the audio characteristics of the audio block can be determined. In this case, the hash value can be signed with a private key. Afterward, the audio signal and the signed hash value can be transmitted together. Optionally, the hash value can be transmitted over different channels.
[0031] The signing unit can sign at least one audio block, at least one audio feature of at least one audio block, the hash value of at least one audio block, or the hash value of at least one audio feature of at least one audio block using a private key.
[0032] To further identify the source of the audio signal (e.g., the speaker), this information can be included in the watermark as metadata (e.g., time, date, location, speaker ID, etc.). Alternatively, this information can be stored on a (central) server or a distributed ledger. This metadata can also be signed with a private key.
[0033] These and other aspects of the invention will be described in more detail with reference to the following accompanying drawings.
[0034] Figure 1 A block diagram of a microphone according to an embodiment is shown.
[0035] Figure 2 A block diagram of a microphone according to another embodiment is shown.
[0036] Figure 3 A schematic representation of the authentication of audio signals detected by a microphone is shown.
[0037] Figure 4 A block diagram of a method for providing authenticated audio signals is shown.
[0038] Figure 5 A block diagram of a method for authenticating audio signals is shown.
[0039] Figure 6 A block diagram of the audio signal signature module is shown.
[0040] Figure 7 A block diagram of the audio signal signature module is shown.
[0041] Figure 8 A block diagram of the audio signal signature module is shown, and
[0042] Figure 9 A block diagram of the audio signal signature module is shown.
[0043] Figure 1A block diagram of a microphone according to an embodiment is shown. Microphone 100 includes a microphone capsule 110 that captures audio signals and outputs detected audio signals 111. Microphone 100 also includes an analog-to-digital converter (ADC) 120 that digitizes the captured audio signals 111 and outputs a digital audio signal 121. Furthermore, a block divider 130 is provided that divides the digital audio signal 121 into multiple audio blocks 131. The multiple audio blocks 131 are input to a hash generator 150, which generates a hash value 151 for each audio block 131. The hash values 151 are input to a signature unit 140, where each hash value 151 is signed with the private key 142 of microphone 100. The output 141 of signature unit 140 corresponds to the hash value signed with the private key 142. At audio output 190, multiple audio blocks 131 and a signed hash value 141 are output together, and the hash value can be embedded in the output audio signal 191, for example, as metadata in a file or stream. Alternatively, the hash value can be transmitted in a separate stream.
[0044] Therefore, according to Figure 1 For example, authentication is performed based on the hash value of the audio signal. In other words, a signed hash value can be used for authentication on the receiver side, for example, based on a public key. This achieves accurate bit-level verification.
[0045] according to Figure 1 In different examples, multiple audio blocks 131 can be directly input to the private key signing unit, where a signature of at least one audio block 131 is generated based on the private key. The generated signature 141 can be output to the output, where the audio blocks are output together with the generated signature 141. In this example, instead of generating a hash value, the signature of the audio blocks is generated directly.
[0046] according to Figure 1 In another example, the audio characteristics of an audio block can be determined, and the private key signing unit 140 generates a signature based on the video characteristics of the audio block and outputs the generated signature 141 via the output unit 190.
[0047] Figure 2A block diagram of a microphone according to another embodiment is shown. Microphone 100 includes a microphone capsule 110 that captures an audio signal and outputs a captured audio signal 111. The captured audio signal 111 is forwarded to an analog-to-digital converter (ADC) 120, which digitizes the audio signal 111 and outputs a digital audio signal 121. The digital audio signal 121 is input to a block divider 130, which outputs multiple audio blocks 131. Optionally, the audio blocks 131 are forwarded to a hash generator 150, which generates a hash value 151 for each audio block 131. In a private key signing unit 140, the hash value 151 of each audio block 131 is signed with a private key 142 and optionally encrypted, and the signed hash value 141 is output to a watermark generator 160, which generates a watermark 161 and embeds the watermark 161 into the multiple audio blocks 131. The hash generator 150 can also generate hash values of the audio features of the audio blocks. The audio output 190 receives a plurality of audio blocks 131 with an embedded watermark 161 and outputs the audio blocks 131 together with the embedded watermark 161 as an output audio signal 191.
[0048] Preferably, the watermark 161 is embedded in each audio block 131. However, the watermark may also be embedded in only some audio blocks.
[0049] Therefore, at the output of audio output 190 (i.e., at the output of microphone 100), an audio block 131 with an embedded watermark 161 is output. This audio signal (with the embedded watermark) can be stored or transmitted. Since the watermark 161 is embedded in, for example, essentially every audio block 131, each audio block 131 can be authenticated individually. If a watermark with the hash value of the audio block is embedded in a subsequent audio block, then the current block and the subsequent audio blocks are requested for authentication. The watermark 161 may also be embedded only in some audio blocks within the audio block. The watermark may also include additional information, such as block boundaries. This information can be used when authenticating the audio signal.
[0050] according to Figure 2 In this example, the generation of the hash value can be omitted, and the audio block or the audio features of the audio block can be received by the private key signing unit 140. The private key signing unit 140 generates a signature based on the private key 142 and outputs the generated signature 141 to the watermark generator 160, which generates a watermark based on the generated signature. This example is advantageous because it does not require the generation of the hash value of the audio block or the hash value of the audio features of the audio block.
[0051] Therefore, audio signal authentication can be performed by verifying the signatures of hash values or audio features of multiple audio blocks based on the public key associated with private key 142. Furthermore, the verified hash values or audio features are compared with the hash values or audio features of multiple received audio blocks to determine if the audio signal has been altered. If the hash values or audio features correspond to each other, then the audio signal can be authenticated.
[0052] Audio watermarks can be prominent identifiers embedded in audio signals, such as copyright information previously used to identify audio signals. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy it. Preferably, if the audio signal with the embedded watermark is copied, stored, or transmitted, the embedded watermark will not be altered.
[0053] Therefore, according to Figure 2 For example, authentication is performed based on the hash value of the audio signal or the hash value of audio features, along with the use of a watermark. This achieves bit-level accurate verification. Figure 3 A schematic representation of the authentication of an audio signal is shown. For example, a microphone 100 is used to capture an audio signal. The audio signal may also be a pre-captured audio signal. The captured audio signal is divided into multiple audio blocks, and at least one signature is generated based on at least one audio block (or an audio feature of at least one audio block). Optionally, a hash value of the audio block or a hash value of the audio feature of the audio block may be generated and signed with the private key 142 of the microphone 100. In particular, a watermark 161 may be embedded in the audio signal. The audio signal may be transmitted via network 200 or stored. The authenticator 300 may receive the audio signal with the embedded watermark and may verify the signature of the signed audio block or the signed audio feature based on the public key 401 associated with the microphone, which may be stored, for example, on server 400 or a distributed ledger.
[0054] Optionally, if a hash value has already been used, the hash value can be extracted, and the extracted hash value can be compared with a hash value based on the received audio signal to authenticate the audio signal.
[0055] Therefore, according to Figure 3 In this example, authentication is performed based on the signature and watermark of the audio block. This achieves bit-level accurate verification.
[0056] Figure 4A block diagram of a method for generating an authenticable audio signal is shown. In step S11, microphone 100 captures or receives an audio signal. The audio signal is digitized in AD converter 120, and block generator 130 divides the audio signal into multiple audio blocks. If the audio signal is already a digital audio signal, the AD conversion can be omitted. In other words, the audio signal is processed based on audio blocks. The length of an audio block can be, for example, between 1 second and 10 seconds. In step S12, a watermark can be generated based on watermark generator 160 and embedded into the audio signal. The watermark can contain metadata of the previously processed audio blocks. Optionally, step S12 can be omitted. In step S13, audio features of the audio signal can be extracted, which can be related to the audible features or audio characteristics of the audio signal. One example is MFCC (Mel-frequency cepstral coefficients) calculated from the audio signal.
[0057] In step S14, a hash value for an audio block or a hash value for an audio feature can be determined. Such a hash value can be determined, for example, using MD5 or SHA-256 methods. The hash value can be determined based on the audio signal or based on the audio features of at least one audio block. Alternatively, hashing can be skipped and only the audio features of at least one audio block can be extracted.
[0058] In step S15, a signature of the audio block or a signature of the audio features of the audio block is generated based on private key 142. Alternatively, a signature can be generated based on metadata (hash value, audio features, or other data). The private key can be associated with a microphone or a signing entity. A watermark to be embedded in the audio block may include a hash value of the audio features of the audio block or directly include the audio features (without a hash value), and may also include other information such as the audio block's sequence number, optionally the date, time, and location, as well as a digital signature. The obtained audio signal can be identified using the private key associated with the microphone and by signing the metadata with the private key.
[0059] Figure 5 A flowchart of a method for authenticating audio signals is disclosed. In step S21, the received audio blocks are examined to determine block boundaries. This can be done, for example, by means of a synchronization signal. Once the block boundaries are determined, the process continues to step S22, where a watermark 161 is extracted from the audio blocks. The extracted watermark 161 can be associated with the current audio block or a previous audio block.
[0060] In step S23, the signature of the signed metadata embedded in the audio block is verified using public key 401. This is advantageous because any manipulation of the watermark can be detected, and the microphone from which the audio signal was previously captured can be identified.
[0061] In step S24, audio features of the audio block can be extracted.
[0062] In step S25, the audio features of the received audio block are determined and compared with the embedded audio features of the received audio block. If the audio signal has been tampered with, the audio features of the received audio signal will not correspond to the audio features embedded in the received audio signal.
[0063] In step S26, optionally, the block number of the audio block that can be embedded in the watermark is determined and compared with the previous sequence number to determine whether any audio block is missing or whether the order of the blocks has been modified.
[0064] If a watermark containing information about an audio block is embedded into the same audio block, robust audio features must be used, meaning the watermark doesn't significantly affect those features. As an example, such audio features could be related to human hearing ability. Alternatively, if the current audio block is embedded with a watermark associated with a previous audio block, that watermark is known before the audio features of the current audio block are determined. Therefore, a watermark associated with a previous audio block can be embedded into the current audio block before determining its audio features. This allows the algorithm to work on the same, already watermarked audio signal for both signatures and authentication.
[0065] If the microphone is used to detect live audio signals and therefore audio transmission must be performed with low latency (e.g., in a live interview or press conference), then optionally, the hash of the audio block can be embedded as a watermark into subsequent audio blocks. This is advantageous because latency can be reduced since the audio signal can be output without waiting for the current audio block to end.
[0066] In this scenario, any tampering with the audio signal can be detected even if the watermark does not include the audio block's serial number. However, using the serial number as part of the watermark can provide more information about modifications to the audio stream / file.
[0067] The drawback of using hash offsets by including the hash value in the watermark of subsequent audio blocks is that the last audio block at the end of the audio sequence cannot be authenticated. However, this should not be a problem if the length of the audio block is between 1 and 10 seconds.
[0068] According to the example, each audio signal captured by microphone 100 and output by the microphone can include a watermark, which can be used to identify the actual microphone that captured the audio signal. If the microphone is registered, then the corresponding microphone can be identified. By using the sequence number or metadata of the audio block embedded in subsequent audio blocks, it can be determined when a portion of the audio signal was removed.
[0069] It can also determine when an audio signal was introduced between subsequent audio blocks, or when the order of the original audio blocks was changed.
[0070] According to the present invention, a method for authenticating an audio signal already detected by a microphone is provided. The received audio signal is authenticated using a pair of private and public keys. This is advantageous because neither the microphone nor the decoder needs to be online. All necessary information is embedded in the audio signal.
[0071] Alternatively, the audio signal may be part of a video file.
[0072] Alternatively, the signed metadata can be included in the audio file (e.g., in an ADM file format) instead of being included in the watermark. In this case, the audio signal cannot be modified.
[0073] Optionally, the signed metadata can be transmitted independently of the audio signal. In this case, no watermarking is performed. For example, the metadata might be uploaded to a web server or distributed ledger, while the audio is distributed via a separate channel. In this scenario, the metadata, the audio signal, or both might require some form of synchronization information to ensure that the correct metadata is assigned to the correct audio block. This could consist, for example, of audio characteristics or block numbers.
[0074] Optionally, the process of generating the signed audio signal can be performed in a software solution based on a pre-existing audio recording or audio stream. The private key can be entered into the software, for example, via a dongle, text input, and authorized via biometric signals such as fingerprints or facial recognition.
[0075] Optionally, the authentication process can calculate a similarity score between the analyzed signal and the audio features of the signed audio signal. This allows for determining the probability that the signal is still authentic even if slight signal processing, such as gain, has occurred. Here, the hash value is not determined.
[0076] According to the present invention, a method for proving the authenticity of an audio signal through digital signature is provided. This means that for a signed audio signal, it is possible to check whether the assumed creator of the audio signal is the actual creator. Furthermore, it is possible to check whether the audio signal has been altered after its creation (changes to the signal, removal of recorded portions, and addition of other audio recordings).
[0077] Figure 6A block diagram of an audio signal signature module is shown. The audio signal signature module 500 includes a block divider 130 that divides the digital audio signal 121 into multiple audio blocks 131. Optionally, if the audio signal has already been divided into audio blocks, the block divider can be omitted. Additionally, an audio feature unit 180 can be provided, which extracts audio features from the audio blocks 131 and outputs audio-related information 181. A private key signature unit 140 is provided, which can receive the audio-related information 181 and generate a signature based on the audio signal and / or audio features 143 (i.e., audio-related information). The signature unit 140 can output the multiple audio blocks and audio feature information separately.
[0078] According to the example, the digital audio signal 121 can be divided into multiple audio blocks 131 of length, for example, 1 to 10 seconds, in the block divider 130. To enable later authentication of the audio signal, audio-related information that can be generated from the audio signal is obtained during the signing process. This can be the blocks of the digital audio signal 131 themselves or audio features extracted from the audio blocks (optionally, hashes generated based on the blocks of the digital audio signal, or hashes generated based on the audio features of an audio block). Additional information 113, such as location or time, along with this audio signal-related information, can be digitally signed using the private key 142.
[0079] Private key 142 may be part of the software / hardware since production, or it may be provided in other ways, such as via a dongle, keypad, or other means. Additionally, the private key may be unlocked via some form of user authentication, such as a biometric sensor or password. For each private key, there exists a public key used to verify the digital signature. Each user / device / organization should use a unique key pair. The metadata used later for authentication may also include a hash of the audio block or a hash of the characteristics of the audio block, optionally along with additional information such as time or location, and must have a digital signature that signs this information. This information, along with the digital signature, may be added to the audio signal as a separate metadata stream, embedded as metadata in the audio file, or embedded in the audio signal using a watermarking method. Signed metadata and audio recordings may also be distributed through a separate channel. Where the metadata is not embedded in the audio signal itself, the metadata may include references such as absolute or relative timestamps to link it to the audio block.
[0080] Ideally, to reduce the amount of data (embedded or as metadata) that must be provided in addition to the audio signal, only the hash value of the metadata can be signed. In this case, for example, the complete audio characteristics of the block plus additional information such as location, plus a signature of the hash of the aforementioned value, could be provided. This could result in a shorter digital signature.
[0081] Figure 7 A block diagram of the audio signal signature module is shown. According to... Figure 7 The audio signal signature module basically corresponds to Figure 6 In addition to the audio signal signature module, a hash generator 150 is also provided. The hash generator 150 can generate a hash value 151 based on audio block 131 from multiple audio blocks or based on audio features from box 180. The signature unit 140 is capable of signing audio block 131. The signature unit 140 can sign extracted audio features of an audio block and / or it can sign the hash value 151 of an audio block or extracted audio features.
[0082] Provide unit 103, or unit 103 receives audio block 131, audio-related information 181 and hash value 151 and outputs one of them.
[0083] Figure 8 A block diagram of the audio signal signature module is shown. Figure 8 The audio signal signature module basically corresponds to Figure 6 or Figure 7 An audio signal signature module is provided. Here, a block divider 130, a watermark generator 160, an audio feature unit 180, and a hash generator may be provided. The audio feature unit 180 and the hash generator 150 are optional. The block divider 130 receives a digital audio signal and divides the audio signal into multiple audio blocks 131. In the watermark generator 160, metadata can be embedded as a watermark into the audio signal. The output 161 of the watermark generator 160 can be used as an audio output. This is particularly advantageous if low-latency audio transmission is required, such as for live performances. The audio blocks 131 can be processed by the audio feature unit 180 to extract audio features. The hash generator 150 can generate a hash value 151 based on the audio blocks 131, the output 161 of the watermark generator, and / or based on the audio features 181 extracted by the audio feature unit 180. The outputs of the block divider, the watermark generator, the audio feature unit 180, and / or the hash generator 151 can be used as inputs to the signature unit 140. As in the example above, the signing unit 140 uses the private key 142 to sign the signal at its input and any other possible information, that is, the signing unit 140 generates a signature.
[0084] Figure 9 A block diagram of an audio signal signature module is shown. The audio signal signature module can be based on... Figure 6 , Figure 7 or Figure 8 The audio signal signature module 500 specifically includes a block divider 130, a watermark generator 160, and an audio feature unit 180.
[0085] According to the implementation, the sequence of subsequent audio blocks 131 is processed and a watermark is embedded, such that a sequence of audio sub-blocks is output at output 190. In an audio block N having multiple sub-blocks, one or more audio features are extracted from the audio feature unit. The extracted features are digitally signed by the signature unit 140 and embedded as a watermark into the subsequent audio block N+1. Figure 9 In the example, the audio features of an audio block are embedded into the watermark of subsequent audio blocks.
[0086] To reduce latency, for example, in live applications, the signed metadata of the current audio block can be output along with a subsequent audio block, or entirely independently of audio processing. Here, the watermark added to the current audio block contains the signed metadata of the last audio block. Therefore, the algorithm does not need to wait until the entire block is captured before outputting. The watermark embedding can work on smaller sub-blocks of the actual audio block. The latency of the entire process then depends on the processing time of audio feature extraction, signing, watermark embedding, and the size of the sub-block, not on the size of the entire audio block. Since the size of the sub-block can be much smaller than the size of an audio block, for example, 20 ms, this results in lower latency than embedding the audio features of one block into the same block.
[0087] For authentication, an audio signal and signed metadata are required. The metadata can be provided via a side channel or embedded in the audio signal as a watermark. When using a watermark, the watermark must be read from the audio signal. One way to achieve this, where block boundaries are required for watermark readout, is to try different offsets of the block boundaries until successful watermark readout becomes possible. Another approach is to embed some synchronization signal with the watermark into the audio signal.
[0088] To authenticate an audio signal, it must be processed within a block with the same boundaries used for signing. Since audio may be cut, block boundaries are not readily known at authentication time. In one approach, where signed metadata is provided separately from the audio signal—for example, as metadata within an audio recording file—the metadata can be attached with timestamps that can be adjusted by audio processing software that cuts the audio. Block boundaries can then be calculated based on these timestamps. If this is not the case, another approach is to try assuming the recording begins with an entire signed block and attempt to authenticate it using the available metadata. If this doesn't work, it can be repeated, for example, with sample offsets. If even offsets larger than the block size do not result in successful authentication, the signal can be considered unauthenticable. If an offset results in successful authentication, the block boundaries can be calculated based on that offset.
[0089] To also authenticate slightly altered audio signals (e.g., in cases of volume changes or equalizer use), a pre-described method without hashing should be used. Where the audio signal itself or the audio features used are different, a similarity measure can be used in conjunction with a threshold to determine if the signal has been altered too much or if it is still authentic. Acceptable thresholds may vary for different use cases.
[0090] To enhance system security, it is meaningful to establish a process that allows for key revocation, for example, in the event of key theft.
[0091] Public keys used for authentication can be provided in various ways. One option is to store all public keys in a centralized database to make it easy to find the required key. To address potential trust issues in this centralized instance, a database with all public keys can be provided based on distributed ledger technology. Another option is for organizations / individuals using this technology to provide public keys on their own websites.
[0092] List of reference numerals
[0093] 100 microphones
[0094] 101 or unit
[0095] 102 Audio signal metadata
[0096] 103 or unit
[0097] 110 Microphone Cap
[0098] 111 Microphone audio signal
[0099] 112 audio signal
[0100] 113 Metadata
[0101] 120 AD converter
[0102] 121 digital audio signal
[0103] 130-block divider
[0104] 131 audio blocks
[0105] 140 Private Key Signature Unit
[0106] 141 Output
[0107] 142 Private Key
[0108] 143 Signature Output
[0109] 150 Hash Generator
[0110] 151 hash value
[0111] 160 Watermark Generator
[0112] 161 watermarks
[0113] 180 audio feature units
[0114] 181 Audio-related information
[0115] 190 audio output
[0116] 191 First audio output signal
[0117] 192 Second audio output signal 200 Network
[0118] 300 Authentication
[0119] 400 server
[0120] 401 Public Key
[0121] 500 audio signal signature module
Claims
1. A microphone (100), comprising: Microphone capsule (110), which is suitable for capturing audio signals, An AD converter (120) is configured to convert a captured audio signal (111) into a digital signal (121). A block generator (130) is configured to divide the digital audio signal (121) into multiple audio blocks (131). A private key signing unit (140) is configured to generate at least one signature (141) based on at least one of the plurality of audio blocks (131) using a private key (142) associated with the microphone (100); and An audio output (190) is configured to output the at least one audio block (131) and at least one generated signature (151).
2. The microphone (100) according to claim 1, wherein, The at least one signature (141) is generated based on at least one audio block (131) or based on the audio features of at least one audio block.
3. The microphone (100) according to claim 1, 2 or 3 further comprises: A hash generator (150) is configured to generate a hash value (151) based on at least one audio block (131) or based on the audio features of at least one audio block (131). The private key signing unit (140) is configured to generate a signature based on the hash value of the audio block or the hash value of the audio feature. The audio output (190) is configured to output at least one audio block (131) and the generated signature (151).
4. The microphone (100) according to any one of claims 1 to 3, further comprising: A watermark generator (160) is configured to generate at least one watermark (161) based on at least one audio block, based on the signature of the at least one audio block, based on the signed hash value of the at least one audio block (131), or based on the signed hash value of the audio features of the at least one audio block, and to incorporate the watermark (161) into the at least one audio block (131) such that an audio signal with the embedded watermark (161) is output via the audio output (190).
5. The microphone (100) according to claim 4, wherein, The watermark (161) of the audio block (131) is embedded in subsequent audio blocks (131).
6. The microphone (100) according to claim 4 or 5, wherein, The at least one watermark (161) also includes metadata (102) of the audio signal (121).
7. A method for generating an authenticable audio signal, the method comprising the following steps: Capture audio signals or receive captured audio signals. Convert the captured or received audio signal into a digital audio signal (121). The digital audio signal (121) is divided into multiple audio blocks (131). Using the private key (142), a signature (141) is generated based on the audio block (131) of the audio signal, and Output at least one audio block (131) and the generated signature.
8. The method according to claim 7, further comprising the following step: In the hash generator (150), a hash value is generated based on at least one audio block (151) or based on the audio features of the at least one audio block (131).
9. The method according to claim 7, comprising the following steps: A watermark (161) is generated by a watermark generator (160) based on at least one audio block, based on the audio features of the at least one audio block, based on a signed hash value of at least one audio block (131), or based on a signed hash value of the audio features of at least one audio block. The watermark (161) is incorporated into at least one audio block (131) such that the audio signal is output together with the embedded watermark (161).
10. The method according to claim 9, wherein, At least one watermark (161) includes metadata (102) of the audio signal.
11. The method according to claim 7 or 8, in, The watermark (161) associated with the audio block is embedded in subsequent audio blocks.
12. A method for authenticating an audio signal generated based on the method according to any one of claims 7 to 11, Detect the audio block (131). Verify the signature of at least one audio block based on the public key.
13. The method of claim 12, further comprising the step of: Extract the hash value (151) from the audio block (131). The signature of the hash value (151) is verified based on the public key associated with the private key of the microphone. The extracted hash value (151) is compared with the hash value determined based on the received audio block.
14. An audio signal signature module (500), comprising: A private key signing unit (140) is configured to sign a plurality of audio blocks (131) or at least one audio feature of the plurality of audio blocks (131) using a private key associated with the audio signal signing module (500).
15. A method for generating an authenticable audio signal, the method comprising the following steps: Receives digital audio signals captured by a microphone or receives audio signals. The digital audio signal (121) is divided into multiple audio blocks (131). Perform at least one of the following signature operations: Use the private key (142) to generate at least one signature for at least one audio block (131). Use the private key (142) to generate at least one signature of at least one audio feature of the audio block. At least one signature is generated using the private key to generate at least one hash value of the audio block (131), and Use the private key (142) to generate at least one signature hash value for at least one audio feature of the audio block.
16. The method of claim 15, comprising the following steps: A watermark (161) is generated by a watermark generator (160) based on at least one audio block, based on the audio features of the at least one audio block, based on a signed hash value of at least one audio block (131), or based on a signed hash value of the audio features of at least one audio block. The watermark (161) is incorporated into at least one audio block (131) such that the audio signal is output together with the embedded watermark (161).