Audio signal signature module and method of generating authenticatable audio signals
By segmenting, extracting features, and embedding authentication information through an audio signal signature module, the complexity of audio signal authentication in existing technologies is solved, achieving simple and efficient audio signal authentication and improving security and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SENNHEISER ELECTRONICS GMBH & CO KG
- Filing Date
- 2024-09-23
- Publication Date
- 2026-06-09
AI Technical Summary
Existing technologies are not easy to use for robustly authenticating audio signals, especially when dealing with deepfake videos and audios. Extensive manual research is required to verify the authenticity and source of the audio signals.
The audio signal signature module divides the audio signal into multiple blocks, extracts features and generates a signature, embeds authentication access information, and uses private key signing and public key verification for authentication, reducing storage requirements and improving security.
It enables simple and robust authentication of audio signals, capable of detecting minor tampering, improving the security of audio signals and the reliability of authentication, while reducing storage and bandwidth requirements.
Smart Images

Figure CN122180959A_ABST
Abstract
Description
[0001] This invention relates to an audio signal signature module, a microphone and a method for generating authenticable audio signals, an audio signal authenticator, and a method for authenticating audio signals.
[0002] Given the modern ability to create deepfake videos and audio, there is an increasing need to verify or authenticate audio signals such as those used in politicians' speeches.
[0003] Authenticating human audio signals has been extremely difficult until now. Typically, extensive manual research is required to verify or authenticate such audio signals and to verify the source of the audio signal.
[0004] Therefore, the object of the present invention is to provide an apparatus that enables the generation of authenticable audio signals and an apparatus for authenticating such audio signals, which enables the audio signals to be authenticated in a simple and robust manner.
[0005] This objective is achieved by the audio signal signature module according to claim 1, the microphone according to claim 4, the method for generating an authenticable audio signal according to claim 5, the audio signal authenticator according to claim 6, and the method for authenticating an audio signal according to claim 10.
[0006] This invention relates to generating an authenticable audio signal using an audio signal signature module or a method for generating an authenticable audio signal, and authenticating an audio signal using an authenticator or a method for authenticating an audio signal. The audio signal must be processed by the audio signal signature module to include information that enables subsequent authentication. During the authentication phase, the authenticator authenticates the received or stored audio signal based on information associated with the audio signal.
[0007] According to one aspect, an audio signal signing module is provided, comprising: an input terminal for receiving a digital signal; a block segmenter for segmenting the received digital audio signal into a sequence of multiple audio blocks; an audio feature extractor for extracting audio features from the current audio block; and a signing unit for generating a signature associated with the current audio block by applying a private key to the audio features. The signing unit also provides signed authentication information including the extracted audio features and the signature, wherein the signed authentication information is output to a memory. The audio signal signing module further includes an information embedder for generating an audio block with embedded information by embedding authentication access information into the current audio block. The authentication access information includes information about how the signed authentication information can be retrieved from the memory and / or where the signed authentication information can be retrieved from the memory. Furthermore, an audio output terminal is provided for outputting an authenticable audio signal comprising a sequence of audio blocks with embedded authentication access information.
[0008] According to one aspect, the information embedder is implemented as a watermark generator, which is configured to generate at least one watermark based on authentication access information and introduce the watermark into at least one audio block, such that an audio signal with the embedded watermark is output via an audio output terminal.
[0009] According to one aspect, the signed authentication information also includes metadata belonging to the current block. The metadata may include at least one of the following: an identifier of the user or entity that has signed the audio signal, and in particular, the location and time of signing or generating the audio signal.
[0010] Therefore, a microphone is provided, comprising: a microphone capsule for capturing audio signals; an A / D converter configured to convert the captured audio signals into digital signals; a block generator configured to segment the digital audio signals into multiple audio blocks; an audio feature extractor for extracting audio features from at least one audio block; a signature unit configured to provide a signature of authentication information including the extracted audio features and / or metadata of at least one audio block; an audio output terminal configured to output the signed authentication information to a memory; and an information embedder configured to embed authentication access information into at least one audio block. The authentication access information includes information about how the signed authentication information can be retrieved from the memory or where the signed authentication information can be retrieved from the memory, wherein the audio output terminal is configured to output at least one audio block with the embedded authentication access information.
[0011] The present invention also relates to a method for generating an authenticable audio signal. A digital audio signal is received and segmented into a sequence of multiple audio blocks. Audio features are extracted from the current audio block, and a signature associated with the current audio block is generated by applying a private key to the extracted audio features. Signed authentication information, including the audio features and the signature, is provided. The signed authentication information is output to a memory. An audio block with embedded authentication access information is generated by embedding authentication access information into the current block. The authentication access information includes information about how the signed authentication information can be retrieved from the memory and / or where the signed authentication information can be retrieved from the memory. An authenticable audio signal is output. The authenticable audio signal comprises a sequence of multiple audio blocks with embedded authentication access information. The present invention also relates to an audio signal authenticator comprising a block boundary detector and a block segmenter for receiving the audio signal and outputting a sequence of multiple audio blocks with embedded authentication access information. Furthermore, the audio signal authenticator includes: an access information extractor for extracting authentication access information from the current block; and an authentication access information parser for extracting access information regarding how the signed authentication information belonging to the current audio block can be retrieved from memory and / or where the signed authentication information belonging to the current audio block can be retrieved from memory. The signed authentication information retrieved from memory includes a signature and audio features belonging to the current audio block. Additionally, a perceptual similarity analyzer is provided for comparing audio blocks and audio features with each other and providing similarity results.
[0012] In a similarity analyzer, audio blocks can be compared to the audio features of the original audio signal. These audio features can be true copies of the original audio signal. However, this would require significant storage space to store the original audio signal. If the audio features relate to other features of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Therefore, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within a similarity analyzer. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.
[0013] This invention also relates to a method for authenticating audio signals by an authenticator. An audio signal is received, and a sequence of multiple audio blocks with embedded authentication access information is output. Authentication access information is extracted from the current audio block, and access information regarding how signed authentication information belonging to the current block can be retrieved from memory and / or where in memory the signed authentication information belonging to the current block can be retrieved is extracted. The signed authentication information is retrieved from memory. The signed authentication information includes a signature and audio features belonging to the current block. The audio blocks and audio features are compared with each other to provide similarity results.
[0014] Audio features involve the audio representation of an audio signal, attempting to capture relevant aspects of human auditory perception of the audio signal. Audio features should be chosen to reduce the required data rate (for transmitting the audio features) and thus reduce the storage space required on the server. While the audio feature may be the original audio signal itself, preferably, it can be a transparently coded version of the original signal that requires far less storage space. Another example of an audio feature can be the output of any model that mimics human auditory perception. An audio feature can be the Mel-frequency power spectrum of an audio signal with Mel-frequency cepstral coefficients (MFCC).
[0015] According to one aspect, the authentication access information parser is configured to determine the identity information of the user or entity that initially signed the audio signal, wherein the public key and the user's identity can be retrieved from an external database based on the determined identity information.
[0016] One approach provides a signature verifier that is configured to verify a combination of signature and authentication information using a public key and provide a signature verification result.
[0017] According to one aspect, the audio signal includes a combiner configured to combine a perceived similarity statement and a verification result into a final authentication result.
[0018] According to one aspect, a method for authenticating an audio signal by an authenticator is provided, the method comprising the steps of: receiving an audio signal; outputting a plurality of audio blocks having embedded authentication access information; extracting authentication access information from the audio blocks indicating where authentication information belonging to the current audio block can be retrieved; retrieving authentication information from a memory, the authentication information including a signature, audio features representing the audio blocks, and / or other information; and perceptually comparing the audio blocks and audio features with each other to provide a similarity result.
[0019] The authentication information is signed with the private key of a microphone or a device that performs audio signal capture or post-processing to generate a unique signature indicating the source of the authentication information. The private key may be associated with a user of the signing module, such as a microphone. On the decoder side, the public key can be used to decode or verify the authentication information. In other words, the signature is used to verify the source of the authentication information.
[0020] Therefore, a signature can be used to verify that the authentication information originates from a specific source belonging to the owner of the private key. By further comparing the audio signal with the authentication information, it can be verified that the audio signal has not been tampered with beyond permissible levels. Only a signature is generated based on, for example, the private key used for a microphone.
[0021] It is advantageous to store authentication information, such as metadata, externally rather than within the audio block, as this does not increase the required bit rate and bandwidth. Only the authentication access information is transmitted with the audio block. If authentication of the audio signal is required, the authentication access information is extracted, and the associated authentication information, such as metadata, is retrieved.
[0022] Therefore, audio blocks, independent of the captured audio signal, provide authentication information. If authentication is not performed, this information is unnecessary, and significant bandwidth and bitrate savings can be achieved compared to attaching authentication information to the audio signal in a suitable container format. Furthermore, when authentication is required, access information can be extracted from the audio blocks, and authentication information can be retrieved. Since the authentication process is not latency-critical, extracting authentication information from an external source before performing authentication is sufficient.
[0023] Authentication access information can be embedded as a watermark into an audio block, or as part of an audio block within a dedicated container, or as part of an audio file within a file-based environment.
[0024] Authentication information and authentication access information can be provided for at least one audio block. Therefore, each audio block can be authenticated individually. This allows for exhaustive authentication. Thus, even minor tampering with the audio signal (where authentication information is available) can be detected. This significantly improves the security of the captured audio signal. Providing authentication information for each audio block is advantageous because it even allows verification of audio signals containing different audio blocks from different audio sources. Any audio signal including different audio blocks can be verified, provided the necessary authentication information is available.
[0025] The following describes an example of improving the security of audio signal transmission or storage. Audio features can be extracted from blocks of digital audio signals, and signatures based on these audio features can be generated using a person's or entity's private key. The captured audio signal and the signed authentication information can be output.
[0026] Therefore, it is possible to verify the combination of the received authentication information and signature using a public key, and to determine whether the transmitted audio signal was captured and signed by a specific microphone or a specific software entity with a specific private key by subsequently comparing the audio signal with the audio features included in the authentication information.
[0027] According to one aspect of the invention, the microphone also includes a watermark generator for generating a watermark based on authentication access information.
[0028] You can use a private key and a public key pair to perform signing on audio blocks and verification of the signatures.
[0029] According to one aspect of the present invention, in addition to audio characteristics, the authentication information may also include microphone identification, microphone location / position information, recording time and date, and / or microphone model, serial number, etc.
[0030] Authentication information comprises information that can be used to authenticate the received or selected audio signal. This authentication information enables the authenticator to perform authentication of the audio signal. Authentication information may include information related to the audio signal or its attributes, i.e., audio characteristics. Authentication information may also include metadata related to information independent of the actual audio signal. This information may be the user's ID, the ID of the microphone used to capture the audio signal, the microphone model, date, time, location or position (GPS positioning), and / or the serial number of the audio block.
[0031] The authenticity check of the received audio signal can be performed by a decoder, which can be implemented, for example, in a cloud service, computer, tablet or smart device.
[0032] A block splitter can be used to divide an audio signal into multiple audio blocks. The length of an audio block can be, for example, between 0.5 s and 20 s. In particular, the block length can also vary over time, which can be indicated, for example, by a watermark. A digital watermark can be embedded in each audio block. The watermark can also be embedded in every nth audio block.
[0033] The sequence of audio blocks in the audio signal can be used to associate a sequence number with each audio block. This is advantageous because later, during the authentication of the audio signal, it can be determined whether an audio block has been removed from the chain of audio blocks by checking its sequence number. This can also detect the order in which changes were made to the original blocks.
[0034] Optionally, the metadata of the audio block can include the microphone user, date, time, GPS location, etc. This increases the likelihood that the authentication result is reliable.
[0035] Authentication information may include audio characteristics representing audio signals or audio blocks of audio signals based on human perception, or a transparently encoded version of the original audio signal, to reduce the storage space required by the server.
[0036] Audio signal authentication can be performed by: detecting block boundaries from special information in the watermark or audio container; and extracting block-specific authentication information (e.g., a link (e.g., consisting of a UUID and a user-based unique block number) from the audio watermark or directly from the audio container). The user's public cryptographic key (e.g., identified by the UUID from the link) can be used to verify the authenticity of the authentication information. The audio signal block is compared with the audio features contained in the authentication information to verify its authenticity.
[0037] Using the audio signal authentication method described above, malicious modification of the audio signal can be detected by comparing the audio block with the authentication information associated with the audio block. Modification of the audio signal to be authenticated and / or modification of authentication information stored externally can also be detected by examining the signature of the authentication information. If the audio block of the audio signal is modified and the authentication access information is also changed, the authentication process will notice this because the signature of the authentication information is the signature of another person's or device's authentication information.
[0038] On the one hand, if the authentication process does not provide a clear indication of the authenticity of the audio signal or audio block, the authentication information stored externally can be used to manually verify the audio signal or audio block.
[0039] As an example, the actual audio features extracted from the audio block and optional additional metadata are stored in a signed manner on, for example, a server. A unique (explicit) link to the metadata stored on, for example, the server is embedded in a watermark or in a suitable audio container. This link may include a unique user identifier (UUID) combined with the user-based unique block number of the corresponding audio block.
[0040] To verify the audio signal of an audio block, block boundaries can be detected from a watermark or specific information in the audio container. Then, a link to block-specific, signed authentication information (e.g., consisting of a UUID and a user-based unique block number) is extracted from the audio watermark or directly from the audio container. This block-specific, signed authentication information consists of audio features, optional metadata, and a signature. The combination of authentication information and signature is verified using the user's public cryptographic key (e.g., identified by the UUID from the link). Thus, it is verified whether the authentication information has been signed by the claimed audio signal source. The audio signal block is compared with audio features to verify perceptual similarity, and therefore, the authenticity of the audio signal block. For perceptual comparison, audio features can be determined based on the received audio signal block, and these features can be directly compared with audio features from signed authentication information stored, for example, on a server.
[0041] The proposed authentication method prevents any malicious modification to the actual audio signal by comparing the audio signal block with the signed authentication information. Authentication may fail if the comparison allows a certain limit on modifications and exceeds the allowed threshold. Furthermore, if the authentication information on the server is also altered in addition to malicious modification of the audio signal block, this can be detected by a mismatch between the authentication information and the signature associated with it. Moreover, if, in addition to malicious modification of the audio signal block, the link (within the watermark or audio container) is changed to point to a different (but appropriate) signed authentication information on the server, signature verification and subsequent comparison may indeed succeed, but the signature will be from a different user, provided the attacker cannot access the original user's private key used for cryptographic signing.
[0042] If the (automated) comparison between the audio signal to be authenticated and the authentication information is performed in a manner that allows for minor modifications (e.g., through perceptual encoding), and the result of the comparison is ambiguous (i.e., no clear statement can be made about the authenticity of the audio signal), then a “manual” perceptual comparison can be performed using a link to the authentication information (where the included audio features correspond to the original audio file or its perceptually encoded version) to increase trust in the method.
[0043] Providing audio block-specific links (e.g., via UUID and user-based unique block numbers) allows for the following advantages (e.g., in watermarking scenarios): audio files can be cut at any time without the cutting tool being aware of authentication processing and embedded links. Alternatively, in cases where audio and links are jointly transmitted using an audio container, file-based links can be embedded once in the included header to reduce data rate. File-based links can consist of a UUID and a location within metadata (e.g., the original file's sampling index) to indicate how the audio file in question is aligned with a reference on the server.
[0044] These and other aspects of the invention will be described in more detail with reference to the following accompanying drawings.
[0045] Figure 1 A block diagram of the audio signal signature module is shown.
[0046] Figure 2 Figure 2 A diagram showing the overall audio signature and authentication workflow is provided.
[0047] Figure 3 A block diagram of an audio signal authenticator is shown.
[0048] Figure 4 A block diagram of the audio signal signature module is shown, and
[0049] Figure 5 A block diagram of an audio signal authenticator is shown.
[0050] This invention relates to the generation and authentication of authenticable audio signals. In other words, the audio signal must be processed to include information that enables subsequent authentication. During the authentication phase, the received or stored audio signal is authenticated based on information associated with it.
[0051] Figure 1A diagram of an audio signal signature module is shown. The audio signal signature module 500 receives an input audio signal 121 from a microphone 100 or another audio source, such as a recorder, via an audio input terminal 502, and outputs an output audio signal 501 via an audio output terminal 190. The output audio signal 501 may be based on the input audio signal 121. The microphone 100 may include at least one microphone pickup head 110 and an analog-to-digital converter (ADC) 120. The at least one microphone pickup head 110 can capture an audio signal and output the captured analog audio signal 111. The captured analog audio signal 111 can be forwarded to the ADC 120, which can digitize the audio signal 111 and output a digital audio signal 121. The audio signal signature module 500 may be implemented as a device separate from the microphone 100, or may be included within the microphone 100.
[0052] The audio signal signature module 500 can receive audio signal 121 from microphone 100 or from another source, which in Figure 1 The input is indicated by selector 101. Digital audio signal 121 is input to block splitter 130, which divides the audio signal 121 into audio blocks 131 and outputs a sequence of multiple audio blocks 131. Optionally, if the audio signal 121 has already been divided into audio blocks 131, block splitter 130 can be omitted. The sequence of audio blocks 131 is forwarded to audio feature extractor 150, which generates audio features 151 for each audio block 131 in the sequence. Private key signing unit 140 receives authentication information including the audio features 151 of each audio block 131 and optional additional information 113 (e.g., metadata), such as location or time, and signs the authentication information with private key 142, thereby providing signed authentication information 141, which can be stored, for example, on memory 600 located on an external server. The signed authentication information 141 may include authentication information 152 containing audio features 151 and optional additional information 113 (see [link to authentication information]). Figure 2 ) and signature 144 generated by private key signing unit 140 ( Figure 2 Alternatively, the authentication information 141 may be stored in internal memory or any other memory, as long as it can be accessed by an authenticator that authenticates the received audio signal.
[0053] The signed authentication information 141 itself includes authentication information, namely (a) audio feature 151, optionally (b) additional information 113, and (c) signature 144. Each audio block 131 is fed in parallel, for example, to an information embedder 160 for embedding, for example, authentication access information 162 into each audio block 131. The authentication access information 162 includes information about where the authentication information 152 (or more precisely, the signed authentication information 141) of that particular audio block 131 can be accessed at, for example, memory 600 on a server, and / or how the authentication information 152 (or more precisely, the signed authentication information 141) of that particular audio block 131 can be accessed at, for example, memory 600 on a server. Embedding can be achieved, for example, by means of watermarking.
[0054] Therefore, the information embedder 160 can be implemented as a watermark generator that generates watermark 163. Another option is to embed the authentication access information 162 into a suitable audio container next to the actual audio block 131. Thus, the information embedder 160 outputs the audio block 161 with the embedded authentication access information 162. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at the audio output terminal 190. The output audio signal 501 can be transmitted, for example, via network 300 (…). Figure 2 The output audio signal 501 can be broadcast or stored. In particular, the output audio signal 501 can even be further modified to a lesser extent, including modification types such as sample rate conversion, perceptual compression, and trimming.
[0055] When the information embedder 160 is a watermark generator, its output 161 can alternatively replace the original audio block 131 and be input to the audio feature extractor 150, which in Figure 2 The selection is indicated by selector 102. This alternative offers the advantage of extracting audio features based on the final signal 161 output by the audio signal signature module 500 instead of the unoutput intermediate signal 131. Since the watermark generator is designed to perceptually alter the audio signal as little as possible, both input options (131 or 161) are reasonable for the audio feature extractor 150.
[0056] Audio features involve attempting to capture relevant aspects of human auditory perception in an audio representation, while preferably reducing the required data rate and thus the storage space required on the server or memory 600. While audio features may be the original audio signal itself, preferably, they may be a transparently coded version of the original signal that requires far less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.
[0057] The audio signal 501 at the output of the audio signal module 500 (with embedded watermarks or authentication access information included in the audio container) can be stored or transmitted. Since the watermark can be embedded in, for example, essentially each audio block 131, each audio block 131 can be authenticated individually. The watermark 103 can also be embedded only in some audio blocks within the audio block set. In addition to the authentication access information, the watermark 163 may also include additional information such as block boundaries. This information can be used to determine the block length when authenticating the audio signal.
[0058] The audio watermark 163 can be a unique identifier embedded in the audio signal and, for example, copyright information previously used to identify the audio signal. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, if the audio signal with the embedded watermark is copied, stored, or transmitted, the embedded watermark will not be altered. The same reasoning applies to authentication access information included in the audio container.
[0059] Figure 2 A diagram illustrating the overall audio signing and authentication workflow is shown. An audio signal is captured by a microphone 100 outputting a digital audio signal 121. Audio features 151 are extracted from this audio signal 121 and signed using, for example, the private key 142 of the audio signal signing module 500. The audio features 151, together with a digital signature 144 generated by the signing unit 140, form signed authentication information 141. The signed authentication information 141 can be forwarded to a server or memory 600, where it is stored for later retrieval. Information regarding where and / or how the signed authentication information 141 can be retrieved, i.e., authentication access information 162, is provided to the audio signal signing module 500. The audio signal signing module 500 uses the authentication access information 162 and embeds it into an audio block 131 of the audio signal. Therefore, the output signal 501 of the audio signal signing module 500 includes the audio block 131 with the embedded authentication access information. The audio block 131 of the audio signal with embedded authentication access information 162 is output by the audio signal signature module 500 and can be stored or distributed, for example, via network 300.
[0060] This distribution may involve modification or tampering with authentication information 152 and / or audio signal 121. The signed authentication information 141 is verified using public key 143. This verification indicates potential modifications to the audio features. The integrity of the audio signal can be verified by extracting audio features from the received signal and comparing these audio features with those contained in the signed authentication information 141. The final authentication result 261 is determined based on these two checks.
[0061] For example, an audio signal signature module 500, implemented as part of microphone 100, is used to capture audio signals. Alternatively, the audio signal signature module 500 can receive audio signals. The audio signal can also be an audio signal pre-captured by other means. See reference... Figure 1 The captured audio signal is segmented into multiple audio blocks, and the authentication information of at least one audio block is signed based on the private key 142 of the audio signal signing module 500. The signed authentication information 141 is output by the audio signal signing module 500 and can be stored on a server or memory 600. Authentication access information 162, i.e., information about where the authentication information can be retrieved from the server or memory 600, can be provided to the audio signal signing module 500. The authentication access information 162 is embedded in the audio block received or captured by the audio signal signing module 500. Therefore, the audio signal signing module 500 performs the following operations: a) outputting an output audio signal 501 including the audio block with the embedded authentication access information to the network 300 for storage or forwarding, and b) outputting the signed authentication information 141 to a server or memory 600 where the information can be stored.
[0062] Optionally, authentication access information can be embedded into the audio block of the audio signal using a watermark. The watermark can be associated with the authentication access information, which is related to the location where the authentication information is stored in memory 600. This authentication information can be referenced as described above. Figure 1 The description is generated by audio feature extractor 150. Audio blocks of the audio signal with embedded authentication access information can be transmitted or stored via network 300. Authenticator 200 can receive audio signal 501 with embedded watermark and extract authentication access information. Authentication information is retrieved from server 600 based on the authentication access information. The signature of the authentication information stored on server or storage 600 can be verified based on public key 143 associated with audio signal signature module 500, which can be stored, for example, on server 400 or a distributed ledger.
[0063] Figure 3 A block diagram of an audio signal authenticator is shown. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. This audio signal 171 may be from... Figure 1The audio signal 501 is output from the audio signal signature module 500. This audio signal 501 is input to a block boundary detector and a block segmenter 205, which determine block boundaries and output multiple audio blocks 201 with embedded authentication access information 212. If the authentication access information 212 has already been embedded using watermarking, it is assumed that the block boundaries were previously encoded into the watermark, and the block boundaries can be detected from the watermark at this stage. In the case where the audio signal is transmitted by an audio container containing the authentication access information 212 as metadata, it is assumed that the authentication access information is available for each block, where the block boundaries are part of the metadata.
[0064] An audio block 201 with embedded authentication access information is input to an access information extractor 210, which extracts authentication access information 212 from the audio block 201 (which, if no tampering has occurred, corresponds to authentication access information 162 from the signature module). An authentication access information parser 230 determines the access information based on the authentication access information 212, i.e., where on the server or memory 600 the signed authentication information 141 belonging to the current audio block 201 is found. The authentication information 141 retrieved from the server or memory 600 first includes audio features 151 representing the audio block 201, secondly includes a signature 221, and thirdly includes optional additional information 113 such as the location and / or time of the signature occurrence. If no tampering has occurred, the signature 221 corresponds to the signature 144 generated by the signature unit 140.
[0065] In the perceptual similarity analyzer 240, audio blocks 201 and audio features 151 are perceptually compared with each other to provide a similarity statement 241.
[0066] In the similarity analyzer 240, audio block 201 can be compared with audio features 151 of the original audio signal. The audio features can be a true copy of the original audio signal. However, this would require a large amount of storage space to store the original audio signal. If the audio features relate to other features of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Therefore, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, in the similarity analyzer 240. In other words, the similarity analyzer 240 should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.
[0067] In its simplest case, this can be a binary decision about whether the quantities being compared are perceived to be sufficiently equal. Alternatively, the similarity statement 241 can be more detailed, for example, by providing a similarity score.
[0068] The authentication access information parser 230 also determines, based on the authentication access information 212, which owner of the private key 142 initially signed the audio signal 171, and this owner is represented, for example, by an ID. To name just a few possibilities, the owner of the private key 142 could be a natural person, an organization, or a registered microphone. Using this ID, the identity 270 of the owner of the public key 143 and the private key can be retrieved from an external database. Subsequently, the public key 143 is used to verify the combination of the signature 221 and the authentication information 152 (audio feature 151 and additional information 113) within the signature verifier 250, thereby providing a signature verification result 251. Finally, in the combiner 260, the perceptual similarity statement 241 and the verification result 251 are combined into a final authentication result 261.
[0069] Note that the processing proposed here for audio content can be similarly applied to any time-related data. In particular, this processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark, or uses a video container instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison.
[0070] Figure 4 A schematic block diagram of an audio signal signing module is shown. The audio signal signing module 500 receives a digital audio signal 121 at its audio input 502. This digital audio signal 121 is input to a block segmenter 130, which outputs a sequence of multiple audio blocks 131. The audio blocks 131 are forwarded to an audio feature extractor 150, which generates an audio feature 151 for each audio block 131. In a private key signing unit 140, the user's or entity's private key 142 is used to sign the audio feature 151 of each audio block 131, along with optional additional information 113 such as location or time authentication information, thereby providing signed authentication information 141. The signed authentication information 141 itself consists of the audio feature 151, the generated signature 144, and optional additional information 113. Figure 5 The authentication method shown requires two components during verification: individual audio blocks 131 and signed authentication information 141. In principle, they can be embedded in the same suitable audio container, but they can also be distributed across different physical channels. Optionally, a user or entity identifier 145 can be added to these two components, which can, for example, enable the retrieval of the appropriate public key for automated verification.
[0071] Figure 5A schematic block diagram of authenticator 200 is shown. Authenticator 200 receives, for example, an audio signal in the form of an audio block 201 to be verified, and signed authentication information 141. The signed authentication information 141 includes audio features 151, additional information 113, and a signature 221. Within a perceptual similarity analyzer 240, the audio features 151 contained in the authentication information 141 are perceptually compared with the audio block 201, thereby providing a similarity statement 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity statement 241 can be more detailed, for example, providing a similarity score. Furthermore, using knowledge about the user or entity that created the signature 221, the corresponding public key 143 and optionally the user identity 270 can be obtained. Optionally, this knowledge can be derived from a user or entity identifier 145 provided along with the authentication information 141 and the audio block 201. Subsequently, public key 143 is used to verify the combination of signature 221 and authentication information 141 within signature verification block 250, thereby providing signature verification result 251. Finally, in combiner 260, perceived similarity statement 241 and verification result 251 are combined into final authentication result 261.
[0072] According to the example, each audio signal captured by microphone 100 and output by the microphone may include a watermark, which can be used to identify the actual microphone that captured the audio signal. If the microphone is used by several users, several user IDs can be associated with the microphone. If the microphone is registered, the corresponding microphone can be identified. User IDs can also be registered. The user ID can be a unique ID or a sub-ID of the microphone ID. The microphone ID and / or user ID can be part of the watermark. Using the sequence number of the audio block, it can be determined when a portion of the audio signal has been removed.
[0073] Alternatively, the audio signal may be part of a video file.
[0074] Alternatively, the signed authentication information can be included in the audio file (e.g., in an ADM file format) instead of in the watermark. In this case, the audio signal does not need to be modified.
[0075] Optionally, the processing to generate the signed audio signal can be performed in a software solution based on a pre-existing audio recording or audio stream. The private key can be entered into the software, for example, via a dongle or text input, and can be authorized, for example, via biometric signals such as fingerprints or facial detection.
[0076] The authentication process can calculate a similarity score between the audio features of the signed audio signal and the analyzed signal. This allows determining the likelihood that the signal is still authentic even if minute signal processing, such as gain adjustments, has occurred.
[0077] Authentication requires an audio signal and a signed authentication message. The signed authentication message can be provided via a side channel. When using a watermark, the authentication access information must be read from the audio signal, and the authentication information must be retrieved accordingly. Where watermark readout requires block boundaries, one method to achieve this is to try different offsets of the block boundaries until the watermark can be successfully read. Another method is to embed some synchronization signal with the watermark into the audio signal.
[0078] To enhance system security, processing is provided that allows for key revocation, for example, in the event of key theft.
[0079] Public keys used for authentication can be provided in various ways. One option is to store all public keys in a centralized database to make it easy to find the required key. To address potential trust issues in this centralized instance, a database with all public keys can be provided based on distributed ledger technology. Another option is for organizations / individuals using this technology to provide public keys on their own websites.
[0080] Torsten Dau, Dirk Pueschel, and Armin Kohlrausch describe another example of audio features in: A quantitative model of the “effective” signal processing in the auditory system. I. model structure, The Journal of the Acoustical Society of America, 99(6):3615-3622, 1996, hereinafter referred to as the Perceptual Model (PEMO). This paper describes a quantitative model designed to describe how the auditory system processes acoustic signals. The focus is on developing a model that mimics the functional processing of auditory stimuli in a way that mimics human perception, particularly in the context of complex sounds such as speech or music.
[0081] The key components of this model can be categorized into peripheral processing, envelope extraction, modulation filtering, and a decision-making stage. The peripheral processing section captures the initial stages of auditory processing, including converting acoustic signals into neural representations through mechanisms such as outer and middle ear filtering and nonlinearities in the cochlea. The envelope extraction section involves extracting the temporal envelope of sound, which is crucial for understanding amplitude modulation and speech processing. The modulation filtering section involves a system for filtering amplitude modulation, simulating how the auditory system is sensitive to different modulation frequencies. In the decision-making stage, the processed auditory signals are used to make decisions about the properties of the sound, representing a higher level of auditory perception.
[0082] The model was designed to align with experimental psychoacoustic data, particularly in its ability to simulate the auditory system's response to amplitude-modulated sounds. This method provides a framework for understanding effective signal processing in human auditory perception, especially for tasks involving the detection and discrimination of complex auditory patterns.
[0083] Performing a comparison of the received audio block with the extracted or retrieved authentication information in the transform domain (i.e., with respect to the audio features perceived by human hearing) ensures that the identified differences are perceptually meaningful. In principle, any similarity metric (e.g., empirical cross-correlation coefficient or relative error relative to a reference) can be used to map the comparison results to a numerical value.
[0084] When a further transformation is applied to the Mel frequency power spectrum of an audio signal, Mel frequency cepstral coefficients (MFCCs) are obtained. MFCCs use the Mel scale, a perceptual scale for pitch, designed to approximate how humans perceive sound. The Mel scale is non-linear, allowing for finer division of low frequencies and coarser division of high frequencies, thus more closely resembling human auditory perception.
[0085] Mel-frequency cepstral coefficients use the cepstral spectrum as a transform, which converts the signal from the frequency domain back to a domain where the rate of change of the signal can be analyzed. The idea is to capture the spectral characteristics of the signal in a way that emphasizes the information-carrying components.
[0086] The process of calculating MFCC from an audio signal involves several steps: pre-emphasis: typically by applying a high-pass filter to enhance the high-frequency components of the signal to balance the spectrum; framing: dividing the audio signal into short, overlapping frames, typically 20 to 40 milliseconds long, because speech signals are non-stationary, but they can be considered quasi-stationary within these short frames; windowing: multiplying each frame by a window function, such as a Hamming window, to reduce edge effects and smooth the signal; Fourier transform: applying a Fast Fourier Transform (FFT) to each windowed frame to convert the time-domain signal to the frequency domain; Mel filter bank: the resulting spectrum is passed through a series of triangular bandpass filters spaced according to the Mel scale, a step that approximates how humans perceive sound frequencies; logarithmic calculation: calculating the logarithm of the power of each Mel-filtered spectrum, which simulates the human ear's response to loudness; Discrete cosine transform (DCT): finally, the logarithmic Mel spectrum is transformed using a DCT. The result is a set of coefficients called MFCC. Typically, only the first 12 to 13 coefficients are retained, as these contain the most important information.
[0087] MFCC can capture the wide spectral shape of an audio signal in a compact form, making it very useful for machine learning algorithms in speech and audio processing. MFCC helps distinguish different phonemes, speaker identity, and other audio features by providing robust representations of variations in pitch, volume, and other factors.
[0088] Audio features may also include:
[0089] Linear predictive coding (LPC) is a method that uses information from a linear predictive model to represent the spectral envelope of a digital speech signal in compressed form. This method estimates the parameters of filters that can be used to reconstruct the signal. LPC is very effective for modeling the formants (resonant frequencies) of speech sounds.
[0090] - Perceptual Linear Prediction (PLP) coefficients are similar to linear predictive coding, but incorporate aspects of human auditory perception, such as critical band spectral resolution, equal-loudness curves, and the intensity-loudness power law. Perceptual Linear Prediction (PLP) coefficients are designed to mimic the non-linear perception of loudness and frequency by the human ear.
[0091] Gammatone filter bank features involve filter banks that simulate the filtering that occurs in the human cochlea. They are similar to Mel filter banks, but use gammatone filters instead of triangular filters. Gammatone filter bank features are useful for capturing the detailed frequency structure of audio, especially in tasks involving environmental sound classification or hearing aid design.
[0092] - Chromaticity features, which represent 12 different pitch categories (C, C#, D, etc.) of a musical octave. They capture the harmonic and melodic characteristics of music. These are particularly useful in music information retrieval, tonality detection, and chord recognition tasks.
[0093] - Mel spectrogram: While similar to MFCC, the Mel spectrogram is the result of applying a Mel filter bank directly to the power spectrogram without further transformations (such as DCT). The Mel spectrogram preserves more detailed frequency information and is commonly used as input to deep learning models, often in conjunction with convolutional neural networks (CNNs) in tasks such as audio event detection, speech recognition, and music genre classification.
[0094] - The Constant Q Transform (CQT) provides a time-frequency representation with a logarithmic frequency scale similar to the Mel scale, but with a variable time resolution that matches the frequency resolution. It is particularly useful for music applications because it provides a better representation of musical pitch than the linear FFT or even the Mel scale.
[0095] - Deep learning-based features: Learned features from deep learning models, such as embeddings from trained neural networks, can also be used as alternatives to traditional handcrafted features such as MFCCs. These features are generally more robust and can capture complex patterns that are difficult to model using traditional methods.
[0096] - Spectral Subband Centroid (SSC) is the centroid of the energy distribution in the different subbands of the captured signal. This centroid can be interpreted as the "centroid" of the spectrum within each subband. It provides information about the energy distribution across frequency bands and is sometimes used as a supplement to MFCC.
[0097] - Relative Spectrum (RASTA) features involve filtering the logarithmic energy of the speech signal to highlight modulation frequencies important for speech recognition. RASTA-PLP is a combination of RASTA filtering and PLP analysis, providing robust features against noise and channel variations.
[0098] Audio features involve the audio representation of an audio signal, attempting to capture relevant aspects of human auditory perception of the audio signal. Audio features should be chosen to reduce the required data rate (for transmitting the audio features) and thus reduce the storage space required on the server. While the audio feature may be the original audio signal itself, preferably, it can be a transparently coded version of the original signal that requires far less storage space. Another example of an audio feature could be the output of any model that mimics human auditory perception. Audio features could be the Mel-frequency power spectrum and Mel-frequency cepstral coefficients (MFCCs) of the audio signal.
[0099] In particular, audio features can be acoustic features, spectral features, and statistical features of the audio signal.
[0100] Acoustic features may include: pitch (fundamental frequency); prosody (rhythm, stress, intonation) and / or duration, and silence patterns (pauses, speech timing, or inconsistencies in breath sounds can be used as audio features).
[0101] Spectral features may include: formant frequencies (such as resonant frequencies in speech); Mel frequency cepstral coefficients (MFCCs) (the short-time power spectrum of sound, which can reveal artifacts introduced during synthesis or tampering); spectrogram analysis (tampering may manifest as unusual or obscured energy distributions in the time-frequency representation); and / or phase information.
[0102] Acoustic features may include: statistical and signal processing features, such as noise and residuals (e.g., subtle background noise or artifacts introduced during synthesis may not match natural recordings); high-frequency content; and / or phase distortion (some tampering may introduce phase anomalies, which can be detected using signal processing techniques).
[0103] Acoustic features may include: temporal features, such as jitter and flicker (e.g., changes in frequency and amplitude); and temporal coherence (e.g., sudden shifts or changes in speech features, such as unnatural interruptions in the signal).
[0104] Acoustic features may include: behavioral or semantic inconsistencies, such as content coherence (logical inconsistencies or unnatural speech flow in spoken content may indicate tampering); affect and naturalness (emotional tone may not match the content, or the voice may lack the natural nuances of human emotion).
[0105] Audio features can be any one of the audio features mentioned above or a combination thereof.
[0106] By combining these features with machine learning or signal processing techniques, models can be trained to detect artifacts and inconsistencies that indicate depth-spoofed audio.
[0107] List of reference numerals in the attached diagram:
[0108] 100 microphones
[0109] 101 Selector
[0110] 102 Selector
[0111] 110 microphone pickup head
[0112] 111 Microphone simulates audio signal
[0113] 113 Metadata
[0114] 120 AD converter
[0115] 121 digital audio signal
[0116] 130-block divider
[0117] 131 audio blocks
[0118] 140 Private Key Signature Unit
[0119] 141 Signed authentication information
[0120] 142 Private Key
[0121] 143 Public Key
[0122] 144 signatures
[0123] 145 Identifier
[0124] 150 Audio Feature Extractors
[0125] 151 Audio Features
[0126] 152 Authentication Information
[0127] 160 Information Embedder / Watermark Generator
[0128] 161 Audio blocks with embedded information
[0129] 162 Authentication Access Information
[0130] 163 Watermark
[0131] 171 audio signal
[0132] 190 audio output terminal
[0133] 200 Audio Signal Authenticator
[0134] 205 Block Boundary Detector and Block Segmenter
[0135] 210 Access Information Extractor
[0136] 212 Authentication Access Information
[0137] 221 signatures
[0138] 230 Authentication Access Message Parser
[0139] 240 Perceptual Similarity Analyzer
[0140] 241 Similarity Results
[0141] 250 signature verifier
[0142] 251 Signature verification result
[0143] 260 combiner
[0144] 261 Authentication Result
[0145] 270 Identity
[0146] 300 Network
[0147] 400 server
[0148] 500 audio signal signature module
[0149] 501 Certified Audio Signal
[0150] 502 Audio Input Terminal
[0151] 600 Server or Storage
Claims
1. An audio signal signature module (500), comprising: The input terminal (502) is configured to receive digital audio signals (121). A block divider (130) is configured to divide the received digital audio signal (121) into a sequence of multiple audio blocks (131). An audio feature extractor (150) is configured to extract audio features (151) from the current audio block (131). The signing unit (140) is configured to generate a signature (144) associated with the current audio block (131) by applying a private key (142) to the audio feature (151), and is configured to provide signed authentication information (141) including the extracted audio feature (151) and the signature (144), wherein the signed authentication information (141) is output to a memory (600). An information embedder (160) is configured to generate an audio block (131) with embedded information (161) by embedding authentication access information (162) into the current audio block (131), wherein the authentication access information (162) includes information about how the signed authentication information (141) can be retrieved from the memory (600) and / or where the signed authentication information (141) can be retrieved from the memory (600), and An audio output terminal (190) is configured to output an authenticable audio signal (501) comprising a sequence of audio blocks having embedded authentication access information (161).
2. The audio signal signature module (500) according to claim 1, wherein, The information embedder (160) is implemented as a watermark generator (160), which is configured to introduce the authentication access information (162) as a watermark (163) into the current audio block (131), such that the authenticateable audio signal (501) includes a sequence of audio blocks with the embedded watermark (163).
3. The audio signal signature module (500) according to claim 1 or 2, wherein, The signed authentication information (141) also includes metadata (113) belonging to the current audio block (131). The metadata (113) includes at least one of the following: the identifier (145) of the user or entity that has signed the audio signal, the location, and the time.
4. A microphone (100), comprising: Microphone pickup head (110) is used to capture audio signals (111). An AD converter (120) is configured to convert the captured audio signal (111) into a digital audio signal (121), and The audio signal signature module (500) according to claim 1, 2 or 3.
5. A method for generating an authenticable audio signal (501), comprising the following steps: Receive digital audio signals (121). The digital audio signal (121) is divided into a sequence of multiple audio blocks (131). Extract audio features (151) from the current audio block (131). A signature (144) associated with the current audio block (131) is generated by applying the private key (142) to the extracted audio features (151). Provide signed authentication information (141) including the audio feature (151) and the signature (144). The signed authentication information (141) is output to the memory (600). An audio block (131) with embedded information (162) is generated by embedding authentication access information (162) into the current audio block (131), wherein the authentication access information (162) includes information about how the signed authentication information (141) can be retrieved from the memory (600) and / or where the signed authentication information (141) can be retrieved from the memory (600), and The authenticated audio signal (501) is output, which includes a sequence of multiple audio blocks with embedded authentication access information (161).
6. An audio signal authenticator (200), comprising: A block boundary detector and a block segmenter (205) are configured to receive an audio signal (171) and output a sequence of multiple audio blocks (201) with embedded authentication access information (212, 162). Access information extractor (210) is configured to extract the authentication access information (212, 162) from the current audio block (201). An authentication access information parser (230) is configured to extract access information about how the signed authentication information (141) belonging to the current audio block (201) can be retrieved from the memory (600) and / or where the signed authentication information (141) belonging to the current audio block (201) can be retrieved from the memory (600). The signed authentication information (141) retrieved from the memory (600) includes a signature (144, 221) and audio features (151) belonging to the current audio block (201), and A perceptual similarity analyzer (240) is configured to compare the audio block (201) of the received audio signal with the retrieved audio features (151), or to compare the audio features extracted from the audio block (201) of the received audio signal with the retrieved audio features (151) to provide a similarity result (241).
7. The audio signal authenticator (200) according to claim 6 further includes: The signature verifier (250) is configured to verify the combination of the signature (144, 221) and the audio feature (151) using a public key (143) and provide a signature verification result (251). The public key (143) is associated with the private key (142), which can be used by the audio signal signature module (500) to generate an authentic audio signal (501).
8. The audio signal authenticator (200) according to claim 7, wherein, The authentication access information parser (230) is configured to determine the identifier (145) of the user or entity that initially signed the audio signal. Among them, the public key (143) and the identity (270) of the user or entity can be retrieved from an external database (400) based on the determined identifier (145).
9. The audio signal authenticator (200) according to claim 7 or 8, further comprising: The combiner (260) is configured to combine the similarity result (241) and the signature verification result (251) into a final authentication result (261).
10. A method for authenticating an audio signal (171) by an authenticator (200), comprising the following steps: Receive audio signal (171). Output a sequence of multiple audio blocks (201) with embedded authentication access information (162). Extract the authentication access information (212, 162) from the current audio block (201). Extract access information (212, 162) about how to retrieve the signed authentication information (141) belonging to the current audio block (201) from the memory (600) and / or where to retrieve the signed authentication information (141) belonging to the current audio block (201) from the memory (600). Retrieve signed authentication information (141) from the memory (600), the signed authentication information (141) including a signature (144, 221) and an audio feature (151) belonging to the current audio block (201), and The audio blocks (201) of the received audio signal and the retrieved audio features (151) are compared with each other to provide a similarity result (241), or The audio features extracted from the audio block (201) of the received audio signal (171) and the retrieved audio features (151) are compared with each other to provide similarity results (241).
11. The method of claim 10, further comprising the step of: The combination of the signature (144, 221) and the audio feature (151) is verified using the public key (143), and Provide a signature verification result (251), wherein the public key (143) is associated with a private key (142), which can be used by the audio signal signature module (500) to generate an authentic audio signal (501).
12. A computer program product for generating authenticated audio signals, in, The computer program product includes a program code device that causes the audio signal signature module according to claim 1 to perform the method according to claim 5.
13. A computer program product for authenticating audio signals. in, The computer program product includes a program code device that causes the audio signal authenticator according to claim 6 to perform the method according to claim 10.