Audio signal signature module and method of generating authenticable audio signal

By segmenting and signing audio signals using an audio signal signature module, embedding authentication access information, and verifying with a public key, the complexity of audio signal authentication in existing technologies is solved, achieving simple and robust audio signal authentication and detection.

CN121889798APending Publication Date: 2026-04-17SENNHEISER ELECTRONICS GMBH & CO KG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SENNHEISER ELECTRONICS GMBH & CO KG
Filing Date
2024-09-23
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies make it difficult to authenticate audio signals simply and robustly, especially for audio signals such as politicians' speeches, which usually requires a lot of manual research.

Method used

The audio signal is segmented into multiple audio blocks by an audio signal signature module, audio features are extracted and a signature is generated, authentication access information is embedded, and authentication is performed using private key signing and public key verification. Watermarking technology is used to embed authentication information into the audio signal to achieve authentication.

Benefits of technology

It achieves simple and robust authentication of audio signals, capable of detecting the authenticity and integrity of audio signals, preventing malicious modification, and reducing storage and transmission bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121889798A_ABST
    Figure CN121889798A_ABST
Patent Text Reader

Abstract

There is provided an audio signal signature module (500) comprising: an input (502) configured to receive a digital audio signal (121); and a block divider (130) configured to divide the received digital audio signal (121) into a sequence of audio blocks (131) having a block length. A signature unit (140) generates a signature (144) associated with the current audio block (131) by applying a private key (142) to the extracted audio feature (151). The combined information is provided as signed authentication information (141) and stored in a predetermined storage location in a memory (600). The predetermined storage location may be used as authentication access information (162) and a current audio block (131) embedded by an information embedder. An audio output (190) provides an authenticable audio signal (501) comprising a sequence of audio blocks with embedded authentication access information (162). A block divider (130) adjusts the block length of the audio block in accordance with the size of authentication access information to be embedded in the audio block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention relates to an audio signal signature module, a microphone and a method for generating authenticable audio signals, an audio signal authenticator, and a method for authenticating audio signals.

[0002] Given the modern ability to create deepfake videos and audio, there is an increasing need to verify and authenticate audio signals such as those used in politicians' speeches.

[0003] Authenticating human audio signals has been extremely difficult until now. Typically, extensive manual research is required to verify or authenticate such audio signals and to verify the source of the audio signal.

[0004] Therefore, the object of the present invention is to provide an apparatus capable of generating authenticable audio signals and an apparatus for authenticating such audio signals, which enables the audio signals to be authenticated in a simple and robust manner.

[0005] This objective is achieved by the audio signal signature module according to claim 1, the microphone according to claim 5, the method for generating an authenticable audio signal according to claim 6, the audio signal authenticator according to claim 8, the method for authenticating an audio signal according to claim 12, and the computer program product according to claims 14 and 15.

[0006] This invention relates to generating an authenticable audio signal using an audio signal signature module or a method for generating an authenticable audio signal, as well as authenticating an audio signal using an authenticator or a method for authenticating an audio signal. The audio signal must be processed by the audio signal signature module to include information that enables subsequent authentication. During the authentication phase, the authenticator authenticates the received or stored audio signal based on information associated with the audio signal.

[0007] Therefore, an audio signal signature module is provided, comprising: an input terminal for receiving a digital signal; a block segmenter for segmenting the received digital audio signal into a sequence of multiple audio blocks having a block length and a block-specific payload capacity; an audio feature extractor for extracting audio features from the current audio block; and a signature unit for generating a signature associated with the current audio block by applying a private key to the audio features. The signature unit also provides signed authentication information including the extracted audio features and the signature. The audio signal signature module further includes an information embedder for generating audio blocks with embedded information by embedding authentication access information into the current audio block. Furthermore, an audio output is provided for outputting an authenticated audio signal comprising a sequence of audio blocks with embedded authentication access information. The block segmenter adjusts the block length of the audio block according to the size of the authentication access information to be embedded in the audio block, so that the authentication access information can be accommodated in the audio block without producing (significant) audible artifacts in the audio signal. Therefore, the block length of the audio block can vary according to the requirements of the information to be embedded. This also allows for easy adjustment of block lengths, taking into account a trade-off between shorter block lengths for finer-grained time-based verification and longer block lengths for embedding authentication access information.

[0008] According to one aspect, the information embedder is implemented as a watermark generator to introduce authentication access information as a watermark into the current audio block, such that the authenticable audio signal includes a sequence of audio blocks with the embedded watermark.

[0009] According to one aspect, at least one watermark has a specific size, and the block generator adjusts the block length of the audio block based on the specific size of at least one watermark.

[0010] On one hand, the block segmenter analyzes the audio characteristics of the audio blocks to determine the block-specific payload capacity for each audio block, and adjusts the length of the audio blocks so that the block-specific payload capacity exceeds the size of the watermark to be embedded.

[0011] The amount of information that can be embedded into an audio block using a watermark (without producing significant acoustic artifacts) is referred to here as the block-specific (watermark) payload capacity, which depends on the signal characteristics of the audio block. To embed a watermark of a given size into an audio block, the block length can be adjusted. For example, if the block-specific payload capacity of the audio block is smaller than the size of the watermark to be embedded, the block length can be appropriately increased, resulting in an increase in the required block-specific payload capacity.

[0012] It should also be noted that block-specific payload capacity depends on the specified robustness of the watermark to manipulation (i.e., the processing steps the signal will undergo after watermarking, such as perceptual coding at a specific bit rate). Typically, the need for greater robustness is closely related to a reduction in block-specific payload capacity.

[0013] For watermarking to be a viable option for authenticating audio signals, the data rate of the information to be embedded should be at least partially lower than the watermark payload capacity rate, which is the average amount of information that can be inserted into the audio signal per unit time while maintaining audibility. Watermarking typically utilizes effects of human perception, such as time masking and frequency masking, to maintain imperceptibility. In particular, the actual amount of information that can be embedded per unit time is highly dependent on short-time signal characteristics and is therefore typically highly time-varying.

[0014] Assuming the authentication method mentioned above uses a consistent block length that varies over time, a problem arises: for a single block, the data size of the authentication access information may be too large compared to the data that can be embedded in the watermark, making the watermarking scheme infeasible. Therefore, to remedy this, the block length can be adjusted according to the watermarking requirements. The block length of a single audio block or a group of audio blocks can be adjusted according to the watermarking requirements. A block segmenter can analyze audio features to determine how much watermark payload can be incorporated and adjust the block length accordingly.

[0015] There is a trade-off between the idea of ​​using short blocks (approximately a few seconds) in the authentication process and the need to include watermarks in the audio blocks.

[0016] The block length can be adjusted by selecting a minimum predefined block length before analyzing the signal characteristics of the audio block to determine if embedding the required information (e.g., a link to block-specific authentication information) is possible. If not, the block length of the audio block under consideration is increased, and the embedding feasibility test is performed again. These two steps are repeated until embedding is feasible. In this case, the actual embedding is completed, and the next audio block is processed. Information about the final block length used must be (possibly encoded) pre-appended to the embedded data to enable correct watermark decoding and thus correctly reconstruct the block length used for authentication.

[0017] According to one aspect, the signed authentication information also includes metadata belonging to the current block. The metadata may include at least one of the following: an identifier of the user or entity that has signed the audio signal, and in particular, the location and time of signing the audio signal or generating the audio signal.

[0018] The present invention also relates to a microphone, which includes a microphone capsule suitable for capturing audio signals and an audio signal signature module as described above.

[0019] This invention also relates to a method for generating an authenticated audio signal. A digital audio signal is received and segmented into a sequence of multiple audio blocks by a block segmenter, each audio block having a block length. Audio features are extracted from the current audio block, and a signature associated with the current audio block is generated by applying a private key to the extracted audio features. Signed authentication information, including the audio features and the signature, is provided. Audio blocks with embedded authentication access information are generated by embedding authentication access information into the current block. An authenticated audio signal is output. The authenticated audio signal comprises a sequence of multiple audio blocks with embedded authentication access information. The block segmenter adjusts the block length of the audio blocks according to the size of the authentication access information to be embedded in the audio blocks.

[0020] According to an advantageous implementation, the method includes analyzing the payload capacity of the audio block to determine whether the payload capacity is sufficient to embed authentication access information, and increasing the block length only when the payload capacity is insufficient to embed the authentication access information. In many cases, the pre-selected block length proves sufficient, and no further steps are required.

[0021] The present invention also relates to an audio signal authenticator, comprising a block boundary detector and a block segmenter for receiving audio signals, detecting block boundaries or block lengths of audio blocks, and outputting a sequence of multiple audio blocks with embedded authentication access information. An authentication access information extractor is provided, configured to extract authentication access information from the audio blocks. The extracted authentication access information includes a signature and audio features belonging to the current audio block. Furthermore, a perceptual similarity analyzer is provided for comparing audio blocks and audio features with each other and providing similarity results.

[0022] One approach provides a signature verifier to verify a combination of signature and audio features using a public key and to provide the signature verification result. The public key is associated with a private key, which can be used by the audio signal signing module to generate an authenticable audio signal.

[0023] On one hand, a combiner is provided to combine similarity results and signature verification results into a final authentication result.

[0024] In a similarity analyzer, audio blocks can be compared to the audio features of the original audio signal. These audio features can be true copies of the original audio signal. However, this would require significant storage space to store the original audio signal. If the audio features are related to other features of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Therefore, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within the similarity analyzer. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.

[0025] Audio features involve attempting to capture an audio representation of relevant aspects of human auditory perception, while preferably reducing the required data rate and thus the storage space required on the server. While it could be the original audio signal itself, preferably, it could be a transparently encoded version of the original signal that requires far less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.

[0026] According to one aspect, a method for authenticating an audio signal via an authenticator is provided, the method comprising the steps of: receiving an audio signal; outputting a plurality of audio blocks having embedded authentication access information; extracting authentication access information from the audio blocks, the authentication access information indicating where authentication information belonging to the current audio block can be retrieved; retrieving authentication information from a memory, the authentication information including a signature, audio features representing the audio blocks and / or other information; and perceptually comparing the audio blocks and audio features with each other to provide similarity results.

[0027] The authentication information is signed using the private key of a microphone or a device that performs audio signal capture or post-processing to generate a unique signature indicating the source of the authentication information. The private key can be associated with the user of the signing module (e.g., the microphone). On the decoder side, the public key can be used to decode or verify the authentication information. In other words, the signature is used to verify the source of the authentication information.

[0028] Therefore, a signature can be used to verify that the authentication information originates from a source who is the owner of the private key. By further comparing the audio signal with the authentication information, it can be verified that the audio signal has not been manipulated beyond permissible levels. Only a signature is generated based on, for example, the private key of a microphone.

[0029] It is advantageous to store authentication information, such as audio features, externally rather than within the audio block, because the required bit rate and bandwidth do not increase. Authentication access information is transmitted only along with the audio block. If authentication of the audio signal is required, the authentication access information is extracted, and the associated authentication information, such as audio features, is retrieved.

[0030] Therefore, authentication information is provided from audio blocks independent of the captured audio signal. If authentication is not performed, no authentication information is needed, and significant bandwidth and bitrate savings can be achieved compared to appending authentication information to the audio signal in a suitable container format. Furthermore, when authentication is required, access information can be extracted from the audio blocks, and authentication information can be retrieved. Since the authentication process is not latency-critical, extracting authentication information from an external source before performing authentication is sufficient.

[0031] Authentication access information can be embedded as a watermark into an audio block, or as part of an audio block into a dedicated container, or as part of an audio file into a file-based environment.

[0032] Authentication information and authentication access information can be provided for at least one audio block. Therefore, each audio block can be authenticated individually. This allows for exhaustive authentication. Thus, even minor tampering with the audio signal (where authentication information is available) can be detected. This significantly improves the security of the captured audio signal. Providing authentication information for each audio block is advantageous because it even allows verification of audio signals comprising different audio blocks from different audio sources. Any audio signal including different audio blocks can be verified, provided the necessary authentication information is available.

[0033] It can extract audio features from blocks of digital audio signals and generate signatures based on these audio features using the private keys of individuals or entities. It can output the captured audio signal and the signed authentication information.

[0034] Therefore, it is possible to verify the combination of the received authentication information and signature using a public key and then determine whether the transmitted audio signal was captured and signed by a specific microphone or a specific software entity with a specific private key by comparing the audio signal with the audio features included in the authentication information.

[0035] According to the example, a microphone is provided, which includes a microphone capsule and an audio signal signature module as described above.

[0036] You can use a private key and a public key pair to sign audio blocks and verify the signatures.

[0037] According to one aspect of the present invention, in addition to audio characteristics, the authentication information may also include microphone identification, microphone location / position information, recording time and data and / or microphone model, serial number, etc.

[0038] Authentication information comprises information that can be used to authenticate received or selected audio signals. This authentication information enables the authenticator to perform authentication of the audio signals. Authentication information may include information related to the audio signal or its characteristics (i.e., audio features). Authentication information may also include metadata related to information independent of the actual audio signal. This information may be the user's ID, the ID of the microphone used to capture the audio signal, the microphone model, date, time, location or position (GPS positioning), and / or the serial number of the audio block.

[0039] The authenticity check of the received audio signal can be performed by a decoder, which can be implemented, for example, on a cloud service, computer, tablet, or smart device.

[0040] A block splitter can be used to divide an audio signal into multiple audio blocks. The length of an audio block can, for example, be between 0.5 s and 20 s. In particular, it is also possible for the block length to vary over time, which may be indicated, for example, by a watermark. A digital watermark can be embedded in each audio block. Alternatively, the watermark can be embedded in every nth audio block.

[0041] The sequence of audio blocks in the audio signal can be used to associate a sequence number with each audio block. This is advantageous because later, during the authentication of the audio signal, it can be determined whether an audio block has been removed from the audio block chain by checking the sequence number of the audio block. This can also detect the order in which changes were made to the original blocks.

[0042] Optionally, the metadata of the audio block can include the microphone user, date, time, GPS location, etc. This increases the probability of reliable authentication results.

[0043] Authentication information may include audio characteristics representing audio signals or audio blocks of audio signals based on human perception, or a transparently encoded version of the original audio signal, to reduce the storage space required by the server.

[0044] Authentication of an audio signal can be performed by detecting block boundaries from specific information in the watermark or audio container and extracting block-specific authentication information (e.g., a link containing a UUID and a user-based unique block number) from the audio watermark or directly from the audio container. The user's public key (e.g., identified by the UUID from the link) can be used to verify the authenticity of the authentication information. The audio signal block is compared with the audio features contained in the authentication information to verify the authenticity of the authentication information.

[0045] Using the above-described method for authenticating audio signals, malicious modification of the audio signal can be detected by comparing the audio block with the authentication information associated with the audio block. Modification of the audio signal to be authenticated and / or modification of authentication information stored externally can also be detected by examining the signature of the authentication information. If the audio block of the audio signal is modified and the authentication access information is also changed, this will be noticed by the authentication process because the signature of the authentication information is that of another person or device.

[0046] According to one perspective, if the authentication process does not provide an explicit indication of the authenticity of an audio signal or audio block, the authentication information stored externally can be used to manually authenticate the audio signal or audio block.

[0047] According to the example, the actual audio features extracted from the audio block and optional additional metadata are stored in a signed manner, for example, on a server. A unique (explicit) link to the metadata stored on the server is embedded in a watermark or in a suitable audio container. The link may include a unique user identifier (UUID) combined with the user-based unique block number of the corresponding audio block. The audio features involve an audio representation of the audio signal that attempts to capture relevant aspects for human auditory perception, while preferably reducing the required data rate and thus the storage space required on the server. While it may be the original audio signal itself, preferably, it can be a transparently encoded version of the original signal that requires much less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.

[0048] To authenticate the audio signal of an audio block, block boundaries can be detected based on a watermark or specific information in the audio container. Then, a link containing block-specific signature authentication information (e.g., a UUID and a user-based unique block number) is extracted from the audio watermark or directly from the audio container. The combination of authentication information and signature is verified using the user's public key (e.g., identified by the UUID from the link). Accordingly, it is verified whether the authentication information has been signed by the claimed audio signal source. The audio signal block is compared with audio features to verify perceptual similarity, and thus the authenticity of the audio signal block. For perceptual comparison, audio features can be determined from the received audio signal block, which can be directly compared with audio features from, for example, signature authentication information stored on a server.

[0049] The proposed authentication method prevents any malicious modification to the actual audio signal by comparing the audio signal block with signed authentication information. Authentication may fail if the comparison allows a limited number of modifications and the allowed modification threshold is exceeded. Furthermore, if the authentication information on the server is altered accordingly along with the malicious modification of the audio signal block, this can be detected by a mismatch between the authentication information and the signature associated with it. Additionally, if a link (within a watermark or audio container) is altered along with the malicious modification of the audio signal block to reference a different (but appropriate) signed authentication information on the server, signature verification and subsequent comparison may indeed succeed, but the signature will be from a different user, assuming the attacker cannot access the original user's private key used for the cryptographic signature.

[0050] Suppose that the (automatic) comparison between the audio signal to be authenticated and the authentication information is performed in a way that allows for minor modifications (e.g., through perceptual coding), and the result of the comparison is unclear (i.e., no explicit statement can be made about the authenticity of the audio signal), then a link to the authentication information (where the audio features contained therein will correspond to the original audio file or its perceptually coded version) can be used for a “manual” perceptual comparison to increase trust in the method.

[0051] By providing audio block-specific links (e.g., via UUID and a user-based unique block number), the following advantages can be achieved: (e.g., in watermarking scenarios) audio files can be cut at any time without the cutting tool being aware of authentication processing and embedded links. Alternatively, in cases where the audio container is used for the joint transmission of audio and links, file-based links can be embedded once in the included header to reduce data rate. This can consist of a UUID and a location within metadata (e.g., a sample index of the original file) indicating how the audio file in question is aligned with a reference on the server.

[0052] The present invention also relates to a computer program product for generating authenticated audio signals. This computer program product includes program code that, when executed by a processor, causes the audio signal signature module to perform the method for generating authenticated audio signals.

[0053] The present invention also relates to a computer program product for authenticating audio signals. The computer program product includes program code that, when executed by a processor, causes the aforementioned audio signal authenticator to perform a method for authenticating audio signals.

[0054] These and other aspects of the invention will be described in more detail with reference to the following accompanying drawings.

[0055] Figure 1A block diagram of the audio signal signature module is shown.

[0056] Figure 2 A flowchart illustrating the process of adjusting the block length of an audio block is provided, along with...

[0057] Figure 3 A block diagram of an audio signal authenticator is shown.

[0058] Figure 1 A diagram is shown of an audio signal signature module 500 that receives digital audio signal 121 from a microphone 100 or other audio source such as a recorder (not shown). A selector switch 101 can selectively forward the audio signal 121 from the microphone 100 or other audio source to an input 502 of the audio signal signature module 500. The audio signal signature module 500 outputs an output audio signal 501 via an audio output 190. The output audio signal 501 is based on the input audio signal 121.

[0059] Microphone 100 includes at least one microphone capsule 110 and an analog-to-digital converter (ADC) 120. The at least one microphone capsule 110 is configured to capture an audio signal and output an analog audio signal 111. The analog audio signal 111 is transmitted to the ADC 120 to convert the analog audio output signal 111 into a digital audio signal 121.

[0060] In one embodiment, the audio signal signature module 500 is integrated into the microphone 100. In an alternative embodiment, the audio signal signature module 500 is implemented as a device separate from the microphone 100.

[0061] A digital audio signal 121 is input to a block divider 130, which divides the audio signal 121 into audio blocks 131 and outputs a sequence of audio blocks 131. Each audio block 131 has a specific block length. The sequence of audio blocks 131 is forwarded to an audio feature extractor 150, which generates audio features 151 for each audio block 131 in the sequence. A private key signing unit 140 receives authentication information including the audio features 151 of each audio block 131 and optional other information 113 (e.g., metadata) (such as location or time), and signs the authentication information with a private key 142 to generate signed authentication information 141, which can be stored in a memory 600 implemented as a storage device (e.g., as an external server). Alternatively, the signed authentication information 141 can be stored in internal memory or any other memory, as long as it can be used for authenticator access to authenticate the received audio signal.

[0062] In all cases, before the corresponding audio block 131 is processed in the audio feature extractor 150, an encoded representation of the storage location where the signed authentication information 141 of the currently analyzed audio block 131 will be stored in the memory 600 can be made available. Such an encoded representation can include a unique user identifier and a user-associated block number, which facilitates explicit identification of the storage location. The actual mapping between the encoded location and the actual storage location can be performed by a dedicated server, which acts similarly to a DNS server that maps website names to IP addresses. Specifically, the authentication access information can be an encoded representation of the storage location. Figure 1 In this context, it is assumed that the private key signing unit 140 has knowledge of the encoded representation of the aforementioned storage location, and the private key signing unit 140 uses this knowledge to generate the corresponding authentication access information 162.

[0063] The authentication access information 162 has a specific size, which is substantially independent of the size of the corresponding / associated audio block, but can still exceed the amount of information that can be integrated as a watermark into the analyzed audio block 131. The authentication access information 162 is forwarded to the information embedder 160, which attempts to embed the authentication access information 162 as a watermark into the current audio block 131. If the current payload capacity of the analyzed audio block 131 is insufficient to allow the attempted embedding, the block segmenter 130 is requested to increase the size of the current audio block to the desired block size 165, and the attempt to embed the authentication access information 162 as a watermark into the current audio block 131 with the increased size 165 is repeated. These two steps—incremental increase of the block size and attempt to embed the watermark—are performed until embedding is successful. Optionally, in addition to the authentication access information, the final selected block size 165 of the current audio block can also be encoded into the watermark to allow for simpler decoding.

[0064] The signed authentication information 141 itself includes authentication information, namely (a) audio features 151, optionally (b) other information 113, and (c) a signature 144. Each audio block 131 is simultaneously provided to an information embedder 160 for embedding authentication access information 162 and an audio feature extractor 150.

[0065] In an alternative implementation, the information embedder 160 embeds the authentication access information 162 into a suitable audio container adjacent to the actual audio block 131. Regardless of whether the authentication access information 162 is stored as a watermark or in an audio container, the information embedder 160 outputs the audio block 161 with the embedded authentication access information 162. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at audio output terminal 190. The output audio signal 501 can be transmitted or stored. In particular, it can even be further processed to perform sample rate conversion, perceptual compression, and segmentation. The embedded authentication access information 162 is unaffected by further processing or at least only partially affected by it.

[0066] When the information embedder 160 embeds a watermark, its output 501 can alternatively be input to the audio feature extractor 150 instead of the original audio block 131, which in Figure 1 The selection is indicated by selector 102. This alternative offers the advantage of extracting audio features based on the final signal 501 output by the audio signal signature module 500, rather than on the unoutput intermediate signal 131. Since the information embedder 160 is designed to perceptually alter the audio signal as little as possible, both input options (131 or 501) of the audio feature extractor 150 are reasonable.

[0067] Audio features involve attempting to capture an audio representation of relevant aspects used for human auditory perception, while preferably reducing the required data rate and thus the storage space required on the server or memory 600. While it could be the original audio signal itself, preferably, it could be a transparently encoded version of the original signal that requires far less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.

[0068] The audio signal 501 at the output of the audio signal module 500 (containing embedded watermarks or authentication access information included in the audio container) can be stored or transmitted. Since the watermark 163 can be embedded in, for example, essentially every audio block 131, each audio block 131 can be authenticated individually. The watermark 163 can also be embedded only in some audio blocks within the audio block set. In addition to authentication access information, the watermark 163 may also include other information such as block boundaries. This information about the block boundaries is useful for determining the adaptive block length, which is required when authenticating the audio signal.

[0069] For watermarking to be a viable option for authenticating audio signals, the amount of information to be embedded should be less than the watermark payload capacity rate, which is the average amount of information that can be inserted into the audio signal per unit time while maintaining imperceptibility. Watermarking typically utilizes effects in human perception, such as temporal and frequency masking, to maintain imperceptibility. In particular, the actual amount of information that can be embedded per unit time is highly dependent on short-time signal characteristics and is therefore often highly time-varying. Assuming the authentication method mentioned above uses a consistent block length that varies over time, a problem arises: for a single block, the metadata size may be too large compared to the data that can be embedded in the watermark, making the watermarking scheme infeasible.

[0070] Therefore, as shown in the example, the block length is adapted to the watermarking requirements. The block length of a single audio block or a group of audio blocks can be adjusted according to the watermarking requirements. The block segmenter can analyze the audio characteristics to determine how much watermark payload can be introduced.

[0071] There is a trade-off between the idea of ​​using short chunks of approximately one second for the authentication process and the need to include watermarks in the audio chunks. Another advantage of the example above is that it is now possible to fully authenticate audio combinations that are already composed of excerpts from various audio files.

[0072] According to the example, the block length varies depending on the block-specific watermark payload capacity that can be embedded into the audio signal.

[0073] The block segmenter 130 can also analyze the input audio signal to determine audio features of the audio signal that relate to determining the available watermark payload (i.e., the amount of data that can be embedded into the audio block). Alternatively, an audio feature extractor can be used to analyze the audio features of the audio signal.

[0074] Audio features involve attempting to capture an audio representation of relevant aspects used for human auditory perception, while preferably reducing the required data rate and thus the storage space required on the server or memory 600. While it could be the original audio signal itself, preferably, it could be a transparently encoded version of the original signal that requires far less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.

[0075] Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, if the audio signal with the embedded watermark is copied, stored, or transmitted, the embedded watermark will not be altered. The same principle applies to authentication access information included in an audio container.

[0076] Figure 2A flowchart illustrating the process of adjusting the block length of an audio block is shown. The process begins in step S1, where the block segmenter 130 selects a minimum block length for the audio block. In step S2, the payload capacity for watermarking is analyzed for the selected block length. In other words, the capacity (i.e., payload capacity) for embedding a watermark in an audio block with that block length is determined. Optionally, the determination of the block-specific payload capacity is based on a predetermined robustness of the watermark to manipulation. If the block length is insufficient, the block length is increased in step S3. If the block length is sufficient, the process continues to step S4, where the block segmenter uses the selected block length at least for that particular audio block. In step S5, the watermark is embedded into the audio block with the specific block length.

[0077] This block length adjustment can be applied to each audio block. Alternatively, the block length can be determined for a set of blocks for which the watermark size is not expected to change.

[0078] First, if the minimum predefined block length allows for the embedding of the required information (e.g., a link to block-specific metadata as mentioned above), then the minimum predefined block length is selected before analyzing the signal characteristics. If this condition is not met, the length of the audio block under consideration is increased, and the embedding feasibility test is performed again. These two steps are repeated until embedding is feasible. In this case, the actual embedding is completed, and the next signal section is processed. Information about the final block length used must be (possibly encoded) pre-attached to the embedded data to enable correct watermark decoding and thus correctly reconstruct the block length used for authentication. Typically, the characteristics of the audio signal are a priori unknown, making it impossible to guarantee the minimum watermark payload capacity. Therefore, it is reasonable to encode the block length (at least partially) into the watermark using a variable bit rate.

[0079] The reason the proposed method works in the described context is that, on the one hand, the size of the data to be embedded does not (or hardly) change with the increase of block length, and on the other hand, the increase of block length increases the block-specific watermark payload capacity for successful embedding.

[0080] Figure 3 A block diagram of an audio signal authenticator 200 is shown. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. The audio signal 171 is, for example, from... Figure 1 The audio signal 171 is output by the audio signal signature module 500. This audio signal 171 is input to the block boundary detector and block segmenter 205, which outputs multiple audio blocks 201 with embedded authentication access information 212.

[0081] The block boundary detector and block segmenter 205 detect block boundaries, for example, by means of information in the block header. Since the block length of an audio block can vary, the block length should be determined for each audio block or for a group of audio blocks with the same block length.

[0082] If authentication access information has already been embedded using a watermark, it is assumed that the block boundaries were previously encoded into the watermark, and the block boundaries can be detected from the watermark at this stage. In the case where the audio signal is transmitted by an audio container containing authentication access information as metadata, it is assumed that the authentication access information is available for each block, where the block boundaries are part of the metadata.

[0083] An audio block 201 with embedded authentication access information is input to an access information extractor 210, which extracts authentication access information 212 from the audio block 201 (or, if no manipulation occurs, authentication access information 162 from the signature module). An authentication access information interpreter 230 determines access information based on the authentication access information 212, i.e., where on the server or memory 600 can the signed authentication information 141 belonging to the current audio block 201 be found. The authentication information 141 retrieved from the server or memory 600 first includes audio characteristics 151 representing the audio block 201, then includes a signature 144, and then includes optional additional information 113, such as the location and / or time of the signature occurrence.

[0084] In the perceptual similarity analyzer 240, audio blocks 201 and audio features 151 are perceptually compared to each other, thereby providing a similarity statement 241.

[0085] The similarity analyzer 240 can compare the audio block 201 with the audio features 151 of the original audio signal. The audio features can be a true copy of the original audio signal. However, this would require a large amount of storage space to store the original audio signal. If the audio features are related to other features of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Therefore, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, in the similarity analyzer 240. In other words, the similarity analyzer 240 should compare the audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.

[0086] In its simplest case, this could be a binary decision about whether the quantities being compared are perceived to be sufficiently equal. Alternatively, the similarity statement 241 could be more detailed, for example, demonstrating a similarity score.

[0087] The authentication access information interpreter 230 also determines, for example, by an ID, which owner of the private key 142 initially signed the audio signal 171 based on the authentication access information 212. To name just a few possibilities, the owner of the private key 142 could be a natural person, an organization, or a registered microphone. Using this ID, the identity 270 of the private key owner and the public key 143 can be retrieved from an external database 400. Optionally, the external database is a distributed ledger. The signature 144 is then authenticated within the signature verifier 250 using the public key 143 to produce a signature verification result 251. The signature verification result can present two states, such as true or false. Finally, the perceived similarity statement 241 and the verification result 251 are combined in the combiner 260 to form the final authentication result 261.

[0088] Note that the processing proposed here for audio content can be similarly applied to any time-related data. Specifically, the processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark, or a video container is used instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison. Instead of repeatedly attempting to embed information access information 162 as a watermark into the current audio block 131 and increasing the block size 165 if embedding fails, the required block length can be directly determined by analyzing the size of potential watermark payloads that can be embedded for different sizes of the current audio block 131 and selecting the minimum block size that allows successful embedding of the required information access information.

[0089] According to the example, each audio signal captured by microphone 100 and output by the microphone may include a watermark, which can be used to identify the actual microphone that captured the audio signal. If the microphone is used by several users, several user IDs can be associated with the microphone. If the microphone is registered, the corresponding microphone can be identified. User IDs can also be registered. The user ID can be a unique ID or a sub-ID of the microphone ID. The microphone ID and / or user ID can be part of the watermark. Using the sequence number of the audio block, it can be determined when a portion of the audio signal has been removed.

[0090] Alternatively, the audio signal may be part of a video file.

[0091] Alternatively, the signed authentication information can be included in the audio file (e.g., in an ADM-like format) instead of in the watermark. In this case, the audio signal cannot be modified.

[0092] Alternatively, the processing of the audio signal used to generate the signature can be performed in a software solution based on a pre-existing audio recording or audio stream. The private key can be entered into the software, for example, via a dongle or text input, and can be authorized, for example, via biometric signals such as fingerprints or facial detection.

[0093] The authentication process can calculate a similarity score between the audio features of the signed audio signal and the analyzed signal. This allows for the determination of the possibility that the signal is still authentic even if minor signal processing, such as gain adjustment, has occurred.

[0094] Authentication requires an audio signal and signed authentication information. The signed authentication information can be provided via a side channel. When using a watermark, the authentication access information must be read from the audio signal, and the authentication information must be retrieved accordingly. Where watermark readout requires block boundaries, one method to achieve this is to try different offsets of the block boundaries until the watermark can be successfully read. Another method could be to embed some synchronization signal with the watermark into the audio signal.

[0095] To enhance system security, processing is provided that allows for key revocation, for example, in the event of key theft.

[0096] Public keys used for authentication can be provided in various ways. One option is to store all public keys in a centralized database to make it easy to find the required key. To address potential trust issues in this centralized instance, a database with all public keys can be provided based on distributed ledger technology. Another option is for organizations / individuals using this technology to provide public keys on their own websites.

[0097] Torsten Dau, Dirk Pueschel, and Armin Kohlrausch describe other examples of audio features in A quantitative model of the "effective" signal processing in the auditory system. I. model structure, The Journal of the Acoustical Society of America, 99(6):3615-3622, 1996, hereinafter referred to as the Perceptual Model (PEMO). This paper describes a quantitative model designed to describe how the auditory system processes acoustic signals. The focus is on developing models that mimic the functional processing of auditory stimuli perceived by humans, particularly in the context of complex sounds such as speech or music.

[0098] The key components of this model can be categorized into peripheral processing, envelope extraction, modulation filtering, and a decision-making stage. The peripheral processing section captures the initial stages of auditory processing, including converting acoustic signals into neural representations through mechanisms such as outer and middle ear filtering and nonlinearities in the cochlea. The envelope extraction section involves extracting the temporal envelope of the sound, which is crucial for understanding amplitude modulation and speech processing. The modulation filtering section involves a system for filtering amplitude modulation, mimicking how the auditory system is sensitive to different modulation frequencies. In the decision-making stage, the processed auditory signals are used to make decisions about the properties of the sound, representing a higher level of auditory perception.

[0099] The model was designed to align with experimental psychoacoustic data, particularly in its ability to simulate the auditory system's response to amplitude-modulated sounds. This method provides a framework for understanding effective signal processing in human auditory perception, especially for tasks involving the detection and discrimination of complex auditory patterns.

[0100] Performing a comparison of the received audio block with the extracted or retrieved authentication information in the transform domain (i.e., with respect to the audio features perceived by human hearing) ensures that the identified differences are perceptually meaningful. In principle, any similarity measure (e.g., empirical cross-correlation coefficient or relative error relative to a reference) can be used to map the comparison results to a numerical value.

[0101] When a further transformation is applied to the Mel frequency power spectrum of the audio signal, Mel frequency cepstral coefficients (MFCCs) are obtained. MFCCs use the Mel scale, a perceptual scale for pitch, used to approximate how humans perceive sound. The Mel scale is non-linear, making it finer in its division of low frequencies and coarser in its division of high frequencies, thus more closely resembling human auditory perception.

[0102] Mel frequency cepstral coefficients use the cepstral spectrum as a transform, which converts the signal from the frequency domain back to a domain where the rate of change of the signal can be analyzed. This approach captures the spectral characteristics of a signal by emphasizing the information-carrying components.

[0103] The process of calculating MFCC from an audio signal involves several steps: pre-emphasis: typically by applying a high-pass filter to enhance the high-frequency components of the signal to balance the spectrum; framing: dividing the audio signal into short overlapping frames, typically 20 to 40 milliseconds long, because speech signals are non-stationary, but they can be considered quasi-stationary within these short frames; windowing: multiplying each frame by a window function, such as a Hamming window, to reduce edge effects and smooth the signal; Fourier transform: applying a Fast Fourier Transform (FFT) to each windowed frame to convert the time-domain signal to the frequency domain; Mel filter bank: the resulting spectrum is passed through a series of triangular bandpass filters spaced according to Mel scale intervals, a step that approximates how humans perceive sound frequencies; logarithmic calculation: calculating the logarithm of the power of each Mel-filtered spectrum, which simulates the human ear's response to loudness; Discrete cosine transform (DCT): finally, the logarithmic Mel spectrum is transformed using a DCT. The result is a set of coefficients called MFCC. Typically, only the first 12 to 13 coefficients are retained, as these contain the most important information.

[0104] MFCCs can capture the wide spectral shape of audio signals in a compact form, making them extremely useful for machine learning algorithms in speech and audio processing. They help distinguish different phonemes, speaker identities, and other audio features by providing robust representations of variations in pitch, volume, and other factors.

[0105] Perceived audio features may also include:

[0106] Linear predictive coding (LPC) is a method that uses information from a linear predictive model to represent the spectral envelope of a digital speech signal in compressed form. It estimates the parameters of filters that can be used to reconstruct the signal. LPC is very effective for modeling the formants (resonant frequencies) of speech sounds.

[0107] - Perceptual Linear Prediction (PLP) coefficients are similar to linear predictive coding, but incorporate aspects of human auditory perception, such as critical band spectral resolution, equal-loudness curves, and the intensity-loudness power law. Perceptual Linear Prediction (PLP) coefficients are designed to mimic the nonlinear perception of loudness and frequency by the human ear.

[0108] Gammatone filter bank features involve filter banks that simulate the filtering that occurs in the human cochlea. They are similar to Mel filter banks that use gammatone filters instead of triangular filters. Gammatone filter bank features are useful for capturing the detailed frequency structure of audio, especially in tasks involving environmental sound classification or hearing aid design.

[0109] - Chromaticity features, which represent 12 different pitch categories (C, C#, D, etc.) of a musical octave. They capture the harmonic and melodic features of music. These are particularly useful in music information retrieval, tonality detection, and chord recognition tasks.

[0110] - Mel spectrogram: While similar to MFCC, the Mel spectrogram is the result of applying a Mel filter bank directly to the power spectrogram without further transformations (such as DCT). The Mel spectrogram preserves more detailed frequency information and is commonly used as input to deep learning models, often in conjunction with convolutional neural networks (CNNs) in tasks such as audio event detection, speech recognition, and music genre classification.

[0111] - The Constant Q Transform (CQT) provides a time-frequency representation with a logarithmic frequency scale similar to the Mel scale, but with a variable time resolution that matches the frequency resolution. It is particularly useful for music applications because it provides a better representation of musical pitch than linear FFT or even the Mel scale.

[0112] - Deep learning-based features: Learned features from deep learning models, such as embeddings from trained neural networks, can also be used as alternatives to traditional handcrafted features like MFCCs. These features are generally more robust and can capture complex patterns that are difficult to model using traditional methods.

[0113] - Spectral Subband Centroid (SSC) is the centroid of the energy distribution in different subbands of the captured signal. It can be interpreted as the "centroid" of the spectrum within each subband. It provides information about the energy distribution across frequency bands and is sometimes used as a supplement to MFCC.

[0114] - Relative Spectrum (RASTA) features involve filtering the logarithmic energy of the speech signal to highlight modulation frequencies important for speech recognition. RASTA-PLP is a combination of RASTA filtering and PLP analysis, providing robust features against noise and channel variations.

[0115] Audio features involve the audio representation of an audio signal, attempting to capture relevant aspects of human auditory perception of the audio signal. Audio features should be chosen to reduce the required data rate (for transmitting the audio features) and thus reduce the storage space required on the server. While the audio feature may be the original audio signal itself, preferably, it can be a transparently coded version of the original signal that requires far less storage space. Another example of an audio feature could be the output of any model that mimics human auditory perception. Audio features could be the Mel-frequency power spectrum and Mel-frequency cepstral coefficients (MFCCs) of the audio signal.

[0116] In particular, audio features can be acoustic features, spectral features, and statistical features of the audio signal.

[0117] Acoustic features may include pitch (fundamental frequency); prosody (rhythm, stress, intonation) and / or duration and silence patterns (pauses, inconsistencies in speech duration or breath sounds can be used as audio features).

[0118] Spectral features may include formant frequencies (such as resonant frequencies in speech); Mel frequency cepstral coefficients (MFCCs) (the short-time power spectrum of sound, which can reveal artifacts introduced during synthesis or manipulation); spectrogram analysis (manipulation may exhibit unusual or fuzzy energy distributions in a time-frequency representation); and / or phase information.

[0119] Acoustic features can include statistical and signal processing features such as noise and residuals (e.g., subtle background noise or artifacts introduced during synthesis may not match the natural recording); high-frequency content; and / or phase distortion (some manipulations may introduce phase anomalies, which can be detected using signal processing techniques).

[0120] Acoustic features can include temporal features such as jitter and flicker (e.g., variations in frequency and amplitude); and temporal coherence (e.g., abrupt shifts or changes in speech features, such as unnatural interruptions in the signal).

[0121] Acoustic features can include behavioral or semantic inconsistencies, such as content coherence (logical inconsistencies or unnatural speech flow in spoken content may indicate manipulation); affect and naturalness (emotional tone may be inconsistent with content, or the voice may lack the natural nuances of human emotion).

[0122] Audio features can be any one of the audio features mentioned above or a combination thereof.

[0123] By combining these features with machine learning or signal processing techniques, models can be trained to detect artifacts and inconsistencies that indicate depth-spoofed audio.

[0124] List of reference numerals

[0125] 100 microphones

[0126] 101 Selector Switch

[0127] 102 Selector

[0128] 110 Microphone Cap

[0129] 111 Microphone simulates audio signal

[0130] 113 Metadata

[0131] 120 AD converter

[0132] 121 digital audio signal

[0133] 130-block divider

[0134] 131 audio blocks

[0135] 140 Private Key Signature Unit

[0136] 141 Signed authentication information

[0137] 142 Private Key

[0138] 143 Public Key

[0139] 144 signatures

[0140] 150 Audio Feature Extractors

[0141] 151 Audio Features

[0142] 160 Information Embedder / Watermark Generator

[0143] 161 Audio blocks with embedded information

[0144] 162 Authentication Access Information

[0145] 163 Watermark

[0146] 165 audio block size

[0147] 171 audio signal

[0148] 190 audio output

[0149] 200 Audio Signal Authenticator

[0150] 205 Block Boundary Detector and Block Segmenter

[0151] 210 Authentication Access Information Extractor

[0152] 212 Authentication Access Information

[0153] 230 Authentication Access Information Interpreter

[0154] 240 Perceptual Similarity Analyzer

[0155] 241 Similarity Results

[0156] 250 signature verifier

[0157] 251 Signature verification result

[0158] 260 combiner

[0159] 261 Authentication Result

[0160] 270 Identity

[0161] 400 External Database

[0162] 500 audio signal signature module

[0163] 501 Certified Audio Signal

[0164] 502 Audio Input Terminal

[0165] 600 Server or Storage

Claims

1. An audio signal signature module (500), comprising: - Input (502) is configured to receive digital audio signals (121). - A block divider (130) is configured to divide the received digital audio signal (121) into a sequence of audio blocks (131) having a block length and a block-specific payload capacity. - An audio feature extractor (150) is configured to extract audio features (151) from the current audio block (131). - A signature unit (140) is configured to generate a signature (144) associated with the current audio block (131) by applying a private key (142) to the extracted audio features (151), and is configured to provide signed authentication information (141) including the extracted audio features (151) and the signature (144), wherein, The signed authentication information (141) is output to the memory (600). - An information embedder (160) is configured to generate an audio block (131) with embedded information (161) by embedding authentication access information (162) into the current audio block (131), wherein the authentication access information (162) includes information about how the signed authentication information (141) can be retrieved from the memory (600) and / or where the signed authentication information (141) can be retrieved from the memory (600), and - An audio output terminal (190) is configured to output an authenticated audio signal (501), the authenticated audio signal comprising a sequence of the audio blocks (131) having embedded authentication access information (162). The block segmenter (130) is configured to adjust the block length of the audio block (131) according to the block-specific payload capacity and the size of the authentication access information (162) to be embedded in the audio block (131).

2. The audio signal signature module (500) according to claim 1, wherein, The information embedder (160) is configured to introduce the authentication access information (162) as a watermark (163) into the current audio block (131), such that the authenticateable audio signal (501) contains a sequence of audio blocks with the embedded watermark (163).

3. The audio signal signature module (500) according to claim 1 or 2, wherein, The signed authentication information (141) also includes metadata (113) belonging to the current audio block (131). The metadata (113) includes at least one of the following: the identifier (145) of the user or entity whose audio signal has been signed, location, and time.

4. The audio signal signature module (500) according to claim 1, 2 or 3, wherein, The information embedder (160) is also configured to determine the block-specific payload capacity of the current audio block (131) to determine whether the block-specific payload capacity is sufficient to embed the authentication access information (161), and if the block-specific payload capacity is insufficient to embed the authentication access information, to increase the block length.

5. A microphone (100), comprising: - Microphone capsule (110), suitable for capturing audio signals (111). - An AD converter (120) is configured to convert the captured audio signal (111) into a digital audio signal (121), and - The audio signal signature module (500) according to claim 1, 2, 3 or 4.

6. A method for generating an authenticable audio signal (501), the method comprising the following steps: - Receive digital audio signals (121). - The digital audio signal (121) is divided into a sequence of multiple audio blocks (131) by a block divider (130), wherein each audio block (131) has a block length and a block-specific payload capacity. - Extract audio features (151) from the current audio block (131). - A signature (144) associated with the current audio block (131) is generated by applying the private key (142) to the extracted audio features (151). - Provide signed authentication information (141), the signed authentication information (141) including the audio feature (151) and the signature (144). - Output the signed authentication information (141) to the memory (600). - An audio block (131) with embedded information is generated by embedding authentication access information (162) into the current audio block (131), wherein the authentication access information (162) includes information about how the signed authentication information (141) can be retrieved from the memory (600) and / or where the signed authentication information (141) can be retrieved from the memory (600). - Output the authenticable audio signal (501), the authenticable audio signal (501) comprising a sequence of audio blocks with embedded authentication access information (161), and - Adjust the block length of the audio block (131) according to the specific payload capacity of the block and the size of the authentication access information to be embedded in the audio block (131).

7. The method of claim 6, comprising: Determine the block-specific payload capacity of the current audio block (131) to determine whether the block-specific payload capacity is sufficient to embed the authentication access information, and If the block-specific payload capacity is insufficient to embed the authentication access information, the block length is increased.

8. An audio signal authenticator (200), comprising: - A block boundary detector and a block segmenter (205) are configured to receive an audio signal (171) and output a sequence of multiple audio blocks (201) with embedded authentication access information (212, 162). - An access information extractor (210) is configured to extract the authentication access information (212, 162) from the current audio block (201). - An authentication access information interpreter (230) is configured to extract access information about how the signed authentication information (141) belonging to the current audio block (201) can be retrieved from the memory (600) and / or where the signed authentication information (141) belonging to the current audio block (201) can be retrieved from the memory (600). The signed authentication information (141) retrieved from the memory (600) includes a signature (144) and audio features (151) belonging to the current audio block (201), and - A perceptual similarity analyzer (240) is configured to compare an audio block (201) of the received audio signal with a retrieved audio feature (151), or to compare an audio feature extracted from the audio block (201) of the received audio signal with the retrieved audio feature (151) to provide a similarity result (241).

9. The audio signal authenticator (200) according to claim 8 further comprises: - A signature verifier (250) is configured to verify the combination of the signature (144) and the audio feature (151) using a public key (143) and to provide a signature verification result (251). The public key (143) is associated with the private key (142), which can be used by the audio signal signature module (500) to generate an authentic audio signal (501).

10. The audio signal authenticator (200) according to claim 9, wherein, The authentication access information interpreter (230) is configured to determine the identifier of the user or entity that initially signed the audio signal. Among them, the identity (270) of the user or entity and the public key (143) can be retrieved from an external database (400) based on the determined identifier.

11. The audio signal authenticator (200) according to claim 9 or 10, further comprising: The combiner (260) is configured to combine the similarity result (241) and the signature verification result (251) into a final authentication result (261).

12. A method for authenticating an audio signal (171) via an authenticator (200), the method comprising the following steps: - Receive audio signals (171). - Output a sequence of multiple audio blocks (201) with embedded authentication access information (162). - Extract authentication access information (212, 162) from the current audio block (201). - Extract access information (212, 162) about how to retrieve the signed authentication information (141) belonging to the current audio block (201) from the memory (600) and / or where to retrieve the signed authentication information (141) belonging to the current audio block (201) from the memory (600). - Retrieve signed authentication information (141) from the memory (600), the signed authentication information (141) including a signature (144) and audio features (151) belonging to the current audio block (201), and - The audio blocks (201) of the received audio signal are compared with the retrieved audio features (151) to provide similarity results (241), or - The audio features extracted from the audio block (201) of the received audio signal (171) are compared with the retrieved audio features (151) to provide similarity results (241).

13. The method of claim 12, further comprising the step of: - Verify the combination of the signature (144) and the audio feature (151) using the public key (143), and - Provide a signature verification result (251), wherein the public key (143) is associated with a private key (142), which can be used by the audio signal signature module (500) to generate an authentic audio signal (501).

14. A computer program product for generating authenticated audio signals, in, The computer program product includes program code that, when executed by a processor, causes the audio signal signature module according to claim 1 to perform the method according to claims 5 to 7.

15. A computer program product for authenticating audio signals. in, The computer program product includes program code that, when executed by a processor, causes the audio signal authenticator according to claim 8 to perform the method according to claim 12.