Audio signal signing module and method of generating an authenticable audio signal

WO2026175500A1PCT designated stage Publication Date: 2026-08-27SENNHEISER ELECTRONICS GMBH & CO KG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/054511
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2026-08-27

Smart Images

  • Figure EP2025054511_27082026_PF_FP_ABST
    Figure EP2025054511_27082026_PF_FP_ABST
Patent Text Reader

Abstract

An audio signal signing module (500) is provided which, comprises a Turing tester (700) configured to perform a Turing test to differentiate a human from a non-human as user, an input (502) configured to receive a digital audio signal (121), a block divider (130) configured to divide the received digital audio signal (121) into a sequence of a plurality of audio blocks (131), an audio feature extractor (150) configured to extract audio features (151) from a current audio block (131), a signature unit (140) configured to generate a signature (144) associated to the current audio block (131) by applying a private key (142) to the audio features (151) and configured to provide signed authentication information (141), which comprises the extracted audio features (151) and the signature (144), wherein the signed authentication information (141) is outputted to a memory (600), an information embedder (160) configured to generate an audio block (131) with embedded information (161) by embedding authentication access information (162) into the current audio block (131), wherein the authentication access information (162) comprises information on how and / or where the signed authentication information (141) is retrievable from the memory (600), and an audio output (190) configured to output an authenticable audio signal (501) comprising a sequence of the audio blocks with embedded authentication access information (161), In a first operating mode the Turing tester is configured to output a question or prompt which the user must answer, receive an answer in form of an audio signal captured via at least one microphone, analyze the answer captured by the microphone, compare the received answer to a correct answer of the question or prompt, and activate a second operating mode, when the received answer is correct and user has passed the Turing test. The second operating mode is activated after completion of the first operating mode to generate an authenticable audio signal (501) via the input (502), the block divider (130), the audio feature extractor (150), the signature unit (140) and information embedder (160).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Sennheiser electronic SE & Co. KG

[0002] Am Labor 1 , 30900 Wedemark,

[0003] Audio signal signing module and method of generating an authenticable audio signal

[0004] The present invention relates to an audio signal signing module and a method of generating an authenticable audio signal.

[0005] In view of modern capability of creating deep fake video and deep fake audio, the need has risen to be able to verify or authenticate an audio signal like a speech of a politician, etc.

[0006] Until now, it has been very difficult to authenticate an audio signal of a person. Typically, a lot of manual research must be done in order to verify or authenticate such an audio signal or verify the source of the audio signal.

[0007] It is therefore an object of the invention to provide means to enable a generation of an authenticable audio signal with an increased level of safety.

[0008] This object is solved by an audio signal signing module according to claim 1.

[0009] Hence, an audio signal signing module is provided. The module comprises a Turing tester configured to perform a Turing test to differentiate a human from a non-human as user. In a first operating mode the Turing tester is configured to output a question or prompt which the user must answer, to receive an answer in form of an audio signal captured via at least one microphone, to analyze the answer captured by the microphone, to compare the received answer to a correct answer of the question or prompt, and activate a second operating mode, when the received answer is correct and the user has passed the Turing test. The second operating mode is activated after completion of the first operating mode to generate an authenticable audio signal via the input, the block divider, the audio feature extractor, the signature unit and information embedder.

[0010] Moreover, a method of generating an authenticable audio signal is provided. The method is performed in two stages. In a first stage (first operating mode) a Turing test to determine whether the user is human is performed. Only, when the Turing test has determined that the user is a human, the second stage (second operating mode) is released and the actual authenticable audio signal is generated.

[0011] In a first operating mode a Turing test is performed to determine whether a user is a human by generating a question or prompt which the user must answer, receiving an answer in form of an audio signal captured via at least one microphone, analyzing the answer captured by the microphone, comparing the received answer to a correct answer for the question or prompt, and activate a second operating mode, when the captured answer is correct and the user has passed the Turing test. The second operating mode is activated after completion of the first operating mode to generate an authenticable audio signal. The generation of the authenticable audio signal is performed by receiving a digital audio signal, dividing the digital audio signal into a sequence of a plurality of audio blocks, extracting audio features from a current audio block, generating a signature associated to the current audio block by applying a private key to the extracted audio features, providing signed authentication information, which comprises the audio features and the signature, outputting the signed authentication information to a memory, generating an audio block with embedded information by embedding authentication access information into the current audio block. The authentication access information comprises information on how and / or where the signed authentication information is retrievable from the memory, and outputting the authenticable audio signal comprising of a sequence of a plurality of audio blocks with embedded authentication access information.

[0012] Hence, to further improve the security against deep fakes of an audio signal, according to an example, it must be determined whether an actual person (a human) is generating the audio signal. Therefore, a Turing test is performed to determine that no computer is generating the audio signal which is to be picked up by a microphone and on which the authentication proceeding is to be performed.

[0013] In particular, an audio Completely Automated Public Turing test to tell Computers and Humans apart CAPTCHA test is performed. The user is prompted to reply to a question outputted by a speaker to perform the Turing test. The reply of the user is picked up by a microphone and this audio signal is analysed to perform the Turing test. When the Turing test has determined that there is no computer, a generation of an authenticable audio signal by means of an audio signal signing module can be initiated.

[0014] Accordingly, the Turing test is performed before the generation of an authenticable audio signal is started in order to ensure that the audio signal which is picked up by a microphone is indeed not a computer but an actual person (i.e. a human). Thus, the security in view of deepfake audio and deepfake video can be further improved.The Turing test according to an example is in particular helpful when a microphone is used to record an interview or a speech. By the Turing test, it is possible to ensure that the audio signal which is picked up by the microphone is not a fake audio signal for example generated by a computer, etc. In particular, it is possible to avoid that a pre-recorded and potentially modified deepfake audio signal is played back over a loudspeaker and this audio signal is captured by a microphone and that the thus recorded audio signal is used as a basis for generating an authenticable audio signal. To avoid this, the microphone can be used to perform a Turing test to determine whether the source of the audio signal is a human or not.

[0015] In order to perform such a Turing test, an audio CAPTCHA is performed. Here, a question is generated and outputted to the user of the microphone. The user must answer the capture question and the audio signal is captured and analysed. Accordingly, the user must speak out the answer loud. By validating the answer (for example by using speech to text conversion) and the circumstances of the answer (e.g. response time), the recording device can determine whether the answer is given by a human or for example by a conversational Al bot.

[0016] According to an example, the audio signal of the answer is analysed and compared to the audio signal which is to be captured by the recording device. If those two audio signals do not originate from the same person, the capturing of the audio signal can be stopped or a notification can be embedded into the audio signal indicating that the audio signal may be generated by a computer.

[0017] According to an aspect, the result of the Turing test can be embedded into the authenticable audio signal for example as metadata.

[0018] The Turing test can be applied and the audio signal of the answer of the user can be detected by a microphone or another recording device (computer with a sound card), an application processing live audio like a conferencing software (MS Teams) or a Digital Audio Workstation. Optionally, the audio can also be outputted by an independent device like a smartphone.

[0019] According to an aspect, the Turing test can be extended to a video conferencing tool. As an example, before making a call, the conferring tool can ask a user to say a specific word or to perform a specific gesture or movement. The audio signal of the word or the gesture can be analysed to determine whether to perform the Turing test. If the Turing test hasdetermined that a computer is generating this audio signal, this information can be embedded in the video and audio signal which is part of the conference. Thus, the other participants and the video I audio conference can be notified that a computer may be participating in the video conference.

[0020] In addition to the Turing test, a biometric test may be performed to determine whether a human is generating the audio signal as well as to determine which person is generating the audio signal.

[0021] A CAPTCHA test "Completely Automated Public Turing test to tell Computers and Humans Apart" is a security mechanism designed to differentiate between humans and computers like automated bots. It ensures that only human users can access certain services or perform specific actions.

[0022] According to an example, the Turing test can be performed based on an analysis of audio signals: In the configuration phase (prior to recording the interview or a speech) during the Turing test a question or prompt is asked to the user. The user needs to answer correctly and in a reasonable amount of time. For example the Turing test can prompt a question: “please add 5 plus 10”. The user can answer “15”. The question can be asked acoustically via an integrated speaker or e.g., on an integrated display. The user must speak the answer out loud. By validating the answer (by using speech-to-text algorithms) and the circumstances (e.g., response time) the Turing test can determine whether the answer is given by a human or by a conversational Al bot.

[0023] According to an aspect an audio signal signing module is provided, which comprises an input for receiving a digital signal, a block divider for dividing the received digital audio signal into a sequence of a plurality of audio blocks, an audio feature extractor for extracting audio features from a current audio block and a signature unit for generating a signature associated to a current audio block by applying a private key to the audio features. The signature unit also provides signed authentication information which comprises the extracted audio features and the signature, wherein the signed authentication information is outputted to a memory. The audio signal signing module also comprises an information embedder for generating an audio block with embedded information by embedding authentication access information into the current audio block. The authentication access information comprises information on how and / or where the signed authentication information is retrievable from the memory. Furthermore, an audio output is provided for outputting anauthenticable audio signal comprising a sequence of audio blocks with embedded authentication access information.

[0024] According to an aspect the information embedder is implemented as a watermark generator configured to generate at least one watermark based on the authentication access information and introduce the watermark into the at least one audio block such that the audio signal with the embedded watermark is outputted via the audio output.

[0025] According to an aspect, the signed authentication information further comprises metadata belonging to the current block. The metadata can comprise at least one of an identifier of the user or entity who or which has signed the audio signal, a location and a time in particular of the signing of the audio signal or the generation of the audio signal.

[0026] Hence, a microphone is provided which comprises a microphone capsule adapted to capture an audio signal, an AD converter configured to convert the captured audio signal into a digital signal, a block generator configured to divide the digital audio signal into a plurality of audio blocks, an audio feature extractor configured to extract audio features from at least one audio block, a signature unit configured to provide a signature of authentication information, which comprises extracted audio features and / or metadata of the at least one audio block, an audio output configured to output the signed authentication information to a memory, an information embedder configured to embed authentication access information into the at least one audio block. The authentication access information comprises information how or where the signed authentication information is retrievable from the memory wherein the audio output is configured to output the at least one audio block with the embedded authentication access information.

[0027] The invention also relates to a method of generating an authenticable audio signal. A digital audio signal is received and the digital audio signal is divided into a sequence of a plurality of audio blocks. Audio features are extracted from a current audio block and a signature associated to the current audio block is generated by applying a private key to the extracted audio features. Signed authentication information is provided comprising the audio features and the signature. The signed authentication information is outputted to a memory. An audio block with embedded information is generated by embedding authentication access information into the current block. The authentication access information comprises information on how and / or where the signed authentication information is retrievable from the memory. The authenticable audio signal is outputted. The authenticable audio signal com-prises a sequence of a plurality of audio blocks with embedded authentication access information. The invention also relates to an audio signal authenticator which comprises a block bounce detector and block divider for receiving an audio signal and outputting a sequence of a plurality of audio blocks with embedded authentication access information. The audio signal authenticator furthermore comprises an access information extracted for extracting the authentication access information from a current block and an authentication access information interpreter for extracting access information on how and / or where signed authentication information belonging to the current audio block is retrievable from the memory. The signed authentication information retrieved from the memory comprises a signature and audio features belonging to the current audio block. Furthermore, perceptional similarity analyzer is provided for comparing the audio blocks and the audio features to each other and providing a similarity result.

[0028] In the similarity analyzer the audio blocks can be compared to the audio features of the original audio signal. The audio features can be a true copy of the original audio signal. This will however require large storage space to store the original audio signal. If the audio features relate to other characteristics of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Hence, the required audio features can be extracted from the audio signal to be authenticated. This can be performed e.g. in the similarity analyzer. In other words, the similarity analyzer should compare the audio features (as received via the authentication access information, i.e. the audio features of the original audio signal) and corresponding audio features of the audio signal to be authenticated.

[0029] The invention also relates to a method of authenticating an audio signal by an authenticator. An audio signal is received and a sequence of a plurality of audio blocks with embedded authentication access information is outputted. The authentication access information is extracted from a current audio block and access information on how and / or where signed authentication information belonging to the current block is retrievable from the memory is extracted. Signed authentication information is retrieved from the memory. The signed authentication information comprises a signature and audio features belonging to the current block. The audio block and the audio features are compared to each other to provide a similarity result.

[0030] Audio features relate to an audio representation of an audio signal trying to capture relevant aspects for human auditory perception of the audio signal. The audio features should be selected to reduce the required data rate (for the transmission of the audio features) andhence storage space required on a server. Although the audio features could possibly be the original audio signal itself, preferably the audio features can be a transparently coded version of the original signal requiring much less storage space. Another example for audio features could be the output of any model mimicking the human auditory perception. The audio features can be a Mel-frequency power spectrum of an audio signal with Mel Frequency Cepstral Coefficients (MFCCs).

[0031] According to an aspect the authentication access information interpreter is configured to determine identity information of the user or entity who or which has originally signed the audio signal, wherein a public key and an identity of the user can be retrieved from an external database based on the determined identity information.

[0032] According to an aspect a signature verifier is provided which is configured to verify a combination of the signature and the authentication information using the public key and to prove a signature verification result.

[0033] According to an aspect an audio signal comprises a combiner configured to combine the perceptual similarity statement and the verification result to a final authentication result.

[0034] According to an aspect a method of authenticating an audio signal by an authenticator is provided which comprises the steps of receiving an audio signal, outputting a plurality of audio blocks with embedded authentication access information, extracting the authentication access information from the audio blocks indicating where the authentication information belonging to the current audio block is retrievable, retrieving from a memory authentication information comprising a signature, audio features representing the audio block and / or further information and perceptually comparing the audio blocks and the audio features to each other providing a similarity result.

[0035] The authentication information is signed with a private key of the microphone or the device performing the capture or post processing of the audio signals to generate a dedicated signature indicative of the source of the authentication information. The private key can be associated with a user of the signing module, e.g. a microphone. At the decoders side a public key can be used to decode or validate the authentication information. In otherwords, the signature is used to verify the source of the authentication information.

[0036] Hence, by means of the signature it can be verified that the authentication information originates from a certain source that is owner of the private key. By further comparing the audiosignal with the authentication information it can be verified that the audio signal has not been manipulated beyond an allowed level. Only a signature is generated based on the private key e.g. of the microphone.

[0037] Storing the authentication information like the metadata not in the audio block but externally is advantageous as the bit rate and the bandwidth required is not increased. Merely, the authentication access information is transmitted together with the audio blocks. If an audio signal is to be authenticated, the authentication access information is extracted and the associated authentication information e.g. the metadata is retrieved.

[0038] The authentication information is thus provided independently of the audio blocks of the captured audio signal. If no authentication is to be performed, the authentication information is not required, and a significant amount of bandwidth and bitrate can be saved compared to the case if the authentication information was attached to the audio signal in an appropriate container format. Furthermore, when an authentication is to be performed, the access information can be extracted from the audio blocks and the authentication information can be retrieved. As the authentication process is not latency critical it is sufficient to extract the authentication information from an external source before performing the authentication.

[0039] The authentication access information can be embedded into the audio block as watermark or in a dedicated container as part of the audio blocks, or, in a file-based context as part of the audio file.

[0040] The authentication information and the authentication access information can be provided for at least one audio block. Hence, each audio block can be individually authenticated. This allows a fine combed authentication. Accordingly, even a slight manipulation of an audio signal (for which the authentication information is available) can be detected. This significantly increases the security of captured audio signals. Providing the authentication information of each audio block is advantageous as it even allows to verify an audio signal comprising different audio blocks from different audio sources. As long as the necessary authentication information is available any audio signal comprising different audio blocks can be validated.

[0041] In the following an example to improve the security of a transmission or storage of the audio signal is described. Audio features of the digital audio signal blocks can be extracted and signatures based on the audio features can be generated using a private key of a personor entity. The captured audio signal and the signed authentication information can be outputted.

[0042] It becomes thus possible to determine whether a transmitted audio signal was captured and signed by a specific microphone or a specific software entity with a specific private key by verifying the combination of the received authentication information and signature using the public key and by subsequently comparing the audio signal with the audio features included in the authentication information.

[0043] According to an aspect of the invention, the microphone also comprises a watermark generator to generate a watermark based on an authentication access information.

[0044] The signing and verification of a signature of the audio blocks can be performed with a private key and public key pair.

[0045] According to an aspect of the invention, in addition to the audio features, the authentication information can comprise a microphone identification, position / location information of the microphone, time and data of the recording and / or a microphone model, sequence number etc.

[0046] Authentication information relates to information that can be used to authenticate a received or selected audio signal. The authentication information enables an authenticator to perform the authentication of an audio signal. The authentication information can comprise information relating to the audio signal or the properties of the audio signal, namely audio features. The authentication information can also comprise metadata relating to information which are independent of the actual audio signal. This information can be an ID of the user, an ID of the microphone used to capture the audio signal, a microphone model, the date, the time, a position or location (GPS position), and / or a sequence number of the audio blocks.

[0047] The authenticity check of the received audio signal can be performed by a decoder which can, for example, be implemented at a cloud service, a computer, a tablet or a smart device.

[0048] The block divider can be used to divide the audio signals into a plurality of audio blocks. The length of the audio blocks can for example be between 0.5s to 20s. It is in particular also possible that the block length varies overtime, which might for example be indicatedby the watermark. In each audio block, a digital watermark can be embedded. It is also possible to embed the watermark in each nthaudio block.

[0049] The sequence of the audio blocks of the audio signals can be used to associate a sequence number to each audio block. This is advantageous as later on during the authentication of the audio signal it can be determined whether an audio block has been removed from the chain of audio blocks by examining the sequence numbers of the audio blocks. This can also detect a changed order of the original blocks.

[0050] Optionally, metadata of the audio block can be the user of the microphone, the date, the time, a GPS position, etc. Thus, the probability is improved that the result of the authentication is reliable.

[0051] The authentication information can comprise audio features representing the audio signal or audio blocks of the audio signal based on human perception or the original audio signal in a transparently coded version to reduce the storage space required by a server.

[0052] The authentication of the audio signal can be performed by detecting the block boundaries from special information in the watermark or audio container, and extracting the authentication access information (e.g. a link (e.g., consisting of UUID and user based unique block number,) to the block specific authentication information either from the audio watermark or directly from the audio container. The public cryptographic key of the user (e.g. identified by the UUID from link) can be used to verify the authenticity of the authentication information. The audio signal block is compared against the audio features contained in the authentication information to verify its authenticity.

[0053] With the above-described method of authenticating an audio signal it becomes possible to detect malicious modifications of the audio signal by comparing an audio block with the authentication information associated to the audio block. It becomes also possible to detect a modification of the audio signal to be authenticated and / or a modification of the authentication information externally stored by checking the signature of the authentication information. If an audio block of an audio signal is modified and the authentication access information is changed as well, this will be noted by the authentication process as the signature of the authentication information is that of another person or device.According to an aspect, the authentication information stored externally can be used to manually authenticate the audio signal orthe audio block if the authenticating process does not give a clear indication of the authenticity of the audio signal orthe audio blocks.

[0054] According to an example, the actual audio features extracted from the audio blocks and optional additional metadata are stored in a signed way, e.g. on a server. A unique (unambiguous) link to the metadata stored e.g. on the server is embedded into watermark or into a suitable audio container. The link can comprise a unique user identifier (UUID) in combination with a user based unique block number of the respective audio block.

[0055] In order to authenticate the audio signal of an audio block, the block boundaries can be detected from special information in the watermark or audio container. Then the link (e.g., consisting of UUID and user based unique block number,) to the block specific signed authentication information, consisting of the audio features, optional metadata and the signature, is extracted either from the audio watermark or directly from the audio container. By means of a public cryptographic key of the user (e.g., identified by the UUID from link,) the combination of the authentication information and the signature is verified. Accordingly, it is verified whether the authentication information has been signed by the claimed source of the audio signal. The audio signal block is compared against the audio features to verify perceptual similarity, and thus the authenticity of the audio signal block. For the perceptual comparison, audio features can be determined from the received audio signal block, which can be directly compared to the audio features from the signed authentication information stored e.g. on a server.

[0056] The proposed authentication method is able to prevent any malicious modification of the actual audio signal comparing the audio signal block against the signed authentication information. The authentication might fail in case the comparison allows for some limited amount of modification and the allowed modification threshold is exceeded. Moreover, If together with a malicious modification of the audio signal block the authentication information on the server is altered accordingly, this can be detected by a mismatch of the authentication information and the signature associated to the authentication information. Furthermore, if together with a malicious modification of the audio signal block the link (within the watermark or audio container) is changed to refer to different (but suitable) signed authentication information on the server, the verification of the signature and the subsequent comparison may be successful indeed, but the signature will be that of a different user, assuming that the private key of the original user for the cryptographic signing is inaccessible to the attacker.Assuming that the (automated) comparison between the audio signal to be authenticated and the authentication information is carried out in a way to allow some minor modifications (e.g. by perceptual coding) and the result of this comparison is not clear (i.e., no unambiguous statement about the authenticity of the audio signal can be made), the link to the authentication information (in case the contained audio features would correspond to the original audio file or a perceptually coded version of it) could be used for a “manual” perceptual comparison, in order to increase the trust in the method.

[0057] By providing an audio block specific link (e.g., by the UUID and user based unique block number,) allows the advantage (e.g., in the watermarking scenario) that the audio file can be cut at anytime instant without the cutting tool being aware of the authentication process and the embedded links. Alternatively, in case an audio container is used for the joint transport of the audio with the link, a file-based link might be once embedded into the contained header to reduce the data rate. It might be composed of the UUID and a position within the meta data (e.g., a sample index of the original file) to indicate how the audio file in question is aligned with the reference on the server.

[0058] Further aspects are described in the dependent claims.

[0059] These and other aspects of the invention are described in more detail with reference to the following figures.

[0060] Fig. 1 A shows a block diagram of an audio signal signing module,

[0061] Fig. 1B shows a flow chart of a Turing test in a first operating mode, Fig. 2 shows a diagram of an overall audio signing and authentication workflow in a second operating mode,

[0062] Fig. 3. shows a block diagram of an audio signal authenticator,

[0063] Fig. 4 shows a block diagram of an audio signal signing module, and Fig. 5 shows a block diagram of an audio signal authenticator.

[0064] The invention relates to a generation of an authenticable audio signal and to an authentication of an audio signal. In other words, an audio signal must be processed to include information which can enable a later authentication. At the authentication stage, a received or stored audio signal is authenticated based on the information associated to the audio signal.The generation of the authenticable audio signal is performed in two stages: A first operating mode to perform a Turing test and a second operating mode to generate the authenticable audio signal.

[0065] Fig. 1 A shows a diagram of an audio signal signing module. The audio signal signing module 500 comprises a Turing tester 700 for performing a Turing test to determine whether a user is human or non-human (like a computer or a computer bot). In a first operating mode the Turing tester 700 is activated and the Turing test is performed.

[0066] Fig. 1 B shows a flow chart of a Turing test in a first operating mode. The objective of the first operating mode is to perform a Turing test in order to make sure that the user is a human. The first operating mode (Turing test) is preferably performed before the user starts outputting the audio signal (e.g. in an interview or during a speech). Hence, the Turing test can be performed during the step up or configuration phase.

[0067] To perform the Turing test in a first step S1 the Turing test is activated. In a second step S2 a question or prompt which the user must answer is generated. In a third step S3 an answer of the user in form of an audio signal is captured via at least one microphone. In step S4 the answer captured by the microphone is analyzed. In step S5, the received answer is compared to a correct answer of the question or prompt. In step S6, a second operating mode is activated, when the captured answer is correct and the user has passed the Turing test.

[0068] The Turing test based on audio signals as described can also be used for video conferencing tools (e.g. Zoom, MS Teams etc.). For example, before an audio / video call is initiated, a Turing test is activated. E.g. the Turing test can ask the userto answer a question and / or perform a gesture. The response is captured as audio signal and / or video signal. The audio and / or video signal are analyzed and the audio response is compared to the right answer to the question or prompt. The video can be analyzed to determine whether the user has performed the requested gesture. As an example, the user can be prompted to say “potato” and to look to the top left.

[0069] In case the visual or acoustic response is not as expected an indication thereof can be shown to the other participants. Hence, the other participants of the video conference are notified that they might speak to an Al bot.Only when the Turing test has determined that the user is human, a second operating mode is activated and the rest of the audio signaling module is activated. The operation of the audio signing module is described in detail below.

[0070] In the second operating mode, the audio signal signing module 500 receives an input audio signal 121 via an audio input 502 from a microphone 100 or another audio source such as an audio recorder and outputs an output audio signal 501 via an audio output 190. The output audio signal 501 can be based on the input audio signal 121. The microphone 100 can comprise at least one microphone capsule 110 and an analog-to-digital AD converter 120. The at least one microphone capsule 110 can capture an audio signal and can output a captured analogue audio signal 111. The captured analogue audio signal 111 can be forwarded to the analog-to-digital AD converter 120 which can digitize the audio signal 111 and can output a digital audio signal 121 .The audio signal signing module 500 can be implemented as a device separate from the microphone 100 or can be comprised in the microphone 100.

[0071] The audio signal signing module 500 can receive the audio signal 121 from the microphone 100. The digital audio signal 121 is inputted to a block divider 130 which divides the audio signal 121 into audio blocks 131 and outputs a sequence of a plurality of audio blocks 131. Optionally, if the audio signal 121 is already divided into audio blocks 131 , the block divider 130 can be omitted. The sequence of audio blocks 131 are forwarded to an audio feature extractor 150 which generates audio features 151 for each audio block 131 in the sequence of audio blocks. A private key signature unit 140 receives authentication information comprising audio features 151 of each audio block 131 togetherwith optional further information 113 (e.g. metadata) like location or time and signs the authentication information with the private key 142, providing signed authentication information 141 , which can be stored on a memory 600 e.g. located on an external server. The signed authentication information 141 can comprise the authentication information 152 (see Fig. 2) comprising audio features 151 and optional further information 113 and a signature 144 (Fig. 2) generated by the private key signature unit 140. Alternatively, the authentication information 141 can be stored in an internal memory or any other memory, as long as it can be accessed by an authenticator used to authenticate a received audio signal.

[0072] The signed authentication information 141 itself comprises the authentication information, namely (a) audio features 151 , optionally (b) the further information 113 and (c) the signature 144. The individual audio blocks 131 are e.g. in parallel fed to an information embedder 160 for embedding authentication access information 162 e.g. into each audio block 131.The authentication access information 162 comprises information on where and / or how the authentication information 152 (or more precisely the signed authentication information 141), for that particular audio block 131 , can be accessed at the memory 600 e.g. on a server. The embedding can for instance be accomplished by means of watermarking.

[0073] Hence, the information embedder 160 can be implemented as a watermark generator which generates a watermark 163. Another option is to embed the authentication access information 162 into an appropriate audio container next to the actual audio blocks 131. Accordingly, the information embedder 160 outputs audio blocks 161 with embedded authentication access information 162. The output of the embedder 160 corresponds to the output audio signal 501 of the module 500 at the audio output terminal 190. The output audio signal 501 can be broadcasted or stored e.g. via a network 300 (Fig. 2). In particular, it can be even further modified to a little extent, including modification types like sample rate conversion, perceptual compression and cutting.

[0074] In the case that the information embedder 160 is a watermark generator its output 161 could alternatively be input to the audio feature extractor 150 instead of the original audio blocks 131 , which is indicated in Fig. 2 by the selector 102. This alternative offers the advantage that the audio features are extracted based on the final signal 161 that will be output by the audio signal signing module 500 rather than an intermediate signal 131 that is not outputted. As the watermark generator is designed to perceptually alter an audio signal as little as possible, both input options (131 or 161) for the audio feature extractor 150 are reasonable.

[0075] Audio features relate to an audio representation trying to capture the relevant aspects for human auditory perception and at the same time preferably reduces the required data rate and hence storage space required on the server or memory 600. Although it could possibly be the original audio signal itself, preferably it can be a transparently coded version of the original signal requiring much less storage space. Another example for audio features could be the output of any model mimicking the human auditory perception.

[0076] The audio signal 501 (with the embedded watermark or authentication access information included in an audio container) at the output of the audio signal module 500 can be stored or transmitted. As a watermark can be embedded into e. g. substantially each audio block 131 , each audio block 131 can be individually authenticated. The watermark 103 can also be embedded only in some of the audio blocks. The watermark 163 can comprise apart from the authentication access information further information like the block boundaries.This information can be used when authenticating the audio signal to determine a block length.

[0077] An audio watermark 163 can be a distinct identification which is embedded in an audio signal and which is for example previously used for identifying a copyright information of the audio signal. Preferably, the watermark is embedded into the audio signal such that it becomes very difficult to remove or destroy the watermark. Preferably, if the audio signal with the embedded watermark is copied, stored or transmitted, the embedded watermark will not change. The same reasoning applies to authentication access information included in an audio container.

[0078] Fig. 2 shows a diagram of an overall audio signing and authentication workflow. An audio signal is captured by a microphone 100 outputting a digital audio signal 121. From this audio signal 121 audio features 151 are extracted and signed with the private key 142 of e.g., audio signal signing module 500. The audio features 151 together with the digital signature 144 generated by the signature unit 140 form the signed authentication information 141. The signed authentication information 141 can be forwarded to a server or memory 600, where the signed authentication information 141 is stored for later retrieval. The information on where and / or how the signed authentication information 141 can be retrieved, namely the authentication access information 162 is provided for the audio signal signing module 500. The audio signal signing module 500 uses this authentication access information 162 and embeds this information into the audio blocks 131 of the audio signal. Hence, the output signal 501 of the audio signal signing module 500 comprises the audio blocks 131 with the embedded authentication access information. The audio blocks 131 of the audio signal with the embedded authentication access information 162 are outputted by the audio signal signing module 500 and can be stored or distributed e.g. through a network 300.

[0079] This distribution can involve a modification or manipulation of either the authentication information 152 and / or the audio signal 121. The signed authentication information 141 is validated using the public key 143. This validation indicates a potential modification of the audio features. By extracting audio features from the received signal and comparing them with the features contained in the signed authentication information 141 the integrity of the audio signal can be verified. A final authentication result 261 is determined from both checks.An audio signal signing module 500 e.g. implemented as part of a microphone 100 is used to capture an audio signal. Alternatively, an audio signal signing module 500 can receive the audio signal. The audio signal can also be an audio signal captured beforehand by other means. As described with reference to Fig. 1 , the captured audio signal is divided into a plurality of audio blocks and the authentication information of at least one audio block is signed based on the private key 142 of the audio signal signing module 500. The signed authentication information 141 is outputted by the audio signal signing module 500 and can be stored on a server or memory 600. Authentication access information 162, regarding information where the authentication information can be retrieved from the server or memory 600 can be provided to the audio signal signing module 500. The authentication access information 162 is embedded into the audio block received or captured by the audio signal signing module 500. Thus, the audio signal signing module 500 outputs a) an output audio signal 501 , which comprises audio blocks with embedded authentication access information to a network 300 for storing or forwarding and b) signed authenticating information 141 to the server or memory 600 where this information ca be stored.

[0080] Optionally, the authentication access information can be embedded into the audio blocks of the audio signal by means of a watermark. The watermark can relate to authenticating access information relating to a location of a memory 600 where authenticating information is stored. This authentication information can be generated by an audio feature extractor 150 as described above with reference to Fig. 1. The audio blocks of the audio signal with the embedded authentication access information can be transmitted via a network 300 or stored. An authenticator 200 can receive the audio signal 501 with the embedded watermark and extract the authentication access information. The authentication information is retrieved from the server 600 based on the authentication access information. The signature of the authentication information as stored on the server or memory 600 can be validated based on a public key 143 associated to the audio signal signing module 500 which can be stored, for example, on a server 400 or a distributed ledger.

[0081] Fig. 3. shows a block diagram of an audio signal authenticator. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. This audio signal 171 ca be the output audio signal 501 from the audio signal signing module 500 of Fig. 1. This audio signal 171 is input into a block bounds detector and block divider 205, which determines block boundaries and outputs a plurality of audio blocks 201 with embedded authentication access information 212. If the authentication access information 212 has been embedded by means of watermarking, it is assumed that the blockboundaries have been previously coded into the watermark, from which they can be detected at this stage. In case the audio signal is transported by an audio container that contains the authentication access information 212 as metadata, it is assumed that the authentication access information is available per block with the block boundaries being part of the metadata.

[0082] The audio blocks 201 with embedded authentication access information are input into the access information extractor 210, which extracts the authentication access information 212 (corresponding to the authentication access information 162 from the signing modules if no manipulation has taken place) from the audio blocks 201 . The authentication access information interpreter 230 determines access information from the authentication access information 212 namely, where to find the signed authentication information 141 belonging to the current audio block 201 on a server or memory 600. The authentication information 141 retrieved from the server or memory 600 comprise firstly audio features 151 representing the audio block 201 , secondly the signature 221 and thirdly optional further information 113 like location and / or time where and / or when the signing has taken place. The signature 221 corresponds to the signature 144 generated by the signature unit 140 if no manipulation has taken place.

[0083] In a perceptual similarity analyzer 240 the audio blocks 201 and the audio features 151 are perceptually compared to each other, providing a similarity statement 241.

[0084] In the similarity analyzer 240 the audio blocks 201 can be compared to the audio features 151 of the original audio signal. The audio features can be a true copy of the original audio signal. This will however require large storage space to store the original audio signal. If the audio features relate to other characteristics of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Hence, the required audio features can be extracted from the audio signal to be authenticated. This can be performed e.g. in the similarity analyzer 240. In other words, the similarity analyzer 240 should compare the audio features (as received via the authentication access information, i.e. the audio features of the original audio signal) and corresponding audio features of the audio signal to be authenticated.

[0085] In the simplest case this can be a binary decision about whether the compared quantities are perceptually sufficiently equal or not. Alternatively, the similarity statement 241 can be more detailed, proving for instance a similarity score.An authentication access information interpreter 230 further determines from the authentication access information 212 which owner of the private key 142 has originally signed the audio signal 171 , e.g. represented by an ID. The owner of the private key 142 can be a natural person, an organization, a registered microphone to name only a few possibilities. Using this ID, a public key 143 and the identity 270 of the owner of the private key can be retrieved from an external database. The public key 143 is subsequently employed to verify the combination of the signature 221 and the authentication information 152 (the audio features 151 and the further information 113) within a signature verifier 250, proving a signature verification result 251. The perceptual similarity statement 241 and the verification result 251 are finally combined in a combiner 260 to a final authentication result 261.

[0086] Note that the processing that is here proposed for audio content can be similarly applied to any time dependent data. In particular, it can be also applied to video content, if the following modifications are applied, i.e. the microphone 100 is replaced by a video camera, the audio feature extractor 150 is replaced by a video feature extractor, the authentication access information embedding unit 160 uses a visual watermark instead an audio watermark, or uses a video container instead of an audio container and the auditory perceptual comparison 240 is replaced by a visual perceptual comparison.

[0087] Fig. 4 shows a schematic block diagram of an audio signal signing module. The audio signal signing module 500 receives a digital audio signal 121 at its audio input 502, which is inputted to the block divider 130 which outputs a sequence of plurality of audio blocks 131. The audio blocks 131 are forwarded to an audio feature extractor 150 which generates audio features 151 for each audio block 131 . In a private key signature unit 140, authentication information comprising audio features 151 of each audio block 131 together with optional further information 113 like location or time is signed with the private key 142 of a user or entity, providing signed authentication information 141. The signed authentication information 141 itself consists of the audio features 151 , the generated signature 144 and the optional further information 113. Both components, the individual audio blocks 131 and the signed authentication information 141 are required during the verification of the authentication method shown in Fig. 5. They can be in principle embedded into the same suitable audio container, but can also be distributed over different physical channels. Optionally, a user or entity identifier 145 can be added to both components, which can for instance enable the retrieval of the suitable public key for an automated verification.

[0088] Fig. 5 shows a schematic block diagram of an authenticator 200. The authenticator 200 receives an audio signal e.g. in form of audio blocks 201 which are to be verified as well assigned authentication information 141. The signed authentication information 141 comprises audio features 151 , further information 113 and the signature 221. The audio features 151 contained in the authentication information 141 are perceptually compared to the audio blocks 201 within the perceptual similarity analyzer 240, providing a similarity statement 241. In the simplest case this can be a binary decision about whether the compared quantities are perceptually sufficiently equal or not. Alternatively, the similarity statement 241 can be more detailed, providing for instance a similarity score. Further, with knowledge about the user or entity that created the signature 221 , the corresponding public key 143 and optionally the user identity 270 are obtained. Optionally, this knowledge can be derived from a user or entity identifier 145 provided together with the authentication information 141 and the audio blocks 201. The public key 143 is subsequently employed to verify the combination of the signature 221 and the authentication information 141 within the signature verification block 250, proving a signature verification result 251. The perceptual similarity statement 241 and the verification result 251 are finally combined in combiner 260 to a final authentication result 261.

[0089] According to an example, each audio signal which is captured by the microphone 100 and which is outputted by the microphone can comprise a watermark by means of which it is possible to determine the actual microphone which has captured the audio signal. If the microphone is used by several users, several user IDs can be associated to the microphone. If the microphone is registered, then it is possible to identify the respective microphone. The user IDs can also be registered. The user IDs can be dedicated IDs or sub IDs of the microphone ID. The microphone ID and / or the user IDs can be part of the watermark. By means of the sequence numbers of the audio blocks, it can be determined when part of the audio signal has been removed.

[0090] Optionally, the audio signal could be part of a video file.

[0091] Optionally, the signed authentication information can be included in an audio file (e.g., similar to ADM file format) instead of including it into a watermark. In this case the audio signal may not be modified.

[0092] Optionally, the process of generating a signed audio signal could be performed in a software solution based on a pre-existing audio recording or an audio stream. The private key can be entered to the software e.g., by dongle, text input, and authorized e.g., by biometric signal like fingerprint or face detection.The authentication process can calculate a similarity score between the audio features of the signed audio signal and the analyzed signal. This could allow to determine a likelihood that the signal is still authentic even if minor signal processing like e.g., gain has happened.

[0093] For authentication the audio signal as well as the signed authentication information is needed. The signed authentication information might be provided via a side channel. In case watermarking is used the authentication access information must be read out from the audio signal and the authentication information must be retrieved accordingly. In case the watermarking readout needs the block boundaries one approach to achieve this is to try out different offsets of block boundaries until a successful watermark readout is possible. Another approach could be to embed some synchronization signal with a watermark into the audio signal.

[0094] To enhance the security of the system a process is provided that allows to withdraw keys e.g., in case they are stolen.

[0095] The public keys for authentication can be provided in various ways. One option is to store all public keys in one centralized database to make it easy to find the needed key. To tackle potential trust issues in this central instance the database with all public keys could be provided on a distributed ledger technology. Another option is that the organization / person that uses the technology provides the public key on their own website.

[0096] Further examples of the audio features are described by Torsten Dau, Dirk Pueschel, and Armin Kohlrausch. A quantitative model of the "effective" signal processing in the auditory system. I. model structure. The Journal of the Acoustical Society of America, 99(6):3615-3622, 1996, in the following referred to as the perception model (PEMO). This paper describes a quantitative model aimed at describing how the auditory system processes acoustic signals. The focus is on developing a model that mimics the functional processing of auditory stimuli as perceived by humans, particularly in the context of complex sounds like speech or music.

[0097] Key components of the model can include peripheral processing, envelope extraction, modulation filtering and decision stage. The peripheral processing part of the model captures the initial stages of auditory processing, including the transformation of acoustic signals into neural representations through mechanisms such as the outer and middle ear filtering, as well as nonlinearities in the cochlea. The envelope extraction part relates to the extraction of the temporal envelope of sound, which is crucial for understanding amplitudemodulation and speech processing. The modulation filtering part relates to a system that filters amplitude modulations, mimicking how the auditory system is sensitive to different modulation frequencies. At the decision stage the processed auditory signal is used to make decisions about the nature of the sound, representing higher-level auditory perception.

[0098] The model was designed to be consistent with experimental psychoacoustic data, particularly in its ability to simulate the auditory system's response to amplitude-modulated sounds. This approach provides a framework for understanding the effective signal processing in human auditory perception, especially fortasks involving the detection and discrimination of complex auditory patterns.

[0099] Carrying out the comparison of the received audio blocks and the extracted or retrieved authentication information in a transformed domain (namely regarding the human auditory perception audio features) ensures that the identified differences are perceptually meaningful. In principle any similarity measure, e.g., the empirical cross-correlation coefficient or the relative error with respect to the reference, may be used to map the result of the comparison onto a numerical value.

[0100] When applying further transformations to the Mel-frequency power spectrum of an audio signal, Mel Frequency Cepstral Coefficients (MFCCs) are. MFCCs use the Mel scale, which is a perceptual scale of pitches which serves to approximate the way humans perceive sound. The Mel scale is nonlinear, such that it spaces out lower frequencies more finely and higher frequencies more coarsely, aligning more closely with human auditory perception.

[0101] Mel Frequency Cepstral Coefficients use the cepstrum as a transformation that converts a signal from the frequency domain back into a domain where the signal's rate of change can be analyzed. The idea is to capture the spectral properties of a signal in a way that emphasizes the information-carrying components.

[0102] The process to compute MFCCs from an audio signal can involve several steps: Pre-em-phasis: Boosting the high-frequency components of the signal to balance the spectrum, usually by applying a high-pass filter, Framing: The audio signal is divided into short overlapping frames, typically 20-40 milliseconds long. This is because speech signals are non-stationary, but they can be considered quasi-stationary within these short frames, and Windowing: Each frame is multiplied by a window function, such as a Hamming window, toreduce edge effects and smooth the signal. Fourier Transform: A Fast Fourier Transform (FFT) is applied to each windowed frame to convert the time-domain signal into the frequency domain. Mel Filter Bank: The resulting spectrum is passed through a series of triangular bandpass filters spaced according to the Mel scale. This step approximates how humans perceive sound frequencies. Logarithm: The logarithm of the power of each Mel-filtered spectrum is computed. This simulates the human ear's response to loudness, Discrete Cosine Transform (DCT): Finally, the logarithmic Mel spectrum is transformed using the DCT. The result is a set of coefficients called MFCCs. Typically, only the first 12-13 coefficients are kept, as these contain the most important information.

[0103] MFCCs can capture the broad spectral shape of the audio signal in a compact form, making them highly useful for machine learning algorithms in speech and audio processing. They help in distinguishing different phonemes, speaker identities, and other audio characteristics by providing a representation that is robust to variations in pitch, volume, and other factors.

[0104] The audio feature can also comprise:

[0105] - Linear Predictive Coding (LPC) which is a method that represents the spectral envelope of a digital signal of speech in compressed form, using the information from a linear predictive model. It estimates the parameters of a filter that can be used to recreate the signal. LPCs are highly effective for modeling the formants (resonant frequencies) of speech sounds.

[0106] - Perceptual Linear Predictive (PLP) Coefficients which is similar to Linear Predictive Coding but incorporates aspects of human auditory perception, such as the critical-band spectral resolution, equal-loudness curve, and intensity-loudness power law. Perceptual Linear Predictive (PLP) Coefficients aims to mimic the non-linear perception of loudness and frequency by the human ear.

[0107] - Gammatone Filterbank Features relate to filterbanks which simulate the filtering that occurs in the human cochlea. They are similar to Mel filterbanks but use Gammatone filters instead of triangular filters. The Gammatone Filterbank Features are useful for capturing the detailed frequency structure of audio, particularly in tasks involving environmental sound classification or hearing aid design.- Chroma Features, which represent the 12 different pitch classes (C, C#, D, etc.) of a musical octave. They capture harmonic and melodic characteristics of music. These are particularly useful in music information retrieval, key detection, and chord recognition tasks.

[0108] - Mel Spectrogram: While similar to MFCCs, a Mel spectrogram is the result of directly applying the Mel filterbank to a power spectrogram without further transformation (like DCT). The Mel spectrogram retains more detailed frequency information and is often used as input to deep learning models and is used in conjunction with Convolutional Neural Networks (CNNs) in tasks like audio event detection, speech recognition, and music genre classification.

[0109] - Constant-Q Transform (CQT), which is a time-frequency representation that provides a logarithmic frequency scale, similar to the Mel scale, but with a variable time resolution that matches the frequency resolution. It is particularly useful for musical applications because it provides a better representation of musical pitch than the linear FFT or even the Mel scale.

[0110] - Deep Learning-Based Features: Learned features from deep learning models, such as embeddings from a trained neural network, can also be used as an alternative to traditional hand-crafted features like MFCCs. These features are often more robust and can capture complex patterns that are not easily modeled by traditional methods.

[0111] - Spectral Subband Centroids (SSC) which captures the centroid of energy distribution in different subbands of the signal, which can be interpreted as the "center of mass" of the spectrum within each subband. It provides information about the distribution of energy across frequency bands and is sometimes used as a complement to MFCCs.

[0112] - RelAtive SpecTrAI (RASTA) Features which involves filtering the log energy of the speech signal to emphasize modulation frequencies important for speech recognition. RASTA-PLP is a combination of RASTA filtering and PLP analysis, providing robust features against noise and channel variations.

[0113] Audio features relate to an audio representation of an audio signal trying to capture relevant aspects for human auditory perception of the audio signal. The audio features should be selected to reduce the required data rate (for the transmission of the audio features) and hence storage space required on a server. Although the audio features could possibly be the original audio signal itself, preferably the audio features can be a transparently codedversion of the original signal requiring much less storage space. Another example for audio features could be the output of any model mimicking the human auditory perception. The audio features can be a Mel-frequency power spectrum of an audio signal with Mel Frequency Cepstral Coefficients (MFCCs).

[0114] In particular the audio features can be acoustical characteristics, spectral characteristics, and statistical characteristics of the audio signal.

[0115] Acoustical features can include pitch (Fundamental Frequency); prosody (Rhythm, Stress, Intonation) and / or duration and silence patterns (inconsistencies in pauses, speech timing, or breathing sounds can be used as audio features.

[0116] Spectral Features can include formant frequencies (like resonant frequencies in speech); Mel-Frequency Cepstral Coefficients (MFCCs) (the short-term power spectrum of sound and can reveal artifacts introduced during synthesis or manipulation); spectrogram analysis (a manipulation may exhibit unusual or smeared energy distribution in the time-frequency representation); and / or phase information.

[0117] Acoustical features can include statistical and signal processing features, like noise and residuals (like subtle background noises or artifacts introduced during synthesis may not match natural recordings); high-frequency content; and / or phase distortions (some manipulations may introduce phase anomalies, which can be detected with signal processing techniques).

[0118] Acoustical features can include temporal features, like jitter and shimmer (like frequency and amplitude variations); temporal coherence (like sudden transitions or changes in speech characteristics, like unnatural breaks in the signal).

[0119] Acoustical features can include behavioral or semantic Inconsistencies, like content coherence (logical inconsistencies or unnatural flow in the spoken content might be indicative of manipulation); emotion and naturalness (The emotional tone may not align with the content, or the voice may lack the natural nuances of human emotion).

[0120] The audio features can be any one of the above mentioned audio characteristics of the audio signal or a combination of them.By combining these features with machine learning or signal processing techniques, models can be trained to detect artifacts and inconsistencies indicative of deepfake audio.

[0121] Further examples which can be combined with the above examples are described in the following:

[0122] Example A1 : Audio signal authenticator (200), comprising:

[0123] a block bounds detector and block divider (205) configured to receive an audio signal (171) and to output a sequence of a plurality of audio blocks (201) with embedded authentication access information (212, 162),

[0124] an access information extractor (210) configured to extract the authentication access information (212, 162) from a current audio block (201),

[0125] an authentication access information interpreter (230) configured to extract access information on how and / or where the signed authentication information (141) belonging to the current audio block (201) is retrievable from a memory (600),

[0126] wherein the signed authentication information (141) retrieved from the memory (600) comprises a signature (144, 221) and audio features (151) belonging to the current audio block (201), and

[0127] a perceptual similarity analyzer (240) configured to compare the audio blocks (201) of the received audio signal and the retrieved audio features (151) or to compare audio features extracted from the audio blocks (201) of the received audio signal and the retrieved audio features (151) to each other providing a similarity result (241).

[0128] Example A2: Audio signal authenticator (200) according to Example A1 , further comprising:

[0129] a signature verifier (250) configured to verify a combination of the signature (144, 221) and the audio features (151) using a public key (143) and to provide a signature verification result (251),

[0130] wherein the public key (143) is associated to a private key (142) that can be used by an audio signal signing module (500) for generating an authenticable audio signal (501).

[0131] Example A3: Audio signal authenticator (200) according to Example A2, wherein the authentication access information interpreter (230) is configured to determine an identifier (145) of the user or entity who or which has originally signed the audio signal, wherein the public key (143) and an identity (270) of the user or entity can be retrieved from an external database (400) based on the determined identifier (145).Example A4: Audio signal authenticator (200) according to Example A2 or A3, further comprising

[0132] a combiner (260) configured to combine the similarity result (241) and the signature verification result (251) to a final authentication result (261).

[0133] Example A5: Method of authenticating an audio signal (171) by an authenticator (200), comprising the steps of:

[0134] receiving an audio signal (171),

[0135] outputting a sequence of a plurality of audio blocks (201) with embedded authentication access information (162),

[0136] extracting the authentication access information (212, 162) from a current audio block (201),

[0137] extracting access information (212, 162) on how and / or where signed authentication information (141) belonging to the current audio block (201) is retrievable from a memory (600),

[0138] retrieving from the memory (600) signed authentication information (141) comprising a signature (144, 221) and audio features (151) belonging to the current audio block (201), and

[0139] comparing the audio blocks (201) of the received audio signal and the retrieved audio features (151) to each other providing a similarity result (241), or

[0140] comparing audio features extracted from the audio blocks (201) of the received audio signal (171) and the retrieved audio features (151) to each other providing a similarity result (241).

[0141] Example A6 Method according to Example A5, further comprising the steps of: verifying a combination of the signature (144, 221) and the audio features (151) using a public key (143), and

[0142] providing a signature verification result (251), wherein the public key (143) is associated to a private key (142) that can be used by an audio signal signing module (500) for generating an authenticable audio signal (501)..List of reference signs 100 microphone

[0143] 102 selector

[0144] 110 microphone capsule

[0145] 111 microphone analog audio signal

[0146] 113 metadata

[0147] 120 AD converter

[0148] 121 digital audio signal

[0149] 130 block divider

[0150] 131 audio blocks

[0151] 140 private key signature unit

[0152] 141 signed authentication information

[0153] 142 private key

[0154] 143 public key

[0155] 144 signature

[0156] 145 identifier

[0157] 150 audio feature extractor

[0158] 151 audio features

[0159] 152 authentication information160 information embedder / watermark generator 161 audio blocks with embedded information 162 authentication access information

[0160] 163 watermark

[0161] 171 audio signal

[0162] 190 audio output

[0163] 200 audio signal authenticator

[0164] 205 block bounds detector and block divider 210 access information extractor

[0165] 212 authentication access information

[0166] 221 signature

[0167] 230 authentication access information interpreter 240 perceptual similarity analyzer

[0168] 241 similarity result

[0169] 250 signature verifier

[0170] 251 signature verification result

[0171] 260 combiner

[0172] 261 authentication result

[0173] 270 identity300 network

[0174] 400 server

[0175] 500 audio signal signing module 501 authenticable audio signal 502 audio input

[0176] 600 server or memory

[0177] 700 Turing tester

Claims

Claims1. Audio signal signing module (500), comprisinga Turing tester (700) configured to perform a Turing test to differentiate a human from a non-human as user,an input (502) configured to receive a digital audio signal (121),a block divider (130) configured to divide the received digital audio signal (121) into a sequence of a plurality of audio blocks (131),an audio feature extractor (150) configured to extract audio features (151) from a current audio block (131),a signature unit (140) configured to generate a signature (144) associated to the current audio block (131) by applying a private key (142) to the audio features (151) and configured to provide signed authentication information (141), which comprises the extracted audio features (151) and the signature (144), wherein the signed authentication information (141) is outputted to a memory (600),an information embedder (160) configured to generate an audio block (131) with embedded information (161) by embedding authentication access information (162) into the current audio block (131), wherein the authentication access information (162) comprises information on how and / or where the signed authentication information (141) is retrievable from the memory (600),an audio output (190) configured to output an authenticable audio signal (501) comprising a sequence of the audio blocks with embedded authentication access information (161),wherein in a first operating mode the Turing tester is configured tooutput a question or prompt which the user must answer,receive an answer in form of an audio signal captured via at least one microphone, analyze the answer captured by the microphone,compare the received answer to a correct answer of the question or prompt, and activate a second operating mode, when the received answer is correct and user has passed the Turing test,wherein the second operating mode is activated after completion of the first operating mode to generate an authenticable audio signal (501) via the input (502), the block divider (130), the audio feature extractor (150), the signature unit (140) and information embedder (160).

2. Audio signal signing module (500) according to claim 1 , whereinthe audio signal of the answer of the user is analysed and compared to an audio signal captured in the second operating mode,wherein if those two audio signals do not originate from the same person, the capturing of the audio signal can be stopped or a notification can be embedded into the audio signal indicating that the audio signal may be generated by a computer.

3. Audio signal signing module (500) according to claim 1 or 2, whereinthe result of the Turing test can be embedded into the authenticable audio signal.

4. Audio signal signing module (500) according to claim 1 , 2 or 3, whereinthe question or prompt is outputted by an independent external device.

5. Microphone (100), comprising:a microphone capsule (110) adapted to capture an audio signal (111),an AD converter (120) configured to convert the captured audio signal (111) into a digital audio signal (121), andan audio signal signing module (500) according to claim 1 , 2 or 3.

6. Method of generating an authenticable audio signal (501), comprising the steps of:in a first operating mode performing a Turing test to determine whether a user is a human bygenerating a question or prompt which the user must answer,receiving an answer in form of an audio signal captured via at least one microphone,analyzing the answer captured by the microphone,comparing the received answer to a correct answer to the question or prompt, andactivate a second operating mode, when the answer is correct and user has passed the Turing test,wherein the second operating mode is activated after completion of the first operating mode to generate an authenticable audio signal (501), the generation of the authenticable audio signal is performed byreceiving a digital audio signal (121),dividing the digital audio signal (121) into a sequence of a plurality of audio blocks (131),extracting audio features (151) from a current audio block (131),generating a signature (144) associated to the current audio block (131) by applying a private key (142) to the extracted audio features (151),providing signed authentication information (141), which comprises the audio features (151) and the signature (144),outputting the signed authentication information (141) to a memory (600), generating an audio block (131) with embedded information (161) by embedding authentication access information (162) into the current audio block (131), wherein the authentication access information (162) comprises information on how and / or where the signed authentication information (141) is retrievable from the memory (600), andoutputting the authenticable audio signal (501) comprising of a sequence of a plurality of audio blocks with embedded authentication access information (161).

7. Method of generating an authenticable audio signal (501) according to claim 6, whereinin addition to the Turing test, a biometric test is performed to determine whether a human is generating the audio signal as well as to determine which person is generating the audio signal.

8. Method of controlling an audio and / or video-conference, comprisingactivating a first operating mode before an audio and / or videoconference call to determine whether the user is a human,in the first operating mode a Turing test is performed to determine whether a user is a human bygenerating a question or prompt which the user must answer,receiving an answer in form of an audio signal captured via at least one microphone,analyzing the answer captured by the microphone,comparing the received answer to a correct answer to the question or prompt, andwherein if the Turing test has determined that the user is not a human, then this information is embedded into the audio and / or video conference signal.

9. Computer program product for authenticating an audio signal,wherein the computer program product comprises program code means causing the audio signal authenticator according to claim 1 to carry out the method of claim 6.