Method of authenticating audio signal

By embedding authentication access information and features into the audio signal and combining it with signature verification, the complexity of audio signal authentication in existing technologies is solved, achieving simple and robust audio signal authentication and malicious modification detection.

CN121889791APending Publication Date: 2026-04-17SENNHEISER ELECTRONICS GMBH & CO KG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies make it difficult to easily and robustly authenticate audio signals, especially those such as politicians' speeches, which usually require extensive manual research.

Method used

By analyzing the received audio signal, extracting and comparing the authentication access information and audio features in the audio block, embedding authentication information using watermarks or audio containers, and combining it with signature verification, automatic authentication of the audio signal is achieved.

Benefits of technology

It achieves simple and robust audio signal authentication, capable of identifying the source of audio signals and detecting malicious modifications, thus improving the reliability and efficiency of authentication results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121889791A_ABST
    Figure CN121889791A_ABST
Patent Text Reader

Abstract

Accordingly, a method of authenticating an audio signal is provided. The method comprises the steps of: analyzing a received audio signal to determine an audio block and extracting embedded authentication access information from the received audio block; analyzing the extracted authentication access information to retrieve authentication information including audio features of the original audio signal; analyzing audio blocks of the received audio signal to determine whether there is an irregular audio block in which no authentication access information is available; comparing the irregular audio block to a corresponding set of audio features that can be retrieved from memory; and notifying a user which portions of the irregular audio block are successfully authenticated. Thus, the method also enables authentication of irregular audio blocks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention relates to a method for authenticating audio signals.

[0002] Given the current ability to create deepfake videos and audio, there is an increasing need for the ability to verify or authenticate audio signals such as politicians' speeches.

[0003] To date, authenticating an individual's audio signal has been extremely difficult. Typically, extensive manual research is required to verify or authenticate such an audio signal and to verify its origin.

[0004] Therefore, the object of the present invention is to provide a method for authenticating audio signals, which can authenticate a variety of different audio signals in a simple and robust manner.

[0005] This objective is achieved by the method for authenticating audio signals according to claim 1.

[0006] Therefore, a method for authenticating audio signals is provided. This method includes the following steps: analyzing a received audio signal to determine audio blocks and extracting embedded authentication access information from the received audio blocks; analyzing the extracted authentication access information to retrieve authentication information containing audio features of the original audio signal; analyzing the audio blocks of the received audio signal to determine if there are irregular audio blocks for which no authentication access information is available; comparing the irregular audio blocks with a corresponding set of audio features that can be retrieved from memory; and notifying the user which parts of the irregular audio blocks have been successfully authenticated. Therefore, this method also enables the authentication of irregular audio blocks.

[0007] According to one aspect, if the received audio signal includes a portion having at least two audio signals, and if the irregular audio block contains a change from the first audio signal to the second audio signal, the following steps are performed: determining the last complete audio block based on the first signal, retrieving authentication information from memory corresponding to the audio block immediately following the last complete audio block, and sequentially comparing the audio features contained in the retrieved authentication information with the received irregular audio block in a forward time sequence or in pairs until a change from the first audio signal to the second audio signal is detected.

[0008] According to one aspect, if the received audio signal includes a portion having at least two audio signals, and if the irregular audio block contains a change from the first audio signal to the second audio signal, the following steps are performed: determining a first complete audio block based on the second signal, retrieving authentication information from memory corresponding to the audio block immediately preceding the first complete audio block, and sequentially comparing the audio features contained in the retrieved authentication information with the received irregular audio block in reverse chronological order, segment by segment or in pairs, until a change from the second audio signal to the first audio signal is detected.

[0009] According to one aspect, the comparison between irregular audio blocks and audio features in the authentication information is accomplished by: selecting a suitable type of audio feature for the comparison, extracting features of the selected type from the irregular audio signal block, converting the audio features in the authentication information into audio features of the selected type if the type of audio features does not correspond to the selected type, and performing a comparison based on the audio features of the selected type.

[0010] According to one aspect, a method for authenticating an audio signal is provided. The received audio signal is analyzed to determine audio blocks and information for authenticating the audio signal, such as embedded authentication access information, is extracted from the received audio blocks. The authentication access information is analyzed to determine information for retrieving authentication information from external storage. The audio blocks of the received audio signal are analyzed to determine if there are audio blocks for which no authentication information is available. The portion of the received audio signal without authentication information is compared with the corresponding portion of the original audio signal stored on a server. If authentication cannot be successfully performed, the user is notified.

[0011] The authentication process can be performed as described above, or in any other manner. For this invention, it is only necessary to receive information that authentication cannot be performed on certain audio blocks. The method then processes those audio blocks for which no authentication information is available.

[0012] The authentication information required to authenticate the received audio signal can be extracted directly from the audio block (where the authentication information is embedded in the audio block, for example, via a watermark or a suitable audio container), or indirectly where only the authentication access information is embedded in the audio block (again, for example, via a watermark or a suitable audio container). The authentication access information includes information about where and / or how the authentication information can be retrieved.

[0013] Several methods exist for authenticating audio signals. Authentication information is provided along with the audio signal. For example, authentication information can be embedded in the audio signal using a watermark, or the authentication information can be provided externally, while the information required to access the authentication information can be embedded in the audio signal. The audio signal is divided into audio blocks, and these audio blocks form the basis for further authentication processing. The length of the audio block can be encoded in the block header.

[0014] The received or stored audio signal to be authenticated can be a signal from a single source, or it can be an audio signal composed of portions of audio signals from multiple different sources (e.g., music recordings, podcasts, radio programs, voice recordings, etc.) or different recordings of the same audio signal (e.g., multiple microphones). Therefore, the authentication process must be able to authenticate composite audio signals. To achieve this, authentication information must be provided for each audio block. However, if the audio signal consists of different clips from different audio signals, the clips may begin and end at locations that do not correspond to the block boundaries of the audio block used for authentication. Therefore, authentication information for the entire audio block or for a portion of the audio block may not be available. For example, since the block carrying the watermark is only partially part of the audio composition, it may not be possible to decode the watermark using authentication access information.

[0015] To handle such a scenario based on the example, identify those parts of the audio signal where no authentication information is available.

[0016] According to one aspect, an authenticator is provided that can perform the steps of the authentication method described above.

[0017] According to one aspect, a computer program product for authenticating audio signals is provided, wherein the computer program product includes program code units that cause an audio signal authenticator to perform the above-described authentication method.

[0018] As an example, authentication of an audio signal can be performed by detecting block boundaries based on specific information in a watermark or audio container, and by extracting authentication access information (e.g., a link, including a UUID and a user-based unique block number) pointing to block-specific authentication information (e.g., audio features) from the audio watermark or directly from the audio container. The user's public encryption key (e.g., identified by the UUID in the link) can be used to verify the authenticity of the authentication information (e.g., audio features). The audio signal block is compared to the audio features to verify its authenticity.

[0019] Using the method for authenticating audio signals described above, malicious modification of the audio signal can be detected by comparing the audio block with the authentication information associated with the audio block. Modification of the audio signal to be authenticated and / or modification of authentication information stored externally can also be detected by examining the signature of the authentication information. If the audio block of the audio signal is modified and the authentication access information is also altered, this will be noticed by the authentication process because the signature of the audio signal and / or the authentication information is that of another person or device.

[0020] According to one aspect of the invention, authentication information (e.g., metadata) may include a microphone identifier, microphone location / position information, recording time and date, and / or microphone model, etc. Metadata may also include, for example, audio characteristics, recording time, serial number, and recording location. Optionally, the metadata of an audio block may be the microphone user, date, time, GPS location, etc. Authentication information (e.g., metadata) may include audio characteristics representing audio signals based on human perception or audio blocks of audio signals, or a transparently encoded version of the original audio signal, to reduce the storage space required by the server. Therefore, the probability of reliable authentication results is increased.

[0021] The authenticity check of the received audio signal can be performed by the decoder, which can be implemented, for example, on a cloud service, computer, tablet or smart device.

[0022] Optionally, the metadata of the audio block can include the microphone user, date, time, GPS location, etc. This increases the probability of reliable authentication results.

[0023] Audio features refer to the audio representation of an audio signal that attempts to capture relevant aspects of human auditory perception of the audio signal. Audio features should be selected to reduce the required data rate (for transmitting the audio features) and thus reduce the storage space required on the server. While the audio features may be the original audio signal itself, preferably, they may be a transparently encoded version of the original signal, which requires significantly less storage space. Another example of an audio feature could be the output of any model that mimics human auditory perception. Audio features could be the Mel-frequency power spectrum and Mel-frequency cepstral coefficients (MFCCs) of the audio signal.

[0024] According to one perspective, if the authentication process does not provide an explicit indication of the authenticity of an audio signal or audio block, the authentication information stored externally can be used to manually authenticate the audio signal or audio block.

[0025] According to the example, the actual audio features extracted from the audio block and optional additional metadata are stored in a signed manner, for example, on a server. A unique (explicit) link to the metadata stored on the server is embedded in a watermark or in a suitable audio container. The link may include a user-based unique block number for the corresponding audio block along with a unique user identifier (UUID). The audio features involve an audio representation of an audio signal that attempts to capture relevant aspects of human auditory perception, and preferably at a reduced data rate and thus a smaller storage space required on the server. While the audio features may be the original audio signal itself, preferably, the audio features may be a transparently encoded version of the original signal that requires significantly less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.

[0026] To authenticate the audio signal of an audio block, block boundaries can be detected based on a watermark or specific information in the audio container. Then, a link (e.g., including a UUID and a user-based unique block number) pointing to block-specific, signed authentication information (including audio features, optional metadata, and a signature) is extracted from the audio watermark or directly from the audio container. The combination of authentication information and signature is verified using the user's public cryptographic key (e.g., identified by the UUID in the link). Accordingly, it is verified whether the authentication information has been signed by the claimed audio signal source. The audio signal block is compared with audio features to verify perceptual similarity, and thus the authenticity of the audio signal block. For perceptual comparison, audio features can be determined from the received audio signal block, and these features can be directly compared with audio features from, for example, signed authentication information stored on a server.

[0027] The proposed authentication method prevents any malicious modification to the actual audio signal by comparing the audio signal block with the signed authentication information. Authentication may fail if the comparison allows a limited amount of modification, but the allowed threshold is exceeded. Furthermore, if the authentication information on the server is altered accordingly along with the malicious modification of the audio signal block, this can be detected by a mismatch between the authentication information and the signature associated with it. Additionally, if, along with the malicious modification of the audio signal block, the link (within the watermark or audio container) is changed to point to a different (but appropriate) signed authentication information on the server, signature verification and subsequent comparison may indeed succeed, but the signature will be that of a different user—provided the attacker cannot obtain the original user's private key used for the encrypted signature.

[0028] Assuming that the (automatic) comparison between the audio signal to be authenticated and the authentication information is performed in a manner that allows for minor modifications (e.g., through perceptual encoding), and the result of the comparison is ambiguous (i.e., no explicit statement can be made about the authenticity of the audio signal), then a “manual” perceptual comparison can be performed using a link to the authentication information (where the included audio features will correspond to the original audio file or its perceptually encoded version) to increase trust in the method.

[0029] By providing audio block-specific links (e.g., via UUID and user-based unique block numbers), it becomes possible to have the advantage (e.g., in watermarking scenarios) that audio files can be cut at any time without the cutting tool being aware of authentication processing and the embedded links. Alternatively, in cases where the audio container is used for the joint transmission of audio and links, file-based links may be embedded all at once in the included header to reduce data rate. This link may include a UUID and a location within metadata (e.g., a sample index of the original file) to indicate how the audio file in question is aligned with a reference file on the server.

[0030] These and other aspects of the invention are described in more detail with reference to the following accompanying drawings.

[0031] Figure 1 A block diagram of the audio signal signature module is shown.

[0032] Figure 2 A diagram showing the overall audio signature and authentication workflow is provided.

[0033] Figure 3 A block diagram of an audio signal authenticator is shown.

[0034] Figure 4 A block diagram of the audio signal signature module is shown, and

[0035] Figure 5 A block diagram of an audio signal authenticator is shown.

[0036] Figure 6 A block diagram of the audio signal signature module is shown.

[0037] Figure 7 A block diagram of an audio signal authenticator is shown, and

[0038] Figure 8 A diagram of an exemplary synthesized audio file to be authenticated is shown.

[0039] Based on the example, authentication is performed on an audio signal. The audio signal is authenticated based on information embedded in an audio block. This embedded information can be, for example, authentication information directly embedded in the audio block using a watermark, or authentication access information. Authentication access information includes information about where and how stored authentication information associated with the audio block can be retrieved from storage (such as a server). In this case, the authentication information is not directly embedded in the audio block. Figures 1 to 5 Examples involving the use of authentication information, Figures 6 to 8 Examples involving the use of authentication access information.

[0040] Based on the example, authentication of audio blocks is performed. As described in more detail below, further analysis of unauthenticated audio blocks is only possible if not all audio blocks are authenticated.

[0041] To better understand the authentication of audio signals, we will first describe the generation of authenticable audio signals.

[0042] Figure 1 A diagram of an audio signal signature module is shown. The audio signal signature module 500 receives an input audio signal 121 from a microphone 100 or another audio source, such as an audio recorder, via an audio input terminal 502, and outputs an output audio signal 501 via an audio output terminal 190. The output audio signal 501 may be based on the input audio signal 121. The microphone 100 may include at least one microphone capsule 110 and an analog-to-digital converter (ADC) 120. The at least one microphone capsule 110 can capture an audio signal and output the captured analog audio signal 111. The captured analog audio signal 111 can be forwarded to the ADC 120, which can digitize the audio signal 111 and output a digital audio signal 121. The audio signal signature module 500 may be implemented as a device separate from the microphone 100 or may be included within the microphone 100.

[0043] The audio signal signature module 500 can receive audio signal 121 from microphone 100 or from other sources, which in Figure 1The input is indicated by selector 101. Digital audio signal 121 is input to block divider 130, which divides the audio signal 121 into audio blocks 131 and outputs a sequence of multiple audio blocks 131. Optionally, if the audio signal 121 has already been divided into audio blocks 131, block divider 130 can be omitted. The sequence of audio blocks 131 is forwarded to audio feature extractor 150, which generates audio features 151 for each audio block 131 in the audio block sequence. Private key signing unit 140 receives authentication information (e.g., metadata) including the audio features 151 of each audio block 131 and optional other information 113 such as location or time, and signs the authentication information using private key 142 to provide signed authentication information 141, which can be stored, for example, on memory 600 located on an external server. Signed authentication information 141 may include authentication information 152 (see... Figure 2 The authentication information 152 includes audio features 151 and optional other information 113, as well as a signature 144 generated by the private key signing unit 140. Figure 2 Alternatively, the authentication information 141 may be stored in internal memory or any other memory, as long as it can be accessed by the authenticator that authenticates the received audio signal.

[0044] The signed authentication information 141 itself includes authentication information, namely (a) audio feature 151, optionally (b) other information 113, and (c) signature 144. Individual audio blocks 131 are fed in parallel, for example, to an information embedder 160 for embedding authentication access information 162 into, for example, each audio block 131. Authentication access information 162 includes information about where and / or how the authentication information 152 (or more precisely, the signed authentication information 141) for that particular audio block 131 can be accessed, for example, at a memory 600 located on a server. This embedding can be accomplished, for example, by means of a watermark.

[0045] Therefore, the information embedder 160 can be implemented as a watermark generator that generates watermark 163. Alternatively, the authentication access information 162 can be embedded into a suitable audio container adjacent to the actual audio block 131. Accordingly, the information embedder 160 outputs the audio block 161 with the embedded authentication access information 162. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at the audio output terminal 190. The output audio signal 501 can be transmitted, for example, via network 300 (…). Figure 2 The output audio signal 501 can be broadcast or stored. In particular, the output audio signal 501 can even be further modified to a very small extent, including modification types such as sample rate conversion, perceptual compression, and trimming.

[0046] When the information embedder 160 is a watermark generator, its output 161, instead of the original audio block 131, can alternatively be input to the audio feature extractor 150, which in Figure 2 The selection is indicated by selector 102. This alternative offers the advantage of extracting audio features based on the final signal 161 output by the audio signal signature module 500, rather than the unoutput intermediate signal 131. Since the watermark generator is designed to perceptually alter the audio signal as little as possible, both input options (131 or 161) are reasonable for the audio feature extractor 150.

[0047] Audio features involve an audio representation that attempts to capture relevant aspects of human auditory perception, while preferably reducing the required data rate and thus the storage space required on the server or memory 600. While audio features can be the original audio signal itself, preferably, they can be a transparently encoded version of the original signal that requires far less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.

[0048] The audio signal 501 at the output of the audio signal module 500 (containing authentication access information or an embedded watermark within the audio container) can be stored or transmitted. Because the watermark can be embedded in, for example, essentially each audio block 131, each audio block 131 can be authenticated individually. The watermark 103 can also be embedded only in some audio blocks within the audio block set. In addition to the authentication access information, the watermark 163 can also include other information such as block boundaries. This information can be used when authenticating the audio signal to determine the block length.

[0049] The audio watermark 163 can be a unique identifier embedded in the audio signal and, for example, previously used to identify copyright information of the audio signal. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, the embedded watermark will not be altered when the audio signal with the embedded watermark is copied, stored, or transmitted. The same reasoning applies to authentication access information included in the audio container.

[0050] Figure 2A diagram illustrating the overall audio signing and authentication workflow is provided. An audio signal is captured by microphone 100, which outputs a digital audio signal 121. Audio features 151 are extracted from this audio signal 121 and signed using, for example, the private key 142 of audio signal signing module 500. Audio features 151, together with a digital signature 144 generated by signing unit 140, constitute signed authentication information 141. The signed authentication information 141 can be forwarded to a server or storage 600, where it is stored for later retrieval. Information regarding where and / or how to retrieve the signed authentication information 141, namely authentication access information 162, is provided to audio signal signing module 500. Audio signal signing module 500 uses the authentication access information 162 and embeds it into audio blocks 131 of the audio signal. Therefore, the output signal 501 of audio signal signing module 500 includes audio blocks 131 with the embedded authentication access information. The audio block 131 of the audio signal with embedded authentication access information 162 is output by the audio signal signature module 500 and can be stored or distributed, for example, via network 300.

[0051] This distribution may involve modification or tampering with authentication information 152 and / or audio signal 121. Public key 143 is used to verify the signed authentication information 141. This verification indicates potential modification of the audio features. The integrity of the audio signal can be verified by extracting audio features from the received signal and comparing them with the features included in the signed authentication information 141. The final authentication result 261 is determined jointly by these two checks.

[0052] An audio signal signature module 500 (e.g., implemented as part of microphone 100) is used to capture audio signals. Alternatively, the audio signal signature module 500 can receive audio signals. The audio signal can also be an audio signal pre-captured by other means. See reference... Figure 1The captured audio signal is divided into multiple audio blocks, and authentication information for at least one audio block is signed based on the private key 142 of the audio signal signing module 500. The signed authentication information 141 is output by the audio signal signing module 500 and can be stored on a server or memory 600. Authentication access information 162, which is information about where the authentication information can be retrieved from the server or memory 600, can be provided to the audio signal signing module 500. The authentication access information 162 is embedded in the audio blocks received or captured by the audio signal signing module 500. Therefore, the audio signal signing module 500 performs the following operations: a) outputting an output audio signal 501 including audio blocks with embedded authentication access information to the network 300 for storage or forwarding, and b) outputting the signed authentication information 141 to a server or memory 600 where the information can be stored.

[0053] Alternatively, authentication access information can be embedded into audio blocks of the audio signal using a watermark. The watermark may relate to authentication access information associated with the location in memory 600 where the authentication information is stored. This authentication information can be obtained from the referenced above. Figure 1 The described audio feature extractor 150 generates audio blocks of an audio signal with embedded authentication access information. These blocks can be transmitted or stored via network 300. Authenticator 200 can receive the audio signal 501 with an embedded watermark and extract the authentication access information. Authentication information is retrieved from server 600 based on the authentication access information. The signature of the authentication information stored on server or storage 600 can be verified based on the public key 143 associated with the audio signal signature module 500, which can be stored, for example, on server 400 or a distributed ledger.

[0054] Figure 3 A block diagram of an audio signal authenticator is shown. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. This audio signal 171 may be from... Figure 1 The audio signal signature module 500 outputs an audio signal 501. This audio signal 171 is input to a block boundary detector and a block segmenter 205, which determine block boundaries and output multiple audio blocks 201 with embedded authentication access information 212. If the authentication access information 212 has already been embedded in the watermark, it is assumed that the block boundaries have been previously encoded into the watermark, and the block boundaries can be detected from the watermark at this stage. In the case where the audio signal is transmitted by an audio container that includes the authentication access information 212 as metadata, it is assumed that the authentication access information is available for each block, where the block boundaries are part of the metadata.

[0055] An audio block 201 with embedded authentication access information is input to an access information extractor 210, which extracts authentication access information 212 from the audio block 201 (which, if tampered with, corresponds to authentication access information 162 from the signature module). An authentication access information parser 230 determines the access information based on the authentication access information 212, i.e., the location of the signed authentication information 141 belonging to the current audio block 201 can be found on the server or memory 600. The authentication information 141 retrieved from the server or memory 600 first includes audio features 151 representing the audio block 201, secondly includes a signature 221, and further includes optional other information 113, such as the location and / or time of the signature occurrence. If tampered with, the signature 221 corresponds to the signature 144 generated by the signature unit 140.

[0056] In the perceptual similarity analyzer 240, audio blocks 201 and audio features 151 are perceptually compared to each other, thereby providing a similarity statement 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity statement 241 can be more detailed, such as providing a similarity score.

[0057] In a similarity analyzer, audio blocks can be compared to the audio features of the original audio signal. These audio features can be true copies of the original audio signal. However, this would require significant storage space to store the original audio signal. If the audio features relate to other characteristics of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Thus, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within the similarity analyzer itself. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.

[0058] The authentication access information parser 230 also determines, based on the authentication access information 212, which owner of the private key 142 initially signed the audio signal 171, for example, by identifying the owner through an ID. The owner of the private key 142 could be a natural person, an organization, a registered microphone, and a few other possibilities are listed here. Using this ID, the identity 270 of the public key 143 and the private key owner can be retrieved from an external database. The public key 143 is then used within the signature verifier 250 to verify the combination of the signature 221 and the authentication information 152 (audio feature 151 and other information 113), thereby providing a signature verification result 251. The perceived similarity claim 241 and the verification result 251 are finally combined in the combiner 260 to form the final authentication result 261.

[0059] Note that the processing proposed here for audio content can be similarly applied to any time-related data. In particular, the above processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark or uses a video container instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison.

[0060] Figure 4 A schematic block diagram of an audio signal signing module is shown. The audio signal signing module 500 receives a digital audio signal 121 at its audio input 502. This digital audio signal 121 is input to a block segmenter 130, which outputs a sequence of multiple audio blocks 131. The audio blocks 131 are forwarded to an audio feature extractor 150, which generates an audio feature 151 for each audio block 131. In a private key signing unit 140, the user's or entity's private key 142 is used to sign the audio features 151, including each audio block 131, and optional other information 113 such as location or time authentication information, thereby providing signed authentication information 141. The signed authentication information 141 itself includes the audio features 151, the generated signature 144, and optional other information 113. Figure 5 The authentication method shown requires two components during verification: individual audio blocks 131 and signed authentication information 141. In principle, they can be embedded in the same suitable audio container, but they can also be distributed across different physical channels. Optionally, a user or entity identifier 145 can be added to these two components, which can, for example, enable the retrieval of a suitable public key for automatic verification.

[0061] Figure 5A schematic block diagram of authenticator 200 is shown. Authenticator 200 receives, for example, an audio signal in the form of an audio block 201 to be verified, and signed authentication information 141. The signed authentication information 141 includes audio features 151, other information 113, and a signature 221. Within a perceptual similarity analyzer 240, the audio features 151 contained in the authentication information 141 are perceptually compared with the audio block 201, thereby providing a similarity claim 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity claim 241 can be more detailed, thereby providing, for example, a similarity score. Furthermore, using knowledge about the user or entity that created the signature 221, a corresponding public key 143 and an optional user identity 270 are obtained. Optionally, this knowledge can be derived from a user or entity identifier 145 provided along with the authentication information 141 and the audio block 201. Subsequently, public key 143 is used to verify the combination of signature 221 and authentication information 141 within signature verification block 250, thereby providing signature verification result 251. The perceived similarity claim 241 and verification result 251 are finally combined in combiner 260 to form final authentication result 261.

[0062] In the following Figures 6 to 8 The document describes how to generate an authenticable audio signal by embedding authentication information into an audio block, and how to authenticate the audio signal based on the embedded authentication information. As described below, this authentication can be used for authenticating audio signals.

[0063] Figure 6 A schematic diagram of the audio signal signature module is shown. Figure 6 The audio signature module described is merely an example. The audio signal signature module 500 receives an input audio signal 121 from the microphone 100 or another audio source, such as an audio recorder, via an audio input terminal 502, and outputs an output audio signal 501 via an audio output terminal 190. The output audio signal 501 may be based on the input audio signal 121. The microphone 100 may include at least one microphone capsule 110 and an analog-to-digital converter (ADC) 120. The at least one microphone capsule 110 can capture an audio signal and output the captured analog audio signal 111. The captured analog audio signal 111 can be forwarded to the ADC 120, which can digitize the audio signal 111 and output a digital audio signal 121. The audio signal signature module 500 may be implemented as a device separate from the microphone 100 or may be included within the microphone 100.

[0064] The audio signal signature module 500 can receive audio signal 121 from microphone 100 or from other sources, which in Figure 6The input is indicated by selector 101. A digital audio signal 121 is input to a block divider 130, which divides the audio signal 121 into audio blocks 131, each with a block length, and outputs a sequence of multiple audio blocks 131. The block divider 130 can divide the audio signal 121 into multiple audio blocks. Optionally, each audio block receives a block number, which can be a subsequent block number. Optionally, if the audio signal 121 has already been divided into audio blocks 131, the block divider 130 can be omitted. The sequence of audio blocks 131 is forwarded to an audio feature extractor 150, which generates an audio feature 151 for each audio block 131 in the sequence. A private key signing unit 140 receives authentication information (e.g., metadata) including the audio features 151 of each audio block 131 and optional other information 113 such as location or time, and signs the authentication information with a private key 142, thereby providing signed authentication information 141. The signed authentication information 141 may include authentication information 152, which includes audio features 151 and optional other information 113, as well as a signature 144 generated by the private key signing unit 140.

[0065] The signed authentication information 141 itself includes authentication information, namely (a) audio feature 151, optionally (b) other information 113, and (c) signature 144. Each audio block 131 is fed in parallel, for example, to an information embedder 160 for embedding the signed authentication information 141 into, for example, each audio block 131. Embedding can be accomplished, for example, by means of a watermark.

[0066] Therefore, the information embedder 160 can be implemented as a watermark generator that generates watermark 163. Alternatively, the signed authentication information 141 can be embedded into a suitable audio container next to the actual audio block 131. Accordingly, the information embedder 160 outputs the audio block 161 with the embedded signed authentication information 141. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at the audio output terminal 190. The output audio signal 501 can be transmitted, for example, via network 300 (…). Figure 7 The output audio signal 501 can be broadcast or stored. In particular, the output audio signal 501 can even be further modified to a very small extent, including modification types such as sample rate conversion, perceptual compression, and trimming.

[0067] When the information embedder 160 is a watermark generator, its output 161, instead of the original audio block 131, can alternatively be input to the audio feature extractor 150 via the selector 102. This alternative offers the advantage of extracting audio features based on the final signal 161 output by the audio signal signature module 500, rather than the unoutput intermediate signal 131. Since the watermark generator is designed to perceptually alter the audio signal as little as possible, both input options (131 or 161) are reasonable for the audio feature extractor 150.

[0068] The audio signal 501 at the output of the audio signal module 500 (containing authentication access information or an embedded watermark within the audio container) can be stored or transmitted. Because the watermark can be embedded in, for example, essentially each audio block 131, each audio block 131 can be authenticated individually. The watermark 103 can also be embedded only in some audio blocks within the audio block set. In addition to the authentication access information, the watermark 163 can also include other information such as block boundaries. This information can be used when authenticating the audio signal to determine the block length.

[0069] The audio watermark 163 can be a unique identifier embedded in the audio signal and, for example, previously used to identify copyright information of the audio signal. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, the embedded watermark will not be altered when the audio signal with the embedded watermark is copied, stored, or transmitted. The same reasoning applies to authentication access information included in the audio container.

[0070] Figure 7 It shows the relationship with Figure 6 The diagram below shows the audio signal authenticator 200. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. This audio signal 171 may be from... Figure 6 The audio signal 171 is output by the audio signal signature module 500. This audio signal 171 is input to the block boundary detector and block segmenter 205, which determines the block boundaries and outputs multiple audio blocks 201 with embedded authentication information.

[0071] The block boundary detector and block segmenter 205 detect block boundaries, for example, using information from the block header. Since the block length of audio blocks may vary, the block length should be determined for each audio block or for a group of audio blocks with the same block length.

[0072] Audio block 201 is input to authentication information extractor 206, which extracts signed authentication information 141 embedded in the audio block. The signed authentication information 141 may include audio features 151, a signature 144, and optional other information 113, such as the location and / or time of the signature occurrence. If it has not been tampered with, the signature 221 corresponds to the signature 144 generated by signature unit 140.

[0073] In the perceptual similarity analyzer 240, audio blocks 201 and audio features 151 are perceptually compared to each other, thereby providing a similarity statement 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity statement 241 can be more detailed, thus providing, for example, a similarity score.

[0074] In a similarity analyzer, audio blocks can be compared to the audio features of the original audio signal. These audio features can be true copies of the original audio signal. However, this would require significant storage space to store the original audio signal. If the audio features relate to other characteristics of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Thus, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within the similarity analyzer itself. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.

[0075] The authentication information extractor 206 also determines, based on the authentication information 141, which owner of the private key 142 initially signed the audio signal 171, for example, by identifying the owner through an ID. The owner of the private key 142 could be a natural person, an organization, a registered microphone, and a few other possibilities are listed here. Using this ID, the identity 270 of the public key 143 and the private key owner can be retrieved from an external database. Subsequently, the signature 144 and the authentication information 152 (audio feature 151 and other information 113) are verified within the signature verifier 250 using the public key 143, thus providing a signature verification result 251. The perceived similarity claim 241 and the verification result 251 are finally combined in the combiner 260 to form the final authentication result 261.

[0076] Note that the processing proposed here for audio content can be similarly applied to any time-related data. In particular, the above processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark, or uses a video container instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison.

[0077] Figure 8 A diagram is shown of an exemplary received audio file FA comprising portions having two distinct audio signals A and B. Audio signal B may optionally be a segment of signal A. Specifically, FA includes audio blocks A1 to A3 from the first audio signal A and audio blocks B1 and B2 from the second audio signal B. Since audio blocks A1, A2, and B2 have been received in their entirety along with their corresponding watermarks, the authenticator 300 is able to determine the validity and authenticity of audio blocks A1, A2, and B2. However, the portion C of the received audio signal FA that ends after audio block A2 and begins before audio block B2 does not include any valid extractable watermark, and therefore no authentication information is available for this portion of the audio file FA.

[0078] If the received audio signal includes clips from different audio signal sources, where these clips originate from, correspond to, or do not correspond to audio blocks from the first audio signal source A or the second audio signal source B, then the authentication process will encounter problems. In this case, there are no valid watermarks WM1, WM2, and WM3, and therefore the watermarks cannot be extracted.

[0079] The problem in attempting to verify block C of the synthesized audio signal FA (i.e., the part that begins after audio block A2 and ends before audio block B2) is that it is unclear where audio block A3 ends and where audio block B1 begins. To determine this boundary point, reference audio features included in the authentication information of blocks A3 and B1 are retrieved from server 600 and compared with audio signal block C. Here, block A3 is the successor block to block A2, and therefore the corresponding authentication information can be found on server 600. Similarly, here, block B1 is the preceding block to block B2, and therefore the corresponding authentication information can also be found on server 600.

[0080] For comparison, a suitable type of audio feature can be selected. Then, audio features of the selected type are extracted from audio signal block C. Furthermore, audio features included in the authentication information of blocks A3 and B1 are retrieved from server 600 and transformed into audio features of the selected type. If the type of audio feature included in the authentication information of either block A3 or B1 corresponds to the selected audio feature type, the transformation can be omitted. Typically, audio features are calculated for blocks with a relatively short duration compared to the duration of individual blocks A3 and B1; for example, 20 ms for MFCC, compared to a duration of 1 s for A3 and B1. Therefore, a feature sequence is obtained for the analyzed audio signal block C, where each feature corresponds to a sub-block.

[0081] Therefore, the comparison is performed as follows: Audio features starting from audio signal block C are compared in forward time pairs with reference audio features that may be transformed, starting from block A3. We assume that the pairwise comparisons are successful for all sub-blocks up to the Nth sub-block, where the audio features differ too much from each other. In the next step, the audio features of signal block C are compared in pairs with the reference audio features of block B1; however, this time, the comparisons begin from the end of blocks C and B1 and proceed in reverse time.

[0082] If the comparison fails for any sub-block immediately following the Nth sub-block of audio block C, it suggests that tampering has occurred within audio block C. Otherwise, if the comparison succeeds for all compared sub-blocks immediately following the Nth sub-block of audio block C, it can be stated that the original audio signal block A3 has been cut within the Nth sub-block, and further tampering may exist only within the Nth sub-block of block C. Furthermore, this also means that all sub-blocks in audio block C except the Nth sub-block have been successfully verified. Since the duration of a sub-block is relatively short, tampering can be considered irrelevant to the conclusion that the entire audio block C has been successfully verified. If the cut location is reported, it can be combined with a hint regarding possible tampering within the relevant sub-block.

[0083] According to the example, the authenticator includes an input for receiving a received audio signal, a block information extractor, an authentication information extractor, a verification unit, and a determination unit that determines that no authentication information is available in the received audio signal. The block information extractor may include a block boundary extractor and a watermark extractor.

[0084] A block information extractor analyzes the received audio signal to determine block boundaries, and a watermark extractor determines whether a watermark exists in the received audio signal. Block boundaries are important because a watermark is typically associated with each audio block. Watermark extraction can only be performed if the correct block boundaries have been determined. Once the correct block boundaries have been determined, the watermark extractor can extract the watermark. The watermark may, for example, include authentication access information that can be used to retrieve authentication information from an authentication information server. This authentication information can be used to authenticate or verify the received audio signal. Preferably, the authentication information is available for each audio block.

[0085] The authentication information extractor can extract authentication information from the authentication information server based on access information extracted from the watermark. The verification unit can verify audio blocks based on the associated authentication information. The determination unit can be used to determine whether there are any audio blocks in the received audio signal that do not contain authentication information.

[0086] One reason for the absence of authentication information might be as follows: Figure 8 The described scenario. In this case, the authenticator 300 may include the execution reference. Figure 8 The iterative unit 360 of the described process.

[0087] Therefore, using the authentication process described above, it is also possible to authenticate those portions of the received audio signal that lack direct authentication information. However, this requires the original audio source signal A and the original audio source signal B to be stored in the authentication information server, i.e., stored in the audio signal memory 620. However, it is expected that the number of portions of the received audio signal lacking direct authentication information should be very small.

[0088] The authentication process described above is advantageous because it does not require large bandwidth—this is because only those audio blocks from the audio signal source in question, which lack actual authentication information, are extracted. The authentication process only requires that authentication information (e.g., audio characteristics or the original audio signal) be stored on a server or at a location where it can be extracted.

[0089] The transmission of the entire reference signal is not required.

[0090] Therefore, the authentication process described above is a robust and easy-to-implement authentication process.

[0091] In another example, portions of the received audio signal lacking authentication information may originate from audio sources or signals that do not provide authentication information. If authentication information cannot be extracted by the authenticator 300, the authenticator can output this information to the user. For example, the authenticator can output a representation of these portions of the received audio signal in a different color, such as red. Accordingly, the authenticator can signal to the user that this portion of the received audio signal does not have sufficient authentication information and that the authentication process cannot be effectively performed.

[0092] Since no authentication information is available, the audio block in question may be from a valid audio signal or from a malicious audio signal.

[0093] In a similarity analyzer, audio blocks can be compared to the audio features of the original audio signal. These audio features can be true copies of the original audio signal. However, this would require significant storage space to store the original audio signal. If the audio features relate to other characteristics of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Thus, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within the similarity analyzer itself. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.

[0094] Torsten Dau, Dirk Pueschel, and Armin Kohlrausch, in “A quantitative model of the “effective” signal processing in the auditory system. I. model structure”—The Journal of the Acoustical Society of America, 99(6): 3615-3622, 1996—described other examples of audio features, hereinafter referred to as the Perceptual Model (PEMO). This paper describes a quantitative model designed to describe how the auditory system processes sound signals. The focus is on developing models that mimic the functional processing of auditory stimuli by humans, particularly in contexts involving complex sounds such as speech or music.

[0095] The key components of this model can be categorized into peripheral processing, envelope extraction, modulation filtering, and a decision-making stage. The peripheral processing section captures the initial stages of auditory processing, including converting acoustic signals into neural representations through mechanisms such as outer and middle ear filtering and nonlinearities in the cochlea. The envelope extraction section involves extracting the temporal envelope of the sound, which is crucial for understanding amplitude modulation and speech processing. The modulation filtering section involves a system for filtering amplitude modulation, mimicking the auditory system's sensitivity to different modulation frequencies. In the decision-making stage, the processed auditory signals are used to make decisions about the properties of the sound, thus representing a higher level of auditory perception.

[0096] The model was designed to align with experimental psychoacoustic data, particularly its ability to simulate the auditory system's response to amplitude-modulated sounds. This approach provides a framework for understanding effective signal processing in human auditory perception, especially for tasks involving the detection and discrimination of complex auditory patterns.

[0097] Performing a comparison of the received audio blocks with the extracted or retrieved authentication information in the transform domain (i.e., in relation to the audio features perceived by human hearing) ensures that the identified differences are perceptually meaningful. In principle, any similarity measure (e.g., empirical cross-correlation coefficient or relative error relative to a reference quantity) can be used to map the comparison results to a numerical value.

[0098] When a further transformation is applied to the Mel frequency power spectrum of the audio signal, Mel frequency cepstral coefficients (MFCCs) are obtained. MFCCs use the Mel scale, a perceptual scale for pitch used to approximate how humans perceive sound. The Mel scale is non-linear, allowing for finer division of lower frequencies and coarser division of higher frequencies, thus better aligning with human auditory perception.

[0099] Mel-frequency cepstral coefficients use the cepstral spectrum as a means of transforming a signal from the frequency domain back to a domain where the rate of change of the signal can be analyzed. This idea emphasizes the information-carrying components to capture the spectral characteristics of a signal.

[0100] The process of calculating MFCC from an audio signal can include several steps: pre-emphasis: typically achieved by applying a high-pass filter to enhance the high-frequency components of the signal to balance the spectrum; framing: dividing the audio signal into short overlapping frames, typically 20 to 40 milliseconds long, because speech signals are non-stationary but can be considered quasi-stationary within these short frames; windowing: multiplying each frame by a window function, such as a Hamming window, to reduce edge effects and smooth the signal; Fourier transform: applying a Fast Fourier Transform (FFT) to each windowed frame to convert the time-domain signal to the frequency domain; Mel filter bank: passing the resulting spectrum through a series of triangular bandpass filters spaced according to the Mel scale, a step that approximates how humans perceive sound frequencies; logarithmic calculation: calculating the logarithm of the power of each Mel-filtered spectrum, which simulates the human ear's response to loudness; Discrete cosine transform (DCT): finally, using the DCT to transform the log-Mel spectrum. The result is a set of coefficients called MFCC. Typically, only the first 12 to 13 coefficients are retained, as these contain the most important information.

[0101] MFCC can capture the wide spectral shape of an audio signal in a compact form, making it highly useful for machine learning algorithms in speech and audio processing. By providing robust representations of variations in pitch, volume, and other factors, MFCC helps distinguish different phonemes, speaker identity, and other audio characteristics.

[0102] Audio features may also include:

[0103] 1. Linear Predictive Coding (LPC) is a method that uses information from a linear predictive model to represent the spectral envelope of a digital speech signal in compressed form. LPC estimates can be used to estimate the parameters of filters that can be used to reconstruct the signal. LPC is very effective for modeling the formants (resonant frequencies) of speech sounds.

[0104] 2. Perceptual Linear Prediction (PLP) coefficients, similar to linear predictive coding, but incorporating various aspects of human auditory perception, such as critical band spectral resolution, equal-loudness curves, and intensity-loudness power laws. Perceptual Linear Prediction (PLP) coefficients are designed to mimic the nonlinear perception of loudness and frequency by the human ear.

[0105] 3. Gammatone Filterbank Features involve a filterbank that simulates the filtering that occurs in the human cochlea. It is similar to a Mel filterbank, but uses Gammatone filters instead of triangular filters. Gammatone filterbank features are useful for capturing the detailed frequency structure of audio, especially in tasks involving environmental sound classification or hearing aid design.

[0106] 4. Chromaticity features, which represent the 12 different pitch categories (C, C#, D, etc.) of a musical octave. Chromaticity features capture the harmonic and melodic characteristics of music. These are particularly useful in music information retrieval, key detection, and chord recognition tasks.

[0107] 5. Mel Spectrum: While similar to MFCC, the Mel spectrogram is the result of applying a Mel filter bank directly to the power spectrum without further transformations (such as DCT). The Mel spectrogram preserves more detailed frequency information and is often used as input to deep learning models, often in conjunction with convolutional neural networks (CNNs) in tasks such as audio event detection, speech recognition, and music genre classification.

[0108] 6. The Constant-Q Transform (CQT) provides a time-frequency representation with a logarithmic frequency scale similar to the Mel scale, but with a variable time resolution that matches the frequency resolution. CQT is particularly useful for music applications because it offers a better representation of musical pitch compared to linear FFT or even the Mel scale.

[0109] 7. Deep Learning-Based Features: Features learned from deep learning models, such as embeddings from trained neural networks, can also be used as alternatives to traditionally hand-designed features like MFCCs. These features are generally more robust and can capture complex patterns that are difficult to model using traditional methods.

[0110] 8. Spectral Subband Centroids (SSCs) are the centroids of the energy distribution in different subbands of the captured signal. They can be interpreted as the "centroid" of the spectral lines within each subband. SSCs provide information about the energy distribution across the frequency bands and are sometimes used as a supplement to MFCCs.

[0111] 9. Relative Spectral (RASTA) features involve filtering the logarithmic energy of the speech signal to highlight modulation frequencies important for speech recognition. RASTA-PLP is a combination of RASTA filtering and PLP analysis, providing robust characteristics against noise and channel variations.

[0112] Audio features refer to the audio representation of an audio signal that attempts to capture relevant aspects of human auditory perception of the audio signal. Audio features should be selected to reduce the required data rate (for transmission of the audio features) and thus reduce the storage space required on the server. While the audio features can be the original audio signal itself, preferably, they can be a transparently coded version of the original signal, which requires significantly less storage space. Another example of an audio feature can be the output of any model that mimics human auditory perception. Audio features can be the Mel-frequency power spectrum and Mel-frequency cepstral coefficients (MFCCs) of the audio signal.

[0113] Specifically, audio features can be the acoustic properties, spectral properties, and statistical properties of an audio signal.

[0114] Acoustic features may include pitch (fundamental frequency); prosody (rhythm, stress, intonation) and / or duration and silence patterns (pauses, speech timing, or inconsistencies in breath sounds can be used as audio features).

[0115] Spectral features may include formant frequencies (such as resonant frequencies in speech); Mel frequency cepstral coefficients (MFCCs) (the short-time power spectrum of sound, which can reveal artifacts introduced during synthesis or tampering); spectrogram analysis (tampering may present an anomalous or obscured energy distribution in the time-frequency representation); and / or phase information.

[0116] Acoustic features can include statistical features and signal processing features, such as noise and residuals (e.g., subtle background noise or artifacts introduced during synthesis may not match natural recordings); high-frequency content; and / or phase distortion (some tampering may introduce phase anomalies, which can be detected using signal processing techniques).

[0117] Acoustic features can include temporal features such as jitter and flicker (e.g., frequency and amplitude variations); and temporal coherence (e.g., sudden shifts or changes in speech characteristics, such as unnatural interruptions in the signal).

[0118] Acoustic features can include behavioral or semantic inconsistencies, such as content coherence (logical inconsistencies or unnatural fluency in spoken content may indicate tampering); affect and naturalness (emotional tone may be inconsistent with the content, or the voice may lack the natural nuances of human emotion).

[0119] Audio features can be any one or a combination of the audio characteristics mentioned above for an audio signal.

[0120] By combining these features with machine learning or signal processing techniques, models can be trained to detect artifacts and inconsistencies that indicate deep fake audio.

[0121] In the example, a watermark could include a time-ordered sequence of pilot sequences (indicating the start of the watermark, the user ID, and the block ID). If a watermark is detected and can only be partially decoded (e.g., only the pilot sequence, without a meaningful user ID or block ID), it is still reasonable to notify the user accordingly. Advanced users (e.g., news organizations) can infer from the information that the content was signed at some point in time but appears to have been altered afterward. The partial presence of the watermark can also be used to refute claims that the file was never altered.

[0122] In another example, if the authentication information (source information) is not bound to the content via a watermark but, for example, via a time-related metadata format, the digital audio workstation can use accurate timestamps to reference all relevant source information for all source files. In this case, since all the information is already available, the authentication process described above is not necessary.

[0123] According to the example, each audio signal captured by microphone 100 and output by the microphone can include a watermark that identifies the actual microphone that captured the audio signal. If the microphone is used by several users, several user IDs can be associated with the microphone. If the microphone is registered, the corresponding microphone can be identified. User IDs can also be registered. A user ID can be a dedicated ID or a sub-ID of the microphone ID. The microphone ID and / or user ID can be part of the watermark. The sequence number of the audio block can determine which part of the audio signal has been removed.

[0124] Alternatively, the audio signal may be part of a video file.

[0125] Alternatively, the signed authentication information can be included in the audio file (e.g., in an ADM file format) instead of in the watermark. In this case, the audio signal does not need to be modified.

[0126] Alternatively, the processing to generate the signed audio signal can be performed using a software solution based on a pre-existing audio recording or audio stream. The private key can be entered into the software, for example, via a dongle or text input, and authorization can be performed, for example, via biometric signals such as fingerprints or facial recognition.

[0127] The authentication process can calculate a similarity score between the audio features of the signed audio signal and the audio features of the analyzed signal. This allows it to determine the likelihood that the signal is still authentic, even if slight signal processing such as gain has been applied.

[0128] Authentication requires an audio signal and signed authentication information. The signed authentication information can be provided via a side channel. When using a watermark, the authentication access information must be read from the audio signal, and the authentication information must be retrieved accordingly. Where watermark reading requires block boundaries, one method to achieve this is to try different offsets of the block boundaries until the watermark can be successfully read. Another method could be to embed some watermarked synchronization signal into the audio signal.

[0129] To enhance system security, processing is provided that enables, for example, key revocation in the event of key theft.

[0130] Public keys used for authentication can be provided in various ways. One option is to store all public keys in a centralized database to make it easy to find the required key. To address potential trust issues in this centralized instance, a database with all public keys can be provided based on distributed ledger technology. Another option is for the organization / individual using the technology to provide the public key on their own website. List of reference numerals 100 microphones 101 Selector 102 Selector 110 Microphone Cap 111 Microphone simulates audio signal 113 Metadata 120 AD converter 121 digital audio signal 130-block divider 131 audio blocks 140 Private Key Signature Unit 141 Signed authentication information 142 Private Key 143 Public Key 144 signatures 145 Identifier 150 Audio Feature Extractors 151 Audio Features 152 Authentication Information 160 Information Embedder / Watermark Generator 161 Audio blocks with embedded information 162 Authentication Access Information 163 Watermark 165 Size Information 171 audio signal 190 audio output terminal 200 Audio Signal Authenticator 205 Block Boundary Detector and Block Segmenter 206 Authentication Information Extractor 210 Access Information Extractor 212 Authentication Access Information 221 signatures 230 Authentication Access Message Parser 240 Perceptual Similarity Analyzer 241 Similarity Results 250 signature verifier 251 Signature verification result 260 combiner 261 Authentication Result 270 Identity 300 Network 400 server 500 audio signal signature module 501 Certified Audio Signal 502 Audio Input Terminal 600 Server or Storage

Claims

1. A method for authenticating an audio signal, comprising the following steps: - Analyze the received audio signal to determine audio blocks and extract the embedded authentication access information from the received audio blocks. - Analyze the extracted authentication access information to retrieve authentication information containing audio features of the original audio signal. - Analyze the audio blocks of the received audio signal to determine if there are irregular audio blocks (C) for which no authentication access information is available. - Compare the irregular audio block (C) with a corresponding set of audio features that can be retrieved from memory (600) specifically based on the authentication access information, and - Notify the user which parts of the irregular audio block (C) were successfully authenticated.

2. The method according to claim 1, wherein, If the received audio signal includes a portion having at least two audio signals (A, B), and if the irregular audio block (C) contains a change from the first audio signal (A) to the second audio signal (B), then the following steps are performed: The last complete audio block (A2) is determined based on the first signal. Retrieve from the memory (600) the authentication information corresponding to the audio block (A3) immediately following the last complete audio block (A2), and The audio features contained in the retrieved authentication information are compared sequentially or in pairs in a forward time manner with the received irregular audio blocks (C) until a change from the first audio signal to the second audio signal is detected.

3. The method according to claim 1 or 2, wherein, If the received audio signal includes a portion having at least two audio signals (A, B), and if the irregular audio block (C) contains a change from the first audio signal (A) to the second audio signal (B), then the following steps are performed: The first complete audio block (B2) is determined based on the second signal. Retrieve from the memory (600) the authentication information corresponding to the audio block (B1) immediately preceding the first complete audio block (B2), and The audio features contained in the retrieved authentication information are compared sequentially or in pairs with the received irregular audio blocks (C) in reverse chronological order until a change from the second audio signal to the first audio signal is detected.

4. The method according to any one of claims 1, 2, or 3, wherein, The comparison between the irregular audio block (C) and the audio features within the authentication information is performed in the following manner: - Select the appropriate type of audio feature for the comparison. - Extract the selected type of features from the irregular audio signal block (C). - If the type of audio features in the authentication information does not correspond to the selected type, then these audio features are converted to audio features of the selected type, and The comparison is performed based on the audio features of the selected type.

5. The method for authenticating an audio signal (200) according to any one of claims 1 to 4, wherein, The authentication information includes audio features (151) belonging to the current audio block (201) and a signature (144, 221).

6. An audio signal authenticator (200) configured to perform the method according to any one of claims 1 to 5.

7. A computer program product for authenticating audio signals, in, The computer program product includes program code units that cause the audio signal authenticator according to claim 6 to perform the method according to any one of claims 1 to 5.