Method and web browser extension for authenticating audio signals

By embedding authentication information into the audio signal and using web browser extensions for automatic authentication, the problem of difficult audio signal verification in existing technologies is solved, and a simple and user-friendly audio signal authentication method is realized.

CN121970053APending Publication Date: 2026-05-01SENNHEISER ELECTRONICS GMBH & CO KG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SENNHEISER ELECTRONICS GMBH & CO KG
Filing Date
2024-09-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively authenticate audio signals, especially verifying the authenticity and source of audio signals such as politicians' speeches, which typically requires extensive manual research.

Method used

Automatic authentication is performed by embedding authentication information or authentication access information into the audio signal and utilizing web browser extensions. This includes block boundary detection of audio blocks, extraction and verification of authentication information, and display of authentication results.

Benefits of technology

It enables simple and user-friendly audio signal authentication on web pages, avoiding separate verification for each website, reducing the need to modify websites and browsers, and providing instant authentication result display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970053A_ABST
    Figure CN121970053A_ABST
Patent Text Reader

Abstract

A method of authenticating an audio signal of a media file associated with a web page is provided. The method includes selecting, by a user, a web page via a web browser, where the web page provides content having at least one media file including an audio signal, where the content is stored on a web server. The audio signal includes a plurality of audio blocks, each audio block having authentication information or authentication access information related to information on how or where to retrieve externally stored authentication information. The method further includes receiving the audio signal from the web server, extracting or retrieving authentication information of the audio block, and performing an authentication check for the audio block of the received audio signal based on the authentication information, and displaying a result of the authentication check via the web browser.
Need to check novelty before this filing date? Find Prior Art

Description

Methods for authenticating audio signals and web browser extensions

[0001] This invention relates to a method for authenticating audio signals and web browser extensions.

[0002] Given the current ability to create deepfake videos and audios, there is an increasing need for the ability to verify or authenticate audio or video signals, such as speeches by politicians.

[0003] To date, authenticating an individual's audio signal has been extremely difficult. Typically, extensive manual research is required to verify or authenticate such an audio signal and to verify its origin.

[0004] Therefore, the object of the present invention is to provide a method for authenticating audio signals and a web browser extension that enables the authentication of audio signals and displays the authentication results in a user-friendly manner.

[0005] This objective is achieved by the method for authenticating audio signals according to claim 1 and the web browser extension according to claim 5.

[0006] Therefore, a method for authenticating audio signals of media files associated with a webpage is provided. A user selects a webpage via a web browser. The webpage provides content having at least one media file including an audio signal. The content is stored on a web server. The audio signal comprises multiple audio blocks, each audio block having authentication information or authentication access information, the authentication access information relating to how or where to retrieve externally stored authentication information. The method includes the following steps: receiving the audio signal from the web server, extracting or retrieving authentication information for the audio blocks, and performing an authentication check on the audio blocks of the received audio signal based on the authentication information. The results of the authentication check are displayed via a web browser.

[0007] According to one aspect, the following steps are performed in a web browser, web browser extension, or web browser plugin: receiving an audio signal from a web server, extracting or retrieving authentication information for audio blocks, and performing authentication checks on the audio blocks of the received audio signal based on the authentication information.

[0008] According to one aspect, a method for authenticating an audio signal is provided. The audio signal may be part of a media file associated with a website accessed by a user via a web browser. The media file including the audio signal is stored on a web server. The authenticable audio signal comprises multiple audio blocks, each audio block having authentication information or authentication access information regarding where and how to retrieve authentication information stored externally. Authentication information is extracted from the audio blocks or retrieved via the authentication access information. An authentication check is performed on the retrieved audio signal via a web browser extension, and the result of the authentication check is displayed via the web browser. Therefore, the authentication check is performed within a web browser extension.

[0009] This is advantageous because all authentication processing is performed by the web browser extension, eliminating the need to implement verification separately on each website and requiring no modifications to the website or web browser. Authentication results can be forwarded to the web browser, for example, via an API. The web browser can then display the authentication information on the user's screen or via a user interface.

[0010] According to one approach, a method for authenticating an audio signal includes the following steps: determining the owner of the private key that initially signed the audio signal based on authentication information or authentication access information, and displaying the owner's name to the user of a web browser.

[0011] One method for authenticating audio signals includes displaying the current time-related status of the authentication check results in a web browser.

[0012] One approach to authenticating audio signals includes using a web browser to display the results of authentication checks over time for the entire sequence being analyzed.

[0013] One method for authenticating audio signals includes having the web browser explicitly flag the appearance of several blocks that result in a negative authentication.

[0014] According to one approach, a method for authenticating an audio signal includes the following steps: detecting clips in a received or extracted audio signal, wherein truncated sequences of multiple audio blocks, each containing authentication information or authentication access information, are sequentially spliced ​​together; and displaying the appearance of those clips to a user of a web browser.

[0015] According to one aspect, the method for authenticating an audio signal includes: first checking whether the audio signal contains authentication information or authentication access information, and if no authentication information or authentication access information is detected, then playing back the selected media file without displaying the authentication check result via a web browser.

[0016] According to one aspect, the method for authenticating an audio signal includes: first checking whether the audio signal contains authentication information or authentication access information, and if no authentication information or authentication access information is detected, then playing back the selected media file via a web browser and displaying information about the lack of authentication to the user.

[0017] According to one aspect of the present invention, a web browser or web browser extension is provided, the web browser or web browser extension being configured to: access a webpage providing content having at least one media file including an audio signal, wherein the content is stored on a web server, wherein the audio signal comprises a plurality of audio blocks, each audio block having authentication information or authentication access information, the authentication access information relating to how or where to retrieve externally stored authentication information; receive the audio signal from the web server; extract or retrieve the authentication information of the audio blocks of the received audio signal; and perform an authentication check on the audio blocks of the received audio signal based on the authentication information. The results of the authentication check are displayed via the web browser.

[0018] According to one aspect, a web browser extension is provided, which is configured to: receive an audio signal to be authenticated to extract or retrieve authentication information of audio blocks, perform authentication checks on the audio blocks of the received or extracted audio signal based on the authentication information, and output the results of the authentication checks to a web browser associated with the web browser extension.

[0019] Authentication information can be embedded into the corresponding audio block of the audio signal. Alternatively, only the authentication access information can be stored in the audio block. When a web browser extension retrieves the authentication access information, it receives information on how and where the actual authentication information can be retrieved.

[0020] Once the authenticator, in the form of a web browser extension, has received the authentication information for the corresponding audio block, it can perform an authentication check. If all audio blocks of the audio signal have been authenticated, the authenticator can, for example, forward this information via an API to a web browser capable of displaying it on a user interface.

[0021] Alternatively, if there are audio blocks in the audio signal that cannot be authenticated, this information can also be forwarded from the authenticator to a web browser via the API for display on the user interface.

[0022] If the authenticator (i.e., a web browser extension or web browser) has already performed an authentication check on the audio signal, the result of the authentication check can also be displayed via the web browser. Advantageously, the authentication result can be displayed near the location on the display where the media file is located or where a link to the media file points. This is advantageous because the authentication result can be displayed close to the location of the media file or close to the representation of the media file on the webpage. Therefore, a simple and user-friendly method is proposed to verify or authenticate media files that include audio signals in or on a webpage.

[0023] An authenticator is provided, for example, in the form of a browser extension or browser plugin. The authenticator analyzes media files on or within a website to be displayed by a web browser to determine the authenticity of the audio signal in the media file. The authenticator determines the block boundaries of audio blocks, extracts or retrieves authentication information associated with the audio blocks, and performs authentication checks for each audio block.

[0024] If the audio signal or a portion of the audio signal (audio block) has been authenticated, the authenticator can output a notification via a web browser. This can also be done if not all audio blocks have been authenticated. Audio blocks that have not yet been authenticated can be flagged.

[0025] A web browser is an application (like a software application) that enables access to and interaction with websites on the Internet. A web browser retrieves web pages from servers and displays them on your device. To access web pages, URLs, or search for content online, a web browser connects to a web server hosting the site, downloads the content (text, images, videos, audio files), and displays it on your screen.

[0026] Web browser extensions are software modules or programs that can be added to a web browser to extend its functionality. Extensions allow users to customize their browsing experience by adding features, improving usability, or integrating with other services.

[0027] Web browser extensions are available for most major web browsers and can be installed from the browser's extension store or marketplace. Web browser extensions run in the background and interact with the visited web pages to provide enhanced functionality. They enhance or customize the user experience by adding features or changing the behavior of the web browser itself. Extensions work within the browser's user interface and typically modify how web pages are displayed or how additional functionality is added to them.

[0028] Web browser plugins are software components that add specific functionality to larger software applications. Historically, web browser plugins were used to handle specific types of content that browsers couldn't natively support, such as playing videos, viewing PDFs, or running Flash or Java content. Current browsers, by natively supporting these features, have largely eliminated the need for plugins.

[0029] Web browser extensions can be built using standard web technologies such as HTML, CSS, and JavaScript, and can be browser-specific. Plugins are typically third-party software, usually installed to provide functionality that the browser itself cannot support independently.

[0030] Plugins are becoming obsolete. Browsers like Chrome and Firefox have stopped supporting most plugins, except for some basic ones like PDF readers, because current browsers natively support most multimedia formats.

[0031] Depending on your needs, there are several alternatives that use web browser extensions to achieve similar functionality. These alternatives can operate outside the browser or through different integration methods.

[0032] Instead of using browser extensions, standalone applications can provide the same functionality.

[0033] As an alternative to web browser extensions, web services can be used. Websites can provide online tools or functions that work similarly to extensions, without requiring any installation.

[0034] Audio features are related to the audio representation of an audio signal and attempt to capture relevant aspects of human auditory perception of the audio signal. Audio features should be selected to reduce the required data rate (for transmitting the audio features) and thus reduce the storage space required on the server. While the audio features may be the original audio signal itself, preferably, they can be a transparently encoded version of the original signal, which requires significantly less storage space. Another example of an audio feature can be the output of any model that mimics human auditory perception. Audio features can be the Mel-frequency power spectrum and Mel-frequency cepstral coefficients (MFCCs) of the audio signal.

[0035] As an example, a computer program product for authenticating audio signals is provided. This computer program product includes program code units that cause the web browser extension described above to perform the methods described above.

[0036] These and other aspects of the invention are described in more detail with reference to the following accompanying drawings.

[0037] Figure 1 shows a block diagram of the audio signal signature module.

[0038] Figure 2 shows a diagram of the overall audio signature and authentication workflow.

[0039] Figure 3 shows a block diagram of the audio signal authenticator.

[0040] Figure 4 shows a block diagram of the audio signal signature module.

[0041] Figure 5 shows a block diagram of the audio signal authenticator.

[0042] Figure 6 shows a block diagram of the audio signal signature module.

[0043] Figure 7 shows a block diagram of the audio signal authenticator, and

[0044] Figure 8 shows a flowchart of the method for authenticating audio signals.

[0045] According to the example, authentication is performed on an audio signal. The audio signal is authenticated based on information embedded in an audio block. This embedded information can be, for example, authentication information directly embedded in the audio block using a watermark, or authentication access information. Authentication access information includes information about where and how to retrieve stored authentication information associated with the audio block from storage (such as a server). In this case, the authentication information is not directly embedded in the audio block. Figures 1 through 5 illustrate examples using authentication access information, and Figures 6 and 7 illustrate examples using authentication information.

[0046] Based on the example, authentication of audio blocks is performed. As described in more detail below, further analysis of unauthenticated audio blocks is only possible if not all audio blocks are authenticated.

[0047] To better understand the authentication of audio signals, we will first describe the generation of authenticable audio signals.

[0048] Figure 1 illustrates an audio signal signature module. The audio signal signature module 500 receives an input audio signal 121 from a microphone 100 or another audio source, such as an audio recorder, via an audio input terminal 502, and outputs an output audio signal 501 via an audio output terminal 190. The output audio signal 501 may be based on the input audio signal 121. The microphone 100 may include at least one microphone capsule 110 and an analog-to-digital converter (ADC) 120. The at least one microphone capsule 110 can capture an audio signal and output the captured analog audio signal 111. The captured analog audio signal 111 can be forwarded to the ADC 120, which can digitize the audio signal 111 and output a digital audio signal 121. The audio signal signature module 500 may be implemented as a device separate from the microphone 100 or may be included within the microphone 100.

[0049] The audio signal signing module 500 can receive an audio signal 121 from a microphone 100 or from another source, as indicated by selector 101 in FIG. 1. The digital audio signal 121 is input to a block divider 130, which divides the audio signal 121 into audio blocks 131 and outputs a sequence of multiple audio blocks 131. Optionally, if the audio signal 121 has already been divided into audio blocks 131, the block divider 130 can be omitted. The sequence of audio blocks 131 is forwarded to an audio feature extractor 150, which generates an audio feature 151 for each audio block 131 in the audio block sequence. The private key signing unit 140 receives authentication information (e.g., metadata) including the audio feature 151 of each audio block 131 and optional other information 113 such as location or time, and signs the authentication information using a private key 142, thereby providing signed authentication information 141, which can be stored, for example, on a memory 600 located on an external server. The signed authentication information 141 may include authentication information 152 (see Figure 2), which includes audio features 151 and optional other information 113, as well as a signature 144 (Figure 2) generated by the private key signing unit 140. Alternatively, the authentication information 141 may be stored in internal memory or any other memory, as long as it can be accessed by the authenticator used to authenticate the received audio signal.

[0050] The signed authentication information 141 itself includes authentication information, namely (a) audio feature 151, optionally (b) other information 113, and (c) signature 144. Individual audio blocks 131 are fed in parallel, for example, to an information embedder 160 for embedding authentication access information 162 into, for example, each audio block 131. Authentication access information 162 includes information about where and / or how the authentication information 152 (or more precisely, the signed authentication information 141) for that particular audio block 131 can be accessed, for example, at a memory 600 located on a server. This embedding can be accomplished, for example, by means of a watermark.

[0051] Therefore, the information embedder 160 can be implemented as a watermark generator that generates watermark 163. Alternatively, the authentication access information 162 can be embedded into a suitable audio container adjacent to the actual audio block 131. Accordingly, the information embedder 160 outputs the audio block 161 with the embedded authentication access information 162. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at audio output terminal 190. The output audio signal 501 can be broadcast or stored, for example, via network 300 (FIG. 2). In particular, the output audio signal 501 can even be further modified to a very small extent, including modification types such as sample rate conversion, perceptual compression, and clipping.

[0052] In the case where the information embedder 160 is a watermark generator, its output 161, instead of the original audio block 131, can alternatively be input to the audio feature extractor 150, as indicated by selector 102 in FIG. 1. This alternative offers the advantage of extracting audio features based on the final signal 161 output by the audio signal signature module 500, rather than the unoutput intermediate signal 131. Since the watermark generator is designed to perceptually alter the audio signal as little as possible, both input options (131 or 161) are reasonable for the audio feature extractor 150.

[0053] Audio features involve an audio representation that attempts to capture relevant aspects of human auditory perception, while preferably reducing the required data rate and thus the storage space required on the server or memory 600. While audio features can be the original audio signal itself, preferably, they can be a transparently encoded version of the original signal that requires far less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.

[0054] The audio signal 501 at the output of the audio signal module 500 (containing authentication access information or an embedded watermark within the audio container) can be stored or transmitted. Because the watermark can be embedded in, for example, essentially each audio block 131, each audio block 131 can be authenticated individually. The watermark 103 can also be embedded only in some audio blocks within the audio block set. In addition to the authentication access information, the watermark 163 can also include other information such as block boundaries. This information can be used when authenticating the audio signal to determine the block length.

[0055] The audio watermark 163 can be a unique identifier embedded in the audio signal and, for example, previously used to identify copyright information of the audio signal. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, the embedded watermark will not be altered when the audio signal with the embedded watermark is copied, stored, or transmitted. The same reasoning applies to authentication access information included in the audio container.

[0056] Figure 2 illustrates the overall audio signing and authentication workflow. An audio signal is captured by microphone 100, which outputs a digital audio signal 121. Audio features 151 are extracted from this audio signal 121 and signed using, for example, the private key 142 of audio signal signing module 500. Audio features 151, together with a digital signature 144 generated by signing unit 140, constitute signed authentication information 141. The signed authentication information 141 can be forwarded to a server or storage 600, where it is stored for later retrieval. Information regarding where and / or how to retrieve the signed authentication information 141, namely authentication access information 162, is provided to audio signal signing module 500. Audio signal signing module 500 uses the authentication access information 162 and embeds it into audio blocks 131 of the audio signal. Therefore, the output signal 501 of audio signal signing module 500 includes audio blocks 131 with the embedded authentication access information. The audio block 131 of the audio signal with embedded authentication access information 162 is output by the audio signal signature module 500 and can be stored or distributed, for example, via network 300.

[0057] This distribution may involve modification or tampering with authentication information 152 and / or audio signal 121. Public key 143 is used to verify the signed authentication information 141. This verification indicates potential modification of the audio features. The integrity of the audio signal can be verified by extracting audio features from the received signal and comparing them with the features included in the signed authentication information 141. The final authentication result 261 is determined jointly by these two checks.

[0058] An audio signal signing module 500 (e.g., implemented as part of microphone 100) is used to capture audio signals. Alternatively, the audio signal signing module 500 can receive audio signals. The audio signals can also be audio signals pre-captured by other means. As described with reference to FIG1, the captured audio signal is divided into multiple audio blocks, and authentication information of at least one audio block is signed based on the private key 142 of the audio signal signing module 500. The signed authentication information 141 is output by the audio signal signing module 500 and can be stored on a server or memory 600. Authentication access information 162 can be provided to the audio signal signing module 500; the authentication access information 162 is information about where the authentication information can be retrieved from the server or memory 600. The authentication access information 162 is embedded in the audio blocks received or captured by the audio signal signing module 500. Therefore, the audio signal signature module 500 performs the following operations: a) outputting an output audio signal 501, including an audio block with embedded authentication access information, to the network 300 for storage or forwarding, and b) outputting signed authentication information 141 to a server or memory 600 that can store the information.

[0059] Optionally, authentication access information can be embedded into audio blocks of the audio signal using a watermark. The watermark may relate to authentication access information associated with the location in memory 600 where the authentication information is stored. This authentication information can be generated by the audio feature extractor 150 as described above with reference to FIG1. ​​Audio blocks of the audio signal with embedded authentication access information can be transmitted or stored via network 300. Authenticator 200 can receive the audio signal 501 with the embedded watermark and extract the authentication access information. Authentication information is retrieved from server 600 based on the authentication access information. The signature of the authentication information stored on server or memory 600 can be verified based on the public key 143 associated with audio signal signature module 500, which can be stored, for example, on server 400 or a distributed ledger.

[0060] Figure 3 shows a block diagram of an audio signal authenticator. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. This audio signal 171 can be the output audio signal 501 from the audio signal signature module 500 of Figure 1. The audio signal 171 is input to a block boundary detector and a block segmenter 205, which determine block boundaries and output multiple audio blocks 201 with embedded authentication access information 212. If the authentication access information 212 has already been embedded using a watermark, it is assumed that the block boundaries have been previously encoded into the watermark, and the block boundaries can be detected from the watermark at this stage. In the case where the audio signal is transmitted by an audio container including the authentication access information 212 as metadata, it is assumed that the authentication access information is available for each block, where the block boundaries are part of the metadata.

[0061] An audio block 201 with embedded authentication access information is input to an access information extractor 210, which extracts authentication access information 212 from the audio block 201 (which, if tampered with, corresponds to authentication access information 162 from the signature module). An authentication access information parser 230 determines the access information based on the authentication access information 212, i.e., the location of the signed authentication information 141 belonging to the current audio block 201 can be found on the server or memory 600. The authentication information 141 retrieved from the server or memory 600 first includes audio features 151 representing the audio block 201, secondly includes a signature 221, and further includes optional other information 113, such as the location and / or time of the signature occurrence. If tampered with, the signature 221 corresponds to the signature 144 generated by the signature unit 140.

[0062] In the perceptual similarity analyzer 240, audio blocks 201 and audio features 151 are perceptually compared to each other, thereby providing a similarity statement 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity statement 241 can be more detailed, such as providing a similarity score.

[0063] In the similarity analyzer 240, audio blocks can be compared with audio features of the original audio signal. The audio features can be a true copy of the original audio signal. However, this would require a large amount of storage space to store the original audio signal. If the audio features are related to other characteristics of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Thus, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within the similarity analyzer. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.

[0064] The authentication access information parser 230 also determines, based on the authentication access information 212, which owner of the private key 142 initially signed the audio signal 171, for example, by identifying the owner through an ID. The owner of the private key 142 could be a natural person, an organization, a registered microphone, and a few other possibilities are listed here. Using this ID, the identity 270 of the public key 143 and the private key owner can be retrieved from an external database. The public key 143 is then used within the signature verifier 250 to verify the combination of the signature 221 and the authentication information 152 (audio feature 151 and other information 113), thereby providing a signature verification result 251. The perceived similarity claim 241 and the verification result 251 are finally combined in the combiner 260 to form the final authentication result 261.

[0065] Note that the processing proposed here for audio content can be similarly applied to any time-related data. In particular, the above processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark or uses a video container instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison.

[0066] Figure 4 shows a schematic block diagram of a more general audio signal signature module compared to the audio signal signature module shown in Figure 1. In particular, unlike the diagram shown in Figure 4, no method is specified here for distributing the individual audio blocks and their corresponding authentication information—these two components can even be distributed independently of each other, as long as they are accessible at the verification step.

[0067] The audio signal signature module 500 receives a digital audio signal 121 at its audio input terminal 502. This digital audio signal 121 is input to a block segmenter 130, which outputs a sequence of multiple audio blocks 131. The audio blocks 131 are forwarded to an audio feature extractor 150, which generates audio features 151 for each audio block 131. In the private key signature unit 140, the user's or entity's private key 142 is used to sign the audio features 151 of each audio block 131, along with optional other information 113 such as location or time authentication information, thereby providing signed authentication information 141. The signed authentication information 141 itself includes the audio features 151, the generated signature 144, and optional other information 113. During the verification process of the authentication method shown in Figure 5, both components—the individual audio blocks 131 and the signed authentication information 141—are required simultaneously. In principle, they can be embedded in the same suitable audio container, but they can also be distributed across different physical channels. Optionally, a user or entity identifier 145 can be added to both components, which can, for example, enable the retrieval of the appropriate public key for automatic verification.

[0068] Figure 5 shows a schematic block diagram of a more general authenticator 200 corresponding to the signal signature module depicted in Figure 4. Authenticator 200 receives, for example, an audio signal in the form of an audio block 201 to be verified, and signed authentication information 141. The signed authentication information 141 includes audio features 151, other information 113, and a signature 221. Within a perceptual similarity analyzer 240, the audio features 151 contained in the authentication information 141 are perceptually compared with the audio block 201, thereby providing a similarity claim 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity claim 241 can be more detailed, thereby providing, for example, a similarity score. Furthermore, using knowledge about the user or entity that created the signature 221, a corresponding public key 143 and an optional user identity 270 are obtained. Optionally, this knowledge can be derived from a user or entity identifier 145 provided along with the authentication information 141 and the audio block 201. Subsequently, public key 143 is used to verify the combination of signature 221 and authentication information 141 within signature verification block 250, thereby providing signature verification result 251. The perceived similarity claim 241 and verification result 251 are finally combined in combiner 260 to form final authentication result 261.

[0069] Figure 6 below illustrates the generation of an authenticable audio signal by embedding authentication information into an audio block, and the authentication of the audio signal based on the embedded authentication information. This authentication can be used for authenticating audio signals, as described below.

[0070] Figure 6 illustrates a schematic diagram of an audio signal signature module. The audio signature module in Figure 6 is merely an example. The audio signal signature module 500 receives an input audio signal 121 from a microphone 100 or another audio source, such as an audio recorder, via an audio input terminal 502, and outputs an output audio signal 501 via an audio output terminal 190. The output audio signal 501 may be based on the input audio signal 121. The microphone 100 may include at least one microphone capsule 110 and an analog-to-digital converter (ADC) 120. The at least one microphone capsule 110 can capture an audio signal and output the captured analog audio signal 111. The captured analog audio signal 111 can be forwarded to the ADC 120, which can digitize the audio signal 111 and output a digital audio signal 121. The audio signal signature module 500 may be implemented as a device separate from the microphone 100 or may be included within the microphone 100.

[0071] The audio signal signature module 500 can receive the audio signal 121 from the microphone 100 or from another source, as indicated by the selector 101 in FIG. 6. The digital audio signal 121 is input to the block divider 130, which divides the audio signal 121 into audio blocks 131, each with a block length, and outputs a sequence of multiple audio blocks 131. The block divider 130 can divide the audio signal 121 into multiple audio blocks. Optionally, each audio block receives a block number, which can be a subsequent block number. Optionally, if the audio signal 121 has already been divided into audio blocks 131, the block divider 130 can be omitted. The sequence of audio blocks 131 is forwarded to the audio feature extractor 150, which generates an audio feature 151 for each audio block 131 in the sequence of audio blocks 131. The private key signing unit 140 receives authentication information (e.g., metadata) including audio features 151 for each audio block 131 and optional other information 113 such as location or time, and signs the authentication information with a private key 142 to provide signed authentication information 141. The signed authentication information 141 may include authentication information 152, which includes the audio features 151 and optional other information 113, as well as a signature 144 generated by the private key signing unit 140.

[0072] The signed authentication information 141 itself includes authentication information, namely (a) audio feature 151, optionally (b) other information 113, and (c) signature 144. Each audio block 131 is fed in parallel, for example, to an information embedder 160 for embedding the signed authentication information 141 into, for example, each audio block 131. Embedding can be accomplished, for example, by means of a watermark.

[0073] Therefore, the information embedder 160 can be implemented as a watermark generator that generates watermark 163. Alternatively, the signed authentication information 141 can be embedded into a suitable audio container adjacent to the actual audio block 131. Accordingly, the information embedder 160 outputs the audio block 161 with the embedded signed authentication information 141. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at audio output terminal 190. The output audio signal 501 can be broadcast or stored, for example, via network 300 (FIG. 7). In particular, the output audio signal 501 can even be further modified to a very small extent, including modifications such as sample rate conversion, perceptual compression, and clipping.

[0074] When the information embedder 160 is a watermark generator, its output 161, instead of the original audio block 131, can alternatively be input to the audio feature extractor 150 via the selector 102. This alternative offers the advantage of extracting audio features based on the final signal 161 output by the audio signal signature module 500, rather than the unoutput intermediate signal 131. Since the watermark generator is designed to perceptually alter the audio signal as little as possible, both input options (131 or 161) are reasonable for the audio feature extractor 150.

[0075] The audio signal 501 at the output of the audio signal module 500 (containing authentication access information or an embedded watermark included in the audio container) can be stored or transmitted. Because the watermark can be embedded in, for example, essentially every audio block 131, each audio block 131 can be authenticated individually. The watermark 103 can also be embedded only in some audio blocks within the audio block set. In addition to authentication information, the watermark 163 can also include other information such as block boundaries. This information can be used when authenticating the audio signal to determine the block length.

[0076] The audio watermark 163 can be a unique identifier embedded in the audio signal and, for example, previously used to identify copyright information of the audio signal. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, the embedded watermark will not be altered when the audio signal with the embedded watermark is copied, stored, or transmitted. The same reasoning applies to authentication access information included in the audio container.

[0077] Figure 7 shows a block diagram of an audio signal authenticator corresponding to the audio signal authenticator of Figure 6. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication information. This audio signal 171 can be the output audio signal 501 from the audio signal signature module 500 of Figure 6. The audio signal 171 is input to a block boundary detector and a block segmenter 205, which determine block boundaries and output multiple audio blocks 201 with embedded authentication information.

[0078] The block boundary detector and block segmenter 205 detect block boundaries, for example, using information from the block header. Since the block length of audio blocks may vary, the block length should be determined for each audio block or for a group of audio blocks with the same block length.

[0079] Audio block 201 is input to authentication information extractor 206, which extracts signed authentication information 141 embedded in the audio block. The signed authentication information 141 may include audio features 151, a signature 144, and optional other information 113, such as the location and / or time of the signature occurrence. If it has not been tampered with, the signature 221 corresponds to the signature 144 generated by signature unit 140.

[0080] In the perceptual similarity analyzer 240, audio blocks 201 and audio features 151 are perceptually compared to each other, thereby providing a similarity statement 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity statement 241 can be more detailed, thus providing, for example, a similarity score.

[0081] The authentication information extractor 206 also determines, based on authentication information 113 / 141, which owner of the private key 142 initially signed the audio signal 171, for example, by identifying the owner through an ID. The owner of the private key 142 could be a natural person, an organization, a registered microphone, and a few other possibilities are listed here. Using this ID, the identity 270 of the public key 143 and the private key owner can be retrieved from an external database. Subsequently, the signature 144 and the authentication information 152 (audio feature 151 and other information 113) are verified within the signature verifier 250 using the public key 143, thus providing a signature verification result 251. The perceived similarity claim 241 and the verification result 251 are finally combined in the combiner 260 to form the final authentication result 261.

[0082] Note that the processing proposed here for audio content can be similarly applied to any time-related data. In particular, the above processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark, or uses a video container instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison.

[0083] Figure 8 illustrates a flowchart of a method for authenticating an audio signal. In step S1, an audio signal to be authenticated is selected. The audio signal may be part of a media file associated with a webpage of a website accessed by the user via a web browser. The media file including the audio signal is stored on a web server. The authenticable audio signal comprises multiple audio blocks, each containing authentication information or authentication access information regarding where and how to retrieve externally stored authentication information. In step S2, authentication information is extracted from the audio blocks or retrieved via the authentication access information. In step S3, an authentication check is performed on the retrieved audio signal via a web browser extension. In step S4, the results of the authentication check are displayed via a web browser.

[0084] The authentication result can be forwarded from the web browser extension to the web browser, for example, via an API. The web browser can then display the authentication information on the user's screen or through a user interface.

[0085] Authentication information can be embedded into the corresponding audio block of the audio signal. Alternatively, only the authentication access information can be stored in the audio block. When the web browser extension retrieves the authentication access information, it receives information on where the actual authentication information can be retrieved.

[0086] Once the authenticator, in the form of a web browser extension, has received the authentication information for the corresponding audio block, it can perform an authentication check. If all audio blocks of the audio signal have been authenticated, the authenticator can forward this information via API to a web browser that can display it on the user interface. If there are some audio blocks in the audio signal that cannot be authenticated, this information can also be forwarded from the authenticator to the web browser via API for display on the user interface.

[0087] Once the authenticator (i.e., a web browser extension or web browser) has performed the authentication check on the audio signal, the result can also be displayed via the web browser. Advantageously, the authentication result can be displayed near the location on the display where the media file is located or where a link to the media file points. This is advantageous because the authentication result can be displayed close to the location of the media file or close to its representation on the webpage. Therefore, a simple and user-friendly method is proposed to verify or authenticate media files that include audio signals in or on a webpage.

[0088] Web browsers or web browser extensions can implement authenticators based on the examples above, and in particular Figures 3, 5, and 7.

[0089] A web browser is an application used to access websites. When a user requests a webpage from a specific website, the browser retrieves the file from the web server and displays the page on the user's monitor. The purpose of a web browser is to retrieve content from the web or local storage and display that content on the user's device. A web browser can retrieve media files and other information from the web server based on information on the webpage to display the webpage on the user interface or monitor.

[0090] Browser extensions can be software modules used to customize web browsers.

[0091] The authenticator is provided, for example, as a browser extension or browser plugin. The authenticator analyzes media files on or within a website to be displayed by a web browser to determine the authenticity of the audio signal of that media file. The authenticator determines the block boundaries of audio blocks, and extracts or retrieves authentication information associated with each audio block and performs authentication checks for each audio block.

[0092] If the audio signal or a portion of the audio signal (audio block) has been authenticated, the authenticator can output a notification via a web browser. This can also be done if not all audio blocks have been authenticated. Audio blocks that have not yet been authenticated can be flagged.

[0093] Instead of implementing it in a browser extension, the authentication mechanism can also be supported natively by the browser. This retains the advantage of not requiring modifications to the website providing authenticated content while allowing users to see the verification results.

[0094] One aspect of the present invention is to conveniently present to web browser users the results of authentication checks performed according to one of the embodiments of Figures 3, 5 and 7.

[0095] Optionally, in step S1, the web browser or web browser extension may first check whether the selected media file includes authentication information or authentication access information. If no authentication information or authentication access information is detected, the web browser or web browser extension may play the selected media file without displaying the authentication information via the web browser. Alternatively, in this case, the web browser may display information to the user regarding the lack of authentication. If authentication information or authentication access information is detected, the web browser may display the authentication check result along with the selected media file to the user of the web browser.

[0096] If authentication information or authentication access information is detected, the web browser or web browser extension may optionally determine the owner of the private key 142 that initially signed the audio signal based on the authentication information, and display the owner name registered for that owner to the web browser user.

[0097] Optionally, a web browser or web browser extension can display the current time-relevant status of the authentication check result at a given moment. For example, it could display a green marker if the authentication result is positive, a red marker if the result is negative, and a yellow or no marker if no authentication information is available. Alternatively, the web browser can display the time-relevant authentication status of the media file in or near a symbolic representation of the media file (e.g., a player's progress bar / timeline), showing the authentication check result over time for the entire analyzed sequence. In other words, for time-relevant media (video and audio), this visualization could include a representation of the verification metric over time (e.g., the file was modified between seconds 3 and 5, but the rest remained unchanged). This makes it easy to jump to the questionable part of the file. If the audio does not have a UI representation (“background sound”), the browser might display the button or representation somewhere in the browser's control panel.

[0098] Another visualization option to consider is that if malicious signal tampering occurs, several consecutive blocks are more likely to be affected than a single block. Therefore, it is reasonable to specifically mark the occurrence of several blocks with negative authentication results. This could happen, for example, by using a darker red hue whenever the length of a suspected modified segment exceeds a certain length (e.g., 5 seconds). Another option would be to increase the darkness of the red hue as the length of the suspected modified segment increases.

[0099] Furthermore, a website can use only truncated sequences of the initially signed media files, and these segments can be spliced ​​together sequentially. These cuts may or may not alter the meaning of the original recording. To handle these cases, a web browser can additionally display the presence of those cuts to its user. Thus, if the metadata contains the numbers of the authenticated blocks, the presence of any non-contiguous block numbers can be specifically indicated by the authentication result, for example, via flashing or broken chain links, suggesting the presence of cuts in the original signal. If the audio signal to be authenticated is created by combining excerpts of different audio signals, then for each excerpt, the associated signer (e.g., the owner, organization, or individual of the private key) can be listed along with the authentication result. In a simpler version, it might be possible to display only the list of signers responsible for any excerpt in a given audio signal.

[0100] If the ownership and authentication information of a media file is available before playback begins, the web browser can start displaying the results of the authentication check on the entire media file before playback begins. For streaming media such as audio or video, ownership information (or a link to ownership information) is provided in a way that is available from the very beginning of the stream (e.g., via metadata), or ownership information can be provided along with individual chunks / blocks of the stream (e.g., via watermarks). In the latter case, the information displayed can change over time depending on how much information is available. This might mean that at the beginning of the stream, only a link to the authentication service is available, while a more detailed analysis is displayed once the browser has received the entire media stream.

[0101] In the similarity analyzer described above, audio blocks can be compared with audio features of the original audio signal. These audio features can be true copies of the original audio signal. However, this would require significant storage space to store the original audio signal. If the audio features relate to other characteristics of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Thus, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within the similarity analyzer. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.

[0102] In the example, the watermark could include a temporal sequence of pilot signals (indicating the start of the watermark, the user ID, and the block ID). If a watermark is detected and can only be partially decoded (e.g., only the pilot sequence, without a meaningful user ID or block ID), it is still reasonable to notify the user accordingly. Advanced users (e.g., news organizations) can infer from the information that the content was signed at some point in time but appears to have been altered afterward. The partial presence of the watermark can also be used to refute claims that the file was never altered.

[0103] In another example, if the authentication information (source information) is not bound to the content via a watermark but, for example, via a time-related metadata format, the digital audio workstation can use accurate timestamps to reference all relevant source information for all source files. In this case, since all the information is already available, the authentication process described above is not necessary.

[0104] According to the example, each audio signal captured by microphone 100 and output by the microphone can include a watermark, which can be used to identify the actual microphone that captured the audio signal. If the microphone is used by several users, several user IDs can be associated with the microphone. If the microphone is registered, the corresponding microphone can be identified. User IDs can also be registered. The user ID can be a dedicated ID or a sub-ID of the microphone ID. The microphone ID and / or user ID can be part of the watermark. The sequence number of the audio block can be used to determine which part of the audio signal has been removed.

[0105] Alternatively, the audio signal may be part of a video file.

[0106] Alternatively, the signed authentication information can be included in the audio file (e.g., in an ADM file format) instead of in the watermark. In this case, the audio signal does not need to be modified.

[0107] Alternatively, the processing to generate the signed audio signal can be performed using a software solution based on a pre-existing audio recording or audio stream. The private key can be entered into the software, for example, via a dongle or text input, and authorization can be performed, for example, via biometric signals such as fingerprints or facial detection.

[0108] The authentication process can calculate a similarity score between the audio features of the signed audio signal and the audio features of the analyzed signal. This allows it to determine the likelihood that the signal is still authentic, even if slight signal processing such as gain has been applied.

[0109] Authentication requires an audio signal and signed authentication information. The signed authentication information can be provided via a side channel. When using a watermark, the authentication access information must be read from the audio signal, and the authentication information must be retrieved accordingly. Where watermark reading requires block boundaries, one method to achieve this is to try different offsets of the block boundaries until the watermark can be successfully read. Another method is to embed some watermarked synchronization signal into the audio signal.

[0110] To enhance system security, processing is provided that enables, for example, key revocation in the event of key theft.

[0111] Public keys used for authentication can be provided in various ways. One option is to store all public keys in a centralized database to make it easy to find the required key. To address potential trust issues in this centralized instance, a database with all public keys can be provided on a distributed ledger technology. Another option is for the organization / individual using the technology to provide the public key on their own website.

[0112] Torsten Dau, Dirk Pueschel, and Armin Kohlrausch, in “A quantitative model of the “effective” signal processing in the auditory system. I. model structure”—published in The Journal of the Acoustical Society of America, 99(6): 3615-3622, 1996—described other examples of audio features, hereinafter referred to as the Perceptual Model (PEMO). This paper describes a quantitative model designed to describe how the auditory system processes acoustic signals. The focus is on developing models that mimic the functional processing of auditory stimuli by humans, particularly in contexts involving complex sounds such as speech or music.

[0113] The key components of this model can be categorized into peripheral processing, envelope extraction, modulation filtering, and a decision-making stage. The peripheral processing section captures the initial stages of auditory processing, including converting acoustic signals into neural representations through mechanisms such as outer and middle ear filtering and nonlinearities in the cochlea. The envelope extraction section involves extracting the temporal envelope of the sound, which is crucial for understanding amplitude modulation and speech processing. The modulation filtering section involves a system for filtering amplitude modulation, mimicking the auditory system's sensitivity to different modulation frequencies. In the decision-making stage, the processed auditory signals are used to make decisions about the properties of the sound, thus representing a higher level of auditory perception.

[0114] The model was designed to align with experimental psychoacoustic data, particularly its ability to simulate the auditory system's response to amplitude-modulated sounds. This approach provides a framework for understanding effective signal processing in human auditory perception, especially for tasks involving the detection and discrimination of complex auditory patterns.

[0115] Performing a comparison of the received audio block with the extracted or retrieved authentication information in the transform domain (i.e., in relation to the audio features perceived by human hearing) ensures that the identified differences are perceptually meaningful. In principle, any similarity measure (e.g., empirical cross-correlation coefficient or relative error relative to a reference quantity) can be used to map the comparison results to a numerical value.

[0116] When a further transformation is applied to the Mel frequency power spectrum of the audio signal, Mel frequency cepstral coefficients (MFCCs) are obtained. MFCCs use the Mel scale, a perceptual scale for pitch used to approximate how humans perceive sound. The Mel scale is non-linear, allowing for finer division of lower frequencies and coarser division of higher frequencies, thus better aligning with human auditory perception.

[0117] Mel-frequency cepstral coefficients use the cepstral spectrum as a means of transforming a signal from the frequency domain back to a domain where the rate of change of the signal can be analyzed. This idea is to capture the spectral characteristics of a signal by emphasizing the information-carrying components.

[0118] The process of calculating MFCC from an audio signal can include several steps: pre-emphasis: typically achieved by applying a high-pass filter to enhance the high-frequency components of the signal to balance the spectrum; framing: dividing the audio signal into short overlapping frames, typically 20 to 40 milliseconds long, because speech signals are non-stationary but can be considered quasi-stationary within these short frames; windowing: multiplying each frame by a window function, such as a Hamming window, to reduce edge effects and smooth the signal; Fourier transform: applying a Fast Fourier Transform (FFT) to each windowed frame to convert the time-domain signal to the frequency domain; Mel filter bank: passing the resulting spectrum through a series of triangular bandpass filters spaced according to the Mel scale, a step that approximates how humans perceive sound frequencies; logarithmic calculation: calculating the logarithm of the power of each Mel-filtered spectrum, which simulates the human ear's response to loudness; Discrete cosine transform (DCT): finally, using the DCT to transform the log-Mel spectrum. The result is a set of coefficients called MFCC. Typically, only the first 12 to 13 coefficients are retained, as these contain the most important information.

[0119] MFCC can capture the wide spectral shape of an audio signal in a compact form, making it highly useful for machine learning algorithms in speech and audio processing. By providing robust representations of variations in pitch, volume, and other factors, MFCC helps distinguish different phonemes, speaker identity, and other audio characteristics.

[0120] Audio features may also include:

[0121] 1. Linear Predictive Coding (LPC) is a method that uses information from a linear predictive model to represent the spectral envelope of a digital speech signal in compressed form. LPC estimates can be used to reconstruct the parameters of filters used to reconstruct the signal. LPC is very effective for modeling the formants (resonant frequencies) of speech sounds.

[0122] 2. Perceptual Linear Prediction (PLP) coefficients, which are similar to linear predictive coding but incorporate various aspects of human auditory perception, such as critical band spectral resolution, equal-loudness curves, and intensity-loudness power laws. Perceptual Linear Prediction (PLP) coefficients are designed to mimic the nonlinear perception of loudness and frequency by the human ear.

[0123] 3. Gammatone Filterbank Features involve a filterbank that simulates the filtering that occurs in the human cochlea. It is similar to a Mel filterbank, but uses Gammatone filters instead of triangular filters. Gammatone filterbank features are useful for capturing the detailed frequency structure of audio, especially in tasks involving environmental sound classification or hearing aid design.

[0124] 4. Chromaticity features, which represent the 12 different pitch categories (C, C#, D, etc.) of a musical octave. Chromaticity features capture the harmonic and melodic characteristics of music. These are particularly useful in music information retrieval, key detection, and chord recognition tasks.

[0125] 5. Mel Spectrum: While similar to MFCC, the Mel spectrogram is the result of applying a Mel filter bank directly to the power spectrum without further transformations (such as DCT). The Mel spectrogram preserves more detailed frequency information and is often used as input to deep learning models, often in conjunction with convolutional neural networks (CNNs) in tasks such as audio event detection, speech recognition, and music genre classification.

[0126] 6. The Constant-Q Transform (CQT) provides a time-frequency representation with a logarithmic frequency scale similar to the Mel scale, but with a variable time resolution that matches the frequency resolution. CQT is particularly useful for music applications because it provides a better representation of musical pitch compared to linear FFT or even the Mel scale.

[0127] 7. Deep Learning-Based Features: Features learned from deep learning models, such as embeddings from trained neural networks, can also be used as alternatives to traditionally hand-designed features like MFCCs. These features are generally more robust and can capture complex patterns that are difficult to model using traditional methods.

[0128] 8. Spectral Subband Centroids (SSCs) are the centroids of the energy distribution in different subbands of the captured signal. They can be interpreted as the "centroid" of the spectral lines within each subband. SSCs provide information about the energy distribution across the frequency bands and are sometimes used as a supplement to MFCCs.

[0129] 9. Relative Spectral (RASTA) features involve filtering the logarithmic energy of the speech signal to highlight modulation frequencies important for speech recognition. RASTA-PLP is a combination of RASTA filtering and PLP analysis, providing robust characteristics against noise and channel variations.

[0130] Audio features are related to the audio representation of an audio signal and attempt to capture relevant aspects of human auditory perception of the audio signal. Audio features should be selected to reduce the required data rate (for transmission of the audio features) and thus reduce the storage space required on the server. While audio features can be the original audio signal itself, preferably, they can be a transparently coded version of the original signal, which requires significantly less storage space. Another example of an audio feature can be the output of any model that mimics human auditory perception. Audio features can be the Mel-frequency power spectrum and Mel-frequency cepstral coefficients (MFCCs) of the audio signal.

[0131] Specifically, audio features can be the acoustic properties, spectral properties, and statistical properties of an audio signal.

[0132] Acoustic features may include pitch (fundamental frequency); prosody (rhythm, stress, intonation) and / or duration and silence patterns (pauses, speech timing, or inconsistencies in breath sounds can be used as audio features).

[0133] Spectral features may include formant frequencies (such as resonant frequencies in speech); Mel frequency cepstral coefficients (MFCCs) (the short-time power spectrum of sound, which can reveal artifacts introduced during synthesis or tampering); spectrogram analysis (tampering may present an anomalous or obscured energy distribution in the time-frequency representation); and / or phase information.

[0134] Acoustic features can include statistical features and signal processing features, such as noise and residuals (e.g., subtle background noise or artifacts introduced during synthesis may not match natural recordings); high-frequency content; and / or phase distortion (some tampering may introduce phase anomalies, which can be detected using signal processing techniques).

[0135] Acoustic features can include temporal features such as jitter and flicker (e.g., frequency and amplitude variations); and temporal coherence (e.g., sudden shifts or changes in speech characteristics, such as unnatural interruptions in the signal).

[0136] Acoustic features can include behavioral or semantic inconsistencies, such as content coherence (logical inconsistencies or unnatural fluency in spoken content may indicate tampering); affect and naturalness (emotional tone may be inconsistent with the content, or the voice may lack the natural nuances of human emotion).

[0137] Audio features can be any one or a combination of the audio characteristics mentioned above for an audio signal.

[0138] By combining these features with machine learning or signal processing techniques, models can be trained to detect artifacts and inconsistencies that indicate deep fake audio.

Claims

1. A method for authenticating audio signals of media files associated with a webpage, comprising the following steps: The method involves a user selecting a webpage via a web browser, wherein the webpage provides content having at least one media file including an audio signal, wherein the content is stored on a web server, and wherein the audio signal comprises multiple audio blocks, each having authentication information or authentication access information related to how or where to retrieve externally stored authentication information. The method further includes the steps of: receiving the audio signal from the web server, extracting or retrieving authentication information from the audio blocks, performing an authentication check on the audio blocks of the received audio signal based on the authentication information, and displaying the result of the authentication check via the web browser.

2. The method according to claim 1, wherein, The following steps are performed in a web browser, web browser extension, or web browser plugin: receiving the audio signal from the web server, extracting or retrieving authentication information of the audio block, and performing authentication checks on the audio block of the received audio signal based on the authentication information.

3. The method according to claim 1 or 2, wherein, The display of the authentication check results is arranged in or near a symbolic representation of the media file containing the audio signal to be authenticated.

4. The method according to claim 1, 2 or 3, wherein, Those audio blocks or groups of audio blocks that failed the authentication check were highlighted on the web browser.

5. The method according to any one of claims 1 to 4, further comprising the following step: The owner of the private key (142) that initially signed the audio signal is determined based on the authentication information or the authentication access information, and the name of the owner is displayed to the user of the web browser.

6. The method according to any one of claims 1 to 5, wherein, The current time-related status of the authentication check result is displayed via the web browser.

7. The method according to any one of claims 1 to 6, wherein, The web browser displays the results of the authentication check over time for the entire analyzed sequence of the audio signal.

8. The method according to any one of claims 1 to 7, wherein, Explicitly mark the occurrence of several blocks with negative authentication results.

9. The method according to any one of claims 1 to 8, further comprising the following step: The system detects clips in the received or extracted audio signal, wherein truncated sequences of multiple audio blocks, each containing authentication information or authentication access information, are successively spliced ​​together, and displays the appearance of these clips to the user of the web browser.

10. The method according to any one of claims 1 to 9, further comprising the following step: First, check whether the audio signal includes authentication information or authentication access information. If no authentication information or authentication access information is detected, play back the selected media file without displaying the authentication check result via the web browser.

11. The method according to any one of claims 1 to 8, further comprising the following step: First, check whether the audio signal includes authentication information or authentication access information. If no authentication information or authentication access information is detected, then play the selected media file via the web browser and display information about the lack of authentication to the user.

12. A web browser or web browser extension configured to: access a webpage providing content having at least one media file including an audio signal, wherein, The content is stored on a web server, wherein the audio signal comprises multiple audio blocks, each having authentication information or authentication access information, the authentication access information relating to how or where to retrieve externally stored authentication information, receiving the audio signal from the web server, extracting or retrieving authentication information of the audio blocks of the received audio signal, performing authentication checks on the audio blocks of the received audio signal based on the authentication information, and displaying the results of the authentication checks via the web browser.

13. A web browser extension configured to: receive an audio signal to be authenticated, extract or retrieve authentication information of audio blocks of the received audio signal, perform an authentication check on the audio blocks of the received audio signal based on the extracted or received authentication information, and output the result of the authentication check to a web browser associated with the web browser extension.

14. A computer program product for authenticating audio signals, wherein, The computer program product includes program code units that cause the web browser or web browser extension according to claim 12 to perform the method according to any one of claims 1 to 11.