Audio signal authenticator and method of authenticating audio signal
By using an audio signal authenticator and authentication method, utilizing block boundary detection and perceptual similarity analysis, and combining audio features from a human auditory model, the problem of difficulty in authenticating the authenticity and source of audio signals is solved, achieving efficient and robust audio signal authentication and enhanced security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SENNHEISER ELECTRONICS GMBH & CO KG
- Filing Date
- 2024-09-23
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to effectively verify the authenticity and origin of audio signals, especially when faced with the challenge of deepfake audio and video, where traditional methods are prone to failure.
An audio signal authenticator, including a block boundary detector and a block segmenter, is employed. By embedding authentication information into audio blocks, a perceptual similarity analyzer and signature verification technology are used, combined with a human auditory perception model, to calculate audio features such as Mel frequency power spectrum and Mel frequency cepstral coefficients, thereby achieving robust audio signal authentication.
It achieves robust authentication of audio signals, can detect minor tampering, improves the security and reliability of audio signal authenticity verification, and reduces the data rate required for storage and transmission.
Smart Images

Figure CN121889792A_ABST
Abstract
Description
[0001] This invention relates to an audio signal authenticator and a method for authenticating audio signals.
[0002] Given the current ability to create deepfake videos and audio, there is an increasing need for the ability to verify or authenticate audio signals such as politicians' speeches.
[0003] To date, authenticating an individual's audio signal has been extremely difficult. Typically, extensive manual research is required to verify or authenticate such an audio signal and to verify its origin.
[0004] One method for authenticating audio signals is to analyze their characteristics and search for irregularities that suggest tampering. However, this method often fails due to the ever-increasing ability of AI to generate manipulated signals.
[0005] Therefore, the object of the present invention is to provide an audio signal authenticator and a method for authenticating audio signals that enable the authentication of various different audio signals in a simple and robust manner.
[0006] This objective is achieved by the audio signal authenticator according to claim 1 and the method for authenticating audio signals according to claim 10.
[0007] Therefore, an audio signal authenticator is provided, comprising a block boundary detector and a block segmenter configured to receive an audio signal and output a sequence of multiple audio blocks, each audio block having embedded authentication information or embedded authentication access information relating to where the authentication information can be retrieved. The authentication information includes at least audio features belonging to the current audio block. The authenticator also includes a perceptual similarity analyzer configured to: compare audio blocks of the received audio signal with audio features of extracted or retrieved authentication information, or compare audio features extracted from audio blocks of the received audio signal with audio features of extracted or retrieved authentication information, thereby providing similarity results.
[0008] Audio features are the audio features of human auditory perception that are related to human auditory perception.
[0009] In a similarity analyzer, audio blocks can be compared to the audio features of the original audio signal. These audio features can be true copies of the original audio signal. However, this would require significant storage space to store the original audio signal. If the audio features are related to other features of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Thus, the desired audio features can be extracted from the audio signal to be authenticated. This can be performed, for example, within a similarity analyzer. In other words, the similarity analyzer should compare audio features (such as those received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.
[0010] When authenticating audio signals, audio blocks are analyzed to directly extract authentication information or to extract authentication access information from the audio blocks. Authentication access information includes information about where or how authentication information can be retrieved. Audio features are related to audio representations that attempt to capture relevant aspects of human auditory perception.
[0011] According to the example, human auditory perception audio features are calculated by mimicking models of human auditory perception, particularly perceptual models. Alternatively, human auditory perception audio features are the Mel-frequency power spectrum of an audio block.
[0012] According to the example, the perceptual similarity analyzer is configured to perform soft authentication decisions, thereby providing the probability of successful authentication as a result.
[0013] According to the example, the audio signal includes authentication access information about where or how authentication information can be retrieved. The audio signal is received and a sequence of multiple audio blocks with embedded authentication access information is output. Authentication access information is extracted from the current audio block, and access information about how and / or where signed authentication information belonging to the current block can be retrieved from memory is also extracted. The signed authentication information is retrieved from memory. The signed authentication information includes audio features belonging to the current block and a signature. The audio blocks and audio features are compared to each other to provide similarity results. The audio features are human auditory perception audio features related to human auditory perception.
[0014] Each audio block has specific audio characteristics. For robust authentication of the audio signal, and to achieve perceptually transparent coding at a sufficiently high bit rate, audio characteristics that are insensitive to typical audio processing such as resampling or MP3 or AAC conversion are required. Preferably, such audio characteristics will not be affected by such audio processing.
[0015] Acoustic stimuli can be translated into an internal audio representation for the user. In other words, the audio signal received by the user will elicit a specific response in their auditory system. The audio features of an audio block can be correlated with such an internal representation. The internal representation is correlated with auditory system-specific features of the audio signal. In other words, the internal representation is correlated with audio features relevant to human auditory perception. Human auditory perception audio features can be correlated with audio representations that attempt to capture relevant aspects of human auditory perception.
[0016] During the generation of an authentic audio signal, audio blocks of the audio signal can be analyzed to determine corresponding audio features such as corresponding internal representations, i.e., human auditory perception audio features.
[0017] Audio features are related to the audio representation of audio signals that attempt to capture relevant aspects of human auditory perception. Another example of audio features could be the output of any model that mimics human auditory perception.
[0018] An example of the characteristics of audio perceived by human hearing is the Mel frequency power spectrum of an audio signal.
[0019] Torsten Dau, Dirk Pueschel, and Armin Kohlrausch, in “A quantitative model of the “effective” signal processing in the auditory system. I. model structure”—The Journal of the Acoustical Society of America, 99(6): 3615-3622, 1996—described another example of the characteristics of human auditory perception of audio, hereinafter referred to as the Perceptual Model (PEMO). This paper describes a quantitative model designed to describe how the auditory system processes acoustic signals. The focus is on developing a model that mimics the functional processing of auditory stimuli by humans, particularly in contexts involving complex sounds such as speech or music.
[0020] The key components of this model can be categorized into peripheral processing, envelope extraction, modulation filtering, and a decision-making stage. The peripheral processing section captures the initial stages of auditory processing, including converting acoustic signals into neural representations through mechanisms such as outer and middle ear filtering and nonlinearities in the cochlea. The envelope extraction section involves extracting the temporal envelope of the sound, which is crucial for understanding amplitude modulation and speech processing. The modulation filtering section involves a system for filtering amplitude modulation, mimicking the auditory system's sensitivity to different modulation frequencies. In the decision-making stage, the processed auditory signals are used to make decisions about the properties of the sound, thus representing a higher level of auditory perception.
[0021] The model was designed to align with experimental psychoacoustic data, particularly its ability to simulate the auditory system's response to amplitude-modulated sounds. This approach provides a framework for understanding effective signal processing in human auditory perception, especially for tasks involving the detection and discrimination of complex auditory patterns.
[0022] According to one aspect, the authentication access information parser is configured to determine the identity information of the user or entity that initially signed the audio signal, wherein the user's public key and identity can be retrieved from an external database based on the determined identity information.
[0023] One approach provides a signature verifier, which is configured to use a public key to verify a combination of signature and authentication information and provide a signature verification result.
[0024] According to one aspect, the audio signal includes a combiner configured to combine a perceived similarity claim and a verification result into a final authentication result.
[0025] Performing a comparison between the received audio block and the extracted or retrieved authentication information in the transform domain (i.e., in relation to the audio features perceived by human hearing) ensures that the identified differences are perceptually meaningful. In principle, any similarity measure (e.g., empirical cross-correlation coefficient or relative error relative to a reference quantity) can be used to map the comparison results to a numerical value.
[0026] When a further transformation is applied to the Mel frequency power spectrum of the audio signal, Mel frequency cepstral coefficients (MFCCs) are obtained. MFCCs use the Mel scale, a perceptual scale for pitch used to approximate how humans perceive sound. The Mel scale is non-linear, allowing for finer division of lower frequencies and coarser division of higher frequencies, thus better aligning with human auditory perception.
[0027] Mel-frequency cepstral coefficients use the cepstral spectrum as a means of transforming a signal from the frequency domain back to a domain where the rate of change of the signal can be analyzed. The idea is to capture the spectral characteristics of a signal by emphasizing the information-carrying components.
[0028] The process of calculating MFCC from an audio signal can include several steps: pre-emphasis: typically achieved by applying a high-pass filter to enhance the high-frequency components of the signal to balance the spectrum; framing: dividing the audio signal into short overlapping frames, typically 20 to 40 milliseconds long, because speech signals are non-stationary but can be considered quasi-stationary within these short frames; windowing: multiplying each frame by a window function, such as a Hamming window, to reduce edge effects and smooth the signal; Fourier transform: applying a Fast Fourier Transform (FFT) to each windowed frame to convert the time-domain signal to the frequency domain; Mel filter bank: passing the resulting spectrum through a series of triangular bandpass filters spaced according to the Mel scale, a step that approximates how humans perceive sound frequencies; logarithmic calculation: calculating the logarithm of the power of each Mel-filtered spectrum, which simulates the human ear's response to loudness; Discrete cosine transform (DCT): finally, using the DCT to transform the log-Mel spectrum. The result is a set of coefficients called MFCC. Typically, only the first 12 to 13 coefficients are retained, as these contain the most important information.
[0029] MFCC can capture the wide spectral shape of an audio signal in a compact form, making it highly useful for machine learning algorithms in speech and audio processing. By providing robust representations of variations in pitch, volume, and other factors, MFCC helps distinguish different phonemes, speaker identity, and other audio characteristics.
[0030] Human auditory perception of audio features may also include:
[0031] - Linear Predictive Coding (LPC): This method uses information from a linear predictive model to represent the spectral envelope of a digital speech signal in compressed form. LPC estimates can be used to estimate the parameters of filters that can be used to reconstruct the signal. LPC is very effective for modeling the formants (resonant frequencies) of speech sounds.
[0032] - Perceptual Linear Prediction (PLP) Coefficients: These are similar to linear prediction coding, but incorporate various aspects of human auditory perception, such as critical band spectral resolution, equal-loudness curves, and intensity-loudness power laws. Perceptual Linear Prediction (PLP) coefficients are designed to mimic the nonlinear perception of loudness and frequency by the human ear.
[0033] - Gammatone Filterbank Features: These are filter banks that simulate the filtering that occurs in the human cochlea. They are similar to Mel filter banks, but use gammatone filters instead of triangular filters. Gammatone filterbank features are used to capture the detailed frequency structure of audio, especially in tasks involving environmental sound classification or hearing aid design.
[0034] - Chromaticity features: These represent 12 different pitch categories (C, C#, D, etc.) of a musical octave. They capture the harmonic and melodic features of music. These are particularly useful in music information retrieval, key detection, and chord recognition tasks.
[0035] - Mel spectrogram: While similar to MFCC, the Mel spectrogram is the result of applying a Mel filter bank directly to the power spectrogram without further transformations (such as DCT). The Mel spectrogram preserves more detailed frequency information and is often used as input to deep learning models, often in conjunction with convolutional neural networks (CNNs) in tasks such as audio event detection, speech recognition, and music genre classification.
[0036] - Constant-Q Transform (CQT): This provides a time-frequency representation with a logarithmic frequency scale similar to the Mel scale, but with a variable time resolution that matches the frequency resolution. It is particularly useful for music applications because it provides a better representation of musical pitch compared to linear FFT or even the Mel scale.
[0037] - Deep learning-based features: Features learned from deep learning models, such as embeddings from trained neural networks, can also be used as alternatives to traditionally hand-designed features like MFCCs. These features are generally more robust and can capture complex patterns that are difficult to model using traditional methods.
[0038] - Spectral Subband Centroids (SSCs) are the centroids of the energy distribution in different subbands of the captured signal. They can be interpreted as the "centroid" of the spectral lines within each subband. SSCs provide information about the energy distribution across the frequency bands and are sometimes used as a supplement to MFCCs.
[0039] - Reactive Spectral Analysis (RASTA) features: These involve filtering the logarithmic energy of the speech signal to highlight modulation frequencies important for speech recognition. RASTA-PLP is a combination of RASTA filtering and PLP analysis, providing robust features against noise and channel variations.
[0040] Audio features refer to the audio representation of an audio signal that attempts to capture relevant aspects of human auditory perception of the audio signal. Audio features should be selected to reduce the required data rate (for transmission of the audio features) and thus reduce the storage space required on the server. While audio features can be the original audio signal itself, preferably, they can be a transparently coded version of the original signal, which requires significantly less storage space. Another example of an audio feature can be the output of any model that mimics human auditory perception. Audio features can be the Mel-frequency power spectrum and Mel-frequency cepstral coefficients (MFCCs) of the audio signal.
[0041] Specifically, audio features can be the acoustic properties, spectral properties, and statistical properties of an audio signal.
[0042] Acoustic features may include pitch (fundamental frequency); prosody (rhythm, stress, intonation) and / or duration and silence patterns (pauses, speech timing, or inconsistencies in breath sounds can be used as audio features).
[0043] Spectral features may include formant frequencies (such as resonant frequencies in speech); Mel frequency cepstral coefficients (MFCCs) (the short-term power spectrum of sound can reveal artifacts introduced during synthesis or tampering); spectrogram analysis (tampering may manifest as unusual or obscured energy distributions in the time-frequency representation); and / or phase information.
[0044] Acoustic features can include statistical and signal processing features such as noise and residuals (e.g., subtle background noise or artifacts introduced during synthesis may not match natural recordings); high-frequency content; and / or phase distortion (some tampering may introduce phase anomalies, which can be detected using signal processing techniques).
[0045] Acoustic features can include temporal features such as jitter and flicker (e.g., frequency and amplitude variations); and temporal coherence (e.g., sudden transitions or changes in speech features, such as unnatural interruptions in the signal).
[0046] Acoustic features can include behavioral or semantic inconsistencies, such as content coherence (logical inconsistencies or unnatural flow in spoken content may indicate tampering); emotion and naturalness (emotional tone may be inconsistent with content, or the voice may lack the natural nuances of human emotion).
[0047] Audio features can be any one of the above-mentioned audio characteristics of an audio signal or a combination thereof.
[0048] By combining these features with machine learning or signal processing techniques, models can be trained to detect artifacts and inconsistencies that indicate deep fake audio.
[0049] According to one aspect, a method for authenticating an audio signal by an authenticator is provided, the method comprising the steps of: receiving an audio signal; outputting a plurality of audio blocks having embedded authentication access information; extracting the authentication access information from the audio blocks, the authentication access information indicating where authentication information belonging to the current audio block can be retrieved; retrieving the authentication information from a memory, including a signature, audio features representing the audio block, and / or other information; and perceptually comparing the audio blocks and audio features to provide similarity results.
[0050] The authentication information can be signed using the private key of a microphone or a device that performs audio signal capture or post-processing to generate a unique signature indicating the source of the authentication information. The private key can be associated with the user of the signing module (e.g., the microphone). On the decoder side, the public key can be used to decode or verify the authentication information. In other words, the signature is used to verify the source of the authentication information.
[0051] Therefore, a signature can be used to verify that the authentication information originates from a source owned by the private key. Further comparison of the audio signal with the authentication information can verify that the audio signal has not been tampered with beyond permissible levels. The signature is generated solely based on, for example, the private key of a microphone.
[0052] It is advantageous to store authentication information, such as metadata, externally rather than within the audio block, as this does not increase the required bit rate and bandwidth. Only the authentication access information is transmitted with the audio block. If authentication of the audio signal is required, the authentication access information is extracted, and the associated authentication information, such as metadata, is retrieved.
[0053] Therefore, authentication information is provided in audio blocks independent of the captured audio signal. If authentication is not performed, no authentication information is needed, and significant bandwidth and bitrate savings are achieved compared to attaching authentication information to the audio signal in a suitable container format. Furthermore, when authentication is required, access information can be extracted from the audio blocks, and authentication information can be retrieved. Since the authentication process is latency-insensitive, extracting authentication information from an external source before performing authentication is sufficient.
[0054] Authentication access information can be embedded as a watermark into an audio block, or as part of an audio block into a dedicated container, or as part of an audio file into a file-based environment.
[0055] Authentication information and authentication access information can be provided for at least one audio block. Thus, each audio block can be authenticated individually. This enables granular authentication. Therefore, even slight tampering with the audio signal (with its authentication information available) can be detected. This significantly improves the security of the captured audio signal. Providing authentication information for each audio block is advantageous because it even makes it possible to verify audio signals including different audio blocks from different audio sources. Any audio signal, including different audio blocks, can be verified, provided the necessary authentication information is available.
[0056] The following describes examples of enhancing the security of audio signal transmission or storage. Audio features can be extracted from blocks of digital audio signals, and signatures based on these audio features can be generated using a person's or entity's private key. The captured audio signal and the signed authentication information can be output.
[0057] Therefore, it is possible to determine whether the transmitted audio signal was captured and signed by a specific software entity or a specific microphone with a specific private key by verifying the combination of the received authentication information and signature using a public key and by subsequently comparing the audio signal with the audio features included in the authentication information.
[0058] According to one aspect of the invention, the microphone also includes a watermark generator that generates a watermark based on authentication access information.
[0059] The signing of audio blocks and the verification of the signatures can be performed using a private key and public key pair.
[0060] According to one aspect of the present invention, in addition to audio characteristics, authentication information may include microphone identifier, microphone location / positioning information, recording time and date and / or microphone model, serial number, etc.
[0061] Authentication information comprises information that can be used to authenticate the received or selected audio signal. Authentication information enables the authenticator to authenticate the audio signal. Authentication information may include information related to the audio signal or its attributes (i.e., audio characteristics). Authentication information may also include metadata related to information independent of the actual audio signal. This information may be the user's ID, the ID of the microphone used to capture the audio signal, the microphone model, date, time, location or positioning (GPS location), and / or the serial number of the audio block.
[0062] The authenticity check of the received audio signal can be performed by a decoder, which can be implemented on, for example, a cloud service, a computer, a tablet, or a smart device.
[0063] A block splitter can be used to divide an audio signal into multiple audio blocks. The length of an audio block can be, for example, between 0.5 seconds and 20 seconds. In particular, the block length can vary over time, for example, this can be indicated by a watermark. A digital watermark can be embedded in each audio block. A watermark can also be embedded in every nth audio block.
[0064] The sequence of audio blocks in the audio signal can be used to associate a sequence number with each audio block. This is advantageous because later, during the authentication of the audio signal, it can be determined whether an audio block has been removed from the audio block chain by checking the sequence number of the audio block. This can also detect changes in the order of the original blocks.
[0065] Optionally, the metadata of the audio block can include the microphone user, date, time, GPS location, etc. This increases the probability of reliable authentication results.
[0066] Authentication information may include audio features based on human perception, representing audio signals or audio blocks of audio signals, or a transparently encoded version of the original audio signal, in order to reduce the storage space required by the server.
[0067] Authentication of an audio signal can be performed by detecting block boundaries based on specific information in the watermark or audio container, and by extracting authentication access information (e.g., links (e.g., including a UUID and a user-based unique block number) pointing to block-specific authentication information from the audio watermark or directly from the audio container). The user's public encryption key (e.g., identified by the UUID in the link) can be used to verify the authenticity of the authentication information. The audio signal block is compared with the audio features contained in the authentication information to verify its authenticity.
[0068] Using the method for authenticating audio signals described above, malicious modification of the audio signal can be detected by comparing the audio block with the authentication information associated with the audio block. Modification of the audio signal to be authenticated and / or modification of authentication information stored externally can also be detected by examining the signature of the authentication information. If the audio block of the audio signal is modified and the authentication access information is also changed, this will be noticed by the authentication process because the signature of the authentication information is that of another person or device.
[0069] According to one perspective, if the authentication process does not provide an explicit indication of the authenticity of an audio signal or audio block, the authentication information stored externally can be used to manually authenticate the audio signal or audio block.
[0070] According to the example, the actual audio features extracted from the audio block and optional additional metadata are stored, for example, on a server in a signed manner. A unique (explicit) link to the metadata stored on the server is embedded in a watermark or in a suitable audio container. The link may include a user-based unique block number for the corresponding audio block along with a unique user identifier (UUID). The audio features involve an audio representation of an audio signal that attempts to capture relevant aspects of human auditory perception, and preferably at a reduced data rate and thus a smaller storage space required on the server. While the audio features may be the original audio signal itself, preferably, the audio features may be a transparently encoded version of the original signal that requires significantly less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.
[0071] To authenticate the audio signal of an audio block, block boundaries can be detected based on a watermark or specific information in the audio container. Then, a link (e.g., including a UUID and a user-based unique block number) pointing to block-specific, signed authentication information (including audio features, optional metadata, and a signature) is extracted from the audio watermark or directly from the audio container. The combination of authentication information and signature is verified using the user's public cryptographic key (e.g., identified by the UUID in the link). Accordingly, it is verified whether the authentication information has been signed by the claimed audio signal source. The audio signal block is compared with audio features to verify perceptual similarity, and thus the authenticity of the audio signal block. For perceptual comparison, audio features can be determined from the received audio signal block, and these features can be directly compared with audio features from, for example, signed authentication information stored on a server.
[0072] The proposed authentication method prevents any malicious modification to the actual audio signal by comparing the audio signal block with the signed authentication information. Authentication may fail if the comparison allows a limited amount of modification, but the allowed threshold is exceeded. Furthermore, if the authentication information on the server is altered accordingly along with the malicious modification of the audio signal block, this can be detected by a mismatch between the authentication information and the signature associated with it. Additionally, if, along with the malicious modification of the audio signal block, the link (within the watermark or audio container) is changed to point to a different (but appropriate) signed authentication information on the server, signature verification and subsequent comparison may indeed succeed, but the signature will be that of a different user—provided the attacker cannot obtain the original user's private key used for the encrypted signature.
[0073] If the (automatic) comparison between the audio signal to be authenticated and the authentication information is performed in a manner that allows for minor modifications (e.g., through perceptual encoding), and the result of the comparison is ambiguous (i.e., no explicit statement can be made about the authenticity of the audio signal), then a "manual" perceptual comparison can be performed using a link to the authentication information (where the included audio features will correspond to the original audio file or its perceptually encoded version) to increase trust in the method.
[0074] By providing audio block-specific links (e.g., via UUID and user-based unique block numbers), it becomes possible to have the advantage (e.g., in watermarking scenarios) that audio files can be cut at any time without the cutting tool being aware of authentication processing and the embedded links. Alternatively, in cases where the audio container is used for the joint delivery of audio and links, file-based links may be embedded all at once in the included header to reduce data rate. The link may include a UUID and a location within metadata (e.g., a sample index of the original file) to indicate how the audio file in question is aligned with a reference file on the server.
[0075] These and other aspects of the invention are described in more detail with reference to the following accompanying drawings.
[0076] Figure 1 A block diagram of the audio signal signature module is shown.
[0077] Figure 2 A diagram showing the overall audio signature and authentication workflow is provided.
[0078] Figure 3 A block diagram of an audio signal authenticator is shown.
[0079] Figure 4 A block diagram of the audio signal signature module is shown.
[0080] Figure 5 A block diagram of an audio signal authenticator is shown.
[0081] Figure 6 A block diagram of the audio signal signature module is shown.
[0082] Figure 7 A block diagram of an audio signal authenticator is shown, and
[0083] Figure 8 An example of how to visualize authentication results is shown.
[0084] According to the example, authentication is performed on an audio signal. The audio signal is authenticated based on information embedded in an audio block. This embedded information can be, for example, authentication information directly embedded in the audio block using a watermark, or alternatively, authentication access information. The authentication access information includes information about where and how the stored authentication information associated with the audio block can be retrieved from memory (such as a server). In this case, the authentication information is not directly embedded in the audio block.
[0085] To better understand the authentication of audio signals, we will first describe the generation of authenticable audio signals.
[0086] Figure 1 A diagram of an audio signal signature module is shown. The audio signal signature module 500 receives an input audio signal 121 from a microphone 100 or another audio source, such as an audio recorder, via an audio input terminal 502, and outputs an output audio signal 501 via an audio output terminal 190. The output audio signal 501 may be based on the input audio signal 121. The microphone 100 may include at least one microphone capsule 110 and an analog-to-digital converter (ADC) 120. The at least one microphone capsule 110 can capture an audio signal and output the captured analog audio signal 111. The captured analog audio signal 111 can be forwarded to the ADC 120, which can digitize the audio signal 111 and output a digital audio signal 121. The audio signal signature module 500 may be implemented as a device separate from the microphone 100 or may be included within the microphone 100.
[0087] The audio signal signature module 500 can receive audio signal 121 from microphone 100 or from other sources, which in Figure 1 The input is indicated by selector 101. Digital audio signal 121 is input to block divider 130, which divides the audio signal 121 into audio blocks 131 and outputs a sequence of multiple audio blocks 131. Optionally, if the audio signal 121 has already been divided into audio blocks 131, block divider 130 can be omitted. The sequence of audio blocks 131 is forwarded to audio feature extractor 150, which generates audio features 151 for each audio block 131 in the audio block sequence. Private key signing unit 140 receives authentication information (e.g., metadata) including the audio features 151 of each audio block 131 and optional other information 113 such as location or time, and signs the authentication information using private key 142 to provide signed authentication information 141, which can be stored, for example, on memory 600 located on an external server. Signed authentication information 141 may include authentication information 152 (see... Figure 2The authentication information 152 includes audio features 151 and optional other information 113, as well as a signature 144 generated by the private key signing unit 140. Figure 2 Alternatively, the authentication information 141 may be stored in internal memory or any other memory, as long as it can be accessed by the authenticator that authenticates the received audio signal.
[0088] The signed authentication information 141 itself includes authentication information, namely (a) audio feature 151, optionally (b) other information 113, and (c) signature 144. Individual audio blocks 131 are fed in parallel, for example, to an information embedder 160 for embedding authentication access information 162 into, for example, each audio block 131. Authentication access information 162 includes information about where and / or how the authentication information 152 (or more precisely, the signed authentication information 141) for that particular audio block 131 can be accessed, for example, at a memory 600 located on a server. This embedding can be accomplished, for example, by means of a watermark.
[0089] Therefore, the information embedder 160 can be implemented as a watermark generator that generates watermark 163. Alternatively, the authentication access information 162 can be embedded into a suitable audio container adjacent to the actual audio block 131. Accordingly, the information embedder 160 outputs the audio block 161 with the embedded authentication access information 162. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at the audio output terminal 190. The output audio signal 501 can be transmitted, for example, via network 300 (…). Figure 2 The output audio signal 501 can be broadcast or stored. In particular, the output audio signal 501 can even be further modified to a very small extent, including modification types such as sample rate conversion, perceptual compression, and trimming.
[0090] When the information embedder 160 is a watermark generator, its output 161, instead of the original audio block 131, can alternatively be input to the audio feature extractor 150, which in Figure 1 The selection is indicated by selector 102. This alternative offers the advantage of extracting audio features based on the final signal 161 output by the audio signal signature module 500, rather than the unoutput intermediate signal 131. Since the watermark generator is designed to perceptually alter the audio signal as little as possible, both input options (131 or 161) are reasonable for the audio feature extractor 150.
[0091] Audio features involve an audio representation that attempts to capture relevant aspects of human auditory perception, while preferably reducing the required data rate and thus the storage space required on the server or memory 600. While audio features can be the original audio signal itself, preferably, they can be a transparently encoded version of the original signal that requires far less storage space. Another example of audio features could be the output of any model that mimics human auditory perception.
[0092] The audio signal 501 (with embedded watermark or authentication access information) at the output of the audio signal module 500 can be stored or transmitted. Because the watermark can be embedded in, for example, essentially every audio block 131, each audio block 131 can be authenticated individually. The watermark 103 can also be embedded only in some audio blocks within the audio block set. In addition to authentication access information, the watermark 163 can also include other information such as block boundaries. This information can be used when authenticating the audio signal to determine the block length.
[0093] The audio watermark 163 can be a unique identifier embedded in the audio signal and, for example, previously used to identify copyright information of the audio signal. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, the embedded watermark will not be altered when the audio signal with the embedded watermark is copied, stored, or transmitted. The same reasoning applies to authentication access information included in the audio container.
[0094] Figure 2 A diagram illustrating the overall audio signing and authentication workflow is provided. An audio signal is captured by, for example, a microphone 100, which outputs a digital audio signal 121. Audio features 151 are extracted from this audio signal 121 and signed using, for example, a private key 142 of an audio signal signing module 500. The audio features 151, together with a digital signature 144 generated by a signing unit 140, constitute signed authentication information 141. The signed authentication information 141 can be forwarded to a server or storage 600, where it is stored for later retrieval. Information regarding where and / or how to retrieve the signed authentication information 141, namely authentication access information 162, is provided to the audio signal signing module 500. The audio signal signing module 500 uses the authentication access information 162 and embeds it into an audio block 131 of the audio signal. Therefore, the output signal 501 of the audio signal signing module 500 includes the audio block 131 with the embedded authentication access information. The audio block 131 of the audio signal with embedded authentication access information 162 is output by the audio signal signature module 500 and can be stored or distributed, for example, via network 300.
[0095] This distribution may involve modification or tampering with authentication information 152 and / or audio signal 121. Public key 143 is used to verify the signed authentication information 141. This verification indicates potential modifications to the audio features. The integrity of the audio signal can be verified by extracting audio features from the received signal and comparing them with the features included in the signed authentication information 141. The final authentication result 261 is determined jointly by these two checks.
[0096] An audio signal signature module 500 (e.g., implemented as part of microphone 100) is used to capture audio signals. Alternatively, the audio signal signature module 500 can receive audio signals. The audio signal can also be an audio signal pre-captured by other means. See reference... Figure 1 The captured audio signal is divided into multiple audio blocks, and authentication information for at least one audio block is signed based on the private key 142 of the audio signal signing module 500. The signed authentication information 141 is output by the audio signal signing module 500 and can be stored on a server or memory 600. Authentication access information 162, which is information about where the authentication information can be retrieved from the server or memory 600, can be provided to the audio signal signing module 500. The authentication access information 162 is embedded in the audio blocks received or captured by the audio signal signing module 500. Therefore, the audio signal signing module 500 performs the following operations: a) outputting an output audio signal 501 including audio blocks with embedded authentication access information to the network 300 for storage or forwarding, and b) outputting the signed authentication information 141 to a server or memory 600 where the information can be stored.
[0097] Alternatively, authentication access information can be embedded into audio blocks of the audio signal using a watermark. The watermark may relate to authentication access information associated with the location in memory 600 where the authentication information is stored. This authentication information can be obtained from the referenced above. Figure 1 The described audio feature extractor 150 generates audio blocks of an audio signal with embedded authentication access information. These blocks can be transmitted or stored via network 300. Authenticator 200 can receive the audio signal 501 with an embedded watermark and extract the authentication access information. Authentication information is retrieved from server 600 based on the authentication access information. The signature of the authentication information stored on server or storage 600 can be verified based on a public key 143 associated with the audio signal signature module 500, which can be stored, for example, on server 400 or a distributed ledger.
[0098] Figure 3 A block diagram of an audio signal authenticator is shown. Figure 3 The signal authenticator can be with Figure 1The audio signal signature module operates together. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. This audio signal 171 can be from... Figure 1 The audio signal signature module 500 outputs an audio signal 501. This audio signal 171 is input to a block boundary detector and a block segmenter 205, which determine block boundaries and output multiple audio blocks 201 with embedded authentication access information 212. If the authentication access information 212 has already been embedded in the watermark, it is assumed that the block boundaries have been previously encoded into the watermark, and the block boundaries can be detected from the watermark at this stage. In the case where the audio signal is transmitted by an audio container that includes the authentication access information 212 as metadata, it is assumed that the authentication access information is available for each block, where the block boundaries are part of the metadata.
[0099] An audio block 201 with embedded authentication access information is input to an access information extractor 210, which extracts authentication access information 212 from the audio block 201 (which, if tampered with, corresponds to authentication access information 162 from the signature module). An authentication access information parser 230 determines access information based on the authentication access information 212, i.e., where and how to find the signed authentication information 141 belonging to the current audio block 201 on the server or memory 600. The authentication information 141 retrieved from the server or memory 600 first includes audio features 151 representing the audio block 201, secondly includes a signature 221, and further includes optional other information 113, such as the location and / or time of the signature occurrence. If tampered with, the signature 221 corresponds to the signature 144 generated by the signature unit 140.
[0100] In the perceptual similarity analyzer 240, audio block 201 and audio feature 151 are perceptually compared with each other, thereby providing a similarity statement 241.
[0101] In the similarity analyzer 240, audio block 201 can be compared with audio features 151 of the original audio signal. The audio features can be a true copy of the original audio signal. However, this would require a large amount of storage space to store the original audio signal. If the audio features are related to other features of the original audio signal, the similarity analyzer should compare the received audio features with the audio features of the audio signal to be authenticated. Therefore, the desired audio features can be extracted from the audio signal to be authenticated. For example, this can be performed in, for example, the similarity analyzer 240. In other words, the similarity analyzer 240 should compare audio features (such as audio features received via authentication access information, i.e., the audio features of the original audio signal) with the corresponding audio features of the audio signal to be authenticated.
[0102] In its simplest case, this could be a binary decision about whether the compared quantities are perceptibly equal enough. Alternatively, similarity claim 241 could be more detailed, such as providing a similarity score. Authentication access information parser 230 also determines, based on authentication access information 212, which owner of private key 142 initially signed audio signal 171, for example, by identifying the owner through an ID. The owner of private key 142 could be a natural person, an organization, a registered microphone, and a few other possibilities are listed here. Using this ID, the identity 270 of public key 143 and the private key owner can be retrieved from an external database. The signature 221 and authentication information 152 (audio feature 151 and other information 113) are then verified within signature verifier 250 using public key 143, providing signature verification result 251. Perceptual similarity claim 241 and verification result 251 are finally combined in combiner 260 to form final authentication result 261.
[0103] Note that the processing proposed here for audio content can be similarly applied to any time-related data. In particular, the above processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark or uses a video container instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison.
[0104] Figure 4 A schematic block diagram of an audio signal signing module is shown. The audio signal signing module 500 receives a digital audio signal 121 at its audio input 502. This digital audio signal 121 is input to a block segmenter 130, which outputs a sequence of multiple audio blocks 131. The audio blocks 131 are forwarded to an audio feature extractor 150, which generates an audio feature 151 for each audio block 131. In a private key signing unit 140, the user's or entity's private key 142 is used to sign the audio features 151, including each audio block 131, and optional other information 113 such as location or time authentication information, thereby providing signed authentication information 141. The signed authentication information 141 itself includes the audio features 151, the generated signature 144, and optional other information 113. Figure 5 The authentication method shown requires two components during verification: individual audio blocks 131 and signed authentication information 141. In principle, they can be embedded in the same suitable audio container, but they can also be distributed across different physical channels. Optionally, a user or entity identifier 145 can be added to these two components, which can, for example, enable the retrieval of a suitable public key for automatic verification.
[0105] Figure 5 A schematic block diagram of authenticator 200 is shown. Authenticator 200 receives, for example, an audio signal in the form of an audio block 201 to be verified, and signed authentication information 141. The signed authentication information 141 includes audio features 151, other information 113, and a signature 221. Within a perceptual similarity analyzer 240, the audio features 151 contained in the authentication information 141 are perceptually compared with the audio block 201, thereby providing a similarity claim 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity claim 241 can be more detailed, thereby providing, for example, a similarity score. Furthermore, using knowledge about the user or entity that created the signature 221, a corresponding public key 143 and an optional user identity 270 are obtained. Optionally, this knowledge can be derived from a user or entity identifier 145 provided along with the authentication information 141 and the audio block 201. Subsequently, public key 143 is used to verify the combination of signature 221 and authentication information 141 within signature verification block 250, thereby providing signature verification result 251. The perceived similarity claim 241 and verification result 251 are finally combined in combiner 260 to form final authentication result 261.
[0106] In the following Figure 6 The document describes how to generate an authenticable audio signal by embedding authentication information into an audio block, and how to authenticate the audio signal based on the embedded authentication information. As described below, this authentication can be used for authenticating audio signals.
[0107] Figure 6 A schematic diagram of the audio signal signature module is shown. Figure 6 The audio signature module described is merely an example. The audio signal signature module 500 receives an input audio signal 121 from the microphone 100 or another audio source, such as an audio recorder, via an audio input terminal 502, and outputs an output audio signal 501 via an audio output terminal 190. The output audio signal 501 may be based on the input audio signal 121. The microphone 100 may include at least one microphone capsule 110 and an analog-to-digital converter (ADC) 120. The at least one microphone capsule 110 can capture an audio signal and output the captured analog audio signal 111. The captured analog audio signal 111 can be forwarded to the ADC 120, which can digitize the audio signal 111 and output a digital audio signal 121. The audio signal signature module 500 may be implemented as a device separate from the microphone 100 or may be included within the microphone 100.
[0108] The audio signal signature module 500 can receive audio signal 121 from microphone 100 or from other sources, which in Figure 6The input is indicated by selector 101. A digital audio signal 121 is input to a block divider 130, which divides the audio signal 121 into audio blocks 131, each with a block length, and outputs a sequence of multiple audio blocks 131. The block divider 130 can divide the audio signal 121 into multiple audio blocks. Optionally, each audio block receives a block number, which can be a subsequent block number. Optionally, if the audio signal 121 has already been divided into audio blocks 131, the block divider 130 can be omitted. The sequence of audio blocks 131 is forwarded to an audio feature extractor 150, which generates an audio feature 151 for each audio block 131 in the sequence. A private key signing unit 140 receives authentication information (e.g., metadata) including the audio features 151 of each audio block 131 and optional other information 113 such as location or time, and signs the authentication information with a private key 142, thereby providing signed authentication information 141. The signed authentication information 141 may include authentication information 152, which includes audio features 151 and optional other information 113, as well as a signature 144 generated by the private key signing unit 140.
[0109] The signed authentication information 141 itself includes authentication information, namely (a) audio feature 151, optionally (b) other information 113, and (c) signature 144. Each audio block 131 is fed in parallel, for example, to an information embedder 160 for embedding the signed authentication information 141 into, for example, each audio block 131. Embedding can be accomplished, for example, by means of a watermark.
[0110] Therefore, the information embedder 160 can be implemented as a watermark generator that generates watermark 163. Alternatively, the signed authentication information 141 can be embedded into a suitable audio container next to the actual audio block 131. Accordingly, the information embedder 160 outputs the audio block 161 with the embedded signed authentication information 141. The output of the embedder 160 corresponds to the output audio signal 501 of module 500 at the audio output terminal 190. The output audio signal 501 can be transmitted, for example, via network 300 (…). Figure 7 The output audio signal 501 can be broadcast or stored. In particular, the output audio signal 501 can even be further modified to a very small extent, including modification types such as sample rate conversion, perceptual compression, and trimming.
[0111] When the information embedder 160 is a watermark generator, its output 161, instead of the original audio block 131, can alternatively be input to the audio feature extractor 150 via the selector 102. This alternative offers the advantage of extracting audio features based on the final signal 161 output by the audio signal signature module 500, rather than the unoutput intermediate signal 131. Since the watermark generator is designed to perceptually alter the audio signal as little as possible, both input options (131 or 161) are reasonable for the audio feature extractor 150.
[0112] The audio signal 501 at the output of the audio signal module 500 (containing authentication access information or an embedded watermark within the audio container) can be stored or transmitted. Because the watermark can be embedded in, for example, essentially each audio block 131, each audio block 131 can be authenticated individually. The watermark 103 can also be embedded only in some audio blocks within the audio block set. In addition to the authentication access information, the watermark 163 can also include other information such as block boundaries. This information can be used when authenticating the audio signal to determine the block length.
[0113] The audio watermark 163 can be a unique identifier embedded in the audio signal and, for example, previously used to identify copyright information of the audio signal. Preferably, the watermark is embedded in the audio signal, making it very difficult to remove or destroy the watermark. Preferably, the embedded watermark will not be altered when the audio signal with the embedded watermark is copied, stored, or transmitted. The same reasoning applies to authentication access information included in the audio container.
[0114] Figure 7 It shows the relationship with Figure 6 The block diagram of the audio signal authenticator corresponds to the audio signal authenticator. The audio signal authenticator 200 receives an audio signal 171 with embedded authentication access information. This audio signal 171 may be from... Figure 6 The audio signal 171 is output by the audio signal signature module 500. This audio signal 171 is input to the block boundary detector and block segmenter 205, which determines the block boundaries and outputs multiple audio blocks 201 with embedded authentication information.
[0115] Block boundary detectors and block segmenters 205 detect block boundaries, for example, using information in the block header.
[0116] Audio block 201 is input to authentication information extractor 206, which extracts signed authentication information 141 embedded in the audio block. Signed authentication information 141 may include audio features 151, a signature 144, and optional other information 113, such as the location and / or time of the signature occurrence. If no tampering has occurred, the signature 221 corresponds to the signature 144 generated by signature unit 140.
[0117] In the perceptual similarity analyzer 240, audio blocks 201 and audio features 151 are perceptually compared to each other, thereby providing a similarity statement 241. In its simplest case, this can be a binary decision about whether the compared quantities are perceptually sufficiently equal. Alternatively, the similarity statement 241 can be more detailed, thus providing, for example, a similarity score.
[0118] The authentication information extractor 206 also determines, based on the authentication information 141, which owner of the private key 142 initially signed the audio signal 171, for example, by identifying the owner through an ID. The owner of the private key 142 could be a natural person, an organization, a registered microphone, and a few other possibilities are listed here. Using this ID, the identity 270 of the public key 143 and the private key owner can be retrieved from an external database. Subsequently, the signature 144 and the authentication information 152 (audio feature 151 and other information 113) are verified within the signature verifier 250 using the public key 143, thus providing a signature verification result 251. The perceived similarity claim 241 and the verification result 251 are finally combined in the combiner 260 to form the final authentication result 261.
[0119] Note that the processing proposed here for audio content can be similarly applied to any time-related data. In particular, the above processing can also be applied to video content if the following modifications are applied: microphone 100 is replaced by a video camera, audio feature extractor 150 is replaced by a video feature extractor, authentication access information embedding unit 160 uses a visual watermark instead of an audio watermark, or uses a video container instead of an audio container, and auditory perception comparison 240 is replaced by visual perception comparison.
[0120] Figure 3 , Figure 5 and Figure 7A comparison between the audio block 201 and audio feature 151 can be achieved, for example, by extracting features of the same type from the audio block 201 and performing a comparison against these features. Alternatively, different types of features can be selected for comparison. In this case, this type of feature must be extracted from the audio block 201 before performing the actual comparison, and the audio features in the authentication information must be converted into this type of feature. The idea is that the type of features available within the authentication information does not necessarily correspond to the type of features used for perceptual comparison. In particular, it may be advantageous to select the type of features used for comparison based on the importance of specific perceptual aspects captured by different types of features.
[0121] Figure 8 An exemplary illustration of the authentication result is shown. Specifically, the upper half depicts an exemplary original speech signal, subdivided into multiple frames for which authentication is to be performed individually. An exemplary modified signal is shown in the lower half, where specific modified areas are marked by arrows, indicating that consonants have been replaced compared to the original speech signal. Multiple individual bars are displayed below the waveform, using color to indicate the authentication result, with green representing a positive authentication result and red representing a negative authentication result.
[0122] In the perceptual similarity analyzer, received audio blocks are compared with extracted or retrieved audio features. Audio features are human auditory perceptual audio characteristics related to human auditory perception. Therefore, the received audio signal is analyzed to determine its human auditory perceptual audio characteristics. These audio features are then compared with the received or extracted audio features.
[0123] As an example, the audio characteristics perceived by human hearing can be the Mel frequency power spectrum of an audio block.
[0124] Audio features are related to the audio representation of audio signals that attempt to capture relevant aspects of human auditory perception. Another example of audio features could be the output of any model that mimics human auditory perception.
[0125] Performing a comparison of the received audio block with the extracted or retrieved authentication information in the transform domain (i.e., the domain associated with the audio features perceived by human hearing) will ensure that the identified differences are perceptually meaningful. In principle, any similarity measure (e.g., empirical cross-correlation coefficient or relative error relative to a reference quantity) can be used to map the comparison results to numerical values.
[0126] On one hand, the comparison between the received audio signal to be authenticated and the audio features from the original audio signal can be based on soft decision-making. Such probabilities can be directly derived from the similarity values by mapping an interval of possible similarity values to an interval of possible probabilities (i.e., the interval [0, 1]). This mapping should generally be monotonically increasing, as a higher similarity value should imply a higher authentication probability.
[0127] To generate such a mapping:
[0128] 1. Collect a set of L test signals that are considered to represent the target use case.
[0129] 2. Define a representative set of signal manipulation methods consisting of G benign methods and M malicious methods.
[0130] 3. Apply each of the G+M manipulation methods to each of the L test signals, and calculate the similarity between the original and modified test signals based on a similarity metric. Therefore, a total of (G+M)×L similarity values are provided (assuming the similarity is calculated based on the total length of each signal rather than on signal sub-intervals).
[0131] 4. Discretize the entire interval of possible similarity values into N separate subintervals In, n=1, ..., N.
[0132] 5. For each sub-interval In, count the number of benevolent modifications Gn and the number of malicious modifications Mn that generate similarity values within that sub-interval.
[0133] 6. Using these numbers, relative frequency And for each similarity value interval I k The expected authentication probability P for k=1, ..., N k pass To calculate.
[0134] Since Hn≥0 and PN=1, therefore P k As expected, it increases monotonically.
[0135] The proposed soft authentication decision can also be applied to individual audio blocks of the audio signal, rather than considering the entire audio signal. Similarly, the example of how to calculate the mapping from similarity values to authentication probabilities in step 3 can be modified as follows: similarity values are calculated for all consecutive audio blocks of the audio signal instead of just for the entire signal. Therefore, the resolution of the resulting mapping can be increased by correspondingly increasing the number N of sub-intervals used to discretize the possible similarity value intervals.
[0136] Audio block-by-block authentication is advantageous because it allows for the precise identification of any tampered audio blocks. Furthermore, it enables a simple method to visualize where tampering has occurred.
[0137] The modified audio signal (i.e., an audio signal that does not correspond to the original audio signal) can be displayed by marking the modified area or the modified audio block (e.g., by using arrows). Additionally, the waveform of the audio block can be displayed. To further improve the display of the tampered audio block, individual bars indicating the authentication result can be provided using color. Green can represent a positive authentication result, and red can represent a negative authentication result.
[0138] As an alternative to displaying authentication decisions using two colors (such as green and red), color scales can be used to display the probability of successful authentication, with different probabilities displayed by different colors.
[0139] Another visualization option to consider is that, in the event of malicious signal tampering, several consecutive blocks are more likely to be affected than a single block. Therefore, it is reasonable to intentionally flag the occurrence of several blocks with negative authentication results. This could be done by using a darker shade of red whenever the length of a suspected modified segment exceeds a certain length (e.g., 5 seconds). Another option is to increase the darkness of the red hue as the length of the suspected modified segment increases.
[0140] Additionally, if the metadata contains the number of the authenticated block, the presence of any non-contiguous block number can be specifically indicated by the authentication result, such as a flashing or broken chain, thus suggesting a cut in the original signal.
[0141] Torsten Dau, Dirk Pueschel, and Armin Kohlrausch, in “A quantitative model of the “effective” signal processing in the auditory system. I. model structure”—The Journal of the Acoustical Society of America, 99(6): 3615-3622, 1996—described an example of human auditory perception of audio characteristics, hereinafter referred to as the Perceptual Model PEMO. This paper describes a quantitative model designed to describe how the auditory system processes sound signals. The focus is on developing a model that mimics the functional processing of auditory stimuli by humans, particularly in contexts involving complex sounds such as speech or music.
[0142] The key components of this model can be categorized into peripheral processing, envelope extraction, modulation filtering, and a decision-making stage. The peripheral processing section captures the initial stages of auditory processing, including converting acoustic signals into neural representations through mechanisms such as outer and middle ear filtering and nonlinearities in the cochlea. The envelope extraction section involves extracting the temporal envelope of the sound, which is crucial for understanding amplitude modulation and speech processing. The modulation filtering section involves a system for filtering amplitude modulation, mimicking the auditory system's sensitivity to different modulation frequencies. In the decision-making stage, the processed auditory signals are used to make decisions about the properties of the sound, representing a higher level of auditory perception.
[0143] The model was designed to align with experimental psychoacoustic data, particularly its ability to simulate the auditory system's response to amplitude-modulated sounds. This approach provides a framework for understanding effective signal processing in human auditory perception, especially tasks involving the detection and discrimination of complex auditory patterns.
[0144] This model is applicable to masking conditions with both stochastic and deterministic masking sounds, independent of the temporal relationship between the masking sound and the signal. As part of the peripheral processing section, the model includes stages that simulate various aspects of temporal adaptation. This part is responsible for the model's ability to predict threshold decay under forward masking conditions. The model's decision stage is implemented as an optimal detection process, enabling the evaluation of the effectiveness of the preprocessing stage. A key feature of this model is that it processes real-time signals, and the threshold is derived using an adaptive tracking process within the simulated three-interval forced-choice process. Therefore, the same stimuli used in psychoacoustic experiments can be used in the simulation, and the results from the optimal processor model can be directly compared to thresholds obtained by human observers.
[0145] Since the model's detection performance depends on signal preprocessing, the different processing stages in the auditory system and their implementation in the model will be described in more detail. Sine waves of different frequencies produce peak responses at different locations along the basilar membrane. This frequency-location conversion implies that each location has bandpass filter characteristics, which is simulated using a linear basilar membrane model. This model can be implemented as a wave-digital filter, calculating, for example, the filtered signal in 120 output segments. Using off-frequency information is of no benefit to the subject as long as broadband noise is used to mask the sound. The signal at the output of a specific basilar membrane segment is half-wave rectified and then low-pass filtered at 1 kHz. This stage roughly simulates the conversion of the mechanical oscillations of the basilar membrane into receptor potentials in the inner hair cells. For high carrier frequencies, the low-pass filtering essentially preserves the signal envelope. The effect of adaptation is simulated through a feedback loop. This stage compresses stable signals almost logarithmically, while rapidly fluctuating inputs are converted in a more linear manner. These feedback loops were studied to explain the time masking effect and will be described in more detail below. In the stage immediately following the feedback loop, the signal undergoes a low-pass filter at 8 Hz, corresponding to a time constant of 20 ms. The combination of the nonlinear adaptive model and optimal detector processing has a significant impact on the results. This model incorporates adaptive characteristics of the auditory periphery. The adaptive stage in this model consists of five feedback loops connected in series with different time constants. Within each individual element, the low-pass filtered output is fed back, forming the denominator of the division element. The divisor is the instantaneous charging state of the low-pass filter, determining the attenuation applied to the input. The time constant can vary between 5 ms and 500 ms. The model has the following property: after the arrival of a stationary signal, the low-pass filter is charged according to its time constant. Because the charging state of the low-pass filter participates as a divisor, the more it is charged, the greater the attenuation of the input signal. For a stationary signal, the input value I produces a value OI at the output of the first feedback loop, derived from the stationarity condition I / O. For a chain of n loops, we obtain an output of O(2nI). For n=5, this approximates a logarithmic transformation. The range of output values obtained for a stationary input with levels between 0 dB and 100 dB is linearly mapped to the range of 0 to 100 model units (MU). For a perfect logarithmic transformation, a change of 1 MU will correspond to a 1-dB change at the input, independent of the input level. Using five feedback loops, we obtain the input-output characteristics, i.e., the transition from the input signal level in dB SPL to the model unit (MU). There is a small deviation compared to the straight line that would represent the logarithmic transformation at this scale. When the input signal is low, a specific increment in the input level results in a slightly smaller increment at the output compared to when the input signal is high.Since we assume constant sensitivity with respect to the output value MU, this small deviation from the logarithmic transformation will predict that level changes that are just perceptible at high input levels will be smaller than those at low input levels. Rapid input changes are linearly transformed compared to the time constant of the low-pass filter. If these changes are slow enough for the capacitor's state of charge to change accordingly, the attenuation gain will also change. Therefore, each element combines static compression nonlinearity with high sensitivity to rapid time changes. When the signal is turned off, the capacitor's charge, and thus the attenuation applied to the input, decreases in each stage according to its time constant. Therefore, the time constant determines how quickly the system recovers to its resting state. The charge of each low-pass element decays exponentially, meaning linear decay on a dB scale. By combining several linear attenuations with different slopes, a piecewise approximation of the forward masking curve can be achieved. In this way, a given signal can reduce the response to a subsequently presented signal. This makes it possible to simulate the nonlinear dependence of the forward masking threshold function on the masking sound level and masking sound duration. When the masking sound is short, those stages with long time constants cannot be fully charged, so only the attenuation caused by the fast stages affects the forward masking threshold. To establish the absolute threshold of the auditory signal, the minimum value at the input of the adaptive stages is limited by a constant value. This can be interpreted as simulating "physiological noise" added before the logarithmic transformation. If the current output of the basilar membrane filter signal is less than the threshold, it is replaced with the threshold. In the case of simultaneous masking, this scaling is not important. When simulating primarily at the absolute threshold, random noise should be used in the model instead of a constant threshold. The model is calibrated based on the 1-dB criterion in the intensity discrimination task. This value is just noticeable in the first step of adjusting the model parameters.
[0146] Alternatively, the audio signal may be part of a video file.
[0147] Alternatively, the signed authentication information can be included in the audio file (e.g., in an ADM file format) instead of in the watermark. In this case, the audio signal does not need to be modified.
[0148] Alternatively, the processing to generate the signed audio signal can be performed using a software solution based on a pre-existing audio recording or audio stream. The private key can be entered into the software, for example, via a dongle or text input, and authorization can be performed, for example, via biometric signals such as fingerprints or facial recognition.
[0149] The authentication process can calculate a similarity score between the audio features of the signed audio signal and the audio features of the analyzed signal. This allows it to determine the likelihood that the signal is still authentic, even if slight signal processing such as gain has been applied.
[0150] Authentication requires an audio signal and signed authentication information. The signed authentication information can be provided via a side channel. When using a watermark, the authentication access information must be read from the audio signal, and the authentication information must be retrieved accordingly. Where watermark reading requires block boundaries, one method to achieve this is to try different offsets of the block boundaries until the watermark can be successfully read. Another method could be to embed some synchronization signal with the watermark into the audio signal.
[0151] To enhance system security, processing is provided that enables, for example, key revocation in the event of key theft.
[0152] Public keys used for authentication can be provided in various ways. One option is to store all public keys in a centralized database to make it easy to find the required key. To address potential trust issues in this centralized instance, a database with all public keys can be provided based on distributed ledger technology. Another option is for the organization / individual using the technology to provide the public key on their own website.
[0153] In the example, a watermark could include a time-ordered sequence of pilot sequences (indicating the start of the watermark, the user ID, and the block ID). If a watermark is detected and can only be partially decoded (e.g., only the pilot sequence, without a meaningful user ID or block ID), it is still reasonable to notify the user accordingly. Advanced users (e.g., news organizations) can infer from the information that the content was signed at some point in time but appears to have been altered afterward. The partial presence of the watermark can also be used to refute claims that the file was never altered.
[0154] In another example, if the authentication information (source information) is not bound to the content via a watermark but, for example, through a time-related metadata format, the digital audio workstation can use accurate timestamps to reference all relevant source information for all source files. In this case, since all the information is already available, the authentication process described above is not necessary.
[0155] According to the example, each audio signal captured by microphone 100 and output by the microphone can include a watermark that identifies the actual microphone that captured the audio signal. If the microphone is used by several users, several user IDs can be associated with the microphone. If the microphone is registered, the corresponding microphone can be identified. User IDs can also be registered. A user ID can be a dedicated ID or a sub-ID of the microphone ID. The microphone ID and / or user ID can be part of the watermark. The sequence number of the audio block can determine which part of the audio signal has been removed.
[0156] Alternatively, the audio signal may be part of a video file.
[0157] Alternatively, the signed authentication information can be included in the audio file (e.g., in an ADM file format) instead of in the watermark. In this case, the audio signal does not need to be modified.
[0158] Alternatively, the processing to generate the signed audio signal can be performed using a software solution based on a pre-existing audio recording or audio stream. The private key can be entered into the software, for example, via a dongle or text input, and authorization can be performed, for example, via biometric signals such as fingerprints or facial recognition.
[0159] The authentication process can calculate a similarity score between the audio features of the signed audio signal and the audio features of the analyzed signal. This allows it to determine the likelihood that the signal is still authentic, even if slight signal processing such as gain has been applied.
[0160] Authentication requires an audio signal and signed authentication information. The signed authentication information can be provided via a side channel. When using a watermark, the authentication access information must be read from the audio signal, and the authentication information must be retrieved accordingly. Where watermark reading requires block boundaries, one method to achieve this is to try different offsets of the block boundaries until the watermark can be successfully read. Another method could be to embed some watermarked synchronization signal into the audio signal.
[0161] To enhance system security, processing is provided that enables, for example, key revocation in the event of key theft.
[0162] Public keys used for authentication can be provided in various ways. One option is to store all public keys in a centralized database to make it easy to find the required key. To address potential trust issues in this centralized instance, a database with all public keys can be provided based on distributed ledger technology. Another option is for the organization / individual using the technology to provide the public key on their own website.
[0163] Examples are also discussed below:
[0164] Example 1: An audio signal authenticator (200) includes: A block boundary detector and block segmenter (205) are configured to receive an audio signal (171) and output a sequence of multiple audio blocks (201), each audio block having embedded authentication information or embedded authentication access information (212, 162), the authentication access information relating to where the authentication information can be retrieved. The authentication information (141) includes at least the audio features (151) belonging to the current audio block (201), and A perceptual similarity analyzer (240) is configured to compare audio blocks (201) and audio features (151) with each other to provide similarity results (241). The audio features are calculated using a model that mimics human auditory perception.
[0165] Example 2: The audio signal authenticator (200) according to Example 1, wherein, The model that mimics human auditory perception is the Perceptual Model (PEMO).
[0166] Example 3: An audio signal authenticator (200) according to Example 1 or 2, wherein, The perceptual similarity analyzer (240) is configured to perform soft authentication decisions, thereby providing the probability of successful authentication as a result.
[0167] Example 4: The audio signal authenticator (200) according to any one of Examples 1 to 3 further includes: Access information extractor (210) is configured to extract authentication access information (212, 162) from the current audio block (201). An authentication access information parser (230) is configured to extract access information about how and / or where the signed authentication information (141) belonging to the current audio block (201) can be retrieved from memory (600). The signed authentication information (141) retrieved from the memory (600) includes audio features (151) belonging to the current audio block (201) and a signature (144, 221).
[0168] Example 5: The audio signal authenticator (200) according to Example 4 further includes: A signature verifier (250) is configured to use a public key (143) to verify a combination of a signature (144, 221) and human auditory perception audio features (151) and provide a signature verification result (251). The public key (143) is associated with the private key (142) that can be used by the audio signal signature module (500) to generate an authentic audio signal (501).
[0169] Example 6: An audio signal authenticator (200) according to Example 4 or 5, wherein, The authentication access information parser (230) is configured to determine the identifier (145) of the user or entity that initially signed the audio signal. Among them, it is possible to retrieve the identity (270) and public key (143) of a user or entity from an external database (400) based on the determined identifier (145).
[0170] Example 7: The audio signal authenticator (200) according to Example 5 further includes:
[0171] Combiner (260) is configured to combine the similarity result (241) and the signature verification result (251) into the final authentication result (261).
[0172] Example 8: A method for authenticating an audio signal (171) by an authenticator (200), comprising the following steps: Receive audio signal (171). Output a sequence of multiple audio blocks (201), each audio block (201) having embedded authentication information or embedded authentication access information (162), the authentication access information being related to where the authentication information can be retrieved. The authentication information (141) includes at least the audio features (151) belonging to the current audio block (201), and The received audio block (201) is compared with the extracted or retrieved audio features (151) to provide similarity results (241). The audio features are calculated using a model that mimics human auditory perception.
[0173] Example 9: A computer program product for authenticating audio signals. The computer program product includes program code units that cause the audio signal authenticator according to claim 1 to execute the method according to claim 8.
[0174] Example 10: An audio signal authenticator (200) according to any one of Examples 1 to 7, wherein the perceptual similarity analyzer (240) is further configured to: calculate audio features of a selected type based on an audio block (201), and wherein the comparison is based on the audio features of the selected type.
[0175] Example 11: According to the audio signal authenticator (200) of Example 11, the perceptual similarity analyzer (240) is further configured to convert the audio features in the authentication information (151) into audio features of a selected type.
[0176] Example 12: An audio signal signature module (500), including The input terminal is configured to receive digital audio signals (121). A block divider (130) is configured to divide a digital audio signal (121) into a sequence of audio blocks (131), wherein the sequence includes at least one audio block. An audio feature extractor (150) is configured to extract audio features (151) from the current audio block (131). The signing unit (140) is configured to: generate a signature (144) belonging to the current audio block (131) by applying a private key (142) to an audio feature (151), and provide signed authentication information (141) including the audio feature (151) and the signature (144). The audio features are calculated using a model that mimics human auditory perception, and the audio features (151) are represented in a format with a data rate lower than that of the audio block (131). Furthermore, the audio signal signature module (500) outputs signed authentication information (141) and audio block (131). List of reference numerals 100 microphones 101 Selector 102 Selector 110 Microphone Cap 111 Microphone simulates audio signal 113 Metadata 120 AD converter 121 digital audio signal 130-block divider 131 audio blocks 140 Private Key Signature Unit 141 Signed authentication information 142 Private Key 143 Public Key 144 signatures 145 Identifier 150 Audio Feature Extractors 151 Audio Features 152 Authentication Information 160 Information Embedder / Watermark Generator 161 Audio blocks with embedded information 162 Authentication Access Information 163 Watermark 165 Size Information 171 audio signal 190 audio output terminal 200 Audio Signal Authenticator 205 Block Boundary Detector and Block Segmenter 206 Authentication Information Extractor 210 Access Information Extractor 212 Authentication Access Information 221 signatures 230 Authentication Access Message Resolver 240 Perceptual Similarity Analyzer 241 Similarity Results 250 signature verifier 251 Signature verification result 260 combiner 261 Authentication Result 270 Identity 300 Network 400 server 500 audio signal signature module 501 Certified Audio Signal 502 Audio Input Terminal 600 Server or Storage
Claims
1. An audio signal authenticator (200), comprising: A block boundary detector and a block segmenter (205) are configured to receive an audio signal (171) and output a sequence of multiple audio blocks (201), the audio blocks having embedded authentication information or embedded authentication access information (212, 162), the authentication access information relating to where the authentication information can be retrieved. The authentication information (141) includes audio features (151) belonging to the current audio block (201), and A perceptual similarity analyzer (240) is configured to: compare audio blocks (201) of a received audio signal (171) with audio features (151) of extracted or retrieved authentication information (141), or compare audio features extracted from audio blocks (201) of a received audio signal (171) with audio features (151) of extracted or retrieved authentication information (141), thereby providing a similarity result (241). The audio features mentioned above are human auditory perception audio features related to human auditory perception.
2. The audio signal authenticator (200) according to claim 1, wherein, The human auditory perception audio features are calculated by mimicking a model of human auditory perception, specifically, a perceptual model (PEMO).
3. The audio signal authenticator (200) according to claim 1 or 2, wherein, The perceptual similarity analyzer (240) is configured to perform soft authentication decisions, thereby providing the probability of successful authentication as a result.
4. The audio signal authenticator (200) according to any one of claims 1 to 3, further comprising: Access information extractor (210), which is configured to extract the authentication access information (212, 162) from the current audio block (201). An authentication access information parser (230) is configured to extract access information about how and / or where the signed authentication information (141) belonging to the current audio block (201) can be retrieved from the memory (600). The signed authentication information (141) retrieved from the memory (600) includes audio features (151) belonging to the current audio block (201) and a signature (144, 221).
5. The audio signal authenticator (200) according to claim 4 further includes: A signature verifier (250) is configured to use a public key (143) to verify the combination of the signature (144, 221) and the human auditory perception audio feature (151) and provide a signature verification result (251). The public key (143) is associated with the private key (142) that can be used by the audio signal signature module (500) to generate an authentic audio signal (501).
6. The audio signal authenticator (200) according to claim 4 or 5, wherein, The authentication access information parser (230) is configured to determine the identifier (145) of the user or entity that initially signed the audio signal. Among them, the identity (270) of the user or entity and the public key (143) can be retrieved from an external database (400) based on the determined identifier (145).
7. The audio signal authenticator (200) according to claim 5 or 6 further comprises: Combiner (260), which is configured to combine the similarity result (241) and the signature verification result (251) into a final authentication result (261).
8. The audio signal authenticator (200) according to any one of claims 1 to 7, wherein The perceptual similarity analyzer (240) is further configured to calculate audio features based on audio blocks (201) of the received audio signal (171), wherein the comparison is based on audio features of a selected type.
9. The audio signal authenticator (200) of claim 8, wherein The perceptual similarity analyzer (240) is also configured to convert audio features within the authentication information (151) into audio features of the selected type.
10. A method for authenticating an audio signal (171) by an authenticator (200), comprising the following steps: Receive audio signal (171). Output a sequence of multiple audio blocks (201), each audio block (201) having embedded authentication information or embedded authentication access information (162), the authentication access information (162) being related to where the authentication information can be retrieved. The authentication information (141) includes at least the audio features (151) belonging to the current audio block (201), and The audio blocks (201) of the received audio signal (171) are compared with the extracted or retrieved audio features (151) to provide similarity results (241), or The audio features extracted from the audio block (201) of the received audio signal (171) are compared with the retrieved audio features (151) to provide similarity results (241). The audio features mentioned above are human auditory perception audio features related to human auditory perception.
11. An audio signal signature module (500), comprising: The input terminal is configured to receive digital audio signals (121). a block divider (130) configured to divide the digital audio signal (121) into a sequence of audio blocks (131), wherein The sequence includes at least one audio block. An audio feature extractor (150) is configured to extract audio features (151) from the current audio block (131). The signing unit (140) is configured to: generate a signature (144) belonging to the current audio block (131) by applying a private key (142) to the audio feature (151), and provide signed authentication information (141) including the audio feature (151) and the signature (144). The audio features are calculated using a model that mimics human auditory perception, and the audio features (151) are represented in a format with a data rate lower than that of the audio block (131). Furthermore, the audio signal signature module (500) outputs the signed authentication information.
12. A computer program product for authenticating audio signals, wherein The computer program product includes program code units that cause the audio signal authenticator of claim 1 to execute the method of claim 10.