Spectrogram-visible audio watermarking

WO2026199055A1PCT designated stage Publication Date: 2026-10-01AL-WARD TAREK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2025/050424
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2026-10-01

Smart Images

  • Figure CA2025050424_01102026_PF_FP_ABST
    Figure CA2025050424_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a forensic watermarking technique for embedding textual information into audio files, creating visible text only when analyzed with a spectrogram while remaining inaudible during normal playback. Unlike traditional watermarking that alters content or degrades audio quality, this innovation enables content creators to distribute media with invisible attribution that can be forensically verified without compromising user experience. The invention generates text-based audio watermarks by: rendering text on a digital canvas; mapping the rendered text to specific high-frequency components in an audio signal; and generating an audio file containing these mapped frequency components. This creates a dual-layer media where primary content remains pristine while a hidden forensic identifier exists in the frequency domain, revealed only through spectrogram analysis. The system allows customization of parameters including character spacing, frequency range, and duration for different applications while maintaining integrity of content and audio quality.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Patent Application — Spectrogram-Visible Audio Watermarking

[0002] Text-to-Audio Watermark System and Method

[0003] Field Of The Invention

[0004] The present Invention relates to advanced digital forensic watermarking technology, specifically to systems and methods for embedding textual information into audio files in a manner that renders the text visible only when the audio is analyzed with a spectrogram while remaining substantially inaudible during normal playback. This invention represents a significant advancement in content attribution and ownership verification that does not significantly compromise content quality or user experience.

[0005] Background Of The Invention

[0006] Digital watermarking is a technique used to embed information into digital media, such as images, audio, or video, in a way that is difficult to remove. Digital watermarking approaches generally fall into distinct categories: (1) visible watermarking, which overlays logos or text directly onto visual content, creating obstructions that degrade the viewing experience; and (2) audio watermarking, which attempts to embed information in sound files. Traditional audio watermarking techniques typically focus on embedding copyright information or other metadata in a manner that is robust against various transformations while attempting to remain imperceptible to listeners. However, these methods often struggle to balance robustness with imperceptibility - the more robust the watermark is against removal or transformation, the more likely it is to introduce audible artifacts that degrade the listening experience.

[0007] Spectrograms are visual representations of the spectrum of frequencies in a sound signal as they vary with time. They are commonly used in audio analysis to visualize the frequency content of audio signals. A spectrogram displays the amplitude of a particular frequency at a particular time as a color or brightness value.

[0008] Existing content protection approaches generally fall into two problematic categories: (1) visible watermarks that obstruct and degrade the visual content, or (2) audible watermarks that interfere with the listening experience. Both approaches compromise the integrity ofthe original content and diminish user experience. Furthermore, existing audio watermarking techniques generally focus on embedding copyright information or identification codes that are not meant to be visually interpreted and are often detectable or extractable by automated systems rather than providing human-readable forensic evidence.

[0009] Conventional audio watermarking methods are not designed to create visually recognizable text patterns when viewed through a spectrogram while remaining substantially inaudible. Existing methods that attempt to create visual patterns in spectrograms often lack the precision and clarity needed for legible text, particularly when the text needs to be customizable by end-users, or they create audible artifacts that compromise audio quality.

[0010] Furthermore, in scenarios involving confidential or pre-release content distribution, there is a critical need for personalized watermarking solutions that can uniquely identify each recipient of the content. Traditional watermarking approaches typically embed generic ownership information or non-human-readable codes that require specialized software to decode. This creates significant challenges when investigating content leaks, as the source of the leak cannot be immediately identified without proprietary detection tools.

[0011] There exists a need for an innovative forensic watermarking approach that can embed attribution and ownership information into content without visibly altering the content or degrading audio quality, while still providing verifiable proof of ownership when needed.

[0012] Summary Of The Invention

[0013] The present invention provides a novel forensic watermarking technique for embedding textual information into audio files in a manner that creates visible text only when the audio is analyzed with a spectrogram, while remaining substantially inaudible during normal playback. The term "substantially inaudible" acknowledges that while the watermark is designed to be imperceptible to most listeners during normal playback conditions, the degree of inaudibility may vary based on factors such as the listener's hearing sensitivity, playback equipment, and environmental conditions.

[0014] This innovation represents a paradigm shift in digital watermarking by creating a dual-layer media experience: the primary layer delivers pristine, unobstructed content to users, while a secondary forensic layer contains ownership or attribution information that can be revealed through spectrogram analysis when needed.In one aspect, the invention provides a method for generating a forensic text-based audio watermark, comprising: rendering the text on a digital canvas; mapping the rendered text to frequency components in an audio signal within a specific high-frequency range; and generating an audio file containing the mapped frequency components.

[0015] In another aspect, the invention provides a system for generating a forensic text-based audio watermark, comprising: a processor; and a memory storing instructions that, when executed by the processor, cause the system to: receive text input from a user; render the text on a digital canvas; map the rendered text to frequency components in an audio signal within a specific high-frequency range; and generate an audio file containing the mapped frequency components.

[0016] The invention provides several significant advantages over prior art techniques:

[0017] 1. The watermarked audio files maintain pristine audio quality while containing embedded text that is clearly visible when analyzed with a spectrogram.

[0018] 2. The watermark remains substantially inaudible during normal playback, preserving the intended listening experience.

[0019] 3. The approach enables content to be shared without visible watermarks that obstruct or degrade the visual experience.

[0020] 4. The forensic nature of the watermark provides verifiable proof of ownership or attribution that can be revealed when needed.

[0021] 5. The ability to embed personalized text enables the creation of unique watermarks for each recipient of confidential or pre-release content, facilitating the tracking of unauthorized distribution to specific individuals.

[0022] Furthermore, the invention includes a temporal pattern design that repeats the watermark throughout the audio file, ensuring visibility throughout the entire duration while maintaining the substantially inaudible nature of the watermark.Detailed Description Of The Invention

[0023] Method Overview

[0024] FIG. 2 illustrates the complete method for generating an audio watermark from text according to an embodiment of the present invention. The process begins with receiving text input and configuring parameters such as duration, character spacing, and frequency range. The text is then rendered on a digital canvas, converted to uppercase for better visibility, and flipped vertically to account for spectrogram orientation.

[0025] The rendered text is mapped to frequency components by dividing the canvas into time slices, analyzing pixel brightness in each slice, and mapping pixel positions to frequency values. Sine waves are generated for each frequency and combined for each time slice. Character spacing adjustments are applied if specified, and the signal is normalized to avoid clipping. Fade-in and fade-out effects are applied, and the pattern is repeated throughout the audio duration. Finally, the audio is converted to WAV format and saved as a file.

[0026] System Architecture

[0027] FIG. 1 illustrates the overall architecture of a text-to-audio watermark system according to an embodiment of the present invention. The system includes:

[0028] - A user interface component that receives input from users, including the text to be embedded, the desired duration of the audio file, character spacing parameters, and other configuration options.

[0029] - A text rendering component that converts the input text into a visual representation on a digital canvas.

[0030] - A frequency mapping component that maps the visual representation to frequency components in an audio signal.

[0031] - An audio generation component that creates the audio signal containing the mapped frequency components.

[0032] - A file output component that saves the audio signal as a digital audio file, typically in WAV format.

[0033] Text Rendering Process

[0034] FIG. 3 illustrates the text rendering process according to an embodiment of the present invention. The process begins by receiving text input from a user. The text is then converted to uppercase to ensure consistency and improve visibility in the spectrogram.The system creates a digital canvas with dimensions based on the text length and desired character height. The text is rendered on the canvas using a bold font, typically Arial or a similar sans-serif font, with a specified character height (e.g., 24 pixels). The system supports adjustable character spacing to optimize visibility in the spectrogram.

[0035] The rendered text is displayed against a black background with white text to maximize contrast. The canvas is then flipped vertically to account for the orientation in spectrogram displays, where higher frequencies appear at the top.

[0036] Frequency Mapping Process

[0037] FIG. 4 illustrates the frequency mapping process according to an embodiment of the present invention. The process maps the visual representation of the text to frequency components in an audio signal.

[0038] The system divides the canvas into time slices, with each vertical column in the image corresponding to a specific time segment in the audio. For each time slice, the system analyzes the brightness values of pixels in the column.

[0039] For each pixel with sufficient brightness (typically brightness > 0.1 on a scale of 0 to 1), the system maps the vertical position of the pixel to a frequency value within a specified range. In a preferred embodiment, this range is between 7,500 Hz and 17,000 Hz, which balances audibility and visibility in a spectrogram.

[0040] The mapping from vertical position to frequency can be linear or non-linear. In a preferred embodiment, a slighdy non-linear mapping is used to optimize visibility in the spectrogram.

[0041] For each mapped frequency, the system generates a sine wave with amplitude proportional to the pixel brightness. These sine waves are combined to create the audio signal for each time slice.

[0042] Character Spacing Control

[0043] FIG. 5 illustrates the character spacing control mechanism according to an embodiment of the present invention. The invention includes a mechanism for adjusting the spacing between characters in the rendered text. This spacing control allows users to optimize the visibility and legibility of the text in the spectrogram.When character spacing is increased, the system measures each character individually and adds the specified spacing between characters. This can improve legibility in the spectrogram, particularly for complex or lengthy text.

[0044] The character spacing parameter can be adjusted based on the specific requirements of the watermark, allowing for optimization between compact representation and maximum legibility.

[0045] Temporal Pattern Design

[0046] FIG. 6 illustrates the temporal pattern of watermark repetition according to an embodiment of the present invention. To ensure the watermark is visible throughout the audio file, the system repeats the watermark at regular intervals.

[0047] In a preferred embodiment, the base watermark is approximately 4 seconds long, followed by a 1-second gap of silence. This pattern is repeated throughout the duration of the audio file as specified by the user.

[0048] For longer audio files, the system uses an efficient streaming approach to generate the audio without excessive memory usage. This involves generating the audio in chunks and writing them directly to the output file.

[0049] Audio Generation and Output

[0050] The system combines the frequency components for each time slice to create the complete audio signal. The signal is normalized to avoid clipping, and fade-in and fade-out effects are applied at the beginning and end of the audio to prevent clicks or pops.

[0051] The audio signal is then converted to a standard audio format, typically 16-bit PCM WAV format with a sample rate of 44,100 Hz. The resulting file can be played on standard audio players but will reveal the embedded text when analyzed with a spectrogram.

[0052] Frequency Deletion Method for Existing Audio

[0053] FIG. 8 illustrates the frequency deletion method for watermarking existing audio files according to an embodiment of the present invention. In addition to generating new audio files with embedded watermarks, the invention also provides a method for watermarking existing audio files through selective frequency deletion. This approach modifies the spectral content of an existing audio file to create visible text patterns when viewed through aspectrogram analyzer, while maintaining the overall quality and intelligibility of the original content.

[0054] The frequency deletion method operates by selectively reducing the amplitude of specific frequencies in the audio signal that correspond to the text pattern. This creates a "negative" watermark where the text appears as areas of reduced spectral energy when viewed in a spectrogram. While some watermarking methods embed data by shifting blocks in the time domain, our approach preserves the temporal structure of the audio while creating visible text patterns in the frequency domain.

[0055] A key innovation in the frequency deletion method is the use of multiple frequency ranges for watermarking, which we call "tiled frequency ranges." This approach embeds the same text pattern across multiple non-overlapping frequency bands, creating redundant watermarks that enhance visibility and robustness. The system uses predefined frequency ranges, typically:

[0056] - 7,000 Hz to 15,000 Hz (high range)

[0057] - 3,000 Hz to 5,000 Hz (mid-high range)

[0058] - 1,500 Hz to 3,600 Hz (mid-low range)

[0059] - 600 Hz to 1,500 Hz (low range)

[0060] This multi-band approach ensures that the watermark remains visible even if certain frequency ranges are affected by audio processing, compression, or environmental factors.

[0061] The process includes:

[0062] 1. Reading an existing audio file and converting it to a time-domain representation. 2. Generating a text pattern similar to the one used in the text-to-audio watermark method.

[0063] 3. Processing the audio in overlapping chunks using a Fast Fourier Transform (FFT) to convert each chunk from the time domain to the frequency domain.

[0064] 4. For chunks within a watermark area, checking each frequency bin against all tiled frequency ranges.

[0065] 5. If a frequency falls within one of the defined ranges, mapping it to the corresponding position in the text pattern.

[0066] 6. If the corresponding position in the pattern has high brightness (indicating the presence of text), reducing the amplitude of that frequency by a specified factor, typically expressed in decibels.7. Applying an inverse FFT to convert the chunk back to the time domain.

[0067] 8. Combining the processed chunks using an overlap-add method with a window function to minimize artifacts at the boundaries between chunks.

[0068] The system supports two pattern types for frequency deletion watermarks:

[0069] 1. "Single" pattern: Places the watermark at a specific start time in the audio

[0070] 2. "Tiled" pattern: Repeats the watermark throughout the audio file with a specified spread duration between repetitions

[0071] The frequency deletion method provides several advantages:

[0072] It preserves the temporal integrity of the audio

[0073] It allows for the watermarking of existing content without the need to recreate the audio

[0074] It can be applied to complex audio material such as music or speech

[0075] It maintains the overall quality and intelligibility of the original content

[0076] It provides a more subtle watermarking approach that is less likely to be detected during normal listening

[0077] It creates human-readable text patterns in spectrograms

[0078] The deletion strength parameter allows users to control the visibility of the watermark in the spectrogram versus its audibility during playback.

[0079] Amplitude Scaling

[0080] The system includes controls for adjusting the amplitude (volume) of the watermark. This is specified in decibels (dB) and converted to a linear scale for audio processing.

[0081] By default, the watermark is set to a volume level that makes it substantially inaudible during normal playback while ensuring it is clearly visible in a spectrogram. Users can adjust this parameter based on their specific requirements for audibility versus visibility.

[0082] Frequency Range Optimization

[0083] The system is optimized to work within the 7,500 Hz to 17,000 Hz frequency range, which provides an ideal balance between audibility and visibility. This range is high enough to minimize interference with most audio content while remaining within the range that can be effectively captured and displayed by standard spectrogram analyzers.The frequency mapping can be adjusted to emphasize certain parts of the text or to accommodate different spectrogram analysis settings.

[0084] Technical Implementation Details

[0085] The invention encompasses two distinct watermarking approaches, each with its own mathematical foundation and implementation details: (1) the text-to-audio watermark generation method that creates new audio files with embedded text, and (2) the frequency deletion method that modifies existing audio files.

[0086] A. Text-to-Audio Watermark Generation

[0087] The core of the text-to-audio watermarking process relies on precise sine wave generation and combination. For each pixel in the rendered text image with brightness above the threshold, a sine wave is generated according to the following formula:

[0088] Sine Wave Generation

[0089] s(t) = A • sin(2πft)

[0090] • s(t) is the sine wave amplitude at time t

[0091] • A is the amplitude, proportional to pixel brightness (typically scaled to range [0,1]) • f is the frequency in Hz, mapped from the pixel's vertical position

[0092] • t is the time in seconds

[0093] Frequency Mapping

[0094] The frequency mapping from pixel position to frequency value follows this equation:

[0095] /

[0096]

[0097] = / mt.n + ( / max - / mt.n ) • (1 - ( \v n Jf

[0098] • f is the resulting frequency in Hz

[0099] • ^

[0100]

[0101] min is r^efrequency of the range (typically 7,500 Hz)

[0102] • ^

[0103]

[0104] max is r^e maximumfrequency of the range (typically 17,000 Hz)

[0105] • y is the pixel's vertical position from the top of the canvas

[0106] • h is the total height of the canvas in pixels

[0107] • y (gamma) is a non-linearity factor (typically 1.1) that slighdy emphasizes mid-range frequencies for improved visibilityAdditive Synthesis and Signal Normalization

[0108] For each time slice (corresponding to a vertical column in the image), multiple sine waves are combined using simple additive synthesis:

[0109] 5(0 = Zs.(t)

[0110]

[0111] i

[0112] • S(t) is the combined audio signal at time t

[0113] • s.(t) represents each individual sine wave

[0114] Signal Normalization

[0115] The resulting signal is then normalized to prevent clipping:

[0116] s

[0117]

[0118] norm (Kt)7= — maxs((n|S)(t)|) • A target

[0119] • s (t) is the normalized signal

[0120] • max(|S(t)|) is the maximum absolute amplitude of the original signal

[0121] • ^target is tarSetamplitude (typically 0.8 to leave headroom)

[0122] B. Frequency Deletion Method

[0123] Window Functions and Overlap-Add Method

[0124] To avoid artifacts at chunk boundaries during the frequency deletion process, we process the audio in overlapping chunks using a Hann window function:

[0125] Hann Window Function

[0126] w(n) = 0.5 • (1 - cos(2πn / N))

[0127] • w(ri) is the window coefficient at sample index n

[0128] • N is the window size (typically 2048 or 4096 samples)

[0129] The overlapping factor is typically 75%, meaning each chunk overlaps with 75% of the previous and next chunks. This ensures smooth transitions when chunks are recombined.

[0130] Overlap-Add Algorithm

[0131] 1. Initialize an output buffer with zeros

[0132] 2. For each processed chunk:

[0133] a. Multiply the chunk by the window function

[0134] b. Add the windowed chunk to the output buffer at the appropriate position, accounting for the overlap

[0135] 3. Normalize the output buffer to maintain consistent amplitudeFast Fourier Transform (FFT) Implementation

[0136] The frequency deletion method relies on the Fast Fourier Transform (FFT) to convert audio signals between time and frequency domains.

[0137] FFT Definition

[0138] X( / c) = S x(n) ■ e~' "

[0139]

[0140] n=0

[0141] • X(fc) is the complex frequency-domain representation

[0142] • x(n) is the time-domain signal

[0143] • N is the number of samples

[0144] • k is the frequency bin index

[0145] • n is the time sample index

[0146] Inverse FFT (IFFT)

[0147] xm = 7 Z VO •

[0148]

[0149] k=0

[0150] Frequency Bin Modification

[0151] For frequency deletion, we modify the magnitude of specific frequency bins while preserving phase information:

[0152] 7'V¥(fc)

[0153] X

[0154]

[0155] mo „diffi-e „d (k) = |X( / C) | • a(k) • e

[0156] • X modified ('kJ) is the modified frequency-domain representation

[0157] • \X(k) | is the magnitude of the original frequency bin

[0158] • xX( / c) is the phase of the original frequency bin

[0159] • a( / c) is the amplitude scaling factor for bin k, determined by the text pattern

[0160] The amplitude scaling factor a(k) is calculated based on the deletion strength parameter (DS) in decibels:

[0161] 10 20 if bin k corresponds to text

[0162] a{k) —

[0163] otherwise

[0164]

[0165] For example, with a deletion strength of -12 dB, the amplitude of affected frequency bins is reduced to approximately 25% of their original value.Example Implementation

[0166] In an example implementation, a user enters the text " PATENT EXAMPLE" with a duration of 10 seconds, character spacing of 2, and a frequency range of 7,500-17,000 Hz.

[0167] The system renders the text on a canvas with dimensions based on the text length and a character height of 24 pixels. The character spacing of 2 adds additional space between each character to improve legibility.

[0168] The rendered text is mapped to frequency components in the 7,500-17,000 Hz range. The resulting audio file contains the embedded text that is substantially inaudible during normal playback but clearly displays " PATENT EXAMPLE" when viewed with a spectrogram analyzer.

[0169] FIG. 7 shows a screenshot of a spectrogram analysis of the resulting audio file, clearly displaying the text " PATENT EXAMPLE" in the high-frequency range.

[0170] In another example implementation, the frequency deletion method is applied to an existing audio file containing music. The text " WATERMARK" is embedded using the "tiled" pattern type with a spread duration of 5 seconds and a deletion strength of -12 dB. The resulting audio file maintains the musical content of the original while clearly displaying the text " WATERMARK" when viewed with a spectrogram analyzer.

[0171] FIG. 9 shows a screenshot of a spectrogram analysis of an audio file watermarked using the frequency deletion method. The image clearly demonstrates how the text appears as areas of reduced spectral energy (darker regions) in the spectrogram, creating a "negative" watermark that is visible upon analysis but remains substantially inaudible during normal playback. This visual evidence highlights the effectiveness of the frequency deletion method in creating clear, legible text patterns while preserving the audio quality of the original content.Personalized Watermarking for Leak Tracking

[0172] The invention is particularly valuable for scenarios involving the distribution of confidential or pre-release content where tracking the source of potential leaks is critical. In such applications, the system can be used to create uniquely identifiable versions of the same content for different recipients.

[0173] For example, a movie studio distributing pre-release screeners to critics could embed each recipient's name or a unique identifier in the audio track. If " PREVIEW FOR JOHN SMITH -DO NOT DISTRIBUTE" is embedded as the watermark text, this identifier would remain substantially inaudible during normal playback, preserving the viewing experience.

[0174] However, if the content were leaked, a simple spectrogram analysis would immediately reveal the source of the leak without requiring specialized decoding software.

[0175] This personalized watermarking capability provides several benefits:

[0176] 1. It creates a strong deterrent against unauthorized sharing

[0177] 2. It simplifies leak investigations by making the source immediately identifiable 3. It provides clear, human-readable evidence that can be used in legal proceedings 4. It eliminates the need for proprietary detection algorithms to identify the source of leaked content

[0178] The system can be integrated into content distribution workflows to automatically generate personalized watermarks for each recipient, making it practical to implement even for large-scale distribution scenarios.

Claims

Patent Application — Spectrogram-Visible Audio WatermarkingThe Embodiments Of The Invention In Which An Exclusive Property Or Privilege Is Claimed Are Defined As Follows:The Embodiments Of The Invention In Which An Exclusive Property Or Privilege Is Claimed Are Defined As Follows:

1. A method for generating a spectrogram-visible audio watermark from text, comprising:receiving text input;rendering the text on a digital canvas with dimensions proportional to the text length, with a fixed height and width determined by text measurement plus padding;flipping the canvas vertically by applying a scale transformation to align text orientation with spectrogram visualization;- processing the canvas data column-by-column, where each column corresponds to a specific time segment in the resulting audio;identifying pixels in each column having brightness values exceeding 10% of maximum pixel intensity;mapping each identified pixel's vertical position to a frequency value within a predetermined frequency range of 7,500-17,000 Hz using a linear or logarithmic mapping function;generating sine waves at the mapped frequencies with amplitudes proportional to pixel brightness;repeating the generated audio pattern at regular intervals; and generating an audio file containing the resulting frequency-domain watermark with normalized amplitude to prevent clipping.

2. The method of claim 1, further comprising:- providing a user-adjustable character spacing parameter ranging from 0 to 10 units;applying the spacing parameter when rendering text by adding the specified spacing between individual characters; and- previewing the adjusted text to the user before processing.

3. The method of claim 1, wherein:the brightness threshold corresponds to a luminance value of 10% of maximum pixel intensity, calculated as the sum of red, green, and blue color channel values (each ranging from 0-255) divided by 765 (the maximum possible sum), resulting in a normalized brightness value between 0 and 1.

4. The method of claim 1, wherein:the predetermined frequency range comprises at least one of:- a high band (7,000-15,000 Hz),a mid-high band (3,000-5,000 Hz),a mid-low band (1,500-3,600 Hz), ora low band (600-1,500 Hz);and wherein multiple frequency bands may be used simultaneously with decreasing amplitude in lower frequency bands.

5. A system comprising a processor and memory storing instructions that, when executed by the processor, cause the system to perform the method of claim 1.

6. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1.

7. A method for embedding a text-based watermark into an existing audio file, comprising:converting the existing audio file into a frequency-domain representation using Fast Fourier Transform (FFT) with a window size of 4096 samples and 50% overlap between windows;rendering text on a digital canvas and identifying pixels exceeding a brightness threshold of 10%;mapping identified pixels to corresponding frequency components within predefined frequency bands;selectively reducing the amplitude of those frequency components by applying a reduction factor calculated as 10A(deletionStrength / 20), where deletionstrength is a user-configurable parameter measured in decibels; applying the reduction factor to both real and imaginary parts of the complex FFT coefficients and their symmetric counterparts to maintain signal properties;- processing the audio in overlapping chunks with a Hann window function to reduce artifacts; andconverting the modified frequency-domain representation back to a time-domain audio file using inverse FFT and overlap-add synthesis, resulting in a spectrogram where the text appears as a negative image formed by attenuated frequency components.

8. The method of claim 7, further comprising:selecting a watermark pattern type as either:a single occurrence at a specified start time position within the audio, ora tiled repetition throughout the audio with a configurable spread duration between repetitions.

9. The method of claim 7, further comprising:embedding the watermark across multiple non-overlapping frequency bands, specifically:- a high band (7,000-15,000 Hz),a mid-high band (3,000-5,000 Hz),a mid-low band (1,500-3,600 Hz), anda low band (600-1,500 Hz);- wherein each lower frequency band receives progressively reduced amplitude modification.

10. A system comprising a processor and memory storing instructions that, when executed by the processor, cause the system to perform the method of claim 7.

11. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 7.

12. A method for tracking unauthorized distribution of content, comprising:generating multiple versions of an audio file, each containing a unique spectrogram-visible text watermark with recipient-identifying information; distributing each version to its corresponding recipient;upon discovery of unauthorized distribution, analyzing the spectrogram of the distributed audio using a Fast Fourier Transform (FFT) with a window size between 1024-4096 samples to visualize the frequency content; and identifying the source of unauthorized distribution by visually or algorithmically extracting the text watermark from the spectrogram.