Method and apparatus for generating watermarked audio by using encoder trained on basis of multiple determination units

The use of a GAN-trained encoder with multiple discriminators for audio watermarking addresses vulnerabilities in conventional techniques, ensuring robust and high-capacity watermarking that protects copyrighted works and maintains sound quality.

WO2025164836A1PCT designated stage Publication Date: 2025-08-07XINAPSE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/002748
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2024-03-04
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Conventional audio watermarking techniques are vulnerable to attacks, require specialized knowledge, and have limitations in data embedding capacity, making them susceptible to misuse and ineffective in protecting copyrighted works.

Method used

A method and device for generating watermarked audio using an encoder trained with multiple discriminators in a Generative Adversarial Network (GAN), where a first discriminator determines watermark insertion based on a spectrogram and a second discriminator determines it based on an audio waveform, ensuring robustness and resilience against audio modulation.

Benefits of technology

The solution provides an inaudible audio watermarking technique that is resistant to various attacks and maintains sound quality, effectively protecting copyrighted works and preventing misuse of synthetic voices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024002748_07082025_PF_FP_ABST
    Figure KR2024002748_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and an apparatus for generating watermarked audio by using an encoder trained on the basis of multiple determination units. An embodiment of the present disclosure may provide a method for generating watermarked audio by using an encoder trained on the basis of multiple determination units, the method including the steps of: inputting, to a watermark encoder, watermark data and target audio data into which the watermark data is to be inserted; and acquiring, as an output of the watermark encoder, marked audio data obtained by inserting the watermark data into the target audio data, wherein the watermark encoder is trained as a generator of a generative adversarial network (GAN), and a discriminator of the GAN comprises: a first determination unit for determining the insertion of a watermark on the basis of a spectrogram; and a second determination unit for determining the insertion of a watermark on the basis of an audio waveform.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for generating watermark audio using an encoder learned based on a multi-discriminant unit

[0001] The present disclosure relates to a method and apparatus for generating watermarked audio using an encoder trained based on multiple discriminators. More specifically, the present disclosure relates to a method and apparatus for generating watermarked audio using an encoder trained as a generator for the discriminator based on Generative Adversarial Networks (GANs), in a discriminator comprising a first discriminator for determining the insertion of a watermark based on a spectrogram and a second discriminator for determining the insertion of a watermark based on an audio waveform.

[0002] Recent advancements in audio synthesis technology have made it possible to generate voices that mimic real people even with small amounts of audio, increasing the accessibility and usability of high-quality synthesis technology. However, the potential for abuse of audio synthesis technology, such as voice phishing using synthetic voices, is also increasing. In this regard, audio watermarking techniques can be used for identifying synthetic audio, identifying copyright information, and forensics. Audio watermarking techniques can help prevent crime and protect and utilize copyrighted works through synthetic voice detection.

[0003] Conventional audio watermarking techniques include Least Significant Bit (LSB), echo hiding, spread spectrum, patchwork, Quantization Index Modulation (QIM), and steganography. However, these techniques rely heavily on specialized knowledge and experience, making implementation difficult. Furthermore, watermarked audio is vulnerable to various attacks and has limitations in the size of data that can be embedded as a watermark. To address these issues, research into audio watermarking using deep learning is beginning, but it remains vulnerable to tampering.

[0004] Accordingly, there is a need for an audio watermarking technique that can embed sufficient watermark data, maintain quality close to the original, and is robust to audio modulation without relying on specialized knowledge and empirical knowledge.

[0005] The present disclosure provides a method and device for generating watermarked audio using an encoder trained based on multiple discriminators. The problems addressed by the present disclosure are not limited to those mentioned above, and other problems and advantages of the present disclosure not mentioned above can be understood through the following description and will be more clearly understood through embodiments of the present disclosure. Furthermore, it will be appreciated that the problems and advantages addressed by the present disclosure can be realized by the means and combinations thereof set forth in the claims.

[0006] As a technical means for achieving the above-described technical problem, a first aspect of the present disclosure may provide a method for generating watermark audio using an encoder trained based on multiple discriminators, including the steps of: inputting watermark data and target audio data, which is an object of embedding the watermark data, into a watermark encoder; and obtaining marked audio data in which the watermark data is embedded in the target audio data as an output of the watermark encoder, wherein the watermark encoder is trained as a generator of a Generative Adversarial Network (GAN), and a discriminator of the GAN includes a first discriminator for determining the embedding of a watermark based on a spectrogram and a second discriminator for determining the embedding of a watermark based on an audio waveform.

[0007] A second aspect of the present disclosure provides a device for generating watermark audio using an encoder trained based on multiple discriminators, including a memory having at least one program stored therein; and a processor configured to operate by executing the at least one program, wherein the processor inputs watermark data and target audio data to be inserted into the watermark data into a watermark encoder, and obtains marked audio data in which the watermark data is inserted into the target audio data as an output of the watermark encoder, wherein the watermark encoder is trained as a generator of a Generative Adversarial Network (GAN), and a discriminator of the GAN includes a first discriminator for determining the insertion of a watermark based on a spectrogram and a second discriminator for determining the insertion of a watermark based on an audio waveform.

[0008] A third aspect of the present disclosure can provide a non-transitory computer-readable recording medium having recorded thereon a program for executing the method of the first aspect of the present disclosure on a computer.

[0009] Other aspects, features and advantages other than those described above will become apparent from the following drawings, claims and detailed description of the invention.

[0010] According to the above-described problem solving means of the present disclosure, by using a deep learning-based watermark encoder, an audio watermarking technique that is inaudible to humans but is robust to various attacks and has high resilience can be implemented.

[0011] In addition, the problem-solving means of the present disclosure can solve the problem of misuse related to high-quality voice generation models, and can protect and promote the use of copyrighted works.

[0012] In addition, according to the problem solving means of the present disclosure, effective watermark insertion can be performed regardless of the influence of sound quality, unlike a method that uses only a preset frequency range particularly suitable for audio of a specific sound quality.

[0013] FIG. 1 is a schematic diagram of a system including an audio generating device according to one embodiment.

[0014] Figure 2 is an example of how an audio generation device operates.

[0015] Figure 3 is an exemplary drawing for explaining a watermark decoder.

[0016] Figure 4 is an exemplary diagram illustrating a process by which a watermark encoder generates marked audio data.

[0017] Figure 5 is an exemplary diagram for explaining the operations performed by the watermark encoder.

[0018] Figure 6 is an exemplary diagram for explaining the operations performed by the watermark decoder.

[0019] Figure 7 is an exemplary diagram for explaining the learning of a watermark encoder and a watermark decoder.

[0020] Figure 8 is an exemplary diagram for explaining a learning spectrogram used for learning a watermark encoder.

[0021] Figure 9 is an exemplary diagram for explaining multiple unit learning spectrograms.

[0022] Figure 10 is an exemplary diagram for explaining learning using a plurality of first unit learning spectrograms and a plurality of second unit learning spectrograms.

[0023] Figure 11 is an exemplary diagram for explaining a discriminator used in learning a watermark encoder.

[0024] Figure 12 is an exemplary drawing for explaining the first discrimination unit and the second discrimination unit.

[0025] Figure 13 is an exemplary diagram for explaining the process of generating the first judgment result and the second judgment result.

[0026] Fig. 14 is an exemplary diagram for explaining loss factors used in learning a watermark encoder.

[0027] Fig. 15 is a block diagram of an audio generation device according to one embodiment.

[0028] According to one embodiment of the present disclosure, a method for generating watermark audio may be provided using an encoder trained based on multiple discriminators, including the steps of: inputting watermark data and target audio data to be inserted into the watermark data into a watermark encoder; and obtaining marked audio data in which the watermark data is inserted into the target audio data as an output of the watermark encoder, wherein the watermark encoder is trained as a generator of a Generative Adversarial Network (GAN), and a discriminator of the GAN includes a first discriminator for determining the insertion of a watermark based on a spectrogram and a second discriminator for determining the insertion of a watermark based on an audio waveform.

[0029] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments presented below, but can be implemented in various different forms, and it should be understood that it includes all transformations, equivalents, and substitutes included in the spirit and technical scope of the present invention. The embodiments presented below are provided to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the invention of the scope of the invention. In describing the present invention, if a detailed description of a related known technology is judged to obscure the gist of the present invention, the detailed description thereof will be omitted.

[0030] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0031] Some embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a given function. Furthermore, for example, the functional blocks of the present disclosure may be implemented by various programming or scripting languages. The functional blocks may be implemented by algorithms that execute on one or more processors. Furthermore, the present disclosure may employ conventional techniques for electronic configuration, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.

[0032] Additionally, the connecting lines or connecting members between components depicted in the drawings are merely exemplary representations of functional connections and / or physical or circuit connections. In an actual device, connections between components may be represented by various functional connections, physical connections, or circuit connections that may be replaced or added.

[0033] The present disclosure will be described in detail with reference to the attached drawings below.

[0034] FIG. 1 is a schematic diagram of a system including an audio generating device according to one embodiment.

[0035] Referring to FIG. 1, an audio generation device (10) can input watermark data (102) and target audio data (101) that is the insertion target of the watermark data (102) into a watermark encoder (100).

[0036] The audio generation device (10) may include an audio generation system based on an artificial neural network. An artificial neural network refers to a model in which artificial neurons formed by combining synapses form a network and change the strength of the synapses through learning, thereby having problem-solving capabilities.

[0037] The audio generation device (10) can be implemented with various types of devices such as a personal computer (PC), a server device, a mobile device, an embedded device, etc., and as a specific example, it can include a PC, a smartphone, a tablet device, and / or an IoT (Internet of Things) device that performs generation of audio with a watermark inserted using an artificial neural network, but is not limited thereto.

[0038] In one embodiment, the audio generating device (10) may be a user's mobile device. For example, the audio generating device (10) may be implemented as a smartphone, a tablet PC, a PC, a smart TV, a personal digital assistant (PDA), a laptop, a media player, and / or other mobile electronic devices.

[0039] In another embodiment, the audio generation device (10) may be a server distinct from the user's device. The server may be implemented as a computer device or multiple computer devices that communicate over a network to provide commands, codes, files, content, services, etc. The network is a comprehensive data communication network that enables different entities to communicate smoothly with each other, and may include wired Internet, wireless Internet, and mobile wireless communication networks. For example, the network may include a Local Area Network (LAN), a Wide Area Network (WAN), a Value Added Network (VAN), a mobile radio communication network, a satellite communication network, and combinations thereof. In addition, wireless communication may include, but is not limited to, wireless LAN (Wi-Fi), Bluetooth, Bluetooth low energy, ZigBee, Wi-Fi Direct (WFD), ultra-wideband (UWB), infrared communication (IrDA, infrared Data Association), NFC (Near Field Communication), etc.

[0040] Furthermore, the audio generation device (10) according to one embodiment may be implemented as a dedicated hardware accelerator mounted on a personal computer (PC), a server device, a mobile device, an embedded device, etc. Alternatively, the audio generation device (10) may be implemented as a hardware accelerator such as a neural processing unit (NPU), a Tensor Processing Unit (TPU), a Neural Engine, etc., which are dedicated modules for driving an artificial neural network, but is not limited thereto.

[0041] In one embodiment, the target audio data (101) and the watermark data (102) may be received from an external device through a communication unit included in the audio generation device (10). In another embodiment, the target audio data (101) and the watermark data (102) may be obtained according to a user input through a user interface of the audio generation device (10), or may be selected from data pre-stored in a database of the audio generation device (10), but are not limited thereto.

[0042] In the present disclosure, target audio data (101) refers to audio data that is the target of watermark insertion, which is an analog or digital signal representing various sounds such as voice and / or music. In addition, watermark data (102) in the present disclosure refers to various types of data inserted into audio data to prevent misuse of synthetic audio and protect copyrights through synthetic voice detection and / or source verification.

[0043] In one embodiment, the target audio data (101) may include audio data without a watermark inserted. In one embodiment, the target audio data (101) may include artificially synthesized audio data or real audio data that is not synthesized.

[0044] In one embodiment, the watermark data (102) may include a bit string of a predetermined length. For example, the watermark data (102) may include a bit string of 32 bits, 64 bits, or 128 bits, but is not limited thereto.

[0045] In one embodiment, the watermark data (102) may include at least one of a hash value generated based on the target audio data (101) and user input data. For example, the watermark data (102) may include both a hash value and user input data.

[0046] In one embodiment, the audio generation device (10) can generate a hash value for target audio data (101) using a predetermined hash function and use the generated hash value as watermark data (102).

[0047] In another embodiment, the audio generation device (10) may obtain user input data indicating various properties such as the source and / or characteristics of target audio data (101), and use the obtained user input data as watermark data (102).

[0048] In another embodiment, the audio generation device (10) may generate a hash value of a first length for target audio data (101) using a predetermined hash function, and obtain input data of a second length indicating the source of the target audio data (101) as user input data. The audio generation device (10) may combine the hash value of the first length and the input data of the second length to generate watermark data (102) expressed as a bit string of a preset length.

[0049] The audio generation device (10) can obtain marked audio data (111) in which watermark data (102) is inserted into the target audio data (101) as an output of the watermark encoder (100) for the input target audio data (101).

[0050] In the present disclosure, marked audio data (111) refers to audio data generated by a watermark encoder (100). In one embodiment, the marked audio data (111) may be audio data that is subject to identification, detection, extraction, etc. of a watermark inserted into the marked audio data (111). In one embodiment, the marked audio data (111) may include audio data that is difficult to distinguish from target audio data (101) within a human audible range.

[0051] In one embodiment, the marked audio data (111) may include audio data expressed as a signal of the same or similar form as the target audio data (101). For example, the marked audio data (111) may be a digital signal having the same sampling rate, number of channels, or file format as the target audio data (101). As another example, the marked audio data (111) may include a digital signal having a preset sampling rate, number of channels, or file format regardless of the target audio data (101).

[0052] The watermark encoder (100) may be configured to receive audio data such as target audio data (101) and a watermark such as watermark data (102) as inputs and output audio data with a watermark inserted therein. As illustrated in FIG. 1, the watermark encoder (100) according to one embodiment may be implemented as a component included in an audio generation device (10). The watermark encoder (100) according to another embodiment may be implemented as a separate device distinct from the audio generation device (10) that receives audio data and a watermark as inputs from the audio generation device (10) and outputs audio data with a watermark inserted therein, and transmits the audio data with the watermark inserted therein to the audio generation device (10).

[0053] Each component described in this disclosure as performing at least one operation, such as the watermark encoder (100), may be implemented as a functional package within a single program running on a processor, or as multiple individual programs running on a single processor. In another embodiment, each component, such as the watermark encoder (100), may be implemented as individual hardware, but is not limited thereto.

[0054] According to a preferred embodiment, the audio generation device (10) can input target audio data (101) and watermark data (102) into the watermark encoder (100). The watermark encoder (100) can generate marked audio data (111) based on the target audio data and the watermark data (102). The audio generation device (10) can obtain the marked audio data (111) generated from the watermark encoder (100). Here, a detailed process of generating the marked audio data (111) by the watermark encoder (100) will be described below with reference to FIG. 4 and the like.

[0055] Figure 2 is an example of how an audio generation device operates.

[0056] Referring to FIG. 2, in step 210, the audio generation device (10) can input watermark data (102) and target audio data (101) that is the insertion target of the watermark data (102) into the watermark encoder (100).

[0057] In one embodiment, the watermark data (102) may include a bit string of a preset length.

[0058] In one embodiment, the watermark data (102) may include at least one of a hash value generated based on the target audio data (101) and user-arbitrary input data. For example, the watermark data (102) may include both a hash value and user-arbitrary input data.

[0059] In step 220, the audio generation device (10) can obtain marked audio data (111) in which watermark data (102) is inserted into the target audio data (101) as an output of the watermark encoder (100) for the input target audio data (101).

[0060] In one embodiment, the watermark encoder (100) may include a pre-trained deep learning model having an invertible neural network (INN) structure.

[0061] In one embodiment, the watermark encoder (100) may be configured to perform a plurality of reversible operations, and the watermark decoder may be configured to perform a plurality of inverse operations corresponding to each of the plurality of reversible operations. In one embodiment, the watermark decoder may be used for training the watermark encoder (100). In another embodiment, the watermark decoder may be used to extract watermark data for arbitrary audio data.

[0062] In one embodiment, the watermark encoder (100) can generate a first watermark spectrogram based on watermark data (102), generate a first audio spectrogram based on target audio data (101), generate a final spectrogram based on the first watermark spectrogram and the first audio spectrogram using a plurality of reversible operations, and generate marked audio data (111) based on the final spectrogram.

[0063] According to one embodiment, the plurality of reversible operations may include any reversible operation that generates a watermark spectrogram and an audio spectrogram corresponding to a next reversible operation for any reversible operation based on a watermark spectrogram and an audio spectrogram corresponding to any reversible operation. In this case, the watermark encoder (100) may generate a final spectrogram as an operation result of the last reversible operation among the plurality of reversible operations, and any reversible operation may not include the last reversible operation.

[0064] That is, when there are K multiple reversible operations, the lth reversible operation among the multiple reversible operations may be an operation that generates an (1+1)th watermark spectrogram and an (1+1)th audio spectrogram based on the lth watermark spectrogram and the lth audio spectrogram, and the watermark encoder (100) may generate the (K+1)th audio spectrogram as the final spectrogram as a result of the Kth reversible operation. At this time, K may be any one of natural numbers greater than or equal to 2, and l may be a natural number greater than or equal to 1 and less than or equal to K.

[0065] According to one embodiment, a first watermark spectrogram may be generated by generating first preprocessing data having the same time length as target audio data (101) based on watermark data (102) but including watermark data (102) for at least one time section, and generating a first watermark spectrogram by performing a Short Time Fourier Transform (STFT) on the first preprocessing data. In this case, the first audio spectrogram may be generated by performing an STFT on the target audio data (101).

[0066] In one embodiment, the first preprocessing data may include watermark data (102) for all time intervals divided into predetermined lengths.

[0067] In one embodiment, the watermark encoder (100) may be trained using a watermark decoder that has an inverse operation relationship with the watermark encoder (100), and the training of the watermark encoder (100) may be performed based on a loss factor that includes a difference between training watermark data input to the watermark encoder (100) and restored training watermark data output from the watermark decoder.

[0068] Training of a watermark encoder (100) according to one embodiment may be performed by inputting target training audio data and training watermark data into the watermark encoder (100), obtaining marked training audio data output from the watermark encoder (100), inputting the marked training audio data into a watermark decoder, obtaining restored training watermark data output from the watermark decoder, and adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a difference between the training watermark data and the restored training watermark data.

[0069] At this time, the watermark decoder can be configured to perform a plurality of inverse operations corresponding to each of the plurality of reversible operations performed by the watermark encoder (100), so that the parameters of the watermark decoder corresponding to the parameters of the watermark encoder (100) being adjusted can be adjusted together.

[0070] In one embodiment, obtaining restored learning watermark data may include generating preprocessed marked learning audio data by performing at least one of a time axis shift and a preset attack simulation on marked learning audio data, inputting the preprocessed marked learning audio data into a watermark decoder, and obtaining restored learning watermark data output from the watermark decoder.

[0071] Meanwhile, the loss factor may include the difference between the learning watermark data and the restored learning watermark data, and the difference between the target learning audio data and the marked learning audio data.

[0072] In one embodiment, the watermark encoder (100) may be trained by adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a first learning spectrogram and a second learning spectrogram. In this case, the first learning spectrogram may include a spectrogram of target learning audio data input to the watermark encoder (100). In addition, the second learning spectrogram may include a spectrogram of marked learning audio data output from the watermark encoder (100) based on the input of target learning audio data and learning watermark data.

[0073] In one embodiment, the loss factor may further include the magnitude of the difference between the first learning spectrogram and the second learning spectrogram, and the magnitude of the first learning spectrogram.

[0074] The first learning spectrogram may include a plurality of first unit learning spectrograms generated by performing STFT on target learning audio data with different window sizes. Furthermore, the second learning spectrogram may include a plurality of second unit learning spectrograms corresponding to the plurality of first unit learning spectrograms.

[0075] In one embodiment, the loss element may further include a first transformed value that is a log-scale transformed value of the first learning spectrogram and a second transformed value that is a log-scale transformed value of the second learning spectrogram.

[0076] For example, the loss component may include the magnitude of the difference between the first transformed value and the second transformed value, and a predetermined normalization variable.

[0077] At this time, the normalization variable can be set based on the resolution for the first learning spectrogram or the second learning spectrogram.

[0078] In one embodiment, the watermark encoder (100) may be trained as a generator of a Generative Adversarial Network (GAN). In this case, the discriminator of the GAN may include a first discriminator that determines whether a watermark has been inserted based on a spectrogram, and a second discriminator that determines whether a watermark has been inserted based on an audio waveform.

[0079] According to one embodiment, the first determination unit and the second determination unit can output the results of determination in parallel for the marked training audio data. At this time, the watermark encoder (100) can be trained by adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including the first determination result of the first determination unit and the second determination result of the second determination unit. At this time, the marked training audio data can be output from the watermark encoder (100) based on the input of the target training audio data and the training watermark data.

[0080] Training of a watermark encoder (100) according to one embodiment may be performed by inputting target training audio data and training watermark data into the watermark encoder (100) and obtaining marked training audio data output from the watermark encoder (100), inputting the marked training audio data into a first determination unit and a second determination unit in parallel, obtaining a first determination result and a second determination result for the input marked training audio data, and adjusting at least one parameter included in an operation of the watermark encoder (100) based on the first determination result of the first determination unit and the second determination result of the second determination unit.

[0081] In one embodiment, the first discrimination unit may include a plurality of first sub-discrimination units. Each of the plurality of first sub-discrimination units may perform STFT on the marked training audio data with a window size corresponding to each of the plurality of first sub-discrimination units to generate a discrimination spectrogram, and may output a first intermediate discrimination result in parallel based on the discrimination spectrogram. In addition, the first discrimination result may include a plurality of first intermediate discrimination results.

[0082] Furthermore, in one embodiment, the second discrimination unit may include a plurality of second sub-discrimination units. In this case, each of the plurality of second sub-discrimination units may generate two-dimensional audio data by reconstructing the marked learning audio data using a period corresponding to each of the plurality of second sub-discrimination units, and may output a second intermediate discrimination result in parallel based on the two-dimensional audio data. Furthermore, the second discrimination result may include a plurality of second intermediate discrimination results.

[0083] Meanwhile, the discriminator can be learned by adjusting parameters for at least one of the first discriminator and the second discriminator based on the first discriminator result and the second discriminator result.

[0084] Figure 3 is an exemplary drawing for explaining a watermark decoder.

[0085] Referring to FIG. 3, the watermark decoder (200) may be configured to receive audio data with a watermark inserted, such as marked audio data (111), as input and restore the original audio data and the original watermark. In one embodiment, the watermark decoder (200) may receive marked audio data (111) generated from the watermark encoder (100) as input and output restored audio data (121) and restored watermark data (122).

[0086] In one embodiment, the watermark decoder (200) may be used to extract an embedded watermark for any audio data. In another embodiment, the watermark decoder (200) may be used to train the watermark encoder (100).

[0087] In one embodiment, the watermark decoder (200) may receive audio data as input in which the watermark insertion is unclear. At this time, the user may determine whether the audio input to the watermark decoder (200) includes a watermark based on the output of the watermark decoder (200). For example, if the watermark output from the watermark decoder (200) does not follow a predetermined format, the user may determine that the audio data input to the watermark decoder (200) is not marked audio data (111) in which a watermark is inserted.

[0088] In one embodiment, the watermark decoder (200) may be trained to extract restored watermark data (122) from marked audio data (111). For example, the watermark decoder (200) may be trained to extract restored watermark data (122) identically to the watermark data (102). For another example, the watermark decoder (200) may be trained to have the extracted restored watermark data (122) differ from the watermark data (102) within a preset tolerance range.

[0089] Meanwhile, a watermark decoder (200) according to one embodiment can be trained together with a watermark encoder (100), and the detailed training process of the watermark decoder (200) will be described later with reference to FIG. 6, etc.

[0090] Figure 4 is an exemplary diagram illustrating a process by which a watermark encoder generates marked audio data.

[0091] Referring to FIG. 4, a watermark encoder (100) according to one embodiment may include an encoder operation preprocessing unit (110), a reversible operation unit (120), and an encoder operation postprocessing unit (130).

[0092] The encoder operation preprocessing unit (110) of the present disclosure is an element that preprocesses target audio data (101) and / or watermark data (102) so that the target audio data (101) and watermark data (102) have a format that is easy to operate on each other.

[0093] In one embodiment, the watermark encoder (100) can generate a first watermark spectrogram by preprocessing watermark data (102) and can generate a first audio spectrogram by preprocessing target audio data (101). The spectrogram of the present disclosure refers to data that extracts the frequency components and intensity of a signal over time from a signal such as audio data.

[0094] For example, the encoder operation preprocessing unit (110) can generate a first audio spectrogram by performing STFT on target audio data (101). Here, STFT means performing Fourier transform on audio data for each predetermined window size.

[0095] In addition, for example, the encoder operation preprocessing unit (110) may generate first preprocessing data having the same time length as the target audio data (101) based on the watermark data (102) but including the watermark data (102) for at least one time section, and may generate a first watermark spectrogram by performing STFT on the first preprocessing data.

[0096] In one embodiment, the first preprocessing data may be time series data having the same time length as the target audio data (101). The encoder operation preprocessing unit (110) may include an FC layer (Fully Connected layer) that converts a bit string of a predetermined length into time series data having the same time length as the target audio data (101).

[0097] In one embodiment, the first preprocessing data may include watermark data (102) for at least one time interval. For example, if the first preprocessing data has a time length of 100 seconds, the time interval may be a time interval having a preset time length, such as 1 second, 10 seconds, or 100 seconds. According to one embodiment, the time interval may include intervals of 0s-1s, 1s-2s, 0s-10s, 1s-11s, etc. of the first preprocessing data.

[0098] In one embodiment, the first preprocessing data may include watermark data (102) for at least one time interval divided into a predetermined time length. For example, among the time intervals of the first preprocessing data divided into 1-second lengths such as 0s-1s, 1s-2s, and 2s-3s, at least one interval may include watermark data (102).

[0099] In another embodiment, the first preprocessing data may include watermark data (102) for all time intervals divided into predetermined time lengths. For example, all time intervals of the first preprocessing data divided into 1-second lengths, such as 0s-1s, 1s-2s, and 2s-3s, may include watermark data (102). In this case, if the predetermined time length is sufficiently short, the watermark data (102) may be repeatedly reflected for the entire time interval of the marked audio data (111). Through this, restored watermark data (122) may be extracted from only a specific time interval of the marked audio data (111).

[0100] Meanwhile, the first STFT performed in the process of generating the first audio spectrogram and the second STFT performed in the process of generating the first watermark spectrogram may be identical to each other. That is, the parameters of the first STFT, including window type, window size, sampling rate, and / or FFT size, may be identical to the parameters of the second STFT, but are not limited thereto.

[0101] In the present disclosure, the reversible operation unit (120) performs the insertion of watermark data (102) into target audio data (101). According to one embodiment, the reversible operation unit (120) can generate a final spectrogram using the first audio spectrogram and the first watermark spectrogram, which have been preprocessed by the encoder operation preprocessing unit (110).

[0102] In one embodiment, the reversible operation unit (120) can generate a final spectrogram using multiple operations. In this case, a single operation step consisting of at least one unit operation can be understood as a computational layer. Meanwhile, the specific processes of the multiple operations performed by the reversible operation unit (120) will be described below with reference to FIG. 5 and elsewhere.

[0103] In one embodiment, the encoder operation post-processing unit (130) can generate marked audio data (111) based on the final spectrogram. For ease of operation, STFT is performed on target audio data (101) input to the watermark encoder (100), thereby generating a first audio spectrogram. Thereafter, an operation using the generated first audio spectrogram is performed, thereby generating a final spectrogram. In order to convert the final spectrogram into an audio data format, ISTFT (Inverse Short Time Fourier Transform) is performed on the final spectrogram on which all operations have been performed, thereby generating marked audio data (111).

[0104] Meanwhile, the STFT performed in the process of generating the first audio spectrogram and the ISTFT performed in the process of generating marked audio data (111) may have an inverse transformation relationship. That is, the ISTFT may be configured to output the input to the STFT when the output to the STFT is input.

[0105] Figure 5 is an exemplary diagram for explaining the operations performed by the watermark encoder.

[0106] Referring to FIG. 5, the reversible operation unit (120) can generate a final audio spectrogram from the first audio spectrogram and the first watermark spectrogram using multiple operations.

[0107] In one embodiment, the watermark encoder (100) may include a pre-trained deep learning model having a reversible neural network structure. At this time, the reversible operation unit (120) may generate a final spectrogram using multiple reversible operations. In the present disclosure, a reversible operation refers to an operation in which an inverse operation exists for every unit operation included in the operation. At this time, the unit operation refers to arithmetic operations, application of functions, and combinations thereof performed on an audio spectrogram or a watermark spectrogram.

[0108] In one embodiment, the watermark encoder (100) can generate a first watermark spectrogram based on watermark data (102) and a first audio spectrogram based on target audio data (101). Through the process described above with reference to FIG. 4, the watermark encoder (100) can generate the first watermark spectrogram and the first audio spectrogram using the encoder operation preprocessing unit (110).

[0109] The watermark encoder (100) can generate a final spectrogram using multiple reversible operations performed by the reversible operation unit (120).

[0110] A plurality of reversible operations according to one embodiment may include any reversible operation that generates a watermark spectrogram and an audio spectrogram corresponding to a next reversible operation for any reversible operation based on a watermark spectrogram and an audio spectrogram corresponding to any reversible operation. Here, any reversible operation may not include a last reversible operation, since the last reversible operation does not have a next reversible operation.

[0111] The watermark encoder (100) can generate a final spectrogram as the result of the last reversible operation among multiple reversible operations. For example, when there are K multiple reversible operations, the reversible operation unit (120) can generate the (K+1)th audio spectrogram as the result of the Kth reversible operation as the final spectrogram.

[0112] That is, when there are K multiple reversible operations, the lth reversible operation among the multiple reversible operations may be an operation that generates an (l+1)th watermark spectrogram and an (l+1)th audio spectrogram based on the lth watermark spectrogram and the lth audio spectrogram. At this time, K may be a natural number greater than or equal to 2, and l may be a natural number greater than or equal to 1 and less than or equal to K. The operation layer l illustrated in Fig. 5 may be understood as the lth reversible operation stage that includes the first operation and the second operation as unit operations, respectively.

[0113] In a specific embodiment, in the operation layer l, the first operation may generate the first audio spectrogram (311) and the first watermark spectrogram (312) by at least one unit operation using at least one of the first audio spectrogram (301), the first watermark spectrogram (302), the l+1 audio spectrogram (311) and the l+1 watermark spectrogram (312).

[0114] For example, in the operation layer l, the (1+1)-th watermark spectrogram (312) may be generated through a first operation using the (1+1)-th audio spectrogram (301) and the (1)-th watermark spectrogram (302), and the (1+1)-th audio spectrogram (311) may be generated through a second operation using the (1+1)-th watermark spectrogram (312) and the (1)-th audio spectrogram (301). The following mathematical expression 1 is an example of the first operation performed in the operation layer l of the watermark encoder (100).

[0115]

[0116] In the above mathematical formula 1, corresponds to the l+1 audio spectrogram (311), corresponds to the first audio spectrogram (301), can represent the first watermark spectrogram (302). Here, includes a hidden function. The hidden function hides the original shape of the watermark spectrogram and allows the watermark to be embedded in the audio in a hidden form, rather than being publicly embedded in the audio. The hidden function can be selected from any function that has an inverse function.

[0117] Meanwhile, the following mathematical expression 2 is an example of the second operation performed in the operation layer l of the watermark encoder (100).

[0118]

[0119] In the above mathematical expression 2, represents the l+1 watermark spectrogram (312). Here, contains the influence function, contains the sigmoid function, may include an error correction function. The influence function according to one embodiment determines the influence of the first watermark spectrogram (302) through masking of the (l+1)th audio spectrogram (311) and the first audio spectrogram (301). The sigmoid function is used to appropriately reflect the result of applying the influence function. The error correction function is used for learning to extract the watermark spectrogram from the audio spectrogram of the previous layer. The influence function and the error correction function may be selected from any function having an inverse function.

[0120] Figure 6 is an exemplary diagram for explaining the operations performed by the watermark decoder.

[0121] Referring to FIG. 6, the watermark encoder (100) may be configured to perform a plurality of reversible operations, and the watermark decoder (200) may be configured to perform a plurality of inverse operations corresponding to each of the plurality of reversible operations.

[0122] That is, the watermark decoder (200) can perform the inverse operation of all operations performed in the watermark encoder (100) in the reverse order of the operations performed in the watermark encoder (100). Here, when the order of the first operation and the second operation performed in the operation layer l of the watermark encoder (100) is changed, the result of the operation may or may not be changed. If the result of the operation is changed, it should be noted that the operations must be performed in the order of the inverse operation of the second operation and the inverse operation of the first operation in the operation layer n of the watermark decoder (200) corresponding to the operation layer l.

[0123] The following mathematical expression 3 is an example of an inverse operation for the second operation of the watermark encoder (100) exemplified by the above mathematical expression 2, and is an example of the first inverse operation performed in the operation layer n of the watermark decoder (200).

[0124] The following mathematical expression 4 is an example of an inverse operation for the first operation of the watermark encoder (100) exemplified by the above mathematical expression 1, and is an example of a second inverse operation performed in the operation layer n of the watermark decoder (200).

[0125]

[0126]

[0127] In the above mathematical expressions 3 and 4, n may be K-l+1. That is, the inverse operation of the computation layer 1 of the watermark encoder (100) may be performed in the computation layer K of the watermark decoder (100), and the inverse operation of the computation layer 2 of the watermark encoder (100) may be performed in the computation layer K-1 of the watermark decoder (100). Meanwhile, if the operations performed in all computation layers of the watermark encoder (100) are the same, it can be understood that the same inverse operation for the above operations is performed in all computation layers of the watermark decoder (200).

[0128] Meanwhile, only the spectrogram corresponding to the audio data is input to the computation layer 1 of the watermark decoder (200), and the spectrogram corresponding to the watermark is not input. In one embodiment, the computation layer 1 of the watermark decoder (200) may use a randomly generated spectrogram or a preset spectrogram as the first spectrogram corresponding to the watermark.

[0129] As illustrated in FIG. 1, the watermark encoder (100) receives target audio data (101) and watermark data (102) as inputs and outputs marked audio data (111). The watermark encoder (100) is trained to output marked audio data (111) that is difficult to distinguish from the target audio data (101). In addition, the watermark decoder (200), as illustrated in FIG. 3, receives marked audio data (111) as inputs and outputs restored audio data (121) and restored watermark data (122). The watermark decoder (200) is trained to output restored watermark data (122) that is identical to the watermark data (102) or similar to the watermark data (102) within a range that at least satisfies a predetermined criterion.

[0130] The target audio data (101), watermark data (102), marked audio data (111), restored audio data (121), and restored watermark data (122) of the present disclosure can be understood as being used in the context of utilizing a watermark encoder (100) and a watermark decoder (200). Meanwhile, the target learning audio data (131), learning watermark data (132), marked learning audio data (141), restored learning audio data (151), and restored learning watermark data (152) of the present disclosure can be understood as being used in the context of learning of a watermark encoder (100) and a watermark decoder (200).

[0131] In one embodiment, the watermark encoder (100) can be trained using a watermark decoder (200) that has an inverse operation relationship with the watermark encoder (100), and the training of the watermark encoder (100) can be performed based on a loss factor including the difference between training watermark data (131) input to the watermark encoder (100) and restored training watermark data (152) output from the watermark decoder (200).

[0132] Meanwhile, the watermark decoder (200) has an inverse operation relationship with the watermark encoder (100), so that the learning of the watermark decoder (200) is automatically performed according to the learning of the watermark encoder (100). That is, according to one embodiment, the learning of the watermark encoder (100) can be understood to include the learning of the watermark decoder (200).

[0133] In the present disclosure, a loss element refers to a loss function used to train a deep learning-based model, at least some elements constituting the loss function, and / or at least one variable used to derive elements constituting the loss function. In one embodiment, the watermark encoder (100), the watermark decoder (200), and / or the discriminator (800) described below with reference to FIG. 11, etc., may include a deep learning-based model.

[0134] In one embodiment, the deep learning-based model may be a deep learning model implemented with one or a combination of two or more of various artificial neural network models, such as a pre-net, a CBHG module, a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), an Invertible Neural Network (INN), a Long Short-Term Memory Network (LSTM), and a Bidirectional Recurrent Deep Neural Network (BRDNN).

[0135] According to one embodiment, the loss factor may include the difference between the learning watermark data (132) and the restored learning watermark data (152), and / or the difference between the target learning audio data (131) and the marked learning audio data (141). For example, the watermark encoder (100) may be trained to reduce the difference between the learning watermark data (132) and the restored learning watermark data (152), and to reduce the difference between the target learning audio data (131) and the marked learning audio data (141).

[0136] Training of a watermark encoder (100) according to one embodiment may be performed by inputting target training audio data (131) and training watermark data (132) into the watermark encoder (100), obtaining marked training audio data (141) output from the watermark encoder (100), inputting the marked training audio data (141) into the watermark decoder (200), obtaining restored training watermark data (152) output from the watermark decoder (200), and adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a difference between the training watermark data (132) and the restored training watermark data (152).

[0137] Here, at least one parameter includes at least one parameter that constitutes a unit operation, and may include, but is not limited to, various parameters that constitute a coefficient value for a value input to the operation layer, and / or a function applied to the input value.

[0138] At this time, the watermark decoder (200) in an inverse operation relationship with the watermark encoder (100) can be learned by adjusting the parameters of the watermark decoder (200) corresponding to the adjusted parameters of the watermark encoder (100).

[0139] Meanwhile, in one embodiment, the inverse operation for the watermark encoder (100) performed by the watermark decoder (200) may further include inverse operations of not only the reversible operation unit (120), but also the encoder operation preprocessing unit (110) and the encoder operation postprocessing unit (130).

[0140] In one embodiment, the encoder operation preprocessing unit (110) can generate first preprocessing data for the learning watermark data (132) through the FC layer. The encoder operation preprocessing unit (110) can perform STFT on the first preprocessing data and the target learning audio data (131), respectively, to generate a first audio spectrogram and a first watermark spectrogram. In addition, the encoder operation postprocessing unit (110) can perform ISTFT on the final spectrogram to generate marked learning audio data (141).

[0141] The decoder operation preprocessing unit (not shown) of the watermark decoder (200) can perform the inverse operation of the encoder operation postprocessing unit (130). That is, the decoder operation preprocessing unit can receive marked training audio data (141) as input and perform STFT. Thereafter, the inverse operation unit of the watermark decoder (200) performs the inverse operation for the reversible operation unit (120) of the watermark encoder (100).

[0142] In addition, the decoder operation post-processing unit (not shown) of the watermark decoder (200) can perform the inverse operation of the encoder twist pre-processing unit (110). That is, the decoder operation post-processing unit can perform ISTFT on the (K+1)th audio spectrogram and the (K+1)th watermark spectrogram output from the operation layer K of the watermark decoder (200), and generate restored training audio data (151) corresponding to the (K+1)th audio spectrogram and first post-processing data corresponding to the (K+1)th watermark spectrogram. The decoder operation post-processing unit can include a second FC layer that performs the inverse operation on the first FC layer of the encoder operation pre-processing unit (110), that is, converts time series data of the audio data length into the form of a watermark. The decoder operation post-processing unit can convert the first post-processing data into restored training watermark data (152) through the second FC layer.

[0143] Figure 7 is an exemplary diagram for explaining the learning of a watermark encoder and a watermark decoder.

[0144] Referring to FIG. 7, obtaining restored learning watermark data (152) may include generating preprocessed marked learning audio data (161) by performing at least one of a time axis shift and a preset attack simulation on marked learning audio data (141), inputting the preprocessed marked learning audio data (161) into a watermark decoder (200), and obtaining restored learning watermark data (152) output from the watermark decoder (200).

[0145] Preprocessed marked training audio data (161) can be obtained by inputting marked training audio data (141) into an audio preprocessing unit (400). In one embodiment, the audio preprocessing unit (400) can perform an arbitrary parallel translation with respect to the time axis on the marked training audio data (161). Through this, the watermark encoder (100) and the watermark decoder (200) can be trained to extract or detect a watermark regardless of the start and end points of the time interval in which the watermark is inserted.

[0146] In one embodiment, the attack simulation may include at least one of various audio data attacks, such as Pitch Shift, Random Noise, Low-pass Filter, Median Filter, Re-Sampling, Amplitude Scaling, Lossy Compression, Quantization, Time Stretch, Crop, and Echo Addition.

[0147] Here, pitch shift means that the pitch, pitch, or octave has been adjusted. Random noise means adding random noise to audio data. A low-pass filter refers to a filter that removes high-frequency signals and passes only low-frequency signals. A median filter refers to a filter that takes the median value of surrounding samples for each sample of the signal. Resampling refers to changing the sampling rate of audio data. Amplitude scaling refers to adjusting the amplitude of an audio signal, which represents intensity. Lossy compression refers to compressing audio data while losing some of the data information. Quantization refers to converting a continuous analog signal into a discrete digital signal. Time stretching refers to increasing or decreasing the playback time of audio. Truncation refers to removing part of the audio data. Echo addition refers to adding an echo effect to audio data.

[0148] This enables robust watermark audio generation / extraction, even when watermark positioning issues and various audio attacks are addressed. This means watermark extraction is possible even if the segment containing the watermark begins at any arbitrary location. Furthermore, learning using attack simulations demonstrates that the watermark remains stable even under various audio corruptions, enabling robust watermark insertion for information such as copyright information that requires identification even when audio is corrupted.

[0149] Meanwhile, the loss factor may include the difference (520) between the learning watermark data (132) and the restored learning watermark data (152), and the difference (510) between the target learning audio data (131) and the marked learning audio data (141).

[0150] The following mathematical expression 5 is an example of a loss factor including the difference (510) between target learning audio data (131) and marked learning audio data (141).

[0151]

[0152] In the above mathematical expression 5, is the first loss function, is the target learning audio data (131), is marked learning audio data (141), is a watermark encoder (100), may represent learning watermark data (132). The watermark encoder (100) may be trained based on a final loss function including the first loss function. That is, the watermark encoder (100) may be trained by adjusting parameters in a direction that reduces the first loss function.

[0153] This minimizes the impact of watermarks on audio heard by humans, while maintaining audio quality.

[0154] The following mathematical expression 6 is an example of a loss factor including the difference (520) between the learning watermark data (132) and the restored learning watermark data (152).

[0155]

[0156] In the above mathematical expression 6, is the second loss function, is the restored learning watermark data (152), is a watermark decoder (200), is an attack simulation, is the time axis shift, can represent the shift time length. The watermark encoder (100) can be trained based on the final loss function including the second loss function. That is, the watermark encoder (100) can be trained by adjusting the parameters in a direction that reduces the second loss function.

[0157] Figure 8 is an exemplary diagram for explaining a learning spectrogram used for learning a watermark encoder.

[0158] Referring to FIG. 8, a first learning spectrogram (610) and / or a second learning spectrogram (620) may be used in the learning process of the watermark encoder (100). Here, the first learning spectrogram (610) may include a spectrogram of target learning audio data (131) input to the watermark encoder (100). In addition, the second learning spectrogram (620) may include a spectrogram of marked learning audio data (141) output from the watermark encoder (100) based on the input of target learning audio data (131) and learning watermark data (132).

[0159] In one embodiment, the first learning spectrogram (610) may be a spectrogram generated by performing STFT on target learning audio data (131), and the second learning spectrogram (620) may be a spectrogram generated by performing STFT on marked learning audio data (141).

[0160] Compared to an example of the learning process described above through the above mathematical expression 5, the learning process of the watermark encoder (100) using the first learning spectrogram (610) and the second learning spectrogram (620) performs comparison at the spectrogram level, so learning can be performed regardless of the sound quality of the audio data used for learning.

[0161] In one embodiment, in the training of the watermark encoder (100), comparisons between spectrograms as well as comparisons between original audio and watermark audio are further utilized, thereby reducing the speed of training and the number of training data required for training.

[0162] In one embodiment, the watermark encoder (100) can be trained by adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a first learning spectrogram (610) and a second learning spectrogram (620). In one embodiment, the loss factor can include a difference (630) between the first learning spectrogram (610) and the second learning spectrogram (620). The watermark encoder (100) can be trained in a direction that reduces the difference (630) between the first learning spectrogram (610) and the second learning spectrogram (620).

[0163] Meanwhile, in one embodiment, the loss factor for learning the watermark encoder (100) may further include a first conversion value (710), which is a log scale conversion value of the first learning spectrogram (610), and a second conversion value (720), which is a log scale conversion value of the second learning spectrogram (610). Since human hearing responds closer to a log scale than linearly to changes in sound intensity, learning the watermark encoder (100) using the first conversion value (710) and the second conversion value (720) may be effective in terms of the signal intensity of audio data.

[0164] In one embodiment, the watermark encoder (100) can be trained by adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a first transformed value (710) and a second transformed value (720). In one embodiment, the loss factor can include a difference (730) between the first transformed value (710) and a second transformed value (720). The watermark encoder (100) can be trained in a direction that reduces the difference (730) between the first transformed value (710) and the second transformed value (720).

[0165] Figure 9 is an exemplary diagram for explaining multiple unit learning spectrograms.

[0166] Referring to FIG. 9, the first learning spectrogram (610) may include a plurality of first unit learning spectrograms generated by performing STFT with a plurality of different window sizes on target learning audio data (131). In addition, the second learning spectrogram (620) may include a plurality of second unit learning spectrograms corresponding to the plurality of first unit learning spectrograms. Here, the plurality of unit learning spectrograms refers to a plurality of spectrograms generated by performing STFT with a plurality of window sizes.

[0167] Meanwhile, a second unit learning spectrogram corresponding to a specific first unit learning spectrogram may mean a spectrogram generated with an STFT having the same parameters as the STFT used to generate the first unit learning spectrogram.

[0168] In one embodiment, STFT can be performed with different parameters. The parameters of STFT can include a window size, an FFT size, a HOP size, etc. In particular, when STFT is performed with M different window sizes on one audio data, M spectrograms can be generated for different bands in the frequency domain. Depending on the window size, the time resolution and frequency resolution of the generated spectrograms can vary. According to the spectrogram corresponding to a large window size, the difference between the first unit learning spectrogram and the second unit learning spectrogram can be calculated with a high frequency resolution. On the other hand, according to the discriminant spectrogram corresponding to a small window size, the difference between the first unit learning spectrogram and the second unit learning spectrogram can be calculated with a high time resolution.

[0169] For example, when performing STFT with a first window size on target learning audio data (131), a first unit learning spectrogram 1 may be generated, when performing STFT with a second window size, a first unit learning spectrogram 2 may be generated, and when performing STFT with an M-th window size, a first unit learning spectrogram M may be generated. Here, the first learning spectrogram (610) may include a set of first unit learning spectrograms 1 to M.

[0170] Similarly, when STFT is performed with a first window size on the marked learning audio data (141), a second unit learning spectrogram 1 can be generated, when STFT is performed with a second window size, a second unit learning spectrogram 2 can be generated, and when STFT is performed with an M-th window size, a second unit learning spectrogram M can be generated. Here, the second learning spectrogram (620) can include a set of second unit learning spectrograms 1 to M.

[0171] Here, the difference between the first unit learning spectrogram 1 and the second unit learning spectrogram 1 is the difference analyzed with the time resolution and frequency resolution corresponding to the first window size. Similarly, the difference between the second unit learning spectrogram M and the second unit learning spectrogram M is the difference analyzed with the time resolution and frequency resolution corresponding to the second window size.

[0172] Meanwhile, in one embodiment, a log scale transformation value may be generated for each unit learning spectrogram. For example, a first unit transformation value 1, which is a log scale transformation value of a first unit learning spectrogram 1, may be generated, a first unit transformation value 2, which is a log scale transformation value of a first unit learning spectrogram 2, may be generated, and a first unit transformation value M, which is a log scale transformation value of a first unit learning spectrogram M, may be generated. The first transformation value (710) for the first learning spectrogram (610) may include a set of first unit transformation values ​​1 to M. Here, the unit transformation value may be understood as a log scale transformation value corresponding to each unit learning spectrogram.

[0173] Similarly, a second unit transformation value 1, which is a log scale transformation value of the second unit learning spectrogram 1, can be generated, a second unit transformation value 2, which is a log scale transformation value of the second unit learning spectrogram 2, can be generated, and a second unit transformation value M, which is a log scale transformation value of the second unit learning spectrogram M, can be generated. The second transformation value (720) for the second learning spectrogram (620) can include a set of second unit transformation values ​​1 to M.

[0174] The specific learning process of the watermark encoder (100) using multiple unit learning spectrograms and multiple unit conversion values ​​will be described later with reference to FIG. 10, etc.

[0175] Figure 10 is an exemplary diagram for explaining learning using a plurality of first unit learning spectrograms and a plurality of second unit learning spectrograms.

[0176] Referring to FIG. 10, the watermark encoder (100) can be trained based on a plurality of unit learning spectrograms and a plurality of unit conversion values.

[0177] In one embodiment, the watermark encoder (100) may be trained based on the difference between a first unit learning spectrogram and a second unit learning spectrogram on which STFT is performed with the same window size. The first unit difference refers to the difference between the first unit learning spectrogram and the second unit learning spectrogram. For example, the first unit learning spectrogram 1 and the second unit learning spectrogram 1 may be a pair of spectrograms on which STFT is performed with the same first window size. In this case, the difference between the first unit learning spectrogram 1 and the second unit learning spectrogram 1 may be understood as the first unit difference 1. Similarly, the difference between the first unit learning spectrogram k and the second unit learning spectrogram k may be understood as the first unit difference k.

[0178] Additionally, the watermark encoder (100) can be trained based on the difference between the first unit conversion value and the second unit conversion value. The second unit difference refers to the difference between the first unit conversion value and the second unit conversion value.

[0179] In the first unit transformation value 1 corresponding to the first unit learning spectrogram 1 on which STFT is performed with the first window size and the second unit transformation value 1 corresponding to the second unit spectrogram 1 on which STFT is performed with the first window size, the difference between the first unit transformation value 1 and the second unit transformation value 1 can be understood as the second unit difference 1. Similarly, the difference between the first unit transformation value k and the second unit transformation value k can be understood as the second unit difference k.

[0180] Here, the difference (630) between the first learning spectrogram (610) and the second learning spectrogram (620) may include a set of first unit differences 1 to M, and the difference (730) between the first conversion value (710) and the second conversion value (720) may include a set of second unit differences 1 to M.

[0181] In one embodiment, the watermark encoder (100) can be trained by adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a first learning spectrogram (610) and a second learning spectrogram (620). In one embodiment, the loss factor can include a magnitude of a difference (630) between the first learning spectrogram (610) and the second learning spectrogram (620), and a magnitude of the first learning spectrogram (610).

[0182] The following mathematical expression 7 is an example of a loss factor including the magnitude of the difference (630) between the first learning spectrogram (610) and the second learning spectrogram (620), and the magnitude of the first learning spectrogram (610).

[0183]

[0184] In the above mathematical formula 7, is the third loss function, is the size of the difference (630) between the first learning spectrogram (610) and the second learning spectrogram (620) expressed by the Frobenius norm, represents the size of the first learning spectrogram (610) expressed by the Frobenius norm.

[0185] Meanwhile, depending on the STFT parameters such as window size, can have different values, In addition, the loss value of the third loss function can have different values, and can vary depending on the parameters of the STFT that generates the unit learning spectrogram.

[0186] The watermark encoder (100) can be trained based on the final loss function including the third loss function. That is, the watermark encoder (100) can be trained by adjusting parameters in a direction that reduces the third loss function. Meanwhile, the third loss function can relatively significantly reflect the influence of the spectral peak appearing in the training spectrogram. That is, the third loss function can be understood as useful for fitting the watermark encoder (100) to a frequency band with high energy in the training spectrogram.

[0187] In one embodiment, the watermark encoder (100) can be trained by adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a first transformed value (710) and a second transformed value (720). For example, the loss factor can include a magnitude of a difference (730) between the first transformed value (710) and the second transformed value (720), and a predetermined normalization variable. In this case, the normalization variable can be set based on the resolution for the first learning spectrogram (610) or the second learning spectrogram (620). The normalization variable can be introduced to normalize the magnitude of the difference (730) between the first transformed value (710) and the second transformed value (720).

[0188] In one embodiment, the resolution of the first learning spectrogram (610) or the second learning spectrogram (620) may include at least one of time resolution and frequency resolution. A spectrogram has time resolution and frequency resolution, and the two are in a conflicting relationship. Meanwhile, the time resolution and frequency resolution may be determined based on the window size of the STFT performed when generating the spectrogram. Generally, the larger the window size, the lower the frequency resolution, the higher the time resolution, and the more BINs there may be. In a specific embodiment, the normalization variable may be proportional to the number of BINs.

[0189] The following mathematical expression 8 is an example of a loss factor including the size of the difference (730) between the first conversion value (710) and the second conversion value (720) and a predetermined normalization variable.

[0190]

[0191] In the above mathematical expression 8, is the fourth loss function, can represent a normalization variable. In mathematical expression 8, the size of the difference (730) between the first conversion value (710) and the second conversion value (720) is calculated as the L1 norm. The watermark encoder (100) can be trained based on the final loss function including the fourth loss function. That is, the watermark encoder (100) can be trained by adjusting the parameters in the direction of reducing the fourth loss function. Meanwhile, depending on the parameters of STFT such as the window size, can have different values, In addition, the loss value of the fourth loss function can have different values, and can vary depending on the parameters of the STFT that generates the unit learning spectrogram.

[0192] The fourth loss function can relatively significantly reflect the influence of spectral valleys appearing in the learning spectrogram. In other words, the fourth loss function can be understood as useful for fitting the watermark encoder (100) to a frequency band with low energy in the learning spectrogram.

[0193] Meanwhile, the watermark encoder (100) can be trained based on multiple unit learning spectrograms and multiple unit conversion values. Mathematical expression 9 below is an example of a loss factor including multiple first unit differences and multiple second unit differences.

[0194]

[0195] In the above mathematical formula 9, is an auxiliary loss function, can represent the number of preset STFTs. The unit learning spectrogram can be uniquely generated for each parameter of the STFT, including the window size. can represent the total number of generated unit learning spectrograms.

[0196] In one embodiment, when the STFT is performed with only one window size, is 1, and the first learning spectrogram (610) and the second learning spectrogram (620) may each include only one unit learning spectrogram. In another embodiment, the STFT may be performed for multiple window sizes, and the first learning spectrogram (610) and the second learning spectrogram (620) may each include multiple unit learning spectrograms. In this case, may be the number of unit learning spectrograms included in the first learning spectrogram (610) or the second learning spectrogram (620).

[0197] Meanwhile, the watermark encoder (100) can be trained based on the final loss function including the auxiliary loss function. That is, the watermark encoder (100) can be trained by adjusting the parameters in a direction that reduces the auxiliary loss function.

[0198] According to one embodiment, the auxiliary loss function includes the sum of the third and fourth loss functions described above. The third loss function is relatively advantageous for global optimization, and the fourth loss function is relatively advantageous for local optimization. In other words, the third and fourth loss functions are complementary. In one embodiment, the watermark encoder (100) can satisfy both global optimization and local optimization by being trained using the auxiliary loss function.

[0199] In addition, the learning of the watermark encoder (100) according to one embodiment can be performed using a plurality of unit learning spectrograms generated for each of a plurality of STFTs according to various parameters, so that overfitting of the watermark encoder (100) can be prevented.

[0200] Figure 11 is an exemplary diagram for explaining a discriminator used in learning a watermark encoder.

[0201] Referring to FIG. 11, the watermark encoder (100) can be trained as a generator of a Generative Adversarial Network (GAN). At this time, the discriminator (800) of the GAN can include a first discriminator (810) that determines whether a watermark has been inserted based on a spectrogram, and a second discriminator (820) that determines whether a watermark has been inserted based on an audio waveform.

[0202] A generative adversarial network (GAN) is a network that trains a generator to generate specific data. It combines the generator with a discriminator, trained to determine whether specific data was generated by the generator. During GAN training, the generator can be trained to generate data that is difficult to discriminate, while the discriminator can be trained to discriminate the generated data.

[0203] In one embodiment, from a GAN including a generator watermark encoder (100) and a discriminator (800), the watermark encoder (100) is trained to output marked audio data (111) in which it is difficult to determine whether a watermark has been inserted. To this end, the watermark encoder (100) is trained to output marked training audio data (141) in which it is difficult to determine whether a watermark has been inserted, and the discriminator (800) can be trained to correctly determine whether a watermark has been inserted.

[0204] According to one embodiment, the first discrimination unit (810) and the second discrimination unit (820) can output discrimination results in parallel for the marked training audio data (141). At this time, the first discrimination result (811) of the first discrimination unit (810) and the second discrimination result (821) of the second discrimination unit (820) can indicate a probability value that is greater than or equal to 0 and less than or equal to 1. Meanwhile, unlike that illustrated in FIG. 11, the first discrimination unit (810) and the second discrimination unit (820) according to one embodiment can also output discrimination results in parallel for the target training audio data (131) rather than the marked training audio data (141). At this time, the discrimination result for the target training audio data (131) may be more suitable for use in the training of the discriminator (800) than in the training of the watermark encoder (100).

[0205] Training of a watermark encoder (100) according to one embodiment may be performed by inputting target training audio data (131) and training watermark data (132) into the watermark encoder (100) and obtaining marked training audio data (141) output from the watermark encoder (100), inputting the marked training audio data (141) into a first determination unit (810) and a second determination unit (820) in parallel, obtaining a first determination result (811) and a second determination result (821) for the input marked training audio data (141), and adjusting at least one parameter included in an operation of the watermark encoder (100) based on the first determination result (811) and the second determination result (821).

[0206] In one embodiment, the watermark encoder (100) can be trained by adjusting at least one parameter included in the operation of the watermark encoder (100) based on a loss factor including a first determination result (811) and a second determination result (821).

[0207] Figure 12 is an exemplary drawing for explaining the first discrimination unit and the second discrimination unit.

[0208] Referring to Fig. 12, the first determination unit (810) may include a plurality of first sub-determination units, and the second determination unit (820) may include a plurality of second sub-determination units. The plurality of first sub-determination units and the plurality of second sub-determination units may output independent determination results for each sub-determination unit. At this time, the determination results output for each sub-determination unit are defined as intermediate determination results.

[0209] In one embodiment, the first judgment result (811) may include a set of a plurality of first intermediate judgment results output from a plurality of first sub-judgment units of the first judgment unit (810), and the second judgment result (821) may include a set of a plurality of second intermediate judgment results output from a plurality of second sub-judgment units of the second judgment unit (820).

[0210] In a specific embodiment, the first sub-discrimination unit 1 may output a first intermediate discrimination result 1 based on the marked learning audio data (141), the first sub-discrimination unit 2 may output a first intermediate discrimination result 2 based on the marked learning audio data (141), and the first sub-discrimination unit M may output a first intermediate discrimination result M. At this time, the first discrimination result (811) may include first intermediate discrimination results 1 to M. Here, the first intermediate discrimination results 1 to M may represent a plurality of probability values ​​output in parallel from a plurality of first sub-discrimination units.

[0211] In another specific embodiment, the second sub-discrimination unit 1 may output a second intermediate discrimination result 1 based on the marked training audio data (141), the second sub-discrimination unit 2 may output a second intermediate discrimination result 2 based on the marked training audio data (141), and the second sub-discrimination unit N may output a second intermediate discrimination result N. At this time, the second discrimination result (821) may include second intermediate discrimination results 1 to N. Here, the second intermediate discrimination results 1 to M may represent a plurality of probability values ​​output in parallel from a plurality of second sub-discrimination units.

[0212] Meanwhile, the specific process of outputting multiple first intermediate judgment results and multiple second intermediate judgment results will be described later with reference to FIG. 13.

[0213] Figure 13 is an exemplary diagram for explaining the process of generating the first judgment result and the second judgment result.

[0214] Referring to FIG. 13, each of the plurality of first sub-discrimination units can obtain inputs on which different preprocessing has been performed on marked training audio data (141) and output a first intermediate discrimination result. In addition, each of the plurality of second sub-discrimination units can obtain inputs on which different preprocessing has been performed on marked training audio data (141) and output a second intermediate discrimination result.

[0215] In one embodiment, the first determination unit (810) may include a plurality of first sub-determiners that determine the insertion of a watermark based on a spectrogram, and the second determination unit (820) may include a plurality of second sub-determiners that determine the insertion of a watermark based on an audio waveform. Here, each of the sub-determiners may be implemented as an artificial neural network model and may have a multi-perceptron structure having a plurality of layers.

[0216] In one embodiment, each sub-discriminant unit may include a feature extraction unit that extracts latent features for the input and an inference unit that generates a discrimination result from the latent features. In this case, the latent features extracted may vary depending on the input preprocessing process for each sub-discriminant unit.

[0217] In one embodiment, each of the plurality of first sub-discrimination units included in the first discrimination unit (810) can generate a discrimination spectrogram by performing STFT with a window size corresponding to each of the plurality of first sub-discrimination units on the marked learning audio data (141), and output a first intermediate discrimination result based on the discrimination spectrogram.

[0218] For example, the first sub-discrimination unit 1 can output a first intermediate discrimination result 1 by inputting a first discrimination spectrogram on which STFT is performed with a first window size on marked training audio data (141). The first sub-discrimination unit 2 can output a first intermediate discrimination result 2 by inputting a second discrimination spectrogram on which STFT is performed with a second window size on marked training audio data (141). The first sub-discrimination unit M can output a first intermediate discrimination result M by inputting an M-th discrimination spectrogram on which STFT is performed with an M-th window size on marked training audio data (141).

[0219] According to one embodiment, the first discriminator (810) may perform discrimination based on a Multi-Resolution Spectrogram Discriminator (MRSD). In one embodiment, each of the plurality of first sub-discriminators may be implemented as an MRSD. MRSD is an algorithm that determines watermark insertion by analyzing spectrograms for audio at various resolutions. According to one embodiment, the discriminator spectrograms at various resolutions may include multiple discriminator spectrograms according to various window sizes.

[0220] Meanwhile, the temporal and frequency resolutions of the discriminant spectrogram vary depending on the window size. At this time, temporal and frequency resolutions are in a conflicting relationship. Discriminant spectrograms corresponding to large window sizes offer high frequency resolution, making them suitable for extracting detailed features in the frequency domain of audio. Conversely, discriminant spectrograms corresponding to small window sizes offer high temporal resolution, making them suitable for extracting detailed features in the time domain of audio.

[0221] In addition, in one embodiment, each of the plurality of second sub-discrimination units included in the second discrimination unit (820) can generate two-dimensional audio data by reconstructing the marked learning audio data (141) using a period corresponding to each of the plurality of second sub-discrimination units, and can output the second intermediate discrimination result in parallel based on the two-dimensional audio data.

[0222] In the present disclosure, two-dimensional audio data means two-dimensional data in which audio data is divided into predetermined periods, components of the divided pieces are arranged along one axis, and components at the same position in the divided pieces are arranged along another axis.

[0223] For example, for 30 seconds of audio data, when a predetermined period is 10 seconds, the audio data can be divided into three segments each having a length of 10 seconds. At this time, multiple unit-length audios constituting one segment can be defined as one component. If the unit length is 1 second, each segmented segment has 10 components. The 10 components of each audio segment are arranged along the first axis, and the 3 components at corresponding time positions are arranged along the second axis. Through this, 2D audio data of a size of 10 * 3 on a component basis can be generated.

[0224] For example, the second sub-discrimination unit 1 can output the second intermediate discrimination result 1 by inputting the first two-dimensional audio data on which the first period of reconstruction is performed on the marked learning audio data (141). The second sub-discrimination unit 2 can output the second intermediate discrimination result 2 by inputting the second two-dimensional audio data on which the second period of reconstruction is performed on the marked learning audio data (141). The second sub-discrimination unit N can output the second intermediate discrimination result N by inputting the Nth two-dimensional audio data on which the Nth period of reconstruction is performed on the marked learning audio data (141).

[0225] According to one embodiment, the second discriminator (820) may perform the discriminator based on the MPWD (Multi-Period Waveform Discriminator). In one embodiment, each of the plurality of second sub-discriminators may be implemented as an MPWD. MPWD is an algorithm that determines the insertion of a watermark by analyzing the waveform of an audio into periodic components. Using MPWD, the insertion of a watermark can be determined by analyzing the characteristics appearing in the periodic repetition pattern of the audio using the two-dimensional audio data described above.

[0226] The watermark encoder (100) can be trained based on a loss factor including the result of the determination output from the discriminator (800). The following mathematical expression 10 is an example of a loss factor including the result of the determination of the discriminator (800).

[0227]

[0228] In the above mathematical expression 10, is the fifth loss function, is the output of the watermark encoder (100), i.e. the marked training audio data (141), is the sum of the number of sub-discrimination units included in the first discrimination unit (810) and the number of sub-discrimination units included in the second discrimination unit (820). can represent the intermediate judgment result output from the sub-judgment unit. Note that it can represent both the first discrimination result (811) and the second discrimination result (821). In one embodiment, when the number of first sub-discrimination units is M and the number of second sub-discrimination units is N, can be the sum of M and N.

[0229] Meanwhile, in one embodiment, the watermark encoder (100) may be trained based on a final loss function including a fifth loss function. That is, the watermark encoder (100) may be trained by adjusting parameters in a direction that reduces the fifth loss function.

[0230] In one embodiment, the discriminator (800) can be trained by adjusting parameters for at least one of the first discriminator (810) and the second discriminator (820) based on the first discrimination result (811) and the second discrimination result (821). In this case, the discriminator (800) can be trained not only by using the results of the discrimination for the marked training audio data (141), but also by using the results of the discrimination for the target training audio data (131).

[0231] The fifth loss function exemplified by the above mathematical expression 10 may be suitable for learning the watermark encoder (100), which is the generator. The following mathematical expression 11 is an example of a loss function for learning the discriminator (800).

[0232]

[0233] In the above mathematical expression 11, is the discriminant loss function, may represent target learning audio data (131). In one embodiment, the discriminator (800) may be trained in a direction that reduces the discriminator loss function. That is, the discriminator (800) may be trained to determine that watermark insertion has not been performed on target learning audio data (131), and to determine that watermark insertion has been performed on marked learning audio data (141).

[0234] Fig. 14 is an exemplary diagram for explaining loss factors used in learning a watermark encoder.

[0235] Referring to FIG. 14, in one embodiment, the loss factor used for learning the watermark encoder (100) may include at least one of the difference in audio data (510), the difference in watermark data (520), the first difference (630), the second difference (730), the first determination result (811), and the second determination result (821).

[0236] The specific process of deriving the difference in audio data (510), the difference in watermark data (520), the first difference (630), the second difference (730), the first determination result (811), and the second determination result (821) is the same as that described above with reference to FIGS. 1 to 13. Therefore, the specific derivation process of each loss element will be omitted.

[0237] In one embodiment, at least one loss function may be derived from a loss element including at least one of a difference in audio data (510), a difference in watermark data (520), a first difference (630), a second difference (730), a first determination result (811), and a second determination result (821). For example, the derived loss function may include the first to fifth loss functions described above and an auxiliary loss function. In this case, the auxiliary loss function may be calculated as the sum of the third loss function and the fourth loss function, but is not limited thereto.

[0238] In one embodiment, the watermark encoder (100) may be trained using a final loss function that includes at least one calculated loss function. For example, the watermark encoder (100) may be trained using a final loss function that includes a first loss function, a second loss function, an auxiliary loss function, and a fifth loss function. The following mathematical expression 12 is an example of a final loss function that includes a plurality of calculated loss functions.

[0239]

[0240] In the above mathematical expression 12, is the final loss function used for learning the watermark encoder (100). is a first loss function that uses the difference (510) of audio data as a loss factor, is a second loss function that uses the difference (520) of watermark data as a loss factor, is an auxiliary loss function that uses the first difference (630) and the second difference (730) as loss elements, can represent a fifth loss function that uses the first discrimination result (811) and the second discrimination result (821) as loss elements. Meanwhile, can represent the weights for each loss function that constitutes the final loss function, and can be preset by the learning subject or determined through separate optimization learning.

[0241] A watermark encoder (100) according to one embodiment is trained using various loss functions set based on various loss factors, thereby obtaining the benefits obtained from training using each loss factor.

[0242] Fig. 15 is a block diagram of an audio generation device according to one embodiment. The device (900) illustrated in Fig. 15 may correspond to the audio generation device (10) illustrated in Fig. 1, etc.

[0243] Referring to FIG. 15, the device (900) may include a processor (910) and a memory (920). Only components related to the embodiment are illustrated in the device (900) of FIG. 15. Therefore, those skilled in the art will understand that other general components may be included in addition to the components illustrated in FIG. 15.

[0244] As an example, the device (900) may further include a communication module (not shown). The communication module may include at least one component that enables the device (900) to perform wired / wireless communication with other external devices. For example, the communication module may include a short-range communication unit and / or a mobile communication unit.

[0245] The processor (910) controls the overall operation of the device (900). For example, the processor (910) can control the input unit (not shown), the display (not shown), the communication module (not shown), the memory (920), etc., by executing programs stored in the memory (920). The processor (910) can control the operation of the device (900) by executing programs stored in the memory (920).

[0246] The processor (910) can control at least some of the operations of the device (900) described above in FIGS. 1 to 14.

[0247] For example, the processor (910) can input watermark data and target audio data, which is a target of insertion of the watermark data, into a watermark encoder, and obtain marked audio data in which the watermark data is inserted into the target audio data as an output of the watermark encoder.

[0248] Meanwhile, a specific example of how the processor (910) operates is the same as described above with reference to FIGS. 1 to 14. Therefore, a specific description of the operation of the processor (910) is omitted below.

[0249] The processor (910) may be implemented using at least one of application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.

[0250] The memory (920) is hardware that stores various data processed within the device (900), and can store a program for processing and controlling the processor (910).

[0251] The memory (920) may include random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disk storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.

[0252] Embodiments according to the present disclosure may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. In this case, the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories.

[0253] Meanwhile, the computer program may be specifically designed and configured for the present disclosure, or may be known and available to those skilled in the computer software field. Examples of computer programs may include not only machine language code, such as that generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.

[0254] According to one embodiment, the method according to various embodiments of the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0255] Unless the steps constituting the method according to the present disclosure are explicitly described in a specific order or are otherwise described in a different order, the steps may be performed in any appropriate order. The present disclosure is not necessarily limited by the order in which the steps are described. The use of any examples or exemplary terms (e.g., “etc.”) in this disclosure is merely intended to illustrate the present disclosure in more detail, and the scope of the present disclosure is not limited by the examples or exemplary terms, unless otherwise defined by the claims. Furthermore, those skilled in the art will appreciate that various modifications, combinations, and variations can be configured according to design conditions and factors within the scope of the appended claims or their equivalents.

[0256] Therefore, the spirit of the present disclosure should not be limited to the embodiments described above, and all scopes equivalent to or equivalent to the scope of the patent claims described below, as well as the scope of the present disclosure, should be considered to fall within the scope of the spirit of the present disclosure.

Claims

1. A step of inputting watermark data and target audio data, which is the target of insertion of the watermark data, into a watermark encoder; and A step of obtaining marked audio data in which the watermark data is inserted into the target audio data as an output of the watermark encoder; Including, but not limited to, The above watermark encoder is trained as a generator of a Generative Adversarial Network (GAN). A method for generating watermark audio using an encoder trained on the basis of multiple discriminators, wherein the discriminator of the above GAN includes a first discriminator that determines the insertion of a watermark based on a spectrogram and a second discriminator that determines the insertion of a watermark based on an audio waveform.

2. In paragraph 1, The first determination unit and the second determination unit input marked learning audio data generated from the watermark encoder and output the determination result in parallel, A method in which the watermark encoder is trained by adjusting at least one parameter included in the operation of the watermark encoder based on a loss factor including a first determination result of the first determination unit and a second determination result of the second determination unit.

3. In paragraph 2, A method in which the marked learning audio data is output from the watermark encoder based on input of target learning audio data and learning watermark data.

4. In paragraph 2, The above first determination unit includes a plurality of first sub-determination units, Each of the plurality of first sub-discrimination units generates a discrimination spectrogram by performing STFT (Short Time Fourier Transform) on the marked learning audio data with a window size corresponding to each of the plurality of first sub-discrimination units, and outputs a first intermediate discrimination result based on the discrimination spectrogram. A method wherein the first judgment result includes a plurality of first intermediate judgment results output in parallel from each of the plurality of first sub-judgment units.

5. In paragraph 2, The second determination unit includes a plurality of second sub-determination units, Each of the plurality of second sub-determining units generates two-dimensional audio data by reconstructing the marked learning audio data using a period corresponding to each of the plurality of second sub-determining units, and outputs a second intermediate determination result based on the two-dimensional audio data. A method wherein the second judgment result includes a plurality of second intermediate judgment results output in parallel from each of the plurality of second sub-judgment units.

6. In paragraph 2, A method wherein the discriminator is learned by adjusting at least one parameter of the discriminator based on the first discrimination result and the second discrimination result.

7. Memory in which at least one program is stored; and A processor that operates by executing at least one program; Including, but not limited to, The above processor, Input the watermark data and the target audio data into which the watermark data is to be inserted into the watermark encoder, As an output of the watermark encoder, marked audio data in which the watermark data is inserted into the target audio data is obtained, The above watermark encoder is trained as a generator of a Generative Adversarial Network (GAN). A device for generating watermark audio using an encoder learned based on multiple discriminators, wherein the discriminator of the above GAN includes a first discriminator that determines the insertion of a watermark based on a spectrogram and a second discriminator that determines the insertion of a watermark based on an audio waveform.

8. A non-transitory computer-readable recording medium having recorded thereon a program for executing the method of paragraph 1 on a computer.

Citation Information

Patent Citations

  • Decoding device, audio delivering system, decoding method, and program

    JP2022177411A

  • Apparatus and method for embedding audio watermark, and apparatus and method for detecting audio watermark

    KR1020110014871A

  • Integrated production system enabling the co-culture of shellfish larvae and dinoflagellates controlling parasite ciliates

    KR1020230031566A

  • Digital content management system, verification device, programs therefor, and data processing method

    WO2011121928A1

  • KR20200027475A