Audio watermark embedding method and device, computer device and storage medium

By combining discrete wavelet transform and the U-Net network model, the shortcomings of audio watermarking technology in terms of flexibility and robustness are solved, and efficient and covert audio watermarking embedding is achieved, which can adapt to various audio processing and attacks.

CN119400186BActive Publication Date: 2025-11-04PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411531242.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-11-04
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

Existing audio watermarking technologies are insufficient in terms of flexibility, adaptability, and robustness, making it difficult to cope with the challenges of complex network environments and audio processing technologies.

Method used

Using discrete wavelet transform and U-Net network model, the audio is decomposed into sub-bands of multiple frequency bands. Latent space features are extracted by encoder, watermark information is embedded, and audio is reconstructed by decoder to generate target audio with embedded watermark.

Benefits of technology

It improves the flexibility, adaptability and robustness of audio watermark embedding, effectively protecting audio copyright in complex environments and ensuring the concealment and reliability of the watermark.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119400186B_ABST
    Figure CN119400186B_ABST
Patent Text Reader

Abstract

The application relates to the fields of artificial intelligence and financial technology, and discloses an audio watermark embedding method and device, computer equipment and a storage medium, which are characterized in that the method decomposes host audio to be embedded with a watermark into a plurality of subbands of different frequency bands through discrete wavelet transform; an encoder in a U-Net network model is used to perform layer-by-layer down sampling on each of the subbands, and hidden space features of the host audio are extracted; watermark information to be embedded is determined, and the watermark information is encoded to obtain a watermark vector; the watermark vector is embedded in the hidden space features to obtain fusion features; a decoder in the U-Net network model is used to perform layer-by-layer up sampling on the fusion features, and audio information is reconstructed; inverse discrete wavelet transform is performed on the audio information to generate target audio embedded with the watermark information; and therefore, the application can effectively improve the flexibility, adaptability and robustness of audio watermark embedding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and financial technology, specifically to an audio watermark embedding method, apparatus, computer device, and computer-readable storage medium. Background Technology

[0002] Currently, with the development of computer technology, more and more technologies are being applied in the financial field, and the traditional financial industry is gradually transforming into Fintech. Audio watermarking embedding technology is no exception. For example, financial institutions (such as banks) often need to record customer service calls for quality control and compliance checks. Audio watermarking embedding technology can help ensure that the recording files have not been tampered with and provide a means of verifying their authenticity. However, due to the security and real-time requirements of the financial industry, higher demands are also placed on audio watermarking embedding technology.

[0003] Meanwhile, with the rapid development of internet technology and the continuous advancement of artificial intelligence (AI) technology, the production, editing, sharing, and dissemination of digital audio have become more convenient than ever before. Audio files circulate rapidly through the internet, greatly enriching people's daily lives. However, this convenience has also brought challenges to audio copyright protection. Traditional audio copyright protection methods, such as those relying on legal means and technical measures, are no longer sufficient to address copyright infringement issues in the current online environment. Against this backdrop, digital watermarking technology has emerged, providing an effective means of copyright protection. Digital watermarking embeds hidden information into audio files, achieving functions such as copyright identification, content authentication, and tracking of illegal distribution. Although digital watermarking technology shows great potential in copyright protection, most media's digital watermarking implementation still relies on traditional methods, namely watermark embedding algorithms designed based on expert experience. Traditional digital watermarking technologies typically employ rule-based methods, which need to consider the characteristics of audio information and the perceptual characteristics of the human ear during design. While these methods have achieved some success in specific scenarios, they generally lack sufficient flexibility and adaptability, making it difficult to cope with the increasingly complex online environment and continuously advancing audio processing technologies. Furthermore, traditional methods often exhibit limited robustness when faced with complex transformations and attacks on audio information.

[0004] In summary, how to provide an audio watermark embedding method, apparatus, computer device, and computer-readable storage medium that can effectively improve the flexibility, adaptability, and robustness of audio watermark embedding is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide an audio watermark embedding method, apparatus, computer device and computer-readable storage medium, aiming to solve the problem of how to effectively improve the flexibility, adaptability and robustness of audio watermark embedding.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] In a first aspect, the present invention provides an audio watermark embedding method, comprising:

[0008] The host audio to be watermarked is decomposed into multiple sub-bands of different frequency bands using discrete wavelet transform;

[0009] The encoder in the U-Net network model is used to downsample each sub-band layer by layer to extract the latent space features of the host audio;

[0010] The watermark information to be embedded is determined, and the watermark information is encoded to obtain a watermark vector;

[0011] The watermark vector is embedded into the latent space feature to obtain the fused feature;

[0012] The fused features are upsampled layer by layer using the decoder in the U-Net network model to reconstruct the audio information;

[0013] The audio information is subjected to inverse discrete wavelet transform to generate target audio with the watermark information embedded.

[0014] In a second aspect, the present invention provides an audio watermark embedding device, comprising:

[0015] The decomposition module is used to decompose the host audio to be embedded with watermark into multiple sub-bands of different frequency bands through discrete wavelet transform.

[0016] The extraction module is used to downsample each sub-band layer by layer using the encoder in the U-Net network model to extract the latent space features of the host audio.

[0017] An encoding module is used to determine the watermark information to be embedded and to encode the watermark information to obtain a watermark vector;

[0018] An embedding module is used to embed the watermark vector into the latent space features to obtain fused features;

[0019] The reconstruction module is used to upsample the fused features layer by layer using the decoder in the U-Net network model to reconstruct the audio information;

[0020] The generation module is used to perform inverse discrete wavelet transform on the audio information to generate target audio with the watermark information embedded.

[0021] Thirdly, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the audio watermark embedding method as described above.

[0022] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the audio watermark embedding method as described above.

[0023] Compared to existing technologies, this invention provides an audio watermark embedding method, apparatus, computer device, and computer-readable storage medium. The method involves: decomposing the host audio to be watermarked into multiple sub-bands of different frequency bands using discrete wavelet transform; using an encoder in a U-Net network model to downsample each sub-band layer by layer to extract the latent space features of the host audio; determining the watermark information to be embedded and encoding the watermark information to obtain a watermark vector; embedding the watermark vector into the latent space features to obtain a fused feature; using a decoder in the U-Net network model to upsample the fused feature layer by layer to reconstruct the audio information; and performing an inverse discrete wavelet transform on the audio information to generate target audio embedded with the watermark information. Therefore, this invention can effectively improve the flexibility, adaptability, and robustness of audio watermark embedding. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram illustrating the application environment of an audio watermark embedding method according to an embodiment of the present invention.

[0026] Figure 2 This is a flowchart illustrating an audio watermark embedding method according to an embodiment of the present invention.

[0027] Figure 3 This is a schematic diagram of the program module of an audio watermark embedding device according to an embodiment of the present invention.

[0028] Figure 4This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention.

[0029] Figure 5 This is another structural schematic diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0032] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0033] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."

[0034] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0035] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0036] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0037] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0038] An embodiment of the present invention provides an audio watermark embedding method, which can be applied to, for example... Figure 1 In the application environment shown, the client and server communicate via a network. The client includes, but is not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0039] Please see Figure 2 An embodiment of the present invention provides an audio watermark embedding method, wherein the method includes the following steps:

[0040] S100. The host audio to be embedded with watermark is decomposed into multiple sub-bands of different frequency bands by discrete wavelet transform.

[0041] S200. Using the encoder in the U-Net network model, each sub-band is downsampled layer by layer to extract the latent space features of the host audio.

[0042] S300. Determine the watermark information to be embedded and encode the watermark information to obtain a watermark vector;

[0043] S400. Embed the watermark vector into the latent space feature to obtain the fused feature;

[0044] S500: The fused features are upsampled layer by layer using the decoder in the U-Net network model to reconstruct the audio information;

[0045] S600. Perform inverse discrete wavelet transform on the audio information to generate target audio with the watermark information embedded.

[0046] In specific implementation, this embodiment achieves highly flexible, adaptable, and robust audio watermark embedding by combining Discrete Wavelet Transform (DWT) and the U-Net network model. First, the host audio is decomposed into sub-bands of different frequency bands using DWT. This step allows the watermark information to select the most suitable embedding position based on the specific spectral characteristics of the host audio, thereby improving the flexibility of the embedding method. Second, the encoder of the U-Net network model downsamples the sub-bands and extracts latent space features. The deep learning mechanism of the U-Net network model enables the algorithm to adaptively learn audio features, enhancing the adaptability of the embedding method. Then, the encoded watermark information is embedded into the latent space features. This process can be optimized during network training to ensure the concealment of the watermark and automatic adjustment of the embedding strength, further improving the adaptability of the method. Next, the decoder of the U-Net network model is used to upsample and reconstruct the audio information. The network's self-learning characteristics help maintain sound quality during reconstruction, while the embedded watermark information is also effectively preserved. Finally, the execution of inverse DWT restores the signal to the time domain, generating the target audio containing the watermark information. Throughout this implementation, the introduction of deep learning models, especially the use of the U-Net network model, enabled the embedding method to not only cope with various audio characteristics and complex attacks, but also to adaptively adjust the embedding strategy to maintain the concealment and robustness of the watermark, thereby achieving efficient and reliable audio watermark embedding in different scenarios and conditions.

[0047] Furthermore, in one embodiment, the audio watermark embedding method, before step S100, decomposing the host audio to be watermarked into multiple sub-bands of different frequency bands through discrete wavelet transform, specifically includes the following steps:

[0048] Obtain the original audio uploaded by the target user;

[0049] The original audio is preprocessed to obtain the host audio to which the watermark is to be embedded.

[0050] In specific implementation, this embodiment obtains the host audio to be embedded with watermark by acquiring the original audio uploaded by the target user and performing necessary preprocessing, ensuring that the quality of the host audio meets the requirements for watermark embedding. Then, discrete wavelet transform technology is used to decompose the host audio into multiple sub-bands of different frequency bands. This not only enables multi-resolution analysis of the host audio, but also allows for the identification and focus on specific frequency bands for watermark embedding, thereby improving the targeting and effectiveness of watermark embedding. The technical effect of this process is that it lays the foundation for subsequent watermark embedding steps, enabling the watermark information to be embedded into the host audio more covertly and stably, while enhancing the robustness of the watermark against various possible signal processing and attacks.

[0051] The specific implementation process of the steps in this embodiment is roughly as follows:

[0052] 1. Obtain the original audio: First, receive and confirm the audio file (i.e., the original audio) uploaded by the target user, and ensure that its format and quality meet the standards for subsequent processing;

[0053] 2. Audio preprocessing: Preprocess the uploaded raw audio, including noise reduction, format conversion, and sampling rate adjustment, to eliminate impurities in the recording and standardize the audio format to obtain clear host audio;

[0054] 3. Discrete wavelet transform: The discrete wavelet transform method is used to decompose the preprocessed host audio into multiple sub-bands of different frequency bands. Each sub-band contains detailed information about the host audio in a specific frequency band.

[0055] 4. Subband Analysis: Analyze the characteristics of each subband, including frequency range and energy distribution, to provide a basis for determining the optimal location for watermark embedding;

[0056] 5. Watermark Embedding Preparation: Based on the results of sub-band analysis, select suitable sub-bands for watermark embedding, while ensuring that the selected sub-bands can carry watermark information without affecting the perceptual quality of the host audio.

[0057] Through the above process, this embodiment can more accurately locate the frequency band in the host audio that is most suitable for embedding the watermark, thereby achieving efficient and covert watermark embedding and enhancing the covertness and robustness of the watermark.

[0058] Furthermore, in one embodiment, the audio watermark embedding method, wherein step S200, using the encoder in the U-Net network model to downsample each sub-band layer by layer to extract the latent space features of the host audio, specifically includes the following steps:

[0059] The U-Net network model is pre-built;

[0060] Each sub-band is used as input and sequentially input into the encoder for layer-by-layer downsampling to obtain the depth feature representation corresponding to each sub-band;

[0061] All the obtained deep feature representations are combined to generate the latent space features of the host audio.

[0062] In specific implementation, this embodiment achieves efficient feature extraction of the host audio by pre-constructing a U-Net network model containing an encoder and a decoder. Specifically, by sequentially inputting multiple sub-bands of the host audio into the encoder for layer-by-layer downsampling, the U-Net network model can automatically learn and extract the latent space features of the host audio. These features capture the core representative information of the host audio and have a small dimensionality, providing an optimized embedding space for subsequent watermark embedding. This processing not only improves the adaptability and flexibility of the watermark algorithm but also enhances the concealment and robustness of the watermark, ensuring that the watermark information can be stably embedded into the host audio and can be effectively extracted even if the audio is attacked or processed.

[0063] The specific implementation process of the steps in this embodiment is roughly as follows:

[0064] 1. Construct a U-Net network model: Design and construct a U-Net network model, which consists of two parts: an encoder and a decoder; the encoder is used to progressively reduce the spatial dimension of the input data, while the decoder is used to progressively restore the spatial dimension of the data;

[0065] 2. Initialize the encoder: Configure the encoder, which typically includes multiple convolutional and pooling layers, as well as activation functions such as ReLU; these components are responsible for extracting features from the input subbands and downsampling them;

[0066] 3. Input sub-bands: Each sub-band obtained by discrete wavelet transform is sequentially fed into the encoder of the U-Net network model as input; these sub-bands contain detailed information about the host audio in different frequency bands;

[0067] 4. Layer-by-layer downsampling: The subband signal is transmitted layer by layer in the encoder. Each convolution layer may be followed by a pooling operation to gradually reduce the spatial resolution of the signal while increasing the feature depth, thereby extracting latent space features.

[0068] 5. Extract latent space features: In the last layer of the encoder, the deep feature representations corresponding to the sub-bands are obtained (i.e., the latent space features of the sub-bands). These features capture important information of the original audio information and are prepared for embedding watermark information. Then, the deep feature representations corresponding to all the sub-bands are combined to generate the latent space features of the host audio.

[0069] 6. Skip Connections (Optional): U-Net network models typically include skip connections, which combine downsampled features from the encoder with upsampled features from the decoder to help recover contextual information during the reconstruction process.

[0070] Through the above process, the U-Net network model in this embodiment can extract useful latent space features for each sub-band, thereby forming the latent space features of the host audio, which can then be used in the watermark embedding process, making the embedding of watermark information more effective and robust.

[0071] Furthermore, in one embodiment, the audio watermark embedding method, wherein step S300, determining the watermark information to be embedded and encoding the watermark information to obtain a watermark vector, specifically includes the following steps:

[0072] Select the watermark information to be embedded based on the watermark embedding requirements of the host audio;

[0073] The watermark information is encoded to obtain an initial vector;

[0074] Adjust the length and / or intensity of the initial vector to obtain the watermark vector.

[0075] In practical implementation, the technical effect achievable by this embodiment is that it can customize and prepare suitable watermark information according to the characteristics of the host audio and the specific requirements of watermark embedding. First, the watermark information to be embedded is selected according to the watermark embedding requirements, ensuring that the watermark content matches the intended use. Second, the watermark information is encoded to generate an initial vector, so that the watermark information is converted into a processable format. Then, by adjusting the length and strength of the initial vector, the watermark vector is optimized to adapt to the audio characteristics and perception threshold, so as to ensure that the embedded watermark is not only sufficiently concealed and reduces the impact on the sound quality of the original audio, but also ensures the robustness of the watermark, ensuring that the watermark information can be stably extracted when facing various signal processing and potential attacks. This embodiment improves the flexibility and adaptability of watermark embedding, enabling it to be adjusted and optimized for different application scenarios and security requirements.

[0076] The specific implementation process of the steps in this embodiment is roughly as follows:

[0077] 1. Requirements Analysis: Based on the characteristics of the host audio and the purpose of watermark embedding (such as copyright protection, content authentication, etc.), analyze and determine the requirements of the watermark information to be embedded; this may include determining the type of watermark information (text, image, identifier, etc.) and the properties that need to be satisfied (such as robustness, concealment, etc.);

[0078] 2. Watermark Information Selection: Select or generate watermark information that meets the requirements; for example, if the watermark is used for a copyright notice, the copyright owner's identifier or a specific serial number may be selected.

[0079] 3. Watermark Information Encoding: Convert the selected watermark information into binary form; this can be achieved through various encoding techniques, such as ASCII encoding for text and binary encoding for images, to generate the initial watermark vector;

[0080] 4. Initial vector processing: Perform necessary processing on the initial vector, such as adding redundant bits to enhance robustness, or applying specific coding techniques (such as Reed-Solomon coding) to improve error correction capability.

[0081] 5. Vector Adjustment: Adjust the length and strength of the watermark vector based on the characteristics of the host audio and the perceptual model. This may include expanding or compressing the vector length and adjusting the embedding strength to ensure the concealment of the watermark, while ensuring that the watermark can still be reliably extracted under various possible audio processing and attacks.

[0082] 6. Generate watermark vector: After adjusting the vector, the final watermark vector is obtained; this vector will be used in the subsequent embedding process to embed it into the latent space features of the host audio.

[0083] This embodiment ensures that the watermark vector matches the features of the host audio and meets specific embedding requirements through the above process, thereby achieving effective watermark information embedding.

[0084] Furthermore, in one embodiment, the audio watermark embedding method, wherein step S400, embedding the watermark vector into the latent space feature to obtain the fused feature, specifically includes the following steps:

[0085] The watermark embedding strategy is determined based on the content of the watermark information and the characteristics of the host audio.

[0086] The watermark vector is embedded into the latent space feature according to the watermark embedding strategy to obtain the preliminary fused feature;

[0087] The preliminary features are normalized to obtain the fused features.

[0088] In specific implementation, the technical effects achievable in this embodiment are as follows: First, a watermark embedding strategy is determined based on the content of the watermark information and the characteristics of the host audio, ensuring that the embedding of the watermark information meets specific security requirements while adapting to the characteristics of the host audio. Then, the watermark vector is embedded into the latent space features of the host audio according to the watermark embedding strategy, generating preliminary fused features. This process tightly integrates the watermark information with the audio content, improving the concealment of the watermark and the consistency of the overall audio information. Finally, the preliminary features are normalized, and the resulting fused features not only ensure a balanced distribution of the watermark information in the host audio but also help improve the robustness of the watermark against various signal processing and potential attacks, making the watermark extraction process more reliable and accurate. This series of processing steps comprehensively improves the performance of the audio watermark embedding system, making it more effective and secure in practical applications.

[0089] The specific implementation process of the steps in this embodiment is roughly as follows:

[0090] 1. Determine the watermark embedding strategy: Based on the content of the watermark information, the requirements for the watermark's robustness, concealment, and capacity, as well as the characteristics of the host audio (such as frequency range), determine a suitable watermark embedding strategy; this may include selecting an appropriate embedding location, embedding depth, and embedding method.

[0091] 2. Embedding watermark information: Based on the determined watermark embedding strategy, the watermark vector is embedded into the latent space features of the host audio.

[0092] 3. Generate preliminary fusion features: After embedding watermark information, preliminary fusion features are obtained; these features contain information from the original host audio and the embedded watermark information.

[0093] 4. Normalization Processing: The initial features are normalized to ensure that the distribution and intensity of the features are within a reasonable range. This helps to improve the stability and robustness of the watermark embedding, ensuring that the watermark can be reliably extracted under different audio information and different playback environments.

[0094] 5. Obtain the final fusion features: After normalization, the final fusion features are obtained; these features will be used in the audio information reconstruction process to generate target audio with embedded watermark information.

[0095] Through the above process, this embodiment can obtain the final fusion features, providing an input basis for the subsequent audio information reconstruction process.

[0096] Furthermore, in one embodiment, the audio watermark embedding method, wherein step S500, which involves upsampling the fused features layer by layer using the decoder in the U-Net network model to reconstruct the audio information, specifically further includes the following steps:

[0097] The reconstructed audio information is subjected to a simulated attack test, and the test results are generated.

[0098] Based on the test results, the watermark embedding strategy and / or the decoder are adjusted.

[0099] In practical implementation, this embodiment uses a decoder to upsample the fused features layer by layer, effectively reconstructing the audio information with embedded watermarks, ensuring the quality of the audio information and the integrity of the watermark information. Furthermore, this embodiment uses simulated attack tests to evaluate the robustness of the watermark, and adjusts and optimizes the watermark embedding strategy and decoder parameters based on the test results, thereby improving the reliability and practicality of subsequent watermark embedding, ensuring that the watermark can be effectively extracted in various practical application environments, and meeting the needs of application scenarios such as digital media copyright protection and authentication.

[0100] The specific implementation process of the steps in this embodiment is roughly as follows:

[0101] 1. Input fused features to the decoder: The fused features containing watermark information are provided as input to the decoder part of the U-Net network model; this step is the starting point of the audio information reconstruction process, and the task of the decoder is to gradually restore the original attributes of the audio information.

[0102] 2. Layer-by-layer upsampling: The fused features undergo layer-by-layer upsampling in the decoder. Each layer may include convolution operations, upsampling operations, and activation functions. This process gradually increases the spatial resolution of the features, gradually restoring them to a form close to the original audio information.

[0103] 3. Reconstructing audio information: After final processing by the decoder, reconstructed audio information is generated; theoretically, this audio information should be indistinguishable from the original audio information in terms of hearing, and at the same time contains embedded watermark information.

[0104] 4. Conduct simulated attack tests: Perform simulated attack tests on the reconstructed audio information. This may include operations such as adding noise, compression, resampling, and editing to simulate various processing and damage that audio information may encounter in real-world applications.

[0105] 5. Generate and analyze test results: After simulating an attack, test the audio information to generate results and evaluate the robustness of the watermark; analyze the test results to determine whether the watermark information can be accurately extracted and whether the sound quality meets the requirements.

[0106] 6. Adjust the watermark embedding strategy and decoder: Based on the test results, optimize the watermark embedding strategy, such as modifying the strength of the watermark vector and adjusting the embedding position, to improve the concealment and robustness of the watermark; at the same time, the structure and parameters of the decoder can also be adjusted to better reconstruct audio information and improve the accuracy of watermark extraction.

[0107] 7. Iterative optimization (optional): If necessary, repeat the above process for multiple iterations, continuously adjusting the watermark embedding strategy and decoder settings until the watermark performance reaches a satisfactory level.

[0108] This embodiment can reconstruct audio information through the above process, and at the same time, ensure that the reconstructed audio information can adapt to potential attacks, thereby improving its reliability in practical applications.

[0109] Furthermore, in one embodiment, the audio watermark embedding method, wherein step S600, performing inverse discrete wavelet transform on the audio information to generate target audio embedded with the watermark information, specifically includes the following steps:

[0110] Perform inverse discrete wavelet transform on the audio information to generate an initial audio file embedded with the watermark information;

[0111] The initial audio is denoised and equalized to obtain the target audio, and it is then identified whether the watermark information is embedded in the target audio.

[0112] When the watermark information is embedded in the target audio, the target audio is output.

[0113] In specific implementation, the technical effect achievable in this embodiment is to restore the audio information incorporating watermark information to the time domain through inverse discrete wavelet transform, generating an initial audio containing watermark information; then, the initial audio is optimized through denoising and equalization processing to improve sound quality and ensure the concealment of the watermark, ultimately obtaining the target audio; finally, it identifies whether the watermark information is embedded in the target audio. When the watermark information is embedded in the target audio, the target audio is output to a designated position, completing the entire watermark embedding and audio reconstruction process, ensuring the effective embedding of watermark information and the integrity of the target audio, and meeting the needs of copyright protection and content authentication.

[0114] The specific implementation process of the steps in this embodiment is roughly as follows:

[0115] 1. Inverse Discrete Wavelet Transform: The upsampled fused features, i.e. the frequency domain representation containing watermark information, are transformed back to the time domain using the Inverse Discrete Wavelet Transform (IDWT). This step is to obtain a preliminary version of the audio information embedded with watermark information, i.e., the initial audio.

[0116] 2. Perform noise reduction and equalization processing: Perform a series of noise reduction and equalization processing on the initial audio in order to improve the listening quality of the reconstructed audio, ensure the concealment of the watermark, and meet specific technical standards or user needs.

[0117] 3. Quality Assessment: During or after noise reduction and equalization processing, objective and subjective quality assessments are performed. Objective assessments may use audio quality assessment algorithms such as PESQ (Perceptual Evaluation of Speech Quality), while subjective assessments may involve hearing tests. This step is to ensure that the audio information after embedding the watermark maintains high sound quality.

[0118] 4. Generate target audio: Based on the quality assessment results, fine-tune the parameters of noise reduction and equalization processing to obtain the final target audio; the target audio should effectively embed the watermark information without significantly reducing the listening experience.

[0119] 5. Output target audio: Identify whether the watermark information is embedded in the target audio. When the watermark information is embedded in the target audio, output the target audio to a specified location, such as saving it to the local file system, sending it to a remote server, or streaming it to other devices. This step marks the completion of the watermark embedding process, and the target audio is ready for distribution or further processing.

[0120] Through the above process, this embodiment ensures that the target audio with embedded watermark information not only meets the technical requirements but also closely approximates the original audio in terms of sound, providing users with a high-quality experience. At the same time, the watermark information is properly embedded, providing support for the protection of audio content.

[0121] As can be seen from the above method embodiments, the audio watermark embedding method provided by the present invention includes: decomposing the host audio to be watermarked into multiple sub-bands of different frequency bands through discrete wavelet transform; using the encoder in the U-Net network model to downsample each sub-band layer by layer to extract the latent space features of the host audio; determining the watermark information to be embedded and encoding the watermark information to obtain a watermark vector; embedding the watermark vector into the latent space features to obtain a fusion feature; using the decoder in the U-Net network model to upsample the fusion feature layer by layer to reconstruct the audio information; and performing an inverse discrete wavelet transform on the audio information to generate target audio embedded with the watermark information. Thus, the method of the present invention can effectively improve the flexibility, adaptability, and robustness of audio watermark embedding.

[0122] Understandably, the audio watermark embedding method provided in this embodiment of the invention can be applied to audio watermark embedding scenarios in the fintech field. Specifically, the audio watermark embedding function can be integrated into the corresponding quality control and compliance inspection scenarios of financial institutions (such as banks). Since financial institutions often need to record customer service calls for quality control and compliance inspection, the audio watermark embedding method provided in this embodiment of the invention can help financial institutions ensure that the relevant recording files have not been tampered with and provide a means to verify their authenticity.

[0123] It should be understood that although this application provides the method operation steps as described in the embodiments or flowcharts, conventional or non-inventive labor may include more or fewer operation steps, and these operation steps are not necessarily executed sequentially according to the order of the embodiments or flowcharts. The order of steps listed in the embodiments or flowcharts is merely one way of executing many steps and does not represent the only execution order. It should be noted that there is no necessary sequential order between the above steps. Those skilled in the art can understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in exchange, etc. Moreover, at least some steps in the embodiments or flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn, alternately, or synchronously with other steps or at least a part of the sub-steps or stages of other steps.

[0124] Based on the above method embodiments, please refer to Figure 3 Another embodiment of the present invention also provides an audio watermark embedding device, wherein the device includes:

[0125] Decomposition module 11 is used to decompose the host audio to be embedded with watermark into multiple sub-bands of different frequency bands through discrete wavelet transform.

[0126] Extraction module 12 is used to perform layer-by-layer downsampling on each of the sub-bands using the encoder in the U-Net network model to extract the latent space features of the host audio;

[0127] Encoding module 13 is used to determine the watermark information to be embedded and to encode the watermark information to obtain a watermark vector;

[0128] Embedding module 14 is used to embed the watermark vector into the latent space feature to obtain the fused feature;

[0129] Reconstruction module 15 is used to upsample the fused features layer by layer using the decoder in the U-Net network model to reconstruct audio information;

[0130] The generation module 16 is used to perform inverse discrete wavelet transform on the audio information to generate target audio with the watermark information embedded.

[0131] Furthermore, in one embodiment, the audio watermark embedding device further includes:

[0132] The acquisition module is used to acquire the original audio uploaded by the target user;

[0133] The preprocessing module is used to preprocess the original audio to obtain the host audio to which the watermark is to be embedded.

[0134] Furthermore, in one embodiment, the audio watermark embedding device, wherein the extraction module 12 is specifically used for:

[0135] The U-Net network model is pre-built;

[0136] Each sub-band is used as input and sequentially input into the encoder for layer-by-layer downsampling to obtain the depth feature representation corresponding to each sub-band;

[0137] All the obtained deep feature representations are combined to generate the latent space features of the host audio.

[0138] Furthermore, in one embodiment, the audio watermark embedding device, wherein the encoding module 13 is specifically used for:

[0139] Select the watermark information to be embedded based on the watermark embedding requirements of the host audio;

[0140] The watermark information is encoded to obtain an initial vector;

[0141] Adjust the length and / or intensity of the initial vector to obtain the watermark vector.

[0142] Furthermore, in one embodiment, the audio watermark embedding device, wherein the embedding module 14 is specifically used for:

[0143] The watermark embedding strategy is determined based on the content of the watermark information and the characteristics of the host audio.

[0144] The watermark vector is embedded into the latent space feature according to the watermark embedding strategy to obtain the preliminary fused feature;

[0145] The preliminary features are normalized to obtain the fused features.

[0146] Furthermore, in one embodiment, the audio watermark embedding device, wherein the reconstruction module 15 is specifically further used for:

[0147] The reconstructed audio information is subjected to a simulated attack test, and the test results are generated.

[0148] Based on the test results, the watermark embedding strategy and / or the decoder are adjusted.

[0149] Furthermore, in one embodiment, the audio watermark embedding device, wherein the generation module 16 is specifically used for:

[0150] Perform inverse discrete wavelet transform on the audio information to generate an initial audio file embedded with the watermark information;

[0151] The initial audio is denoised and equalized to obtain the target audio, and it is then identified whether the watermark information is embedded in the target audio.

[0152] When the watermark information is embedded in the target audio, the target audio is output.

[0153] It should be noted that, in the device embodiments of the present invention, the information interaction and execution process between the above modules are based on the same concept as in the method embodiments of the present invention. For details on their specific functions and the resulting technical effects, please refer to the aforementioned method embodiments section, which will not be repeated here.

[0154] Based on the above method embodiments, another embodiment of the present invention also provides a computer device, which can be a server, and its internal structure diagram can be as follows. Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the audio watermarking embedding method server-side as described in any of the above method embodiments.

[0155] Based on the above method embodiments, another embodiment of the present invention also provides a computer device, which can be a client, and its internal structure diagram can be as follows. Figure 5As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When executed by the processor, the computer program implements the functions or steps of the audio watermark embedding method on the client side as described in any of the above method embodiments.

[0156] Those skilled in the art will understand that Figure 4 and Figure 5 The structural schematic diagram shown is only a schematic diagram of a part of the structure related to the present invention and does not constitute a limitation on the computer device on which the present invention is applied. The specific computer device may include more components than shown in the figure, or combine certain components, or have different component arrangements.

[0157] The processor referred to herein can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0158] The memory includes readable storage media, internal memory, etc., where internal memory can be the RAM of a computer device. Internal memory provides an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of the computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal storage units and external storage devices of the computer device. The memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. The memory can also be used to temporarily store data that has been output or will be output.

[0159] Based on the above method embodiments, another embodiment of the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the audio watermark embedding method as described in any of the above method embodiments. The computer-readable storage medium may be non-volatile or volatile.

[0160] It should be noted that the functions or steps that can be achieved by the computer-readable storage medium or computer device, and the technical effects brought about by the functions / steps, can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0161] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc. The disclosed memory components or memories of the operating environment described herein are intended to include one or more of these and / or any other suitable types of memory.

[0162] Those skilled in the art will understand that, for the sake of convenience and brevity, the embodiments of the device of the present invention are only illustrated by the division of the above-mentioned functional units and modules. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the present invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0163] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0164] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0165] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0166] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An audio watermark embedding method, characterized in that, include: The host audio to be watermarked is decomposed into multiple sub-bands of different frequency bands using discrete wavelet transform; The encoder in the U-Net network model is used to downsample each sub-band layer by layer to extract the latent space features of the host audio; The watermark information to be embedded is determined, and the watermark information is encoded to obtain a watermark vector; The watermark vector is embedded into the latent space feature to obtain the fused feature; The fused features are upsampled layer by layer using the decoder in the U-Net network model to reconstruct the audio information; Perform inverse discrete wavelet transform on the audio information to generate target audio with the watermark information embedded. The step of embedding the watermark vector into the latent space features to obtain the fused features includes: The watermark embedding strategy is determined based on the content of the watermark information and the characteristics of the host audio. The watermark vector is embedded into the latent space feature according to the watermark embedding strategy to obtain the preliminary fused feature; The preliminary features are normalized to obtain the fused features; The step of using the decoder in the U-Net network model to upsample the fused features layer by layer to reconstruct the audio information further includes: The reconstructed audio information is subjected to a simulated attack test, and the test results are generated. Based on the test results, the watermark embedding strategy and / or the decoder are adjusted.

2. The audio watermark embedding method according to claim 1, characterized in that, Before the step of decomposing the host audio to be watermarked into multiple sub-bands of different frequency bands using discrete wavelet transform, the following steps are included: Obtain the original audio uploaded by the target user; The original audio is preprocessed to obtain the host audio to which the watermark is to be embedded.

3. The audio watermark embedding method according to claim 1, characterized in that, The step of downsampling each sub-band layer by layer using the encoder in the U-Net network model to extract the latent space features of the host audio includes: The U-Net network model is pre-built; Each sub-band is used as input and sequentially input into the encoder for layer-by-layer downsampling to obtain the depth feature representation corresponding to each sub-band; All the obtained deep feature representations are combined to generate the latent space features of the host audio.

4. The audio watermark embedding method according to claim 1, characterized in that, The process of determining the watermark information to be embedded and encoding the watermark information to obtain a watermark vector includes: Select the watermark information to be embedded based on the watermark embedding requirements of the host audio; The watermark information is encoded to obtain an initial vector; Adjust the length and / or intensity of the initial vector to obtain the watermark vector.

5. The audio watermark embedding method according to any one of claims 1-4, characterized in that, The step of performing inverse discrete wavelet transform on the audio information to generate target audio embedded with the watermark information includes: Perform inverse discrete wavelet transform on the audio information to generate an initial audio file embedded with the watermark information; The initial audio is denoised and equalized to obtain the target audio, and it is then identified whether the watermark information is embedded in the target audio. When the watermark information is embedded in the target audio, the target audio is output.

6. An audio watermark embedding device, characterized in that, include: The decomposition module is used to decompose the host audio to be embedded with watermark into multiple sub-bands of different frequency bands through discrete wavelet transform. The extraction module is used to downsample each sub-band layer by layer using the encoder in the U-Net network model to extract the latent space features of the host audio. An encoding module is used to determine the watermark information to be embedded and to encode the watermark information to obtain a watermark vector; An embedding module is used to embed the watermark vector into the latent space features to obtain fused features; The reconstruction module is used to upsample the fused features layer by layer using the decoder in the U-Net network model to reconstruct the audio information; The generation module is used to perform inverse discrete wavelet transform on the audio information to generate target audio with the watermark information embedded. The step of embedding the watermark vector into the latent space features to obtain the fused features includes: The watermark embedding strategy is determined based on the content of the watermark information and the characteristics of the host audio. The watermark vector is embedded into the latent space feature according to the watermark embedding strategy to obtain the preliminary fused feature; The preliminary features are normalized to obtain the fused features; The step of using the decoder in the U-Net network model to upsample the fused features layer by layer to reconstruct the audio information further includes: The reconstructed audio information is subjected to a simulated attack test, and the test results are generated. Based on the test results, the watermark embedding strategy and / or the decoder are adjusted.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio watermark embedding method as described in any one of claims 1-5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the audio watermark embedding method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Self-adaptive audio blind watermark method based on auditory model

    CN106504757A

  • Method and device for embedding and identifying audio watermark based on deep network

    CN113990330A