Audio watermark embedding method and apparatus, and conference system

By dynamically adjusting the embedding parameters and algorithms in audio watermarking technology, the robustness and imperceptibility of audio watermarking under complex environments and attacks are solved, achieving adaptive protection and effective traceability of audio data.

WO2026056323A1PCT designated stage Publication Date: 2026-03-19HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing audio watermarking technologies struggle to balance robustness and imperceptibility during embedding and extraction. In particular, watermark signals are prone to failure when faced with noise interference, cutting attacks, low-pass filtering attacks, desynchronization attacks, and ripping attacks, making it difficult to achieve effective audio data protection.

Method used

By acquiring raw audio in the terminal environment, the watermark embedding parameters and algorithm are dynamically adjusted to adapt to environmental changes and improve the robustness and extractability of the watermark. This includes adjusting the watermark embedding strength and frequency band, and changing the embedding algorithm when necessary to deal with different attack methods.

Benefits of technology

It achieves adaptive adjustment of audio watermarks under different environments and attacks, improves the robustness and imperceptibility of the watermark, enhances the security and traceability of audio data, and can effectively prevent information leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025095191_19032026_PF_FP_ABST
    Figure CN2025095191_19032026_PF_FP_ABST
Patent Text Reader

Abstract

An audio watermark embedding method (600) and apparatus, and a conference system, relating to the technical field of digital watermarking. The method (600) comprises: acquiring first original audio collected, during a first time period, from an environment where a first terminal is located, the first original audio comprising first playback audio played back by the first terminal during the first time period, and the first playback audio being configured to use a first watermark embedding algorithm to embed a first watermark on the basis of a first watermark embedding parameter value (601); performing watermark extraction on the first original audio on the basis of the first watermark embedding algorithm (602); and if the extraction of the first watermark from the first original audio fails, using the first watermark embedding algorithm to embed, on the basis of a second watermark embedding parameter value, the first watermark into first audio to be played back that is to be played back by the first terminal, so as to obtain second playback audio, the second watermark embedding parameter value differing from the first watermark embedding parameter value, and the second playback audio being configured to be played back by the first terminal during a second time period (603).
Need to check novelty before this filing date? Find Prior Art

Description

Audio watermark embedding method and device, and conference system

[0001] The present application claims priority to the Chinese patent application No. 202411295413.7, filed on September 14, 2024, and entitled "Audio watermark embedding method and device, and conference system", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of digital watermark technology, in particular to an audio watermark embedding method and device, and conference system. BACKGROUND

[0003] As an important branch of information hiding technology, audio watermark technology is very effective for copyright protection of audio works, and therefore has been paid more and more attention. Through audio watermark technology, copyright information or user information of a provider can be embedded as watermark information into audio data, so as to realize copyright protection and leakage tracing of audio data. Compared with image watermark technology, the implementation of audio watermark technology is more difficult, mainly because the human auditory system has higher sensitivity than the visual system. How to embed appropriate watermark in audio is a long-term research direction in the application process of audio watermark technology. SUMMARY

[0004] The present application provides an audio watermark embedding method and device, and conference system.

[0005] In a first aspect, an audio watermark embedding method is provided. The method comprises: obtaining a first original audio collected from an environment where a first terminal is located in a first time period, the first original audio including a first played audio played by the first terminal in the first time period. The first played audio is configured to be embedded with a first watermark based on a first watermark embedding parameter value using a first watermark embedding algorithm. According to the first watermark embedding algorithm, the first original audio is subjected to watermark extraction. If the first watermark fails to be extracted from the first original audio, the first watermark is embedded in a first to-be-played audio to be played by the first terminal based on a second watermark embedding parameter value using the first watermark embedding algorithm, to obtain a second played audio. The second watermark embedding parameter value is different from the first watermark embedding parameter value, and the second played audio is used for playing by the first terminal in a second time period, which is located after the first time period in time sequence.

[0006] The present application extracts the watermark in the audio played by the terminal from the original audio collected from the environment where the terminal is located. In the case where the watermark cannot be successfully extracted from the original audio, the watermark embedding parameter value is adjusted, so that the watermark embedded in the audio played by the terminal afterwards can adapt to the environment and be successfully extracted, realizing adaptive environment adjustment of audio watermark, and improving the robustness of audio watermark.

[0007] Optionally, the first watermark embedding parameter value comprises a first watermark embedding strength and a first watermark embedding frequency band, and the second watermark embedding parameter value comprises a second watermark embedding strength and a second watermark embedding frequency band. The second watermark embedding parameter value is different from the first watermark embedding parameter value, comprising: the second watermark embedding strength is different from the first watermark embedding strength, and / or the second watermark embedding frequency band is different from the first watermark embedding frequency band.

[0008] Optionally, the first original audio further comprises environmental audio generated in an environment where the first terminal is located in a first time period. According to the audio signal characteristics of the environmental audio and the audio signal characteristics of the first to-be-played audio, the first watermark embedding parameter value is adjusted to obtain the second watermark embedding parameter value.

[0009] Optionally, the first watermark embedding parameter value comprises a first watermark embedding strength. The implementation manner of adjusting the first watermark embedding parameter value according to the audio signal characteristics of the environmental audio and the audio signal characteristics of the first to-be-played audio comprises: if the difference between the audio signal strength of the first to-be-played audio and the audio signal strength of the environmental audio is greater than a signal strength threshold, the first watermark embedding strength is increased to obtain a second watermark embedding strength, and the second watermark embedding strength is taken as the watermark embedding strength in the second watermark embedding parameter value.

[0010] In the case that the watermark in the audio played by the terminal is failed to be extracted from the original audio collected from the environment where the terminal is located, the watermark embedding strength can be adjusted, and the watermark signal strength of the watermark information in the carrier audio is attempted to be increased by increasing the watermark embedding strength, so as to improve the probability of successful watermark extraction.

[0011] Optionally, the first watermark embedding parameter value further comprises a first watermark embedding frequency band. If the implementation manner of increasing the first watermark embedding strength when the difference between the audio signal strength of the first to-be-played audio and the audio signal strength of the environmental audio is greater than a signal strength threshold comprises: if the difference between the audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and the audio signal strength of the environmental audio in the first watermark embedding frequency band is greater than a signal strength threshold, the first watermark embedding strength is increased.

[0012] The application compares the audio signal strengths of the carrier audio and the environmental audio in the current watermark embedding frequency band to determine whether the currently used watermark embedding frequency band is suitable for embedding the watermark. If the difference between the audio signal strengths of the carrier audio and the environmental audio in the current watermark embedding frequency band is greater than a signal strength threshold, it indicates that the environmental audio has relatively small interference on the carrier audio in the current watermark embedding frequency band, and the possibility of successful watermark extraction can be improved by increasing the watermark embedding strength.

[0013] Optionally, the implementation manner of adjusting the first watermark embedding parameter value according to the audio signal feature of the environmental audio and the audio signal feature of the first to-be-played audio further includes: if a difference between the audio signal intensity of the first to-be-played audio in the first watermark embedding frequency band and the audio signal intensity of the environmental audio in the first watermark embedding frequency band is less than or equal to a signal intensity threshold, taking the second watermark embedding frequency band as the watermark embedding frequency band in the second watermark embedding parameter value, wherein a difference between the audio signal intensity of the first to-be-played audio in the second watermark embedding frequency band and the audio signal intensity of the environmental audio in the second watermark embedding frequency band is greater than the signal intensity threshold.

[0014] If the difference between the audio signal intensity of the carrier audio and the environmental audio in the current watermark embedding frequency band is less than or equal to the signal intensity threshold, it indicates that the environmental audio has relatively large interference on the carrier audio in the current watermark embedding frequency band, and the influence effect of increasing the watermark embedding intensity on watermark extraction is small. At this time, the possibility of successful watermark extraction can be tried to be improved by adjusting the watermark embedding frequency band.

[0015] Optionally, the above method is applied to a conference service platform. The original audio stream and the environmental audio stream sent by the first terminal are received, the environmental audio stream is obtained after eliminating the playing audio stream played by the first terminal in the original audio stream, and the environmental audio stream is used to provide the second terminal for playing. The first original audio is obtained from the original audio stream, and the environmental audio is obtained from the environmental audio stream.

[0016] Optionally, after receiving the environmental audio stream sent by the first terminal, a second watermark is embedded in the environmental audio stream, and the environmental audio stream embedded with the second watermark is sent to the second terminal. Alternatively, the conference service platform can send the environmental audio stream to the second terminal, and the second terminal embeds the second watermark in the received environmental audio stream and plays it.

[0017] Optionally, unknown watermark detection is performed on the environmental audio. If there is an unknown watermark in the environmental audio, first prompt information is output, and the first prompt information is used to indicate that the environment where the first terminal is located has a risk of information leakage.

[0018] Since there may be an internal attacker in the environment of the terminal who uses the audio watermark as a covert information channel for leaking information to the outside, the present application performs unknown watermark detection on the environmental audio to determine whether there is an unknown watermark in the environmental audio, and records and prompts in the case where there is an unknown watermark in the environmental audio, thereby realizing the sensing ability of the air gap attack based on the unknown watermark, so as to be able to perform post-tracing after the information leakage occurs.

[0019] Optionally, the first watermark embedding parameter value comprises a first watermark embedding strength. If the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark exceeds the watermark signal strength range, the first watermark embedding strength is adjusted to obtain a third watermark embedding strength. The first watermark is embedded in the first to-be-played audio based on the third watermark embedding parameter value by using the first watermark embedding algorithm to obtain a third played audio, the third watermark embedding parameter value comprises the third watermark embedding strength, and the third played audio is used for the first terminal to play in a second time period.

[0020] In the case where the watermark in the audio played by the terminal is successfully extracted from the original audio collected from the environment where the terminal is located, whether the watermark signal strength of the extracted watermark exceeds the watermark signal strength range is judged. If the watermark signal strength of the extracted watermark exceeds the watermark signal strength range, the watermark embedding strength is adjusted, so that the dynamic adjustment of the watermark signal strength is realized, and the watermark signal strength of the watermark extracted later can be within the watermark signal strength range, so as to avoid the problems that the watermark robustness is low due to too small watermark signal strength, or the audio playback quality is affected due to too large watermark signal strength, so as to better balance the robustness and imperceptibility of the audio watermark.

[0021] Optionally, if the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark exceeds the watermark signal strength range, the implementation manner of adjusting the first watermark embedding strength comprises: if the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark is less than the minimum value of the watermark signal strength range, the first watermark embedding strength is increased. Or, if the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark is greater than the maximum value of the watermark signal strength range, the first watermark embedding strength is decreased.

[0022] Optionally, second original audio collected from the environment where the first terminal is located in a second time period is obtained, and the second original audio comprises the second played audio. The watermark is extracted from the second original audio according to the first watermark embedding algorithm. If the first watermark fails to be extracted from the second original audio, the first watermark is embedded in second to-be-played audio to be played by the first terminal by using a second watermark embedding algorithm to obtain fourth played audio, the second watermark embedding algorithm is different from the first watermark embedding algorithm, and the fourth played audio is used for the first terminal to play in a third time period, the third time period is located after the second time period in time sequence.

[0023] In the case where the watermark cannot be successfully extracted after only adjusting the watermark embedding parameter value, the watermark embedding algorithm can be further adjusted to try to use different watermark embedding algorithms for watermark embedding to improve the watermark extractability, and the flexibility of audio watermark embedding is improved.

[0024] Optionally, it is detected whether the first original audio is playback audio. If the first original audio is playback audio, second prompt information is output, and the second prompt information is used to indicate that there is a watermark embedding abnormal problem in the audio played by the first terminal.

[0025] The present application detects whether the obtained first original audio is playback audio, to determine the authenticity of the first original audio, that is, to determine whether the first original audio is actual audio that is truly collected from the environment where the first terminal is located in a first time period, and records and prompts in a case where it is determined that the obtained first original audio is playback audio, so as to be able to conduct post-investigation and trace.

[0026] In a second aspect, an audio watermark embedding device is provided. The audio watermark embedding device includes a plurality of function modules that interact with each other to implement the method in the first aspect and each of the implementations thereof. The plurality of function modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of function modules can be arbitrarily combined or divided based on specific implementation.

[0027] In a third aspect, a conference system is provided, including a conference service platform and a plurality of conference terminals, the plurality of conference terminals are communicatively connected through the conference service platform, and the plurality of conference terminals include a first conference terminal. The first conference terminal may, for example, be the first terminal in the method in the first aspect and each of the implementations thereof, and the conference service platform may be used to execute the method in the first aspect and each of the implementations thereof.

[0028] For example, the first conference terminal is configured to send, to the conference service platform, first original audio collected from an environment where the first conference terminal is located in a first time period, the first original audio including first playback audio played by the first conference terminal in the first time period, the first playback audio being configured to have first watermark embedded based on a first watermark embedding parameter value using a first watermark embedding algorithm. The conference service platform is configured to extract the first watermark from the first original audio according to the first watermark embedding algorithm, and if the extraction of the first watermark from the first original audio fails, embed the first watermark in first to-be-played audio to be played by the first conference terminal based on a second watermark embedding parameter value using the first watermark embedding algorithm to obtain second playback audio, the second watermark embedding parameter value being different from the first watermark embedding parameter value. The first conference terminal is configured to play the second playback audio in a second time period, the second time period being located after the first time period in time sequence.

[0029] In a fourth aspect, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method in the first aspect and each of the implementations thereof.

[0030] In a fifth aspect, a computer program product including instructions is provided, which, when executed by a computing device cluster, causes the computing device cluster to perform the method in the first aspect and each of the implementations thereof.

[0031] In a sixth aspect, a computer-readable storage medium is provided, including computer program instructions, which, when executed by a computing device cluster, causes the computing device cluster to perform the method in the first aspect and each of the implementations thereof.

[0032] In a seventh aspect, a chip is provided, including programmable logic circuit and / or program instructions, which, when the chip is executed, implements the method in the first aspect and each of the implementations thereof. BRIEF DESCRIPTION OF DRAWINGS

[0033] FIG. 1 is an implementation schematic diagram of a digital watermarking technology provided by an embodiment of the present application;

[0034] FIG. 2 is a schematic diagram of a video conference scenario provided by an embodiment of the present application;

[0035] FIG. 3 is a schematic diagram of audio watermarking frame embedding provided by a related art;

[0036] FIG. 4 is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0037] FIG. 5 is a structural schematic diagram of a conference system provided by an embodiment of the present application;

[0038] FIG. 6 is a flow schematic diagram of an audio watermark embedding method provided by an embodiment of the present application;

[0039] FIG. 7 is an implementation schematic diagram of an audio watermark embedding and extraction process provided by an embodiment of the present application;

[0040] FIG. 8 is a structural schematic diagram of an audio watermark embedding device provided by an embodiment of the present application;

[0041] FIG. 9 is a structural schematic diagram of a computing device provided by an embodiment of the present application;

[0042] FIG. 10 is a structural schematic diagram of a computing device cluster provided by an embodiment of the present application;

[0043] FIG. 11 is a structural schematic diagram of another computing device cluster provided by an embodiment of the present application. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0045] Digital watermarking is an information hiding technology. Watermark information is embedded in the carrier file without affecting the readability and integrity of the original file. Digital watermarking technology has a wide range of applications, including copyright protection, anti-counterfeiting, digital forensics, and information hiding. Currently, digital watermarking technology has become one of the important means of digital media security protection. For example, Figure 1 is a schematic diagram of the implementation of a digital watermarking technology provided in this application embodiment. As shown in Figure 1, it generates a watermark from specific identification information (which may be the creator of the digital media, user serial number, or copyright information, etc.) and embeds it into the digital media information carrier. By performing watermark detection on the watermarked digital media, the embedded watermark information can be extracted, thereby indicating some information about the digital media itself. Digital watermarks can be embedded in various digital media, such as images, audio, video, and text. Digital watermarking technology mainly includes image watermarking technology and audio watermarking technology. Image watermarking technology can be used to protect the copyright of image or video works. Audio watermarking technology can be used to protect the copyright of audio works.

[0046] Based on their perceptibility, digital watermarks can be categorized into visible and invisible watermarks. Visible watermarks, also known as overt watermarks, are watermarks that are perceptibly visible, such as identifiers inserted or overlaid on images. Visible watermarks typically refer to visible information directly embedded in digital media, such as text or images. They are generally used to visually identify images or videos obtained from video databases or the internet to prevent these images from being used for illicit commercial purposes. Similarly, audible watermarks are watermarks that are perceptibly visible in audio. Invisible watermarks, also known as hidden watermarks or covert watermarks, are watermarks that are not perceptibly visible (difficult to detect after embedding). Invisible watermarks typically refer to invisible information embedded in digital media, such as digital codes or noise. Embedding invisible watermarks in digital media does not significantly damage the media itself and preserves its original playback quality. When needed, the owner can extract the watermark from the digital media using watermark detection algorithms to prove the authenticity or integrity of the digital media.

[0047] The basic characteristics of digital watermark mainly include fidelity, robustness and capacity. The fidelity is used to measure the similarity before and after the signal is processed. The digital media information after embedding the watermark information should meet certain requirements in perception, which is not necessarily visible or invisible watermark, and is determined according to the application occasion. For invisible watermark, the fidelity is the imperceptibility (transparency) of the watermark. Robustness refers to the extractability and detectability of the watermark after the digital media information with watermark undergoes various signal processing or various attacks. The measurement indicators of robustness include fragility, effectiveness and security. The capacity, also known as embedding rate, loading rate or payload, refers to the maximum number of watermark bits that can be embedded per unit time or in a digital media work.

[0048] With the development of information technology, the problem of illegal leakage of media data is becoming increasingly serious. In the related technology, digital watermark is embedded in the media data to realize the identification and tracking of the media data. The media data includes but is not limited to audio data, video data, file data and other digital streaming media data, such as remote shared desktop, remote shared document and remote shared application, etc.

[0049] Taking a video conference scene as an example, FIG. 2 is a schematic diagram of a video conference scene provided by an embodiment of the present application. As shown in FIG. 2, a secure conference is convened by a conference administrator, an audio and video stream is sent by a sending side conference terminal, data transmission is performed through a conference service platform, an audio and video stream is received and played by a receiving side conference terminal, and an end-to-end interaction process is completed. However, during the audio and video playing process, an internal leaker may perform a recording or forward the played audio and video stream to a person without authority, thereby causing data leakage.

[0050] Among them, the recording behavior includes using a screen recording software to transcribe or using a camera to shoot, and it is almost impossible to prevent others from taking pictures or recording in technical means, so the main idea of preventing recording is to trace the source of the recorder to deter the recording behavior. Therefore, the digital watermark technology is particularly important here. As shown in FIG. 2, after the receiving side conference terminal receives the audio and video stream, real-time watermark generation and embedding can be performed when the audio and video stream is played. The embedded watermark information is bound to the user identity or the device identity accessing the conference. Among them, the user identity accessing the conference can be the identity of the user currently logged in the conference terminal, and the device identity accessing the conference can be the identity of the conference terminal playing the audio and video stream. In this way, once a recording behavior (such as recording from a distance) occurs, the conference administrator only needs to obtain a small part of the leaked segment in the audio and video stream, and can extract the watermark information through a watermark extraction algorithm, and identify the device identity and / or user identity in the watermark information, thereby locating the device or user leaking the data, and realizing the tracing of the internal leaker. The security in the transmission process before watermark embedding is guaranteed by end-to-end encryption.

[0051] In order to reduce the impact of the watermark on the quality of the audio, an audio watermarking technology can be used to embed some invisible information in the audio data, that is, to embed a hidden watermark in the audio data. These information can be in the form of digital code, digital signature, digital watermark, etc., which are embedded in the different time domain, frequency domain, phase, amplitude, etc. of the audio signal to ensure that the quality and audibility of the audio signal are not affected. For example, an audio hidden watermarking technology can be used in a video conference system. The audio data is divided into file information and real-time information, and the internal structure of both the file information and the real-time information adopts a framing form, which divides the entire audio data into a plurality of audio data frames Fi. The audio data frame is the smallest unit of an audio segment, and the sending and coding and decoding are performed in the data length of the audio data frame. Therefore, the current audio watermark embedding scheme can embed watermark information between the frequency signal components of each audio data frame by frequency masking, to obtain an audio data frame Di containing watermark information. Wherein i is an integer greater than 1. For example, FIG. 3 is an audio watermark framing embedding schematic diagram provided by the related art. In the related art, in order to exclude interference and meet the real-time transmission scenario, the position of the audio watermark segment in the audio data frame needs to be determined through synchronization information, so that the audio watermark segment can be accurately restored at the extraction end.

[0052] Compared with the image watermarking technology, the implementation of the audio watermarking technology is more difficult, mainly because the human auditory system has higher sensitivity than the visual system. How to balance the imperceptibility and robustness of the audio watermarking is a long-standing problem in the application process of the audio watermarking technology. In the related art, a fixed watermark embedding parameter value is usually used, for example, a fixed watermark embedding strength is used to embed the watermark in a fixed frequency band of the audio data. However, if the watermark signal strength embedded in the audio data is too large, it is easy to be perceived by the human ear, and the imperceptibility of the watermark is poor. If the watermark signal strength embedded in the audio data is too small, the watermark extractability and detectability will be poor after the signal processing and audio watermark attack, and the watermark robustness is low. The current common audio watermark attack methods include noise interference attack, cutting attack, low-pass filtering attack, desynchronization attack and dubbing attack. Among them, the cutting attack, the low-pass filtering attack and the desynchronization attack are realized by tampering with the audio signal embedded with the watermark.

[0053] The noise interference attack refers to adding noise (such as Gaussian noise or white noise, etc.) to the audio signal embedded with the watermark to interfere with the extraction of the watermark signal.

[0054] Cutting attack refers to cutting or clipping the audio signal to affect the extraction or identification of the watermark signal. Cutting attack may be intercepting a certain time period or specific part of the audio signal, or deleting a part of the audio signal, or changing the sampling rate of the audio signal, resulting in changes in the frequency characteristics of the watermark signal, making it difficult to extract or identify the watermark signal.

[0055] Low-pass filtering attack refers to low-pass filtering processing of the audio signal embedded with watermark. The specific implementation process is as follows: first, select the corresponding low-pass filter according to different audio signals, then set the filter parameters, and finally input the audio signal into the low-pass filter. The signal after filtering is the low-frequency signal. In this way, the watermark signal embedded in the high-frequency part will be filtered out, resulting in the watermark signal being unable to be accurately extracted. Low-pass filtering attack usually needs to select appropriate filter parameters according to the characteristics of the watermark embedding algorithm and the frequency characteristics of the watermark signal. Attackers may use digital signal processing tools or software to implement low-pass filtering attack to process the audio signal to destroy the extractability of the watermark.

[0056] Desynchronization attack aims to use the dependence of watermark algorithm on synchronization information to destroy the extraction process of watermark by modifying the synchronization information. Among them, the synchronization information can be a flag or parameter used to locate the position of the watermark. Desynchronization attack is implemented as follows: the attacker processes the attacked audio signal and modifies or disguises the synchronization information in it, such as modifying the time position, frequency characteristics, etc. of the synchronization signal. This attack has a great impact on audio watermark algorithms that require accurate synchronization information, causing the frames of the attacked audio with watermark to be out of sync with the frames of the audio with watermark before the attack, resulting in watermark extraction failure or error.

[0057] Re-recording attack (or resampling attack) usually refers to the attacker recording the audio content containing watermark, and then using means such as re-embedding watermark in the recorded audio or trying to remove the watermark in the recorded audio to modify or destroy the original watermark information, so that the watermark signal in the recorded audio is difficult to extract, achieving the purpose of destroying the robustness of the watermark. This attack mainly occurs in the process of propagation and copying of audio content.

[0058] In the related art, a fixed watermark embedding strength is usually used to embed a watermark in audio data, and the audio watermark signal strength is limited by imperceptibility, and cannot be set too strong. In addition, the quality of the audio peripheral for playing audio and the noise of the playing environment are usually unpredictable or uncontrollable, so that the watermark robustness in the recorded audio is easily affected by factors such as the quality of the audio peripheral and the environmental noise, resulting in that the watermark in the recorded audio cannot be effectively extracted. Therefore, how to embed appropriate watermark information in the audio data so as to achieve a reasonable compromise between watermark robustness and imperceptibility is a difficulty in the application of audio watermark technology.

[0059] Based on this, the application provides a technical solution. An audio pickup device is deployed in an environment where a terminal is located, and the audio pickup device collects actual audio generated in the environment where the terminal is located. The actual audio is audio that has not been subjected to noise reduction processing, which is referred to as original audio in the application. In the case where the terminal plays audio, the original audio collected by the audio pickup device includes the audio played by the terminal. If environmental audio is generated in the environment where the terminal is located, the original audio can also include the environmental audio. The environmental audio includes but is not limited to environmental noise audio and / or environmental human voice audio. The environmental noise audio includes, for example, low-frequency noise generated when devices such as air conditioners and fans in the local environment are running, and / or noise emitted by any sound-emitting device deployed in the local environment. The environmental human voice audio refers to the speech of a speaker in the local environment. In the case where the audio played by the terminal is configured to be embedded with a watermark, the application collects the original audio generated in the environment where the terminal is located, and performs watermark extraction on the original audio according to a watermark embedding algorithm used for embedding the watermark in the audio played by the terminal. If the watermark extraction from the original audio fails, the watermark embedding parameter value is adjusted, and the watermark is embedded in the audio to be played by the terminal based on the adjusted watermark embedding parameter value. The application can dynamically adjust the watermark embedding parameter value according to the audio playing environment. Compared with using a fixed watermark embedding parameter value, the application scheme enables the watermark in the audio played by the terminal to better adapt to the environment, realizes adaptive environmental adjustment of the audio watermark, and thus improves the audio watermark robustness. In the case of air recording, since the recorded audio and the original audio collected by the application are both actual audio generated in the environment where the terminal is located, in the case where the watermark embedded in the audio played by the terminal can be successfully extracted from the original audio, the watermark can usually be extracted from the recorded audio, thereby realizing the tracing of the recorded leaked audio.

[0060] The technical scheme provided in the application can be applied to a computer device with audio watermark detection capability (referred to as a watermark detection device), which can be a conference service platform or a conference management platform, etc. in a conference scenario. The technical scheme provided in the application is implemented as follows: a first original audio collected from an environment where a first terminal is located in a first time period is acquired, the first original audio includes first playing audio played by the first terminal in the first time period, the first playing audio is configured to be embedded with a first watermark based on a first watermark embedding algorithm and a first watermark embedding parameter value. The first watermark is extracted from the first original audio according to the first watermark embedding algorithm. If the first watermark fails to be extracted from the first original audio, the first watermark is embedded in first to-be-played audio to be played by the first terminal based on a second watermark embedding parameter value by using the first watermark embedding algorithm to obtain second playing audio. The second watermark embedding parameter value is different from the first watermark embedding parameter value. The second playing audio is used for playing by the first terminal in a second time period, and the second time period is located after the first time period in time sequence. The watermark embedding parameter refers to a plurality of parameters related to watermark embedding. The first watermark embedding parameter value and the second watermark embedding parameter value refer to different parameter values of one or more watermark embedding parameters related to the first watermark embedding algorithm. The application extracts the watermark in the audio played by the terminal from the original audio collected from the environment where the terminal is located. In the case that the watermark cannot be successfully extracted from the original audio, the watermark embedded in the audio played by the terminal after adjustment of the watermark embedding parameter value can be adapted to the environment and thus successfully extracted, the adaptive environment adjustment of the audio watermark is realized, and the robustness of the audio watermark is improved. In addition, some attacks on the audio watermark may occur on the terminal side, for example, in the case that the watermark embedding is performed by the terminal, the related software on the terminal may be tampered with, so that the terminal cannot perform watermark embedding before playing the audio, or the terminal is subjected to acoustic attack when playing the audio, such as the existence of some acoustic devices in the environment where the terminal is located to make noise interference on the watermark embedding frequency band. By implementing the scheme of the application, it can be determined whether the audio played by the terminal is effectively embedded with the watermark, so as to improve the controllability of the watermark embedding, realize effective protection of the audio played by the terminal, and play a deterrent role on audio recording and leakage behavior.

[0061] Optionally, the audio watermark embedding algorithm includes, but is not limited to, a least significant bit (LSB) embedding algorithm, a spread spectrum algorithm, or a phase coding algorithm. Among them, the LSB algorithm is a relatively simple watermark embedding method, which hides the watermark information in the least significant bit of the audio data, so as to minimize the impact on the original audio quality. The spread spectrum algorithm enhances the concealment and attack resistance of the watermark by distributing the encoded data into as many spectrums as possible. The spread spectrum algorithm combines the m sequence with excellent performance for encoding and decoding, and has certain robustness to moving picture experts group audio layer 3 (MP3) audio encoding, pulse code modulation (PCM) quantization, and additional noise. The phase coding algorithm uses the characteristic that the human ear auditory system is not sensitive to the absolute phase, uses a reference phase representing the watermark information to replace the absolute phase of the original audio segment, while keeping the relative phase unchanged, to realize the embedding of the watermark. In addition, the audio watermark embedding algorithm can also be realized based on Fourier transform and wavelet transform, and the application does not limit the watermark embedding algorithm used. The following mainly takes the first watermark embedding algorithm as an example, and the corresponding watermark embedding parameters related to the first watermark embedding algorithm include, but are not limited to, watermark embedding strength and / or watermark embedding frequency band.

[0062] Optionally, the first watermark embedding parameter value includes a first watermark embedding strength and a first watermark embedding frequency band, and the second watermark embedding parameter value includes a second watermark embedding strength and a second watermark embedding frequency band. The second watermark embedding parameter value is different from the first watermark embedding parameter value, including that the second watermark embedding strength is different from the first watermark embedding strength, and / or the second watermark embedding frequency band is different from the first watermark embedding frequency band.

[0063] In some embodiments, the first watermark embedding parameter value includes a first watermark embedding strength. The application can further include the following implementation: if the first watermark is successfully extracted from the first original audio, and the watermark signal strength of the extracted first watermark exceeds the watermark signal strength range, the first watermark embedding strength is adjusted to obtain a third watermark embedding strength. Then the first watermark is embedded in the first to-be-played audio based on the third watermark embedding parameter value using the first watermark embedding algorithm to obtain a third played audio. The third watermark embedding parameter value includes the third watermark embedding strength. The third played audio is used for the first terminal to play in a second time period. Conversely, if the watermark signal strength of the extracted watermark is within the watermark signal strength range, the first watermark embedding strength can be kept unchanged, and the watermark can be continuously embedded in the audio to be played by the terminal based on the first watermark embedding parameter value. In the case where the watermark in the audio played by the terminal is successfully extracted from the original audio collected from the environment where the terminal is located, the watermark signal strength of the extracted watermark is judged whether it exceeds the watermark signal strength range. If the watermark signal strength of the extracted watermark exceeds the watermark signal strength range, the watermark embedding strength is adjusted, so as to realize the dynamic adjustment of the watermark signal strength, so that the watermark signal strength of the watermark extracted later can be within the watermark signal strength range, so as to avoid the problems that the watermark robustness is low due to too small watermark signal strength, or the audio playback quality is affected due to too large watermark signal strength affecting the watermark transparency, so as to better balance the robustness and imperceptibility of the audio watermark.

[0064] In some embodiments, if the watermark embedding parameter value is adjusted one or more times for the same watermark embedding algorithm, and the watermark in the audio played by the terminal still cannot be successfully extracted from the original audio collected from the environment where the terminal is located, the watermark embedding algorithm can be further adjusted. The application can further include the following implementation: the second original audio collected from the environment where the first terminal is located in a second time period is obtained, and the second played audio is included in the second original audio. The watermark is extracted from the second original audio according to the first watermark embedding algorithm. If the first watermark fails to be extracted from the second original audio, the first watermark is embedded in the second to-be-played audio to be played by the first terminal using a second watermark embedding algorithm to obtain a fourth played audio. The second watermark embedding algorithm is different from the first watermark embedding algorithm. The fourth played audio is used for the first terminal to play in a third time period. The third time period is located after the second time period in time sequence. In the case where the watermark still cannot be successfully extracted after only adjusting the watermark embedding parameter value, the watermark embedding algorithm can be further adjusted to try to use different watermark embedding algorithms for watermark embedding to improve the watermark extractability, thereby improving the flexibility of audio watermark embedding.

[0065] In some embodiments, in the case that the first period of time includes environmental audio generated in the environment where the first terminal is located in the first original audio, the present application can further include the following implementation, performing unknown watermark detection on the environmental audio, and outputting a first prompt information if there is unknown watermark in the environmental audio. The first prompt information is used to indicate that there is a risk of information leakage in the environment where the first terminal is located. Since there may be an internal attacker in the environment where the terminal is located who uses audio watermark as a covert information channel for external information leakage, the present application performs unknown watermark detection on the environmental audio to determine whether there is unknown watermark in the environmental audio, and records and prompts in the case that there is unknown watermark in the environmental audio, so as to realize the sensing capability of the air gap attack based on unknown watermark, so as to be able to conduct post-investigation and traceback after information leakage occurs.

[0066] In some embodiments, the present application can further include the following implementation, detecting whether the first original audio is replayed audio, and outputting a second prompt information if the first original audio is replayed audio. The second prompt information is used to indicate that there is a watermark embedding abnormal problem in the audio played by the first terminal. The present application detects whether the obtained first original audio is replayed audio, to determine the authenticity of the first original audio, that is, to determine whether the first original audio is actually collected from the environment where the first terminal is located in the first period of time, and records and prompts in the case that it is determined that the obtained first original audio is replayed audio, so as to be able to conduct post-investigation and traceback. In this way, it can be detected that the attacker uses the audio recorded in advance and embedded with the first watermark to imitate the actual audio collected from the environment where the first terminal is located, so as to effectively monitor whether the audio played by the first terminal is actually embedded with the watermark, to strengthen the defense means against terminal software tampering and acoustic attack, and to improve the security and reliability of audio playing.

[0067] Optionally, the present application can be applied to a conference scene, and then watermark information associated with conference information can be used, that is, the audio watermark (including the first watermark and / or the second watermark) used in the present application is associated with the conference information. Optionally, the conference information includes one or more of conference room information to which the local environment belongs, conference number, or conference timestamp information. The conference room information can indicate the place where the audio is propagated, and the conference room information is for example a conference room doorplate number. The conference number and / or the conference timestamp information can be used to identify a conference. The conference timestamp information can include the start time and the end time of the conference. The present application uses watermark information associated with conference information to identify the source of the audio, so that when the audio is traced back later, it can be located to which conference the audio specifically comes from according to the watermark information, so as to narrow the tracing range of the audio data.

[0068] For example, the audio watermark adopted by the present application includes conference information and signature information obtained by signing the conference information by using a target key. The target key is obtained based on an identity private key held by an audio playing device and an authentication public key provided by an authenticator. Only the authenticator has the tracing right of the audio data. The conference information in the audio watermark can declare the source of the audio data, and thus plays the role of the identity information of the audio data, thereby narrowing the tracing range of the audio data. The signature information in the audio watermark is used to ensure the authenticity of the conference information corresponding to the audio data, that is, to ensure the authenticity of the audio data. Since the target key is obtained based on the identity private key held by the audio playing device and the authentication public key provided by the authenticator, as long as the identity private key held by the audio playing device is not disclosed, other user terminals cannot calculate the target key, and thus cannot calculate the signature information by using the target key to forge the watermark, and thus malicious forgery of the watermark by others to disclose the audio data for framing can be prevented. In addition, in the present application, the source of the audio data is proved by the signature information, rather than directly carrying the identity information of the user or the device in the audio watermark, and thus the disclosure of the identity information of the user or the device can also be avoided.

[0069] In a first possible implementation, the target key is generated by the audio playing device by using a key agreement algorithm for the identity private key and the authentication public key. Correspondingly, the signature information in the audio watermark can be a symmetric signature value obtained by the audio playing device by signing the conference information based on a symmetric signature algorithm by using the target key, or a truncated result value of the symmetric signature value. In this implementation, only the authenticator and the audio playing device can calculate the same symmetric key, and as long as the identity private key held by the audio playing device and the authentication private key held by the authenticator are not disclosed, any third party other than the audio playing device and the authenticator cannot calculate the target key, and thus cannot calculate the signature information by using the target key to forge the watermark, and thus malicious forgery of the watermark by others to disclose the audio data for framing can be prevented. In addition, the present application can perform truncation processing on the symmetric signature value, so as to embed the truncated result value as watermark information into the audio data, and thus the embedding capacity of the watermark can be reduced, and thus the fidelity and robustness of the audio data after embedding the watermark information can be ensured. In this implementation, the same key is used for generating the audio watermark and detecting the audio watermark, and the audio watermark can be referred to as a symmetric watermark.

[0070] Optionally, in the first possible implementation manner, the authentication public key is provided by a plurality of approval parties, and the authentication public key is a distributed key generation (DKG) public key. The plurality of approval parties jointly hold n private key fragments, and each approval party holds a private key fragment less than t in number, the DKG public key is calculated based on the n private key fragments, and a private key corresponding to the DKG public key is obtained based on at least t private key fragments of the n private key fragments. Wherein, n is an integer greater than 1, and 2≤t≤n. In this implementation manner, the authentication party needs to rely on a plurality of approval parties to recover the private key corresponding to the authentication public key, and further combine the identity public key held by the audio playback device to calculate the same key as the target key used by the audio playback device to generate the audio watermark. Compared with the scheme in which the authentication party directly holds the authentication private key, in this scheme, since the authentication party is supervised by a plurality of approval parties, the risk of malicious impersonation of the authentication party to frame the terminal user by leaking the audio data can be reduced, the reliability of the authentication party and the security of the watermark are improved, and the reliability of the watermark traceability is further improved.

[0071] In the second possible implementation manner, the target key is generated by the audio playback device based on the identity private key and the authentication public key by using an asymmetric key generation algorithm. Correspondingly, the signature information in the audio watermark can be an asymmetric signature value obtained by the audio playback device signing the conference information based on the target key by using an asymmetric signature algorithm. In this implementation manner, only the audio playback device can calculate the target key used to generate the audio watermark, and any third party including the authentication party can be prevented from maliciously impersonating the watermark to leak the audio data for framing. In addition, only the authentication party can calculate the key that can verify the audio watermark generated by the audio playback device, that is, only the authentication party has the traceability right of the audio data played by the audio playback device, and the terminal user privacy can be well protected. In this implementation manner, the keys used to generate the audio watermark and detect the audio watermark are different, and the audio watermark can be referred to as an asymmetric watermark.

[0072] Optionally, in one implementation manner in which the audio playback device generates the target key based on the identity private key and the authentication public key by using an asymmetric key generation algorithm, the audio playback device generates a derived key based on the identity private key and the authentication public key by using a key derivation function. The audio playback device generates the target key based on the derived key and the identity private key. Optionally, the target key is SK, and SK=(K+SKu)mod q. Wherein, SKu is the identity private key, K is the derived key, q is a prime number, mod q represents taking modulo q, and the value range of SK is [1, q).

[0073] Alternatively, the application can also use watermark information associated with the audio playback device, that is, the audio watermark used in the application is associated with the audio playback device, for example, the first watermark is associated with the first terminal, and the second watermark is associated with the second terminal. Alternatively, for a conference scenario, the application can also combine voice recognition technology to collect the voiceprint information of the participants in the conference start stage and bind the identity, dynamically identify different participant voices in the conference stage through the pre-recorded voiceprint information, and embed the identity information of the corresponding participant as the audio watermark in the different human voice audio in the watermark embedding stage. In this way, the specific speaking stage and the specific speaker can be traced back in the audio tracing stage. The application does not limit the audio watermark used.

[0074] The technical solutions of the application will be described in detail from the aspects of application scenarios, systems, method flows, software devices, and hardware devices.

[0075] The application scenarios of the embodiments of the application will be described below.

[0076] The embodiments of the application can be applied to various audio playback scenarios and can realize protection of the played audio. For example, FIG. 4 is a schematic diagram of an application scenario provided by an embodiment of the application. As shown in FIG. 4, the application scenario includes a terminal, a sound pickup device, and a watermark detection device. The terminal is configured to play audio embedded with a watermark. The sound pickup device is configured to collect original audio generated in an environment where the terminal is located and provide the original audio to the watermark detection device. The original audio collected by the sound pickup device includes the audio played by the terminal. Optionally, the original audio collected by the sound pickup device can also include environmental audio generated in the environment where the terminal is located, which includes but is not limited to environmental noise audio and / or environmental human voice audio. The environmental noise audio includes, for example, low-frequency noise generated when a device such as an air conditioner or a fan in the local environment is running, and / or noise emitted by any sound-emitting device deployed in the local environment. The environmental human voice audio refers to the speech of a speaker in the local environment. The watermark detection device is configured to execute the solutions provided by the embodiments of the application, including but not limited to extracting the watermark in the original audio collected by the sound pickup device, adjusting the watermark embedding parameter value and / or the watermark embedding algorithm, performing unknown watermark detection on the original audio collected by the sound pickup device, or detecting whether the original audio collected by the sound pickup device is a replayed audio, and the like.

[0077] Optionally, in the application scenario shown in FIG. 4, the terminal and the sound pickup device are two independent devices. In actual scenarios, the terminal and the sound pickup device can be the same device, that is, the terminal plays audio and collects original audio generated in the environment in which the terminal is located. Optionally, the sound pickup device can collect audio generated in the environment in which the terminal is located in real time or periodically, or the sound pickup device can only collect audio generated in the environment in which the terminal is located when it detects that the terminal plays audio, and the embodiments of the present application do not limit this.

[0078] The application example takes the application scenario shown in FIG. 4 as an example for description, which is a conference scenario, and the application scenario can be an audio / video conference system or a cloud service conference system. Correspondingly, the terminal is a conference terminal, and the watermark detection device can be a conference service platform or a conference management platform. The conference management platform can include a conference management server and a conference management terminal, and the conference management terminal is usually a conference terminal used by a conference manager. In the conference scenario, the sound pickup device is a device with an audio collection function deployed in the conference scenario. The sound pickup device can send the collected original audio to the conference terminal, and the conference terminal sends the original audio to the watermark detection device (conference service platform or conference management platform). In addition, the conference terminal also performs noise removal processing on the original audio collected by the sound pickup device to remove the audio played by the conference terminal itself in the original audio, to obtain environmental audio, and sends the environmental audio to other conference terminals through the conference service platform.

[0079] Optionally, the conference terminal can be a dedicated physical device or a software program with a conference function, which can run on various computing devices such as mobile phones, tablets, computers, and various user terminals. In this case, the computing device running the software program can also be considered as a conference terminal. The conference terminal joins the video conference through the conference service platform. Specifically, the conference terminal can obtain conference data (including audio data) from the conference service platform, and send locally collected conference data to the conference service platform, so that the conference service platform forwards the conference data to other conference terminals. The conference terminals can be connected through a wireless network, so that the conference participants are not limited by geographical location and can smoothly join the video conference. In some cases, a conference participant can include only one conference user, such as a conference user joining the conference through a conference software program running on a personal mobile phone. In some cases, a conference participant can also include multiple conference users, such as multiple conference users in a conference room joining the conference through a conference terminal in the conference room. The conference service platform can be a multipoint control unit (MCU).

[0080] The conference system of the embodiments of the present application is described below.

[0081] For example, FIG. 5 is a structural schematic diagram of a conference system provided by an embodiment of the present application. As shown in FIG. 5, the conference system includes a conference terminal and a conference service platform (MCU). Optionally, the conference system further includes a conference management platform. The conference management platform can include a conference management server and a conference management terminal. The conference terminal is in communication connection with the conference service platform. The conference service platform is in communication connection with the conference management server. The conference management server is in communication connection with the conference management terminal.

[0082] The conference service platform includes an audio / video transceiving module and a service message module. Optionally, the conference service platform further includes one or more of an audio / video codec module, an audio watermark detection module, or an audio watermark embedding module. The audio / video transceiving module is configured to transceive network-encoded audio / video data with other devices (such as the conference terminal). The audio / video codec module is configured to encode the network-encoded audio / video data into audio / video data in an original format, and encode the audio / video data in the original format into network-format audio / video data. The original format is, for example, WAV format (a standard digital audio file format), and the network format is, for example, real-time transport protocol (RTP) format. The audio watermark detection module is configured to perform audio watermark detection on the audio data in the original format from the audio / video codec module, and the watermark detection result includes watermark signal strength and / or watermark information content, etc. The service message module is configured to send service messages, including the watermark detection result, etc., to the conference management server. The audio watermark embedding module is configured to embed audio watermark in the audio data in the original format from the audio / video codec module, to form audio data with watermark and send it to the audio / video codec module to be encoded into network-format audio network data.

[0083] The conference terminal includes an audio / video transceiving module, an audio / video codec module, an audio / video capturing module, and an audio / video playing module. Optionally, the conference terminal further includes an audio watermark embedding module. The audio / video transceiving module is configured to transceive network-encoded audio / video data with other devices (such as the conference service platform). The audio / video codec module is configured to encode the network-encoded audio / video data into audio / video data in an original format, and encode the audio / video data in the original format into network-format audio / video data, and encode the audio / video data in the original format after noise reduction processing such as eliminating playback sound. The audio / video capturing module is configured to capture audio / video data in the original format. The audio / video playing module is configured to play the decoded audio / video data in the original format. The audio watermark embedding module is configured to embed audio watermark in the audio data in the original format from the audio / video codec module, to form audio data with watermark for the audio / video playing module to play.

[0084] The conference management server comprises a service message module, a conference control module and a storage module. Optionally, the conference management server further comprises an audio watermark detection module. The service message module is configured to exchange service messages with the conference service platform, including watermark detection results, etc., and send conference status data to the conference management terminal, wherein the conference status data comprises audio watermark status of the terminal, such as whether the watermark is embedded effectively.

[0085] The conference management terminal comprises a service message module, a conference service module and a display module. Optionally, the conference management terminal further comprises an audio watermark detection module. The service message module is configured to receive conference status data from the conference management server. The conference service module is configured to receive conference status data from the service message module and display the conference status data through the display module. The display module is configured to display the conference status data, including audio watermark status of all terminals in the conference. The audio watermark detection module is configured to perform audio watermark detection on audio data in original format, and the watermark detection result comprises watermark signal strength and / or watermark information content, etc.

[0086] The method flow of the embodiments of the present application is described below.

[0087] FIG. 6 is a flow diagram of an audio watermark embedding method 600 according to an embodiment of the present application. As shown in FIG. 6, the method 600 comprises but is not limited to the following steps 601 to 603. Optionally, the method 600 further comprises the following steps 604 to 605 and / or steps 606 to 608. The method 600 can be applied to various computer devices with audio watermark detection capability, for example, the method 600 can be applied to the watermark detection device in the application scenario shown in FIG. 4, or can be applied to the conference service platform, the conference management server or the conference management terminal shown in FIG. 5.

[0088] Step 601: obtaining first original audio collected from an environment where a first terminal is located in a first time period, wherein the first original audio comprises first playing audio played by the first terminal in the first time period, and the first playing audio is configured to be embedded with a first watermark based on a first watermark embedding parameter value using a first watermark embedding algorithm.

[0089] Optionally, the first watermark in the first played audio is embedded by the first terminal or embedded by other devices. Taking a conference scenario as an example, the watermark in the audio played by the conference terminal can be embedded by the conference terminal side or can be embedded by the conference service platform. In an audio multi-stream scenario, the watermark in the audio played by the conference terminal is usually embedded by the conference terminal side. In addition, in the audio multi-stream scenario, the uplink audio encoding and downlink audio mixing of the conference terminal are both completed on the conference terminal side, and the conference service platform usually does not participate in the mixing and encoding of the conference audio, but only forwards the uplink audio data of the conference terminal to other conference terminals (usually only the audio data of several routes with the largest sound is forwarded).

[0090] The first played audio is configured to embed the first watermark based on the first watermark embedding algorithm and the first watermark embedding parameter value. It can be understood that the first played audio should be embedded with the first watermark based on the first watermark embedding algorithm and the first watermark embedding parameter value, but in actual application scenarios, due to some attacks on the audio watermark on the watermark embedding side, the first watermark cannot be embedded in the first played audio. For example, in the case where the watermark embedding is performed by the conference terminal, the conference software on the conference terminal can be tampered with, so that the conference terminal cannot perform watermark embedding before playing the audio.

[0091] Optionally, the first original audio further includes environmental audio generated in an environment where the first terminal is located. The environmental audio includes but is not limited to environmental noise audio and / or environmental human voice audio. The environmental noise audio includes, for example, low-frequency noise generated when devices such as air conditioners and fans in the local environment are running, and / or noise emitted by any sound-emitting device deployed in the local environment. The environmental human voice audio refers to the speech of a speaker in the local environment. The first original audio is actual audio collected from the environment where the first terminal is located within the first time period without de-noising processing. The environmental audio can be audio obtained by de-noising the first original audio after removing the first played audio in the first original audio. In a conference scenario, the environmental audio is audio that the first terminal needs to transmit to other conference terminals.

[0092] Optionally, the method provided by the embodiments of the present application can be executed by a conference service platform. The conference service platform can receive an original audio stream and an environmental audio stream sent by a first terminal. The environmental audio stream is an audio stream obtained by eliminating a played audio stream played by the first terminal in the original audio stream. The environmental audio stream is used to provide for a second terminal to play. The first terminal and the second terminal are both conference terminals. Then the conference service platform obtains a first original audio from the original audio stream and obtains environmental audio from the environmental audio stream.

[0093] Optionally, after receiving the environmental audio stream sent by the first terminal, the conference service platform embeds a second watermark in the environmental audio stream, and sends the environmental audio stream embedded with the second watermark to the second terminal for playing. Alternatively, the conference service platform sends the environmental audio stream to the second terminal, and the second terminal embeds a second watermark in the received environmental audio stream and plays it. The second watermark can be the same watermark as the first watermark, or can be a different watermark, and the embodiments of the present application do not limit this.

[0094] Optionally, after obtaining the environmental audio, unknown watermark detection can be further performed on the environmental audio. If there is an unknown watermark in the environmental audio, a first prompt information is output. The first prompt information is used to indicate that the environment where the first terminal is located has a risk of information leakage. Optionally, the unknown watermark detection can be performed on the environmental audio based on a known rule, for example, the environmental audio is detected for watermark according to a plurality of known watermark embedding algorithms respectively. Alternatively, the unknown watermark detection can be performed on the environmental audio based on an artificial intelligence (AI) algorithm, for example, the environmental audio is input into a pre-trained AI model to obtain an output result of the AI model, and the output result is used to indicate whether there is an unknown watermark in the environmental audio. The embodiments of the present application do not limit the implementation manner of detecting unknown watermarks.

[0095] Since there can be an internal attacker in the environment where the terminal is located who uses the audio watermark as a covert information channel for leaking information to the outside, in the embodiments of the present application, unknown watermark detection is performed on the environmental audio to determine whether there is an unknown watermark in the environmental audio, and a record prompt is performed in the case that there is an unknown watermark in the environmental audio, so as to realize the sensing capability of the air gap attack based on the unknown watermark, so that after the information leakage occurs, the post investigation and tracing can be performed. In the conference scenario, the conference service platform can perform watermark sampling detection on the environmental audio stream sent by each conference terminal in an uplink direction. If unknown watermark information is found, it can be fed back to the conference management platform for recording for subsequent tracing, thereby providing a mechanism for detecting whether there is a leakage of confidential information through the conference audio channel.

[0096] Optionally, after obtaining the first original audio, it can be further detected whether the first original audio is playback audio, and if the first original audio is playback audio, second prompt information is output. The second prompt information is used to indicate that there is a watermark embedding abnormal problem in the audio played by the first terminal. Optionally, whether the first original audio is playback audio can be determined according to the audio content of the first original audio, for example, it can be judged whether the audio content in the received original audio stream is repeated, and in the case that the original audio stream includes repeated audio content, it can be determined that the original audio stream comes from a counterfeit audio recorded or generated in advance, and in this case, it can be determined that the first original audio is playback audio. Or, the first original audio can be compared with the environmental audio to determine whether the first original audio is real sound, for example, the received original audio stream and the environmental audio stream can be content matched to judge whether the original audio stream includes the audio content in the environmental audio stream.

[0097] In the embodiments of the present application, whether the obtained first original audio is playback audio is detected to judge the authenticity of the first original audio, that is, to judge whether the first original audio is the actual audio collected from the environment where the first terminal is located in the first time period, and in the case that it is determined that the obtained first original audio is playback audio, a record prompt is given to facilitate subsequent investigation and tracing. In this way, the attack means that the attacker uses the audio recorded in advance and embedded with the first watermark to counterfeit the actual audio collected from the environment where the first terminal is located can be detected, so that whether the audio played by the first terminal is truly embedded with the watermark can be effectively monitored, the defense means against terminal software tampering and acoustic attack is strengthened, and the security and reliability of audio playing are improved. In a conference scenario, the conference service platform can perform playback detection on the original audio stream sent by each conference terminal, and if playback audio is found, it can be fed back to the conference management platform for recording for subsequent tracing, which provides a scheme for monitoring whether the conference terminal effectively executes the audio watermark, strengthens the defense means against conference software tampering and acoustic attack, and improves the confidence of the conference management party in the conference security state.

[0098] Step 602, extracting the watermark from the first original audio according to the first watermark embedding algorithm.

[0099] The embodiments of the present application take the first watermark embedding algorithm as a spread spectrum algorithm as an example to simply describe the watermark embedding and extraction process. For example, FIG. 7 is an implementation schematic diagram of an audio watermark embedding and extraction process provided by the embodiments of the present application. As shown in FIG. 7, the implementation process mainly includes four steps of watermark construction, carrier information processing, watermark embedding and carrier information restoration.

[0100] The first step: watermark construction. In the watermark construction process, the original watermark information (usually character information) needs to be encoded and converted, including character encoding, error correction encoding, and spread spectrum, etc. After encoding and converting the original watermark information, the original watermark information is converted into binary bit information. The binary bit information uniquely represents the watermark information for subsequent extraction and restoration of the watermark information.

[0101] The second step: carrier information processing. In the carrier information processing process, the carrier audio needs to be first processed by frame to obtain audio frames, and then fast Fourier transform (FFT) is used for each audio frame to convert the time domain signal of each audio frame into a frequency domain signal for subsequent watermark embedding. The length of the audio frame is determined based on the actual business delay requirement and the ability of the watermark algorithm. If the length of the audio frame is too long, it will affect the real-time performance, and if it is too small, it will affect the watermark embedding effect and robustness. In order to meet the time requirement of the algorithm in the embodiments of the present application, the length of the audio frame can be set to ten milliseconds.

[0102] The third step: watermark embedding. The watermark embedding step is the core step of the entire real-time audio watermark embedding algorithm. In the watermark embedding process, the binary bit information generated by the first step (watermark construction) is embedded bit by bit into the framed carrier (frequency domain signal of the audio frame), and the frequency domain signal of the audio frame is modified according to the value of the binary bit information to achieve the purpose of watermark embedding. This process needs to select a fixed frequency band as the modification content, and the frequency band selection cannot be too high or too low. Taking the carrier audio as human voice audio as an example, watermark information in a high frequency band will be filtered by audio playback equipment, and watermark information in a low frequency band is easy to overlap with noise signal frequency band, thereby affecting the watermark embedding effect. In the case of carrier audio being human voice audio, the watermark embedding frequency band can be selected to be consistent with the human voice frequency band, for example, selecting a frequency band of 5 kHz to 7 kHz for watermark embedding, so that the watermark information is not easy to be erased.

[0103] The fourth step: carrier information restoration. In the carrier information restoration process, each audio frame with embedded watermark information in the frequency domain signal is extracted and restored. First, the audio frame with embedded watermark information in the frequency domain signal needs to be inverse fast Fourier transformed (IFFT) to restore the frequency domain signal to the time domain signal. Then, all the audio frames with embedded watermark information in the time domain signal are merged to restore the audio data stream containing the watermark information.

[0104] Optionally, in the audio watermark extraction process for the illegally recorded audio data containing watermark information in the audio tracing process, the process is the inverse process of the watermark embedding process for the audio data, including four steps of carrier information processing, watermark extraction, watermark information synthesis and watermark information restoration. In the carrier information processing step, the audio data stream containing watermark information needs to be first processed by frame division, and the length of the audio frame obtained by frame division needs to be consistent with the length of the audio frame in which the watermark is embedded in the second step above, and then FFT is performed on a single audio frame to convert the time domain signal into a frequency domain signal for subsequent watermark extraction. In the watermark extraction step, the same frequency band as the embedding frequency band is selected to restore the watermark information, and the restored information is encoded bit information, which has a certain bit error ratio (BER) and needs to be corrected by an error correction algorithm. In the watermark information synthesis step, multiple original bit information obtained by repeating the watermark extraction step above is weighted and counted to synthesize a unique bit information, reducing the overall sample error caused by a single audio frame. In the watermark information restoration step, the bit information synthesized in the watermark information synthesis step is subjected to overall error checking and encoding restoration to obtain the original watermark information.

[0105] Step 603, if the first watermark extraction from the first original audio fails, embedding the first watermark in the first to-be-played audio of the first terminal based on the second watermark embedding parameter value using the first watermark embedding algorithm to obtain a second played audio, the second watermark embedding parameter value is different from the first watermark embedding parameter value, and the second played audio is used for the first terminal to play in a second time period.

[0106] The second time period is located after the first time period in time sequence. The first time period and the second time period are different audio collection time periods, for example, the second time period is the next audio collection time period of the first time period.

[0107] Optionally, the first original audio includes the first played audio played by the first terminal in the first time period and the environmental audio generated in the environment of the first terminal in the first time period. The first watermark embedding parameter value can be adjusted to obtain the second watermark embedding parameter value according to the audio signal characteristics of the environmental audio and the audio signal characteristics of the first to-be-played audio. The first to-be-played audio is a carrier audio. Optionally, the audio signal characteristics include the audio signal intensity in one or more frequency bands. Because the audio signal characteristics of different carrier audios are different, for example, the signal intensity of male audio in the low frequency band is usually greater than that in the high frequency band, and it is generally suitable to select the middle-low frequency band to embed the watermark, and the signal intensity of female audio in the high frequency band is usually greater than that in the low frequency band, and it is generally suitable to select the middle-low frequency band to embed the watermark, so the watermark embedding frequency band can be determined and adjusted according to the audio signal characteristics of the carrier audio.

[0108] In the embodiments of the present application, by extracting the watermark in the audio played by the terminal from the original audio collected from the environment where the terminal is located, in the case that the watermark cannot be successfully extracted from the original audio, the watermark embedded in the audio played by the terminal afterwards can be adjusted to adapt to the environment so as to be successfully extracted, thereby realizing adaptive environment adjustment of the audio watermark and improving the robustness of the audio watermark. In addition, some attacks on the audio watermark may occur on the terminal side, for example, in the case that the watermark embedding is performed by the terminal, the related software on the terminal may be tampered with, resulting in that the terminal cannot perform watermark embedding before playing the audio, or the terminal is subjected to acoustic attack when playing the audio, such as the existence of some acoustic devices in the environment where the terminal is located to make noise interference on the watermark embedding frequency band. By implementing the scheme provided in the embodiments of the present application, it can be perceived whether the audio played by the terminal is effectively embedded with the watermark, thereby improving the controllability of watermark embedding, realizing effective protection of the audio played by the terminal, and playing a deterrent role on the behavior of secretly recording and leaking the audio.

[0109] Optionally, the first watermark embedding parameter value includes a first watermark embedding strength. According to the audio signal feature of the environmental audio and the audio signal feature of the first to-be-played audio, the implementation manner of adjusting the first watermark embedding parameter value can include: if the difference between the audio signal strength of the first to-be-played audio and the audio signal strength of the environmental audio is greater than a signal strength threshold, increasing the first watermark embedding strength to obtain a second watermark embedding strength, and taking the second watermark embedding strength as the watermark embedding strength in the second watermark embedding parameter value.

[0110] The audio signal strength can be expressed in decibels. Optionally, the signal strength threshold can be 20 decibels.

[0111] In the embodiments of the present application, in the case that the watermark extraction from the original audio collected from the environment where the terminal is located fails, the watermark embedding strength can be adjusted first, and the watermark signal strength of the watermark information in the carrier audio is tried to be improved by increasing the watermark embedding strength, thereby improving the possibility of successful watermark extraction.

[0112] Optionally, in the case that the first watermark embedding parameter value further comprises a first watermark embedding frequency band, the above implementation manner can be alternatively implemented as: if a difference between the audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and the audio signal strength of the environment audio in the first watermark embedding frequency band is greater than a signal strength threshold, the first watermark embedding strength is increased. Optionally, if the difference between the audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and the audio signal strength of the environment audio in the first watermark embedding frequency band is less than or equal to the signal strength threshold, a second watermark embedding frequency band can be used as the watermark embedding frequency band in the second watermark embedding parameter value. Wherein, a difference between the audio signal strength of the first to-be-played audio in the second watermark embedding frequency band and the audio signal strength of the environment audio in the second watermark embedding frequency band is greater than the signal strength threshold.

[0113] In the embodiments of the present application, by comparing the audio signal strengths of the carrier audio and the environment audio in the current watermark embedding frequency band, it is determined whether the currently used watermark embedding frequency band is suitable for embedding the watermark. If the difference between the audio signal strengths of the carrier audio and the environment audio in the current watermark embedding frequency band is greater than the signal strength threshold, it indicates that the interference of the environment audio to the carrier audio in the current watermark embedding frequency band is relatively small, and the possibility of successful watermark extraction can be improved by increasing the watermark embedding strength. If the difference between the audio signal strengths of the carrier audio and the environment audio in the current watermark embedding frequency band is less than or equal to the signal strength threshold, it indicates that the interference of the environment audio to the carrier audio in the current watermark embedding frequency band is relatively large, and the effect of increasing the watermark embedding strength on the watermark extraction is small. At this time, the possibility of successful watermark extraction can be tried to improve by adjusting the watermark embedding frequency band. For example, the signal strength difference between the carrier audio and the environment audio in the 5 kilohertz-6 kilohertz frequency band (the first watermark embedding frequency band) is less than 20 decibels, and the signal strength difference between the carrier audio and the environment audio in the 6 kilohertz-7 kilohertz frequency band is greater than 20 decibels. At this time, the 6 kilohertz-7 kilohertz frequency band can be used as the second watermark embedding frequency band.

[0114] Optionally, in the case that the watermark in the audio played by the terminal is successfully extracted from the original audio collected from the environment where the terminal is located, the following steps 604 to 605 can be performed.

[0115] Step 604: If the first watermark is successfully extracted from the first original audio, and the watermark signal strength of the extracted first watermark exceeds the watermark signal strength range, the first watermark embedding strength is adjusted to obtain a third watermark embedding strength.

[0116] Optionally, the implementation of step 604 is as follows: if the first watermark is successfully extracted from the first original audio, and the watermark signal strength of the extracted first watermark is less than the minimum value of the watermark signal strength range, the first watermark embedding strength is increased. Or, if the first watermark is successfully extracted from the first original audio, and the watermark signal strength of the extracted first watermark is greater than the maximum value of the watermark signal strength range, the first watermark embedding strength is decreased. Optionally, the watermark signal strength can be judged according to the error code rate of the extracted watermark information. The lower the error code rate, the greater the watermark signal strength. Conversely, the higher the error code rate, the smaller the watermark signal strength.

[0117] Step 605: embedding the first watermark in the first terminal to-be-played audio based on the first watermark embedding parameter value by using the first watermark embedding algorithm to obtain a third to-be-played audio, the first watermark embedding parameter value including a third watermark embedding strength, and the third to-be-played audio being used for the first terminal to play in a second time period.

[0118] Optionally, the embedding method of the audio watermark can refer to the related content in step 603 described above, and the embodiments of the present application will not be repeated here.

[0119] In addition, if the watermark signal strength of the extracted first watermark is within the watermark signal strength range, the first watermark embedding strength can be kept unchanged, and the watermark can be continuously embedded in the terminal to-be-played audio based on the first watermark embedding parameter value.

[0120] In the embodiments of the present application, in the case that the watermark in the audio played by the terminal is successfully extracted from the original audio collected from the environment where the terminal is located, whether the watermark signal strength of the extracted watermark exceeds the watermark signal strength range is judged. If the watermark signal strength of the extracted watermark exceeds the watermark signal strength range, the watermark embedding strength is adjusted, so as to dynamically adjust the watermark signal strength, so that the watermark signal strength of the watermark extracted later can be within the watermark signal strength range, so as to avoid the problems that the watermark robustness is low due to too small watermark signal strength, or the audio playback quality is affected due to too large watermark signal strength affecting the watermark transparency, so as to better balance the robustness and imperceptibility of the audio watermark.

[0121] Optionally, if the watermark embedding parameter value is adjusted one or more times for the same watermark embedding algorithm, and the watermark in the audio played by the terminal is still not successfully extracted from the original audio collected from the environment where the terminal is located, the watermark embedding algorithm can be further adjusted. For example, after step 603 is performed, steps 606 to 608 can be further performed.

[0122] Step 606: obtaining second original audio collected from the environment where the first terminal is located in a second time period, the second original audio including a second to-be-played audio.

[0123] The implementation of this step 606 can refer to the above step 601, and the embodiments of the present application will not be repeated here.

[0124] Step 607, extracting the first watermark from the second original audio according to the first watermark embedding algorithm.

[0125] The implementation of this step 607 can refer to the above step 602, and the embodiments of the present application will not be repeated here.

[0126] Step 608, if the first watermark fails to be extracted from the second original audio, embedding the first watermark in the second to-be-played audio to be played by the first terminal by using a second watermark embedding algorithm to obtain a fourth to-be-played audio, the second watermark embedding algorithm being different from the first watermark embedding algorithm, and the fourth to-be-played audio being used for the first terminal to play in a third time period.

[0127] The third time period is located after the second time period in time sequence. Optionally, the embedding manner of the audio watermark can refer to the related content in the above step 603, and the embodiments of the present application will not be repeated here.

[0128] In the case that the watermark embedding parameter value is adjusted and the watermark still cannot be successfully extracted, the embodiments of the present application can further adjust the watermark embedding algorithm to attempt to embed the watermark by using different watermark embedding algorithms to improve the watermark extractability, thereby improving the flexibility of audio watermark embedding.

[0129] The implementation process of the above scheme provided by the embodiments of the present application will be exemplarily described below taking a conference scenario as an example.

[0130] 1. The conference management party convenes a security conference, and initially configures a default watermark embedding algorithm and a watermark embedding parameter value; a conference service platform generates watermark information corresponding to each conference terminal (for example, including a conference terminal A and a conference terminal B), and embeds the corresponding watermark information W in the uplink audio stream sent by the conference terminal A by using the default watermark embedding algorithm based on the watermark embedding parameter value, and then encodes the audio stream embedded with the watermark information into a network transmission format and sends it to the conference terminal B.

[0131] 2. The conference terminal B receives the audio stream and decodes and plays the audio stream, and the played audio stream carries the watermark information W.

[0132] 3. The conference terminal B acquires an original audio stream collected from the local environment, and performs noise reduction and anti-howling processing on the collected original audio stream to eliminate the audio stream played by the conference terminal B in the original audio stream, thereby obtaining an environmental audio stream.

[0133] 4. The conference terminal B sends one original audio stream and one environmental audio stream to the conference service platform, wherein the environmental audio stream is used as conference audio data to participate in conference audio mixing and is sent to other conference terminals.

[0134] 5. The conference service platform performs audio watermark sampling detection on the received original audio stream. If the watermark information W cannot be successfully extracted, the watermark embedding strength is increased or the watermark embedding frequency band or the watermark embedding algorithm is adjusted. If the watermark information W can be successfully extracted, if the watermark signal strength is weak, the watermark embedding strength is increased, and if the watermark signal strength is high, the watermark embedding strength is decreased. Then, the conference service platform applies the adjusted watermark embedding strength, watermark embedding frequency band or watermark embedding algorithm to subsequent watermark embedding of the audio stream sent by the conference terminal A to the conference terminal B, and continuously detects the original audio stream uploaded by the conference terminal B until the watermark signal strength reaches the expected effect.

[0135] 6. The conference service platform feeds back the audio watermark detection result to the conference management server, and the conference management server records the audio watermark detection states of all conference terminals and identifies the conference terminals whose audio watermark does not meet the expectation, and sends the identification to the conference management terminal in real time. The conference management terminal presents the real-time execution states of the audio watermark of all conference terminals in the conference.

[0136] In addition, the conference service platform can also perform watermark sampling detection on the environmental audio stream uploaded by each conference terminal. If unknown watermark information is found, it can be fed back to the conference management server for recording for subsequent tracing. And / or, the conference service platform can also perform playback detection on the original audio stream uploaded by each conference terminal. If a playback audio is found, it can be fed back to the conference management server for recording for subsequent tracing.

[0137] The order of the steps of the above-mentioned audio watermark embedding method provided by the embodiments of the present application can be adjusted appropriately, and the steps can also be increased or decreased according to the situation. Any person skilled in the art can easily think of changes within the technical range disclosed in the present application, which should be covered within the protection scope of the present application.

[0138] The software device of the embodiments of the present application is illustrated below.

[0139] For example, FIG. 8 is a structural schematic diagram of an audio watermark embedding device provided by the embodiments of the present application. As shown in FIG. 8, the audio watermark embedding device 800 includes but is not limited to an acquisition module 801, a watermark extraction module 802 and a watermark embedding module 803. Optionally, please continue to refer to FIG. 8, the audio watermark embedding device further includes one or more of an adjustment module 804, a receiving module 805, a watermark detection module 806, an output module 807 or an audio detection module 808.

[0140] The acquisition module 801 is configured to acquire first original audio collected from an environment where the first terminal is located in a first time period, and the first original audio includes first playing audio played by the first terminal in the first time period, and the first playing audio is configured to be embedded with a first watermark based on a first watermark embedding parameter value by using a first watermark embedding algorithm.

[0141] The watermark extraction module 802 is configured to extract the first watermark from the first original audio according to the first watermark embedding algorithm.

[0142] The watermark embedding module 803 is configured to, if the first watermark fails to be extracted from the first original audio, embed the first watermark in first to-be-played audio to be played by the first terminal to obtain second playing audio by using the first watermark embedding algorithm based on a second watermark embedding parameter value, the second watermark embedding parameter value is different from the first watermark embedding parameter value, and the second playing audio is used for playing by the first terminal in a second time period, and the second time period is located after the first time period in time sequence.

[0143] Optionally, the first watermark embedding parameter value includes a first watermark embedding strength and a first watermark embedding frequency band, the second watermark embedding parameter value includes a second watermark embedding strength and a second watermark embedding frequency band, the second watermark embedding parameter value is different from the first watermark embedding parameter value, including that the second watermark embedding strength is different from the first watermark embedding strength, and / or the second watermark embedding frequency band is different from the first watermark embedding frequency band.

[0144] Optionally, the first original audio further includes environmental audio generated in the environment where the first terminal is located in the first time period. The adjustment module 804 is configured to adjust the first watermark embedding parameter value based on an audio signal feature of the environmental audio and an audio signal feature of the first to-be-played audio to obtain the second watermark embedding parameter value.

[0145] Optionally, the first watermark embedding parameter value includes a first watermark embedding strength, and the adjustment module 804 is configured to, if a difference between an audio signal strength of the first to-be-played audio and an audio signal strength of the environmental audio is greater than a signal strength threshold, increase the first watermark embedding strength to obtain a second watermark embedding strength, and take the second watermark embedding strength as a watermark embedding strength in the second watermark embedding parameter value.

[0146] Optionally, the first watermark embedding parameter value further includes a first watermark embedding frequency band, and the adjustment module 804 is configured to, if a difference between an audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and an audio signal strength of the environmental audio in the first watermark embedding frequency band is greater than a signal strength threshold, increase the first watermark embedding strength.

[0147] Optionally, the adjusting module 804 is further configured to, if a difference between the audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and the audio signal strength of the environmental audio in the first watermark embedding frequency band is less than or equal to the signal strength threshold, take the second watermark embedding frequency band as the watermark embedding frequency band in the second watermark embedding parameter value, wherein a difference between the audio signal strength of the first to-be-played audio in the second watermark embedding frequency band and the audio signal strength of the environmental audio in the second watermark embedding frequency band is greater than the signal strength threshold.

[0148] Optionally, the audio watermark embedding apparatus 800 is applied to a conference service platform. The receiving module 805 is configured to receive an original audio stream and an environmental audio stream sent by a first terminal, the environmental audio stream being an audio stream obtained by eliminating a played audio stream played by the first terminal in the original audio stream, and the environmental audio stream being used to be played by a second terminal. The obtaining module 801 is configured to obtain first original audio from the original audio stream and environmental audio from the environmental audio stream.

[0149] Optionally, the watermark embedding module 803 is further configured to, after receiving the environmental audio stream sent by the first terminal, embed a second watermark in the environmental audio stream, and send the environmental audio stream embedded with the second watermark to the second terminal.

[0150] Optionally, the watermark detecting module 806 is configured to perform unknown watermark detection on the environmental audio. The output module 807 is configured to, if the unknown watermark exists in the environmental audio, output first prompt information, the first prompt information being used to indicate that an environment where the first terminal is located has a risk of information leakage.

[0151] Optionally, the adjusting module 804 is configured to, if the first watermark is successfully extracted from the first original audio and a watermark signal strength of the extracted first watermark is out of a watermark signal strength range, adjust the first watermark embedding strength to obtain a third watermark embedding strength. The watermark embedding module 803 is configured to embed the first watermark in the first to-be-played audio based on a third watermark embedding parameter value by using a first watermark embedding algorithm, to obtain third played audio, the third watermark embedding parameter value including the third watermark embedding strength, and the third played audio being used to be played by the first terminal in a second time period.

[0152] Optionally, the adjusting module 804 is configured to: if the first watermark is successfully extracted from the first original audio and a watermark signal strength of the extracted first watermark is less than a minimum value of the watermark signal strength range, increase the first watermark embedding strength; or if the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark is greater than a maximum value of the watermark signal strength range, decrease the first watermark embedding strength.

[0153] Optionally, the obtaining module 801 is further configured to obtain second original audio collected from the environment where the first terminal is located in a second time period, and the second original audio includes second playback audio. The watermark extraction module 802 is further configured to perform watermark extraction on the second original audio according to a first watermark embedding algorithm. The watermark embedding module 803 is further configured to, if the first watermark fails to be extracted from the second original audio, embed the first watermark in second to-be-played audio to be played by the first terminal by using a second watermark embedding algorithm to obtain fourth playback audio, the second watermark embedding algorithm being different from the first watermark embedding algorithm, and the fourth playback audio being used for the first terminal to play in a third time period, the third time period being located after the second time period in time sequence.

[0154] Optionally, the audio detection module 808 is configured to detect whether the first original audio is replayed audio. The output module 807 is configured to output second prompt information if the first original audio is replayed audio, the second prompt information being used to indicate that there is a watermark embedding abnormal problem in the audio played by the first terminal.

[0155] The obtaining module 801, the watermark extraction module 802, the watermark embedding module 803, the adjusting module 804, the receiving module 805, the watermark detection module 806, the output module 807, or the audio detection module 808 can be implemented by software or by hardware. For example, the implementation of the obtaining module 801 is described below. Similarly, the implementation of the watermark extraction module 802, the watermark embedding module 803, the adjusting module 804, the receiving module 805, the watermark detection module 806, the output module 807, or the audio detection module 808 can refer to the implementation of the obtaining module 801.

[0156] As an example of a software functional unit, the obtaining module 801 can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the obtaining module 801 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region, or can be distributed in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ), or can be distributed in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Generally, one region can include multiple AZs.

[0157] Likewise, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, usually one VPC is set in one region, and communication between two VPCs in the same region and between VPCs in different regions needs to set a communication gateway in each VPC to realize the interconnection between VPCs through the communication gateway.

[0158] As an example of a hardware functional unit, the obtaining module 801 can include at least one computing device, such as a server, etc. Alternatively, the obtaining module 801 can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. Among them, the above-mentioned PLD can be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0159] The multiple computing devices included in the obtaining module 801 can be distributed in the same region or in different regions. The multiple computing devices included in the obtaining module 801 can be distributed in the same AZ or in different AZs. Likewise, the multiple computing devices included in the obtaining module 801 can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs, etc.

[0160] It should be noted that in other embodiments, the obtaining module 801 can be used to perform any step of the audio watermark embedding method, the watermark extraction module 802 can be used to perform any step of the audio watermark embedding method, the watermark embedding module 803 can be used to perform any step of the audio watermark embedding method, the adjustment module 804 can be used to perform any step of the audio watermark embedding method, the receiving module 805 can be used to perform any step of the audio watermark embedding method, the watermark detection module 806 can be used to perform any step of the audio watermark embedding method, the output module 807 can be used to perform any step of the audio watermark embedding method, and the audio detection module 808 can be used to perform any step of the audio watermark embedding method.

[0161] The steps implemented by the acquisition module 801, the watermark extraction module 802, the watermark embedding module 803, the adjustment module 804, the receiving module 805, the watermark detection module 806, the output module 807, and the audio detection module 808 can be specified as needed, and the functions of the audio watermark embedding device can be implemented by implementing different steps in the audio watermark embedding method by the acquisition module 801, the watermark extraction module 802, the watermark embedding module 803, the adjustment module 804, the receiving module 805, the watermark detection module 806, the output module 807, or the audio detection module 808.

[0162] The conference system of the embodiments of the present application is described below.

[0163] The embodiments of the present application provide a conference system, which includes a conference service platform and a plurality of conference terminals, and the plurality of conference terminals are communicatively connected through the conference service platform. The plurality of conference terminals include a first conference terminal. For example, the first conference terminal can be the terminal in FIG. 4, and the conference service platform can be the watermark detection device in FIG. 4. Alternatively, the first conference terminal can be the conference terminal in FIG. 5, and the conference service platform can be the conference service platform in FIG. 5. The conference service platform is configured to execute the method 600.

[0164] For example, the first conference terminal is configured to send, to the conference service platform, first original audio collected from an environment where the first conference terminal is located in a first time period, and the first original audio includes first playback audio played by the first conference terminal in the first time period, and the first playback audio is configured to be embedded with a first watermark based on a first watermark embedding algorithm and a first watermark embedding parameter value. The conference service platform is configured to perform watermark extraction on the first original audio according to the first watermark embedding algorithm, and if the first watermark fails to be extracted from the first original audio, embed the first watermark in first to-be-played audio to be played by the first conference terminal in a second time period based on a second watermark embedding parameter value by using the first watermark embedding algorithm to obtain second playback audio, and the second watermark embedding parameter value is different from the first watermark embedding parameter value. The first conference terminal is configured to play the second playback audio in the second time period, and the second time period is located after the first time period in time sequence.

[0165] The hardware device of the embodiments of the present application is described below.

[0166] For example, FIG. 9 is a structural schematic diagram of a computing device 900 provided by the embodiments of the present application. As shown in FIG. 9, the computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate through the bus 902. The computing device 900 can be a server or a terminal device. It should be understood that the number of processors and memories in the computing device 900 is not limited by the embodiments of the present application.

[0167] The bus 902 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, or the like. For ease of representation, only one line is shown in FIG. 9, but it does not mean that there is only one bus or only one type of bus. The bus 902 can include a path for transmitting information between the components (e.g., the memory 906, the processor 904, the communication interface 908) of the computing device 900.

[0168] The processor 904 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), or the like.

[0169] The memory 906 can include a volatile memory (e.g., a random access memory (RAM)) and / or a non-volatile memory (e.g., a read-only memory (ROM), a floppy disk, a hard disk, or a solid state drive (SSD)).

[0170] The memory 906 stores executable program codes, and the processor 904 executes the executable program codes to implement the functions of the aforementioned modules, including but not limited to the obtaining module, the watermark extraction module, and the watermark embedding module, respectively, so as to implement the audio watermark embedding method, such as the method 600 described above. That is, the memory 906 stores instructions for executing the audio watermark embedding method.

[0171] Alternatively, the memory 906 stores executable program codes, and the processor 904 executes the executable program codes to implement the functions of the aforementioned audio watermark embedding apparatus 800, so as to implement the audio watermark embedding method. That is, the memory 906 stores instructions for executing the audio watermark embedding method.

[0172] The communication interface 908 enables communication among the computing device 900 and other devices or communication networks using, for example but not necessarily, a transceiver module such as a network interface card, a transceiver, and the like.

[0173] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.

[0174] For example, FIG. 10 is a structural schematic diagram of a computing device cluster according to an embodiment of the present application. As shown in FIG. 10, the computing device cluster includes at least one computing device 900. The memory 906 in one or more computing devices 900 in the computing device cluster can store the same instructions for performing the audio watermark embedding method.

[0175] In some possible implementations, the memory 906 in one or more computing devices 900 in the computing device cluster can also respectively store partial instructions for performing the audio watermark embedding method. In other words, the combination of one or more computing devices 900 can collectively execute the instructions for performing the audio watermark embedding method.

[0176] It should be noted that the memory 906 in different computing devices 900 in the computing device cluster can store different instructions for respectively performing partial functions of the audio watermark embedding apparatus. That is, the instructions stored in the memory 906 in different computing devices 900 can implement the functions of one or more of the aforementioned modules including but not limited to the obtaining module, the watermark extraction module, and the watermark embedding module.

[0177] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. For example, FIG. 11 is a structural schematic diagram of another computing device cluster according to an embodiment of the present application. As shown in FIG. 11, two computing devices 900A and 900B are connected through a network. Specifically, the communication interface in each computing device is connected to the network. In this type of possible implementation, the memory 906 in the computing device 900A stores instructions for performing the functions of the obtaining module. Meanwhile, the memory 906 in the computing device 900B stores instructions for performing the functions of the watermark extraction module and the watermark embedding module.

[0178] The connection manner between the computing device clusters shown in FIG. 11 can be that the audio watermark embedding method provided by the embodiments of the present application needs to transmit data and calculate data, and therefore the functions implemented by the watermark extraction module and the watermark embedding module are performed by the computing device 900B.

[0179] It should be understood that the functions of the computing device 900A shown in FIG. 11 can also be completed by multiple computing devices 900. Similarly, the functions of the computing device 900B can also be completed by multiple computing devices 900.

[0180] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection manners of the computing device clusters shown in FIG. 10 and FIG. 11. The difference is that the memory 906 in one or more computing devices 900 in the computing device cluster can store the same instructions for performing the audio watermark embedding method.

[0181] In some possible implementation manners, the memory 906 in one or more computing devices 900 in the computing device cluster can also respectively store partial instructions for performing the audio watermark embedding method. In other words, the combination of one or more computing devices 900 can collectively execute the instructions for performing the audio watermark embedding method.

[0182] It should be noted that the memory 906 in different computing devices 900 in the computing device cluster can store different instructions for performing part of the functions of the audio watermark embedding apparatus 800. That is, the instructions stored in the memory 906 in different computing devices 900 can implement the functions of the audio watermark embedding apparatus 800.

[0183] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the audio watermark embedding method.

[0184] The embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium contains instructions, which instruct the computing device to perform the audio watermark embedding method.

[0185] Finally, it should be noted that: the above examples are used to illustrate the technical solutions of the present application, but not limited to them; although the present application is described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio watermark embedding method characterized by, The method comprises: acquiring first original audio collected from an environment where a first terminal is located in a first time period, the first original audio comprising first playing audio played by the first terminal in the first time period, the first playing audio being configured to be embedded with a first watermark based on a first watermark embedding parameter value using a first watermark embedding algorithm; extracting the first watermark from the first original audio according to the first watermark embedding algorithm; if the first watermark fails to be extracted from the first original audio, embedding the first watermark in first to-be-played audio to be played by the first terminal based on a second watermark embedding parameter value using the first watermark embedding algorithm to obtain second playing audio, the second watermark embedding parameter value being different from the first watermark embedding parameter value, and the second playing audio being used for playing by the first terminal in a second time period, the second time period being located after the first time period in time sequence.

2. The method of claim 1, wherein, The first watermark embedding parameter value comprises a first watermark embedding strength and a first watermark embedding frequency band, the second watermark embedding parameter value comprises a second watermark embedding strength and a second watermark embedding frequency band, and the second watermark embedding parameter value is different from the first watermark embedding parameter value, including that the second watermark embedding strength is different from the first watermark embedding strength, and / or the second watermark embedding frequency band is different from the first watermark embedding frequency band.

3. The method according to claim 1 or 2, characterized in that, The first original audio further comprises environmental audio generated in the environment where the first terminal is located in the first time period, and the method further comprises: adjusting the first watermark embedding parameter value based on an audio signal feature of the environmental audio and an audio signal feature of the first to-be-played audio to obtain the second watermark embedding parameter value.

4. The method of claim 3, wherein, The first watermark embedding parameter value comprises a first watermark embedding strength, and adjusting the first watermark embedding parameter value based on the audio signal feature of the environmental audio and the audio signal feature of the first to-be-played audio comprises: if a difference between the audio signal strength of the first to-be-played audio and the audio signal strength of the environmental audio is greater than a signal strength threshold, increasing the first watermark embedding strength to obtain a second watermark embedding strength, and taking the second watermark embedding strength as the watermark embedding strength in the second watermark embedding parameter value.

5. The method of claim 4, wherein, The first watermark embedding parameter value further comprises a first watermark embedding frequency band, and if the difference between the audio signal strength of the first to-be-played audio and the audio signal strength of the environmental audio is greater than the signal strength threshold, increasing the first watermark embedding strength comprises: if a difference between the audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and the audio signal strength of the environmental audio in the first watermark embedding frequency band is greater than the signal strength threshold, increasing the first watermark embedding strength.

6. The method of claim 5, wherein, Adjusting the first watermark embedding parameter value based on the audio signal feature of the environmental audio and the audio signal feature of the first to-be-played audio further comprises: If a difference between the audio signal strength of the first to-be-played audio in the second watermark embedding frequency band and the audio signal strength of the environmental audio in the second watermark embedding frequency band is less than or equal to the signal strength threshold, a second watermark embedding frequency band is taken as a watermark embedding frequency band in the second watermark embedding parameter value, wherein a difference between the audio signal strength of the first to-be-played audio in the second watermark embedding frequency band and the audio signal strength of the environmental audio in the second watermark embedding frequency band is greater than the signal strength threshold.

7. The method according to any one of claims 3 to 6, characterized in that, The method is applied to a conference service platform, and the method further includes: receiving an original audio stream and an environmental audio stream sent by the first terminal, the environmental audio stream being an audio stream obtained by eliminating a played audio stream played by the first terminal in the original audio stream, and the environmental audio stream being used to provide the second terminal for playing; obtaining the first original audio from the original audio stream and obtaining the environmental audio from the environmental audio stream.

8. The method of claim 7, wherein, After receiving the environmental audio stream sent by the first terminal, the method further includes: embedding a second watermark in the environmental audio stream and sending the environmental audio stream embedded with the second watermark to the second terminal.

9. The method according to any one of claims 3 to 8, characterized in that, The method further includes: performing unknown watermark detection on the environmental audio; if there is an unknown watermark in the environmental audio, outputting first prompt information, the first prompt information being used to indicate that there is a risk of information leakage in the environment where the first terminal is located.

10. The method according to any one of claims 1 to 9, characterized in that, The first watermark embedding parameter value includes a first watermark embedding strength, and the method further includes: if the first watermark is successfully extracted from the first original audio and a watermark signal strength of the extracted first watermark exceeds a watermark signal strength range, adjusting the first watermark embedding strength to obtain a third watermark embedding strength; embedding the first watermark in the first to-be-played audio based on a third watermark embedding parameter value by using the first watermark embedding algorithm to obtain a third played audio, the third watermark embedding parameter value including the third watermark embedding strength, and the third played audio being used for the first terminal to play in the second time period.

11. The method of claim 10, wherein, The if the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark exceeds the watermark signal strength range, the first watermark embedding strength is adjusted, including: if the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark is less than a minimum value of the watermark signal strength range, the first watermark embedding strength is increased; or if the first watermark is successfully extracted from the first original audio and the watermark signal strength of the extracted first watermark is greater than a maximum value of the watermark signal strength range, the first watermark embedding strength is decreased.

12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: obtaining second original audio collected from the environment where the first terminal is located in the second time period, the second original audio including the second played audio; performing watermark extraction on the second original audio according to the first watermark embedding algorithm; and outputting the second played audio to the first terminal. If the first watermark fails to be extracted from the second original audio, a second watermark embedding algorithm is used to embed the first watermark in second to-be-played audio to be played by the first terminal, to obtain fourth to-be-played audio, the second watermark embedding algorithm being different from the first watermark embedding algorithm, and the fourth to-be-played audio being used for the first terminal to play in a third time period, the third time period being located after the second time period in time sequence.

13. The method according to any one of claims 1 to 12, characterized in that, The method further comprises: detecting whether the first original audio is playback audio; if the first original audio is playback audio, outputting second prompt information, the second prompt information being used to indicate that there is a watermark embedding abnormal problem in audio played by the first terminal.

14. An audio watermark embedding apparatus characterized by comprising: The device comprises: an acquisition module configured to acquire first original audio collected from an environment in which a first terminal is located in a first time period, the first original audio comprising first played audio played by the first terminal in the first time period, the first played audio being configured to have a first watermark embedded therein by using a first watermark embedding algorithm based on a first watermark embedding parameter value; a watermark extraction module configured to extract a watermark from the first original audio according to the first watermark embedding algorithm; a watermark embedding module configured to, if the first watermark fails to be extracted from the first original audio, embed the first watermark in first to-be-played audio to be played by the first terminal by using the first watermark embedding algorithm based on a second watermark embedding parameter value, to obtain second to-be-played audio, the second watermark embedding parameter value being different from the first watermark embedding parameter value, and the second to-be-played audio being used for the first terminal to play in a second time period, the second time period being located after the first time period in time sequence.

15. The apparatus of claim 14, wherein, The first watermark embedding parameter value comprises a first watermark embedding strength and a first watermark embedding frequency band, the second watermark embedding parameter value comprises a second watermark embedding strength and a second watermark embedding frequency band, and the second watermark embedding parameter value is different from the first watermark embedding parameter value, including that the second watermark embedding strength is different from the first watermark embedding strength, and / or the second watermark embedding frequency band is different from the first watermark embedding frequency band.

16. The apparatus of claim 14 or 15, wherein, The first original audio further comprises environmental audio generated in the environment in which the first terminal is located in the first time period, and the device further comprises an adjustment module. The adjustment module is configured to adjust the first watermark embedding parameter value according to an audio signal feature of the environmental audio and an audio signal feature of the first to-be-played audio, to obtain the second watermark embedding parameter value.

17. The apparatus of claim 16, wherein, The first watermark embedding parameter value comprises a first watermark embedding strength, and the adjustment module is configured to: if a difference between an audio signal strength of the first to-be-played audio and an audio signal strength of the environmental audio is greater than a signal strength threshold, increase the first watermark embedding strength to obtain a second watermark embedding strength, and use the second watermark embedding strength as a watermark embedding strength in the second watermark embedding parameter value.

18. The apparatus of claim 17, wherein, The first watermark embedding parameter value further comprises a first watermark embedding frequency band, and the adjustment module is configured to: if a difference between the audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and the audio signal strength of the environmental audio in the first watermark embedding frequency band is greater than the signal strength threshold, increasing the first watermark embedding strength.

19. The apparatus of claim 18, wherein, The adjustment module is configured to: if a difference between the audio signal strength of the first to-be-played audio in the first watermark embedding frequency band and the audio signal strength of the environmental audio in the first watermark embedding frequency band is less than or equal to the signal strength threshold, taking a second watermark embedding frequency band as a watermark embedding frequency band in the second watermark embedding parameter value, wherein a difference between the audio signal strength of the first to-be-played audio in the second watermark embedding frequency band and the audio signal strength of the environmental audio in the second watermark embedding frequency band is greater than the signal strength threshold.

20. The apparatus of any one of claims 16 to 19, wherein, The device is applied to a conference service platform, and the device further includes a receiving module. The receiving module is configured to receive an original audio stream and an environmental audio stream sent by the first terminal, the environmental audio stream being an audio stream obtained by eliminating a played audio stream played by the first terminal from the original audio stream, and the environmental audio stream being used to provide the second terminal for playing. The obtaining module is configured to obtain the first original audio from the original audio stream and obtain the environmental audio from the environmental audio stream.

21. The device of claim 20, wherein The watermark embedding module is further configured to embed a second watermark in the environmental audio stream after receiving the environmental audio stream sent by the first terminal, and send the environmental audio stream embedded with the second watermark to the second terminal.

22. The apparatus of any one of claims 16 to 21, wherein, The device further includes a watermark detection module and an output module. The watermark detection module is configured to perform unknown watermark detection on the environmental audio. The output module is configured to output first prompt information if the unknown watermark exists in the environmental audio, the first prompt information being used to indicate that the environment where the first terminal is located has a risk of information leakage.

23. The apparatus of any one of claims 14 to 22, wherein, The first watermark embedding parameter value includes a first watermark embedding strength, and the device further includes an adjustment module. The adjustment module is configured to, if the first watermark is successfully extracted from the first original audio and a watermark signal strength of the extracted first watermark exceeds a watermark signal strength range, adjust the first watermark embedding strength to obtain a third watermark embedding strength. The watermark embedding module is configured to embed the first watermark in the first to-be-played audio based on a third watermark embedding parameter value by using the first watermark embedding algorithm to obtain a third played audio, the third watermark embedding parameter value including the third watermark embedding strength, and the third played audio being used for the first terminal to play in the second time period.

24. The apparatus of claim 23, wherein, The adjustment module is configured to: if the first watermark is successfully extracted from the first original audio and a watermark signal strength of the extracted first watermark is less than a minimum value of the watermark signal strength range, increase the first watermark embedding strength; or If the first watermark is successfully extracted from the first original audio, and a watermark signal strength of the extracted first watermark is greater than a maximum value of the watermark signal strength range, the first watermark embedding strength is decreased.

25. The apparatus of any of claims 14 to 24, wherein, The obtaining module is further configured to obtain second original audio collected from an environment in which the first terminal is located in a second time period, the second original audio including the second playback audio; The watermark extracting module is further configured to extract a watermark from the second original audio according to the first watermark embedding algorithm. The watermark embedding module is further configured to, if the first watermark fails to be extracted from the second original audio, embed the first watermark in second to-be-played audio to be played by the first terminal using a second watermark embedding algorithm to obtain fourth playback audio, the second watermark embedding algorithm being different from the first watermark embedding algorithm, and the fourth playback audio being configured to be played by the first terminal in a third time period, the third time period being located after the second time period in a time sequence.

26. The apparatus of any one of claims 14 to 25, wherein, The apparatus further includes an audio detection module and an output module. The audio detection module is configured to detect whether the first original audio is a replayed audio. The output module is configured to output second prompt information if the first original audio is a replayed audio, the second prompt information being configured to indicate that there is a watermark embedding abnormality problem in audio played by the first terminal.

27. A conferencing system, characterized by The apparatus includes: a conference service platform and a plurality of conference terminals, the plurality of conference terminals being communicatively connected through the conference service platform, and the plurality of conference terminals including a first conference terminal; the first conference terminal being configured to send, to the conference service platform, first original audio collected from an environment in which the first conference terminal is located in a first time period, the first original audio including first playback audio played by the first conference terminal in the first time period, the first playback audio being configured to have a first watermark embedded therein based on a first watermark embedding parameter value using a first watermark embedding algorithm; the conference service platform being configured to extract a watermark from the first original audio according to the first watermark embedding algorithm, and, if the first watermark fails to be extracted from the first original audio, embed the first watermark in first to-be-played audio to be played by the first conference terminal based on a second watermark embedding parameter value using the first watermark embedding algorithm to obtain second playback audio, the second watermark embedding parameter value being different from the first watermark embedding parameter value; the first conference terminal being configured to play the second playback audio in a second time period, the second time period being located after the first time period in a time sequence.

28. A cluster of computing devices, characterized in that, The apparatus includes at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device being configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of any of claims 1 to 13.

29. A computer program product comprising instructions, wherein: The instructions, when executed by the cluster of computing devices, cause the cluster of computing devices to perform the method of any of claims 1 to 13.

30. A computer-readable storage medium, characterized in that, comprising computer program instructions which, when executed by a cluster of computing devices, cause the cluster of computing devices to perform a method as claimed in any of claims 1 to 13.

Citation Information

Patent Citations

  • Strong-robustness watermark embedding and extraction method for original remote sensing images

    CN106408497A

  • Audio watermarking method and system capable of resisting random cutting and rerecording

    CN113506580A

  • Watermark audio generation method and device, electronic equipment and storage medium

    CN118098250A

  • Re-embedding of watermarks in multimedia signals

    CN1659652A

  • Device for embedding digital watermark

    JP2004235953A