Methods and devices for identifying sound watermarks
By using a voice watermark recognition device in a remote conferencing system and setting an encoding threshold based on the noise environment, the problem of decreased watermark recognition accuracy caused by noise interference in remote conferencing is solved, and accurate watermark recognition is achieved under different noise environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2026-04-03
AI Technical Summary
In remote meetings, noise interference with the audio signal reduces the accuracy of watermark recognition and affects the user's voice component in the audio signal along the call transmission path.
By using a sound watermark recognition device in conference terminals and cloud servers, different encoding thresholds are set according to the noise of the transmission environment to identify sound watermark signals in synthetic sound signals. The device includes a memory and a processor, and adapts to different noise environments by eliminating noise interference and determining the encoding thresholds.
It improves the accuracy of voice watermark recognition, can adapt to changing noise interference environments, and ensures accurate recognition of watermark signals under different noise conditions.
Smart Images

Figure CN116137152B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a sound signal processing technology, and more particularly to a method and device for identifying sound watermarks. Background Technology
[0002] Remote conferencing allows people in different locations or spaces to converse, and the related equipment, protocols, and applications are quite mature. It is worth noting, however, that some real-time conferencing programs may synthesize audio signals and sound watermarks to identify the callers.
[0003] Inevitably, if the audio signal is interfered with by noise, the accuracy of the receiver in judging the watermark will decrease, which in turn will affect the user's voice component in the audio signal on the call transmission path. Summary of the Invention
[0004] The present invention relates to a method and device for recognizing sound watermarks. The recognized sound watermark signal results can be effectively set with different encoding thresholds according to the noise of the transmission environment to improve the accuracy of sound watermark recognition.
[0005] According to embodiments of the present invention, the sound watermark identification method is applicable to conference terminals. The sound watermark identification method includes (but is not limited to) the following steps: receiving a synthesized sound signal via a network. This synthesized sound signal includes a sound watermark signal. The sound watermark signal is generated by offsetting the phase of a reflected sound signal based on a watermark identifier. This reflected sound signal is a sound signal obtained by recording sound emitted from a simulated sound source after reflection by an external object and recorded by a microphone. Noise interference transmitted via the network in the synthesized sound signal is determined based on a reflection cancellation sound signal. The reflection cancellation sound signal is a sound signal in the synthesized sound signal whose watermark identifier is one or more codes. An encoding threshold is determined based on the noise interference. The encoding threshold includes a first threshold and a second threshold. The noise interference corresponding to the first threshold is lower than the noise interference corresponding to the second threshold. The first threshold is greater than the second threshold. The sound watermark signal in the synthesized sound signal is identified based on the encoding threshold.
[0006] According to an embodiment of the present invention, a sound watermark identification device includes (but is not limited to) a memory and a processor. The memory stores program code. The processor is coupled to the memory. The processor is configured to load and execute the program code to perform the following steps: receiving a synthesized sound signal via a network. This synthesized sound signal includes a sound watermark signal. The sound watermark signal is generated by offsetting the phase of a reflected sound signal based on a watermark identifier. This reflected sound signal is a sound signal obtained by recording sound emitted from a simulated sound source after reflection by an external object and recorded by a microphone. Determining noise interference transmitted via the network in the synthesized sound signal based on a reflection cancellation sound signal. The reflection cancellation sound signal is a sound signal in the synthesized sound signal whose watermark identifier is one or more codes to eliminate the sound watermark signal in the sound watermark signal. Determining an encoding threshold based on noise interference. The encoding threshold includes a first threshold and a second threshold. The noise interference corresponding to the first threshold is lower than the noise interference corresponding to the second threshold. The first threshold is greater than the second threshold. Identifying the sound watermark signal in the synthesized sound signal based on the encoding threshold.
[0007] According to the embodiment of the present invention, the sound watermark identification method and identification device, for a sound watermark signal generated based on a reflected sound signal, determines noise interference by eliminating sound watermark signals with different codes, and determines a corresponding encoding threshold for the estimated noise interference. This allows for adaptation to varying noise interference. Attached Figure Description
[0008] The accompanying drawings are included to further illustrate the invention, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0009] Figure 1 This is a schematic diagram of a conference call system according to an embodiment of the present invention;
[0010] Figure 2 This is a flowchart of a sound watermark recognition method according to an embodiment of the present invention;
[0011] Figure 3 This is a schematic diagram illustrating virtual reflection conditions according to an embodiment of the present invention;
[0012] Figure 4 This is a flowchart of a method for generating an encoding threshold according to an embodiment of the present invention;
[0013] Figure 5 This is a flowchart illustrating the determination of the encoding threshold according to an embodiment of the present invention;
[0014] Figure 6 This is a flowchart illustrating the determination of the encoding threshold according to another embodiment of the present invention;
[0015] Figure 7This is a flowchart of identifying a sound watermark signal according to an embodiment of the present invention.
[0016] Explanation of icon numbers
[0017] 10, 20: Conference terminals;
[0018] 50: Cloud server;
[0019] 11, 21: Radio;
[0020] 13, 23: Speakers;
[0021] 15, 25, 55: Communication transceivers;
[0022] 17, 27, 57: Memory;
[0023] 19, 29, 59: Processors;
[0024] 70: Voice watermark recognition device;
[0025] S210~S240, S410~S450, S510~S530, S610~S660: Steps;
[0026] S Rx : Receive audio signals during calls;
[0027] S Tx : Transmitting audio signals during a call;
[0028] S WM : Audio watermark signal;
[0029] S Rx +S WM Embedded watermark signal;
[0030] S' Rx 、S” Rx :Reflecting sound signals;
[0031] W: wall;
[0032] d s d w :distance;
[0033] SS: sound source;
[0034] W E :Watermark identifier;
[0035] S A Synthesized sound signals;
[0036] Preprocess the audio signal;
[0037] sB- First sound signal;
[0038] s B+ Second sound signal;
[0039] The third sound signal;
[0040] Fourth sound signal;
[0041] s C The fifth sound signal;
[0042] The sixth sound signal;
[0043] Correlation;
[0044] Th D , Encoding threshold. Detailed Implementation
[0045] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same component reference numerals are used in the drawings and description to denote the same or similar parts.
[0046] Figure 1 This is a schematic diagram of a conference call system 1 according to an embodiment of the present invention. Please refer to... Figure 1 The voice communication system 1 includes, but is not limited to, conference terminals 10 and 20 and cloud server 50.
[0047] Conference terminals 10 and 20 can be landline phones, mobile phones, VoIP phones, tablet computers, desktop computers, laptops, or smart speakers.
[0048] The conference terminal 10 includes (but is not limited to) a microphone 11, a speaker 13, a communication transceiver 15, a memory 17, and a processor 19.
[0049] Microphone 11 can be a dynamic, condenser, or electret condenser microphone, or it can be a combination of other electronic components, analog-to-digital converters, filters, and audio processors that receive sound waves (e.g., human voices, ambient sounds, machine sounds, etc.) and convert them into sound signals. In one embodiment, microphone 11 is used to pick up / record the speaker's voice to obtain a received audio signal of the conversation. In some embodiments, this received audio signal may include the speaker's voice, the sound emitted by speaker 13, and / or other ambient sounds.
[0050] The speaker 13 can be a loudspeaker or an amplifier. In one embodiment, the speaker 13 is used to play sound.
[0051] The communication transceiver 15 may be a transceiver supporting wired networks such as Ethernet, fiber optic networks, or cables (which may include (but is not limited to) components such as connection interfaces, signal converters, and communication protocol processing chips), or it may be a transceiver supporting wireless networks such as Wi-Fi, fourth-generation (4G), fifth-generation (5G), or later-generation mobile networks (which may include (but is not limited to) components such as antennas, digital-to-analog / analog-to-digital converters, and communication protocol processing chips). In one embodiment, the communication transceiver 15 is used to transmit or receive data.
[0052] The memory 17 can be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or similar component. In one embodiment, the memory 17 is used to store program code, software modules, configuration settings, data (e.g., sound signals, watermark identifiers, or sound watermark signals), or files.
[0053] Processor 19 is coupled to receiver 11, speaker 13, transceiver 15, and memory 17. Processor 19 may be a central processing unit (CPU), graphics processing unit (GPU), or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), programmable controller, field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or other similar components or combinations thereof. In one embodiment, processor 19 is used to execute all or part of the operations of its associated conference terminal 10, and can load and execute software modules, files, and data stored in memory 17.
[0054] The conference terminal 20 includes (but is not limited to) a microphone 21, a speaker 23, a communication transceiver 25, a memory 27, and a processor 29. The implementation and function of the microphone 21, speaker 23, communication transceiver 25, memory 27, and processor 29 can be found in the foregoing description of the microphone 11, speaker 13, communication transceiver 15, memory 17, and processor 19, and will not be repeated here. The microphone 21 is used to receive reflected sound signals and transmit them via the communication transceiver 25 to the processor 59 of the cloud server 50.
[0055] The cloud server 50 is directly or indirectly connected to the conference terminals 10 and 20 via a network. The cloud server 50 can be a computer system, a server, or a signal processing device. In one embodiment, the conference terminals 10 and 20 can also serve as the cloud server 50. In another embodiment, the cloud server 50 can serve as an independent cloud server, distinct from the conference terminals 10 and 20. In some embodiments, the cloud server 50 includes (but is not limited to) the same or similar communication transceiver 55, memory 57, and processor 59, and the implementation methods and functions of the components will not be described in detail.
[0056] In one embodiment, the sound watermark identification device 70 may be a conference terminal 10, 20 and / or a cloud server 50. The sound watermark identification device 70 is used to identify sound watermark signals, which will be described in detail in subsequent embodiments.
[0057] The methods described in the embodiments of the present invention will be explained below in conjunction with the various devices, components, and modules in the conference communication system 1. The various processes of this method may be adjusted according to the implementation situation, and are not limited thereto.
[0058] It should also be noted that, for ease of explanation, the same components can perform the same or similar operations, and will not be described again. For example, the processor 19 of the conference terminal 10, the processor 29 of the conference terminal 20, and / or the processor 59 of the cloud server 50 can all implement the same or similar methods of the embodiments of the present invention.
[0059] Figure 2 This is a flowchart of a sound watermark recognition method according to an embodiment of the present invention. Please refer to... Figure 2 The processor 19 receives the synthesized sound signal S via the network. A (Step S210). Specifically, assume that conference terminals 10 and 20 establish a conference call. For example, the conference is established through video conferencing software, voice call software, or by making a phone call, and the speaker can begin speaking. After the microphone 21 records / receives the audio, the processor 29 can obtain the received audio signal S. Rx This call receives audio signal S. RxThe audio content (which may also include ambient sound or other noise) of the speaker corresponding to the conference terminal 20. The processor 29 of the conference terminal 20 can transmit and receive the call audio signal S via the communication transceiver 25 (i.e., via a network interface). Rx In some embodiments, the call receives an audio signal S. Rx It may have undergone echo cancellation, noise filtering, and / or other audio signal processing.
[0060] Next, the processor 59 of the cloud server 50 receives the call reception audio signal S from the conference terminal 20 via the communication transceiver 55. Rx Processor 59 determines the virtual reflection conditions and the received audio signal S during the call. Rx Generate reflected sound signal S' Rx Specifically, a typical echo cancellation algorithm can adaptively eliminate components belonging to the reference signal in the audio signals received by receivers 11 and 21 from the outside (e.g., the call received audio signal S in the call receiving path). Rx The sound recorded by the microphones 11 and 21 includes the shortest path from the loudspeakers 13 and 23 to the microphones 11 and 21, as well as different reflection paths from the environment (i.e., the paths formed by sound reflected from external objects). The location of the reflection affects the time delay and attenuation of the sound signal. In addition, the reflected sound signal may also come from different directions, which can lead to phase shift.
[0061] In one embodiment, the processor 59 can determine the reflected sound signal S' based on the positional relationship. Rx Compared to receiving audio signals during a call S Rx The time delay and amplitude decay. For example, Figure 3 This is a schematic diagram illustrating virtual reflection conditions according to an embodiment of the present invention. Please refer to... Figure 3 Assuming the virtual reflection condition is a wall (i.e., an external object), and the distance between the microphone 21 and the sound source SS is d. s (For example, 0.3, 0.5, or 0.8 meters) and the distance between the microphone 21 and the wall W is d. w Under conditions of (e.g., 1, 1.5, or 2 meters), the reflected sound signal S' Rx Receive audio signal S during a call Rx The relationship can be represented as follows:
[0062] s′ Rx (n)=α1·s Rx (nn w1 (1)
[0063] Where α1 is the amplitude attenuation caused by reflection (i.e., the sound signal being reflected by the wall W), and n is the sampling point or time. wThe time delay caused by the reflection distance (i.e., the distance from the sound source SS through the wall W to the receiver 21).
[0064] In this embodiment of the invention, the processor 59 offsets the phase of the reflected sound signal according to the watermark identifier, and generates a sound watermark signal S accordingly. WM Specifically, processor 59 generates an audio watermark signal by offsetting the phase of the reflected sound signal according to the watermark identifier. In general echo cancellation mechanisms, compared to the phase shift of the reflected sound signal, the time delay and amplitude changes of the reflected sound signal have a greater impact on the error of the echo cancellation mechanism. This change is like being in a completely new interference environment, requiring the echo cancellation mechanism to readjust. Therefore, in this embodiment of the invention, the audio watermark signals corresponding to different values in the watermark identifier only differ in phase, but their time delay and amplitude are the same. That is, the audio watermark signal includes one or more phase-shifted reflected sound signals.
[0065] In one embodiment, the watermark identifier is encoded in a multi-base system, where each of one or more bits in the watermark identifier provides multiple values. For example, in binary, the value of each bit in the watermark identifier can be "0" or "1". In hexadecimal, the value of each bit in the watermark identifier can be "0", "1", "2", ..., "E", "F". In another embodiment, the watermark identifier is encoded using letters, words, and / or symbols. For example, the value of each bit in the watermark identifier can be any of the English letters "A" through "Z".
[0066] In one embodiment, the different values for each bit of the watermark identifier correspond to different phase shifts. For example, assuming the watermark identifier W0 is in base N (N is a positive integer), then N values can be provided for each bit. These N different values correspond to different phase shifts. For example, suppose the watermark identifier W O If it's a binary system, then each bit can provide two values (i.e., 1 and 0). These two different values correspond to two phase shifts. For example, phase shift It is 90° and the phase shift is 90°. It is -90° (i.e., -1).
[0067] Processor 59 can offset the phase of the reflected sound signal (with or without high-pass filtering) based on the value of one or more bits in the watermark identifier. For example, in N-ary mode, processor 59 selects the phase shift based on one or more values in the watermark identifier. One or more of them, and using the selected phase shift The phase shift is performed. For example, if the first bit of the watermark identifier is 1, then the output phase-shifted reflected sound signal... offset relative to the reflected sound signal Other reflected sound signals And so on. Phase shifting can be achieved using the Hilbert transform or other phase shifting algorithms.
[0068] The processor 19 of the conference terminal 10 receives the audio watermark signal S via the network through the communication transceiver 15. WM Or embed watermark signal S Rx +S WM To obtain the synthesized sound signal S A (That is, the transmitted sound watermark signal S) WM Or embed watermark signal S Rx +S WM ).
[0069] Please refer to Figure 2 The processor 19 determines the synthesized sound signal S based on the reflection cancellation sound signal. A Noise interference transmitted via the network (step S220). Specifically, canceling the reflection of the sound signal is to cancel the synthesized sound signal S. A Medium sound watermark signal S WM The watermark identifier is a sound signal with one or more codes. These codes refer to the values or symbols provided by the aforementioned multi-level encoding or other encoding mechanisms. The process of eliminating the sound signal through reflection will be detailed in subsequent embodiments.
[0070] During the transmission process from cloud server 50 to conference terminal 10 via network, its output signal (i.e., the transmitted audio watermark signal S) WM Or embed watermark signal S Rx +S WM After amplitude attenuation α T Transformed into attenuated sound signal S T And subject to noise N T Interference. And the sound signal and noise N T The signal-to-noise ratio (SNR) between them is SNR T =20·log(S) T / N T It is worth noting that using a fixed threshold to identify the audio watermark signal may not be applicable to different noise environments.
[0071] Please refer to Figure 2 The processor 19 determines a coding threshold based on the noise interference (step S230). Specifically, this coding threshold includes a first threshold and a second threshold, where the noise interference corresponding to the first threshold is lower than the noise interference corresponding to the second threshold, and the first threshold is greater than the second threshold. For example, the first threshold is 1.9, and the second threshold is 0.3. The signal-to-noise ratio (SNR) of the noise interference corresponding to the first threshold is...T =∞dB (i.e., no noise interference), and the signal-to-noise ratio (SNR) of the noise interference corresponding to the second threshold is SNR. T = -6dB (i.e., high noise interference). In this example, the values of the first and second thresholds mentioned above were obtained experimentally. However, the values of the first and second thresholds can still be changed according to actual needs, and the embodiments of the present invention are not limited thereto.
[0072] Figure 4 This is a flowchart of a method for generating an encoding threshold according to an embodiment of the present invention. Please refer to... Figure 4 In one embodiment, the processor 19 determines the time delay n based on the delay time n. w and synthesized sound signal S A Generate pre-processed sound signals This preprocessed audio signal It is a synthesized sound signal S A Phase shifted (e.g., 90°, -90°) and delayed by a delay time n w The result is (step S410). It should be noted that this embodiment uses a binary-encoded watermark identifier as an example (i.e., only two values are provided), and these two values correspond to, for example, a phase shift of 90° and -90°. However, if other encodings are used, different phase shifts may occur. Regarding the preprocessing of the audio signal... With synthesized sound signal S A The relationship can be represented as follows:
[0073]
[0074] That is, preprocessing the audio signal It is delayed by a time of n w And the synthesized sound signal S with a 90° phase shift A .
[0075] Regarding the synthesized sound signal S A The original voice signal S received during the call Rx The relationship can be represented as follows:
[0076]
[0077] Among them, the call receives the audio signal s. Rx Become via a 90° phase shift N T For noise interference, α w This is due to amplitude attenuation. And the audio signal received during a call... By delaying by a delay time n w become The preprocessed audio signal described above With synthesized sound signal SA From the relationship, we can derive the following about the preprocessed audio signal. Receive audio signal S during a call Rx Relationship:
[0078]
[0079] Where, α w For amplitude attenuation, N T For noise interference, noise interference N T via a 90° phase shift
[0080] Next, the processor 19 synthesizes the sound signal S. A and preprocessing of audio signals Generate the first sound signal s respectively B- and the second sound signal s B+ (Step S420). In one embodiment, at least one code of the watermark identifier includes a first code and a second code (e.g., W0 = 1, W0 = 0), and the aforementioned reflection cancellation sound signal includes a first sound signal s. B- and the second sound signal s B+ First sound signal s B- The audio signal with the watermark identifier as the first code (e.g., W0 = 1) was eliminated, and the second audio signal s B+ The audio signal with the watermark identifier as the second code (e.g., W0=0) was eliminated.
[0081] Regarding the first sound signal s B- With synthesized sound signal S A The relation can be expressed as follows:
[0082]
[0083] Regarding the first sound signal s B- Receive audio signal S during a call Rx The relationship can be represented as follows:
[0084]
[0085] Regarding the second sound signal s B+ With synthesized sound signal S A The relation can be expressed as follows:
[0086]
[0087] Regarding the second sound signal s B+ Receive audio signal S during a call Rx The relationship can be represented as follows:
[0088]
[0089] Please refer to Figure 4 The processor 19, based on the first sound signal s B- Generate a third sound signal And according to the second sound signal s B+ Generate a fourth sound signal (Step S430). Specifically, the first sound signal s B- A third sound signal is generated by shifting the phase and / or delaying the time. Second sound signal s B+ A fourth sound signal is generated by shifting the phase and / or delaying the time. In one embodiment, the first sound signal s B- Phase shifted by 90° and delayed by a delay time n w The third sound signal was obtained Regarding the third sound signal With the first sound signal s B- The relation can be expressed as follows:
[0090]
[0091] In addition, the second sound signal s B+ Phase shifted by 90° and delayed by a delay time n w The fourth sound signal was obtained. Regarding the fourth sound signal With the second sound signal s B+ The relation can be expressed as follows:
[0092]
[0093] Please refer to Figure 4 Processor 19 based on the third audio signal and the fourth sound signal Determine the first correlation respectively and second correlation (Step S440). Specifically, the processor 19 processes the first audio signal s. B- With the third sound signal Calculate the cross-correlation to determine the first correlation. In addition, the processor 19 supports the second audio signal s B+ With the fourth sound signal Calculate the cross-correlation to derive the second correlation.
[0094] It is worth noting that the first correlation With the second correlation The difference between the absolute values corresponds to the magnitude of noise interference. For example, the first correlation... Noise ratio (SNR) corresponding to noise interference T The relationship between the watermark identifier W0 and the watermark identifier W0 can be represented as follows:
[0095]
[0096] Table (1)
[0097] In other words, when the watermark identifier is the first code (e.g., W0 = 1), it only works in noisy environments (e.g., high signal-to-noise ratio SNR). T At -6dB), the first sound signal s B- With the third sound signal In Partially negatively correlated, in a noise-free environment (SNR) T =∞dB) is then uncorrelated (for example, In noisy environments, the correlation is high and negative (e.g., When the watermark identifier is the second code (e.g., W0 = 0), the first sound signal s B- With the third sound signal In s Rx (n-2·n w )and All of these components are negatively correlated, in a noise-free environment (SNR). T Its correlation is high and negative at (=∞dB) (e.g., High noise environment (SNR) T Its correlation is high and negative at -6dB (e.g., When synthesized sound signal S A When there is no watermark identifier (e.g., W0 = N / A, or not any code), the first sound signal s B- With the third sound signal In s Rx (n-2·n w )and All are negatively correlated; in the absence of noise, the correlation is high and negative (e.g., In noisy environments, the correlation is high and negative (e.g., In other words, when the watermark identifier is the first code (W0=1), it can be detected through the first correlation. The noise interference (i.e., SNR) in network transmission is determined. T =∞dB or SNR T = -6dB).
[0098] Next, the second correlation Noise interference SNR T The relationship with the watermark identifier W0 can be represented as follows:
[0099]
[0100] Table (2)
[0101] As can be seen from Table (2), when the watermark identifier is the first code (e.g., W0 = 1), in a high-noise environment (e.g., SNR), T At -6dB), the second sound signal s B+ With the fourth sound signal In s Rx (n-2·n w )and All of them are positively correlated, and in a noise-free environment (e.g., SNR) T At ∞ dB), the second correlation High and positive (e.g., ); Second correlation under high noise environment High and positive (e.g., When the watermark identifier is the second code (e.g., W0 = 0), only the second audio signal s is detected. B+ With the fourth sound signal Noise in The part is positively correlated in a noise-free environment (e.g., SNR). T Its correlation is low at (=∞dB) (e.g., High noise environment (e.g., SNR) T Its correlation is high and positive at -6dB (e.g., When synthesized sound signal S A When there is no watermark identifier (i.e., W0 = N / A, or not any code), the second audio signal s B+ With the fourth sound signal In s Rx (n-2·n w )and All are positively correlated; in the absence of noise, the correlation is high and positive (e.g., ); the correlation is high and positive in noisy environments (e.g., In other words, when the watermark identifier is the second code (e.g., W0 = 0), it can be detected through the second correlation. The noise interference (i.e., SNR) in network transmission is determined. T =∞dB or SNR T = -6dB).
[0102] Please refer to Figure 4 Processor 19 based on the first correlation and second correlation Determine the encoding threshold (Step S450). Specifically, the first correlation... With the second correlation The difference between the absolute values corresponds to the magnitude of the noise interference.
[0103] In one embodiment, the processor 19 determines the encoding threshold based on the correlation ratio. The correlation ratio is related to the first correlation. and second correlation The absolute value of the sum and the first correlation With the second correlation The largest of the absolute values. Furthermore, the encoding threshold in this embodiment... Used to identify synthetic sound signal S A The sound watermark signal S WM Is it at least one code? For example, the sound watermark signal S WM It is either 1 or 0. (Regarding the encoding threshold) First correlation and second correlation The relationship can be represented as follows:
[0104]
[0105] Through the aforementioned first correlation With the second correlation Based on its characteristics, the coding threshold can be derived. Noise interference SNR T The relationship with the watermark identifier W0 is as follows:
[0106]
[0107] Table (3)
[0108] From Tables (1), (2), and (3), it can be seen that when the watermark identifier is the first code or the second code and the network transmission environment is free from noise interference (e.g., SNR), T When (=∞dB), the first correlation With the second correlation The absolute values differ significantly, and the first correlation... With the second correlation These are positive and negative numbers, respectively. Therefore, the coding threshold corresponding to this noise interference... The value is 1.9 (i.e., the first threshold). However, when the network transmission environment is noisy (e.g., SNR...),... T When (=-6dB), the first correlation With the second correlation The differences between the absolute values are small, and the first correlation is small. With the second correlation These are positive and negative numbers, respectively. Therefore, the coding threshold corresponding to this noise interference... The value is 0.3 (i.e., the second threshold). When the synthesized sound signal S A When there is no watermark identifier (i.e., W0 = N / A), due to the first relevance With the second correlation The differences between the absolute values are small. Therefore, regardless of the magnitude of the noise interference, its coding threshold... The value is 0.3.
[0109] Please refer to Figure 5 In another embodiment, the processor 19 determines the first sound signal s based on the first sound signal s. B- Generate a third sound signal And according to the second sound signal s B+ Generate a fourth sound signal (Step S510). With Figure 4 The difference from the corresponding embodiment is that in this embodiment, the first sound signal s B- After a delay of one delay time n w The third sound signal was obtained And the second sound signal s B+ After a delay of one delay time n w The fourth sound signal was obtained. Regarding the third sound signal in this embodiment With the first sound signal s B- The relation can be expressed as follows:
[0110]
[0111] In addition, regarding the four-tone signals With the second sound signal s B+ The relation can be expressed as follows:
[0112]
[0113] Please refer to Figure 5 Processor 19 based on the third audio signal and the fourth sound signal Determine the first correlation respectively and second correlation (Step S520). Specifically, the processor 19 processes the first audio signal s. B- With the third sound signal Calculate cross-correlation to determine the first correlation. And the second sound signal s B+ With the fourth sound signal Calculate cross-correlation to derive the second correlation. First correlation and second correlation The difference between the absolute values corresponds to the magnitude of noise interference. For example, the first correlation... Or second correlation Signal-to-noise ratio (SNR) corresponding to noise interference T The relationship between the watermark identifier W0 and the watermark identifier W0 can be represented as follows:
[0114]
[0115] Table (4)
[0116] In other words, when the watermark identifier is the first code (e.g., W0=1) or the second code (e.g., W0=0), the first correlation... and second correlation The result is unrelated. That is to say, the first sound signal s B- With the third sound signal They are unrelated to each other, and the second sound signal s B+ With the fourth sound signal They are also unrelated. It is worth noting that only when the synthesized sound signal S... A When there is no watermark identifier (i.e., W0 = N / A), the s in the audio signal Rx (nn w )and The signal is positively correlated, while the noise component is uncorrelated. Therefore, when the synthesized audio signal SA has no watermark identifier (i.e., W0 = N / A) and the transmission environment is noise-free (SNR), the signal is positively correlated. T When the value is ∞dB, the correlation is high and positive. The transmission environment has a high noise level (SNR). T When the value is -6dB, the correlation is low and positive.
[0117] Please refer to Figure 5 Next, processor 19, based on the first correlation... and second correlation The sum of the values determines the encoding threshold Th. D (Step S530). It is worth noting that the encoding threshold Th in this embodiment... D Used to identify synthetic sound signal SA Does the audio watermark signal contain at least one code? For example, is the audio watermark signal N / A? Regarding the encoding threshold Th... D First correlation and second correlation The relationship can be represented as follows:
[0118]
[0119] Next, based on Table (4) and the aforementioned first correlation... and second correlation Based on its characteristics, the encoding threshold Th can be derived. D Noise interference SNR T The relationship with the watermark identifier W0 can be represented as follows:
[0120] <![CDATA[Th D ]]> <![CDATA[W0=1]]> <![CDATA[W0=0]]> <![CDATA[W0=N / A]]> <![CDATA[SNR T =∞dB]]> ±0.3 ±0.3 10 <![CDATA[SNR T =-6dB]]> ±0.3 ±0.3 0.5
[0121] Table (5)
[0122] As shown in Table (5) and the first correlation mentioned above and second correlation Based on the characteristics, it can be concluded that in the absence of a watermark identifier, the first relevance... and second correlation It can be used to determine noise interference in network transmission (i.e., SNR) T =∞dB or SNR T = -6dB). Based on this, the encoding threshold Th can be used to determine the threshold value. D Identify whether there is at least one code in the audio watermark signal.
[0123] Figure 6 This is a flowchart illustrating the determination of the encoding threshold according to another embodiment of the present invention. Please refer to... Figure 6 In one embodiment, the encoding threshold includes a first noise threshold and a second noise threshold. The processor 19 determines the noise threshold based on the delay time n. w and synthesized sound signal S A Generate pre-processed sound signals (Step S610). Specifically, the audio signal is preprocessed. It is a synthesized sound signal S A After a delay of one delay time n w The results obtained. Regarding the preprocessing of audio signals. With synthesized sound signal S A The relationship can be represented as follows:
[0124]
[0125] Regarding the preprocessing of audio signals Receive audio signal S during a call Rx The relationship can be represented as follows:
[0126]
[0127] Next, the processor 19 synthesizes the sound signal S. A and preprocessing of audio signals Generate the fifth sound signal s C (Step S620). Regarding the fifth sound signal s C With synthesized sound signal S A The relation can be expressed as follows:
[0128]
[0129] Regarding the fifth sound signal s C Receive audio signal S during a call Rx The relationship can be represented as follows:
[0130]
[0131]
[0132] In this embodiment, the reflection cancellation sound signal includes a fifth sound signal s. C The fifth sound signal s C Synthetic audio signals that have not been identified by any code (e.g., W0 = N / A) have been eliminated.
[0133] Please refer to Figure 6 Processor 19 based on the fifth sound signal s C Generate the sixth sound signal (Step S630). In this embodiment, the fifth sound signal s C After a delay of one delay time n w To generate a sixth sound signal Regarding the sixth sound signal With the fifth sound signal s C The relation can be expressed as follows:
[0134]
[0135] Processor 19 based on the fifth sound signal s C and the sixth sound signal Determine the third correlation (Step S640). Specifically, processor 19 responds to the fifth audio signal s C and the sixth sound signal Calculate cross-correlation to derive the third correlation. Third correlation This corresponds to the magnitude of noise interference. For example, the third correlation... Signal-to-noise ratio (SNR) corresponding to noise interference T The relationship between the watermark identifier W0 and the watermark identifier W0 can be represented as follows:
[0136]
[0137] Table (6)
[0138] In other words, when the watermark identifier is the first code (i.e., W0 = 1), the fifth sound signal s C With the s in the sound signal Rx (nn w ), and N T (nn w The third correlation between ) The results showed a negative correlation, and the transmission environment was noise-free (SNR). T When the value is ∞dB, the correlation is high and negative (e.g., ); while the transmission environment has a high noise level (SNR) T When the value is -6dB, the correlation is high and negative (e.g., Furthermore, the characteristics of the watermark identifier when it is the second code (i.e., W0 = 1) are the same as those of the first code. It is worth noting that only when the synthesized sound signal S... A When there is no watermark identifier (i.e., W0 = N / A), the noise component in the audio signal The correlation is negative. Therefore, when the synthesized audio signal SA has no watermark identifier (i.e., W0 = N / A) and the transmission environment is noise-free (SNR), the correlation is negative. T When =∞dB), the correlation is low (e.g., ); while the transmission environment has a high noise level (SNR) T When the value is -6dB, the correlation is high (e.g., ).
[0139] Processor 19 based on third correlation Determine the first noise threshold For example, regarding the first noise threshold With the third correlation The relationship can be represented as follows:
[0140]
[0141] Next, based on Table (6) and the aforementioned third correlation... Based on the characteristics, the first noise threshold can be determined. Signal-to-noise ratio (SNR) corresponding to noise interferenceT The relationship with the watermark identifier W0 can be represented as follows:
[0142]
[0143] Table (7)
[0144] As shown in Table (7) and the third correlation mentioned above Based on the characteristics, it can be concluded that in the absence of a watermark identifier (e.g., W0 = N / A), and without noise interference (e.g., SNR), T =∞dB), then the third correlation Smaller and first noise threshold Larger; if there is significant noise interference (e.g., SNR) T =-6dB), then the third correlation Larger and first noise threshold Relatively small. First noise threshold. Used to identify whether there is at least one code in the audio watermark signal in a synthesized audio signal.
[0145] On the other hand, processor 19 determines the second noise threshold based on the correlation ratio. (Step S650). For a detailed explanation of step S650, please refer to [link / reference needed]. Figure 4 And this will not be elaborated further here. That is, the second noise threshold determined in this embodiment. The encoding threshold determined in step S450
[0146] Next, the processor 19 determines the noise threshold based on the first noise threshold. and the second noise threshold Determine the final encoding threshold (Step S660). In one embodiment, the encoding threshold... Related to the first noise threshold With the second noise threshold The difference and the second noise threshold The largest among them. Regarding the encoding threshold. First noise threshold With the second noise threshold The relationship can be represented as follows:
[0147]
[0148] Encoding threshold Used to identify synthetic sound signal S AThe key is whether the audio watermark signal contains at least one code and whether it contains at least one code (e.g., W0 = N / A, W0 = 1, or W0 = 1). Based on the characteristics in Tables (5) and (7), the encoding threshold can be determined. Signal-to-noise ratio (SNR) corresponding to noise interference T The relationship with the watermark identifier W0 can be represented as follows:
[0149]
[0150] Table (8)
[0151] As shown in Table (8), regardless of the value of the watermark identifier (e.g., W0 = N / A, 0 or 1), if there is no noise interference (e.g., SNR), T =∞dB), then the encoding threshold Larger (e.g., If there is significant noise interference (e.g., SNR) T =-6dB), then the encoding threshold Smaller (e.g., This allows us to conform to the characteristics and range of noise variations in the environment.
[0152] Please refer to Figure 2 The processor 19 identifies the synthesized sound signal S based on the encoding threshold. A The sound watermark signal S WM (Step S240). Specifically, processor 19 generates a synthesized sound signal with a 90° phase shift. Figure 7 This is a flowchart illustrating the identification of a sound watermark signal according to an embodiment of the present invention. The processor 19 can identify the synthesized sound signal S... A and the synthesized sound signal after phase shifting Correlation between Identify watermark identifier W E (Step S710). For example, processor 19 synthesizes the sound signal S. A With synthesized sound signals Calculate orthogonal cross correlation and Processor 19 defines the encoding threshold and Th D Then the watermark identifier W E It can be represented as:
[0153]
[0154]
[0155] That is, if the correlation The absolute value is lower than the encoding threshold and Th D Then processor 19 determines that the value of this bit is not any code (e.g., N / A); if the correlation... Above the encoding threshold or Th D Then processor 19 further determines the correlation. This determines whether the value of this bit corresponds to a phase shift of -90° (e.g., 0) or a phase shift of 90° (e.g., 1). In other words, the encoding threshold Th... D This can be used to assist in confirming whether the sound signal is any of the codes in the watermark identifier. Furthermore, to avoid being affected by noise, another part of the identification process involves determining the encoding threshold based on the characteristics of changes in noise interference. Finally, processor 19 can convert these two encoding thresholds or Th D Correlation By comparison, a more accurate watermark identifier can be determined.
[0156] In another embodiment, the processor 19 can identify the synthetic sound signal S using a deep learning-based classifier. A The corresponding values at different time units.
[0157] Regarding varying noise interference, for example, based on experimental experience, the synthesized sound signal S A The transmission process is in a high-noise interference environment (e.g., SNR). T When the signal density is -6dB, a coding threshold of 1.9 is used to identify the audio watermark signal S. WM Watermark identifiers can improve the accuracy of recognition. On the other hand, synthesized sound signals S A The transmission process is conducted in a noise-free environment (e.g., SNR). T When the value is ∞dB, a coding threshold of 0.3 can correctly identify the audio watermark signal S. WM The watermark identifier in the image.
[0158] In summary, in the sound watermark recognition method and device of this invention, noise interference in the transmission environment is determined based on the characteristics of virtual reflected sound signals and reflection-cancelled sound signals in the synthesized sound signal. Furthermore, the encoding threshold for the desired watermark identifier is determined based on the noise interference. Therefore, corresponding encoding thresholds can be used according to different transmission environments to improve the accuracy of watermark identifier recognition.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for recognizing audio watermarks, applicable to conference terminals, characterized in that, The method for recognizing the sound watermark includes: The synthesized sound signal is received via a network, wherein the synthesized sound signal includes a sound watermark signal, the sound watermark signal is generated by offsetting the phase of the reflected sound signal according to the watermark identifier, and the reflected sound signal is a sound signal obtained by recording the sound emitted by a simulated sound source after being reflected by an external object and recorded by a microphone. The noise interference transmitted via the network to the synthesized sound signal is determined based on at least one reflection cancellation sound signal, wherein the at least one code is a value provided by a multi-base encoding, the at least one code includes a first code and a second code, the reflection cancellation sound signal is a sound signal that cancels the watermark identifier of the sound watermark signal in the synthesized sound signal, which is the first code or the second code, the reflection cancellation sound signal includes a first sound signal and a second sound signal, the first sound signal is obtained by subtracting the synthesized sound signal from the sound signal in the case where the watermark identifier is the first code, and the second sound signal is obtained by subtracting the synthesized sound signal from the sound signal in the case where the watermark identifier is the second code; A coding threshold is determined based on the noise interference, wherein the coding threshold includes a first threshold and a second threshold, the noise interference corresponding to the first threshold is lower than the noise interference corresponding to the second threshold, and the first threshold is greater than the second threshold, including: When the noise interference is present, the second threshold is used as the encoding threshold; and When there is no noise interference, the first threshold is used as the encoding threshold; and Identifying the sound watermark signal in the synthesized sound signal according to the encoding threshold includes: Cross-correlation is calculated between the synthesized sound signal and the phase-shifted synthesized sound signal to obtain the correlation. If the absolute value of the correlation is less than the encoding threshold, then it is determined that the value of one bit in the audio watermark signal is not the first code or the second code; and If the absolute value of the correlation is not less than the encoding threshold, then the value of the bit in the sound watermark signal is determined to be either the first code or the second code.
2. The method for recognizing audio watermarks according to claim 1, characterized in that, The steps for determining the noise interference include: A preprocessed sound signal is generated based on the delay time and the synthesized sound signal, wherein the preprocessed sound signal is obtained by phase-shifting and delaying the synthesized sound signal by the delay time; The first sound signal and the second sound signal are generated based on the synthesized sound signal and the preprocessed sound signal, respectively. A third sound signal is generated based on the first sound signal, and a fourth sound signal is generated based on the second sound signal, wherein the first sound signal is phase-shifted and delayed by the delay time to generate the third sound signal, and the second sound signal is phase-shifted and delayed by the delay time to generate the fourth sound signal; A first correlation and a second correlation are determined based on the third sound signal and the fourth sound signal, respectively, wherein the first correlation is the correlation between the first sound signal and the third sound signal, the second correlation is the correlation between the second sound signal and the fourth sound signal, and the difference between the absolute values of the first correlation and the second correlation corresponds to the magnitude of the noise interference.
3. The method for recognizing audio watermarks according to claim 2, characterized in that, The step of determining the coding threshold based on the noise interference includes: The encoding threshold is determined based on the correlation ratio, wherein the correlation ratio is related to the maximum of the sum of the first correlation and the second correlation, and the maximum of the absolute values of the first correlation and the second correlation, and the encoding threshold is used to identify whether the sound watermark signal in the synthesized sound signal is the first code or the second code.
4. The method for identifying audio watermarks according to claim 2, characterized in that, The step of determining the coding threshold based on the noise interference includes: The encoding threshold is determined based on the sum of the first correlation and the second correlation, wherein the encoding threshold is used to identify whether the sound watermark signal in the synthesized sound signal contains the first code or the second code.
5. The method for identifying audio watermarks according to claim 2, characterized in that, The encoding threshold includes a first noise threshold and a second noise threshold, and the step of determining the encoding threshold based on the noise interference includes: The first noise threshold is determined based on the third correlation, wherein the third correlation is related to the correlation between the fifth sound signal and the sixth sound signal, the reflection cancellation sound signal includes the fifth sound signal, the fifth sound signal cancels the synthesized sound signal when the watermark identifier is not the first code or the second code, the sixth sound signal is the sound signal of the fifth sound signal delayed by the delay time, and the first noise threshold is used to identify whether the sound watermark signal in the synthesized sound signal contains the first code or the second code; The second noise threshold is determined based on a correlation ratio, wherein the correlation ratio is related to the maximum of the sum of the first and second correlations and the absolute values of the first and second correlations, and the second noise threshold is used to identify whether the sound watermark signal in the synthesized sound signal is the first code or the second code; and The encoding threshold is determined based on the first noise threshold and the second noise threshold, wherein the encoding threshold is related to the difference between the first noise threshold and the second noise threshold, and the maximum of the second noise threshold, and the encoding threshold is used to identify whether the sound watermark signal in the synthesized sound signal contains the first code or the second code and whether it is the first code or the second code.
6. A device for identifying sound watermarks, comprising: Memory, used to store program code; as well as A processor, coupled to the memory, is characterized in that the processor is configured to load and execute the program code to: The synthesized sound signal is received via a network, wherein the synthesized sound signal includes a sound watermark signal, the sound watermark signal is generated by offsetting the phase of the reflected sound signal according to the watermark identifier, and the reflected sound signal is a sound signal obtained by recording the sound emitted by a simulated sound source after being reflected by an external object and recorded by a microphone. The noise interference transmitted via the network to the synthesized sound signal is determined based on at least one reflection cancellation sound signal, wherein the at least one code is a value provided by a multi-base encoding, the at least one code includes a first code and a second code, the reflection cancellation sound signal is a sound signal that cancels the watermark identifier of the sound watermark signal in the synthesized sound signal, which is the first code or the second code, the reflection cancellation sound signal includes a first sound signal and a second sound signal, the first sound signal is obtained by subtracting the synthesized sound signal from the sound signal in the case where the watermark identifier is the first code, and the second sound signal is obtained by subtracting the synthesized sound signal from the sound signal in the case where the watermark identifier is the second code; A coding threshold is determined based on the noise interference, wherein the coding threshold includes a first threshold and a second threshold, the noise interference corresponding to the first threshold is lower than the noise interference corresponding to the second threshold, and the first threshold is greater than the second threshold, including: When the noise interference is present, the second threshold is used as the encoding threshold; and When there is no noise interference, the first threshold is used as the encoding threshold; and Identifying the sound watermark signal in the synthesized sound signal according to the encoding threshold includes: Cross-correlation is calculated between the synthesized sound signal and the phase-shifted synthesized sound signal to obtain the correlation. If the absolute value of the correlation is less than the encoding threshold, then it is determined that the value of one bit in the audio watermark signal is not the first code or the second code; and If the absolute value of the correlation is not less than the encoding threshold, then the value of the bit in the sound watermark signal is determined to be either the first code or the second code.
7. The sound watermark identification device according to claim 6, characterized in that, The processor is further configured to: A preprocessed sound signal is generated based on the delay time and the synthesized sound signal, wherein the preprocessed sound signal is obtained by phase-shifting and delaying the synthesized sound signal by the delay time; The first sound signal and the second sound signal are generated based on the synthesized sound signal and the preprocessed sound signal, respectively. A third sound signal is generated based on the first sound signal, and a fourth sound signal is generated based on the second sound signal, wherein the first sound signal is phase-shifted and / or delayed by the delay time to generate the third sound signal, and the second sound signal is phase-shifted and / or delayed by the delay time to generate the fourth sound signal; A first correlation and a second correlation are determined based on the third sound signal and the fourth sound signal, respectively, wherein the first correlation is the correlation between the first sound signal and the third sound signal, the second correlation is the correlation between the second sound signal and the fourth sound signal, and the difference between the absolute values of the first correlation and the second correlation corresponds to the magnitude of the noise interference.
8. The sound watermark identification device according to claim 7, characterized in that, The processor is further configured to: The encoding threshold is determined based on the correlation ratio, wherein the correlation ratio is related to the maximum of the sum of the first correlation and the second correlation, and the maximum of the absolute values of the first correlation and the second correlation, and the encoding threshold is used to identify whether the sound watermark signal in the synthesized sound signal is the first code or the second code.
9. The sound watermark identification device according to claim 7, characterized in that, The processor is further configured to: The encoding threshold is determined based on the sum of the first correlation and the second correlation, wherein the encoding threshold is used to identify whether the sound watermark signal in the synthesized sound signal contains the first code or the second code.
10. The sound watermark identification device according to claim 7, characterized in that, The encoding threshold includes a first noise threshold and a second noise threshold, and the processor is further configured to: The first noise threshold is determined based on a third correlation, wherein the third correlation is related to the correlation between the fifth sound signal and the sixth sound signal, the reflection cancellation sound signal includes the fifth sound signal, the fifth sound signal cancels the synthesized sound signal when the watermark identifier is not the first code or the second code, the sixth sound signal is the fifth sound signal delayed by the delay time, and the first noise threshold is used to identify whether the sound watermark signal in the synthesized sound signal contains the first code or the second code; The second noise threshold is determined based on the correlation ratio, wherein the correlation ratio is related to the maximum of the sum of the first correlation and the second correlation and the absolute value of the first correlation and the second correlation, and the second noise threshold is used to identify whether the sound watermark signal in the synthesized sound signal is the first code or the second code. as well as The encoding threshold is determined based on the first noise threshold and the second noise threshold, wherein the encoding threshold is related to the difference between the first noise threshold and the second noise threshold, and the maximum of the second noise threshold, and the encoding threshold is used to identify whether the sound watermark signal in the synthesized sound signal contains the first code or the second code and whether it is the first code or the second code.
Citation Information
Patent Citations
Sound system calibration
EP2899997A1
Voice watermarking system
TW200621026A