Sound watermark processing method and voice communication system

By generating high-frequency sine wave signals of different frequencies to synthesize watermark signals and combining them with speech signals, the problems of time-consuming watermark signal embedding and noise influence are solved, and real-time embedding and anti-noise recognition of watermark signals are achieved.

CN115691520BActive Publication Date: 2025-09-30ACER INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110835016.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-23
Publication Date
2025-09-30
Estimated Expiration
2041-07-23

AI Technical Summary

Technical Problem

In the prior art, the watermark embedding process takes too long and is difficult to meet real-time requirements. In addition, the watermark signal is easily affected by noise during transmission, resulting in distortion and difficulty in recognition.

Method used

By generating high-frequency sine wave signals of different frequencies, synthesizing watermark sound signals and combining them with speech signals in the time domain, the watermark signal is embedded in real time, and the noise impact is reduced by identifying pulse signals at the receiving end.

Benefits of technology

The real-time embedding and noise resistance of the watermark signal are realized, ensuring that the watermark signal can be accurately identified during the transmission process to meet the real-time call needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691520B_ABST
    Figure CN115691520B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method for processing sound watermarks and a voice communication system. Several sine wave signals are generated. These sine wave signals have different frequencies and are high-frequency sound signals. A watermark pattern is mapped onto a time-frequency graph to form a watermarked sound signal. The two dimensions of the watermark pattern in a two-dimensional coordinate system correspond to the time axis and frequency axis of the time-frequency graph, respectively. Each of the several sound frames on the time axis corresponds to a sine wave signal of a different frequency on the frequency axis. The speech signal and the watermarked sound signal are synthesized in the time domain to generate an embedded watermark signal. This allows for real-time sound watermarking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice processing technology, and in particular to a sound watermark processing method and a voice communication system. Background Art

[0002] Remote conferencing allows people in different locations or spaces to engage in conversation, and the related equipment, protocols, and applications are quite mature. However, some real-time conferencing programs may synthesize voice signals and watermarks. However, embedding the watermark can be time-consuming and unsuitable for the real-time nature of conference calls. Furthermore, the voice signal may be distorted by noise during transmission, and the embedded watermark may also be affected, making it difficult to recognize. Summary of the Invention

[0003] The present invention is directed to a sound watermark processing method and a voice communication system, which can embed watermark sound signals in real time and also has an anti-noise function.

[0004] According to an embodiment of the present invention, a method for processing an audio watermark includes (but is not limited to) the following steps: generating multiple sine wave signals. These sine wave signals have different frequencies and are high-frequency audio signals. Mapping a watermark pattern onto a time-frequency graph to form a watermarked audio signal. The two dimensions of the watermark pattern in a two-dimensional coordinate system correspond to the time axis and the frequency axis of the time-frequency graph, respectively. Each of the multiple audio frames on the time axis corresponds to a sine wave signal of a different frequency on the frequency axis. The speech signal and the watermarked audio signal are synthesized in the time domain to generate an embedded watermark signal.

[0005] According to an embodiment of the present invention, a voice communication system includes (but is not limited to) a transmitting device. The transmitting device is configured to generate a plurality of sine wave signals, map a watermark pattern onto a time-frequency graph to form a watermarked audio signal, and synthesize the audio signal and the watermarked audio signal in the time domain to generate an embedded watermarked signal. These sine wave signals have different frequencies and are high-frequency audio signals. The two dimensions of the watermark pattern in a two-dimensional coordinate system correspond to the time axis and the frequency axis of the time-frequency graph, respectively. Each of the plurality of audio frames on the time axis corresponds to a sine wave signal of a different frequency on the frequency axis.

[0006] Based on the foregoing, the voice communication system and audio watermarking method according to embodiments of the present invention use multiple high-frequency sine wave signals of varying frequencies to synthesize a watermarked audio signal corresponding to the watermark pattern. The watermarked audio signal is then combined with the speech signal in the time domain. This allows for real-time embedding of the watermarked audio signal while mitigating the effects of pulse signal noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The accompanying drawings are included to provide a further understanding of the present invention and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present invention and together with the description serve to explain the principles of the present invention.

[0008] Figure 1 is a block diagram of components of a voice communication system according to an embodiment of the present invention;

[0009] Figure 2 is a flow chart of a method for processing a sound watermark according to an embodiment of the present invention;

[0010] Figure 3A and Figure 3B It is a waveform diagram illustrating sine wave signals of different frequencies;

[0011] Figure 4A and Figure 4B yes Figure 3A and Figure 3B The waveform of the windowed sine wave signal;

[0012] Figure 5A It is an example to illustrate the watermark pattern;

[0013] Figure 5B This is an example illustrating a watermark pattern in a two-dimensional coordinate system;

[0014] Figure 5C This is an example Figure 5B The watermark pattern is mapped to a time-frequency graph;

[0015] Figure 5D This is a diagram illustrating the superposition of several sound frames;

[0016] Figure 6 is an example illustrating a watermarked sound signal in a time-frequency graph;

[0017] Figure 7 is an example illustrating a transmitted sound signal in a time-frequency diagram;

[0018] Figure 8 is a flow chart of watermark pattern recognition according to an embodiment of the present invention;

[0019] Figure 9 is a schematic diagram illustrating an example of modification of a default watermark signal.

[0020] Explanation of Figure Numbers

[0021] 1: Voice communication system;

[0022] 10: conveying device;

[0023] 11, 51: communication transceiver;

[0024] 13, 53: memory;

[0025] 15, 55: processor;

[0026] 30: Network;

[0027] 50: receiving device;

[0028] S210-S260, S810-S850: steps;

[0029] S f1 w ,…,S fN w 、S f1 、S f2 : sine wave signal;

[0030] W I :watermark pattern;

[0031] S W : Watermark sound signal;

[0032] X, Y: axis;

[0033] S' H : voice signal;

[0034] S H Wed : Embed watermark signal;

[0035] S A : Transmit sound signals;

[0036] W1,…,W M :Default watermark signal;

[0037] CS: two-dimensional coordinate system;

[0038] TFD: time-frequency diagram;

[0039] W'1,…,W' M : Modified default watermark signal. DETAILED DESCRIPTION

[0040] Reference will now be made in detail to exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.

[0041] Figure 1 is a block diagram of components of a voice communication system 1 according to an embodiment of the present invention. Figure 1 The voice communication system 1 includes but is not limited to one or more transmitting devices 10 and one or more receiving devices 50 .

[0042] The transmitting device 10 and the receiving device 50 may be a wired telephone, a mobile phone, an Internet phone, a tablet computer, a desktop computer, a laptop computer, or a smart speaker.

[0043] The transmitting device 10 includes (but is not limited to) a communication transceiver 11 , a memory 13 , and a processor 15 .

[0044] The communication transceiver 11 may be, for example, a transceiver that supports a wired network such as Ethernet, a fiber optic network, or a cable (which may include, but is not limited to, a connection interface, a signal converter, a communication protocol processing chip, and other components). It may also be a transceiver that supports a wireless network such as Wi-Fi, fourth-generation (4G), fifth-generation (5G), or later-generation mobile networks (which may include, but is not limited to, an antenna, a digital-to-analog / analog-to-digital converter, a communication protocol processing chip, and other components). In one embodiment, the communication transceiver 11 is used to transmit or receive data via a network 30 (e.g., the Internet, a local area network, or other type of network).

[0045] The memory 13 can be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or similar device. In one embodiment, the memory 13 is used to store program code, software modules, configurations, data (e.g., audio signals, watermark patterns, watermarked audio signals, etc.), or files.

[0046] The processor 15 is coupled to the communication transceiver 11 and the memory 13. The processor 15 can be a central processing unit (CPU), a graphics processing unit (GPU), or other programmable general-purpose or special-purpose microprocessor, a digital signal processor (DSP), a programmable controller, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other similar components or a combination of these components. In one embodiment, the processor 15 is configured to perform all or part of the operations of the transmission device 10 and can load and execute various software modules, program codes, files, and data stored in the memory 13.

[0047] The receiving device 50 includes (but is not limited to) a communication transceiver 51, a memory 53, and a processor 55. The implementation and functions of the communication transceiver 51, the memory 53, and the processor 55 can be referred to the description of the communication transceiver 11, the memory 13, and the processor 15, respectively, and will not be repeated here.

[0048] In some embodiments, the transmitting device 10 and / or the receiving device 50 further include a microphone and / or a speaker (not shown). The microphone may be a dynamic, condenser, or electret condenser microphone. The microphone may also be a combination of other electronic components that can receive sound waves (e.g., human voice, ambient sound, machine operation sound, etc.) and convert them into sound signals, an analog-to-digital converter, a filter, and an audio processor. In one embodiment, the microphone is used to receive / record the speaker to obtain a voice signal. In some embodiments, the voice signal may include the speaker's voice, the sound emitted by the speaker, and / or other ambient sounds. The speaker may be a loudspeaker or an amplifier. In one embodiment, the speaker is used to play sound.

[0049] Hereinafter, the method of the embodiment of the present invention will be described with reference to various devices, components and modules in the voice communication system 1. Each process of the method can be adjusted according to the implementation situation and is not limited thereto.

[0050] Figure 2 This is a flow chart of a method for processing a sound watermark according to an embodiment of the present invention. Figure 2 , the processor 15 of the transmitting device 10 generates one or more sine wave signals S f1 ,…,S fN (Step S210). Specifically, the frequencies of these sine wave signals (eg, sine wave or cosine wave) are different. For example, Figure 3A and Figure 3B It is a sine wave signal S with different frequencies. f1 、S f2 Please refer to the waveform diagram of Figure 3A and Figure 3B , sine wave signal S f2 The frequency is higher than the sine wave signal S f1 . Assume there are N sine wave signals S f1 ,…,S fN , that is, N sine wave signals S with different frequencies f1 ,…,S fN N is, for example, 32, 64, 128 or other positive integers.

[0051] In one embodiment, the processor 15 may determine the sine wave signal S at specific frequency intervals (Spacing). f1 ,…,S fN For example, the sine wave signal S f1 The frequency of the sine wave signal S is 16 kHz. f2 The frequency of the sine wave signal S is 16.5kHz. f2 The frequency of is 17kHz, that is, the frequency interval is 500Hz, and the rest are similar. f1 ,…,S fN5 The frequency intervals between them may not be fixed.

[0052] The processor 15 converts these sine wave signals S f1 ,…,S fN The time length is set to the number of samples of one audio frame (time unit) (e.g., 512, 1024, or 2028). In addition, these sine wave signals are high-frequency sound signals (e.g., their frequencies are between 16kHz and 20kHz, but may vary depending on the capabilities of the speaker).

[0053] In one embodiment, the processor 15 further windows the sine wave signals S based on a windowing function (eg, a Hamming window, a rectangular window, or a Gaussian window). f1 ,…,S fN , to generate a windowed sine wave signal S f1 w ,…,S fN w In this way, time intervals are generated between adjacent sound frames in the time domain, and pulses are avoided from being generated between sound frames.

[0054] For example, Figure 4A and Figure 4B yes Figure 3A and Figure 3B The waveform of the windowed sine wave signal. Figure 4A , sine wave signal S f1 After windowing, it becomes S f1 w Please refer to Figure 4B , sine wave signal S f2 After windowing, it becomes S f2 w .

[0055] The processor 15 converts the watermark pattern W I Mapping to the time-frequency graph to form the watermarked sound signal S W (Step S220). Specifically, the watermark pattern W IIt can be designed according to the needs of the user, and the embodiment of the present invention is not limited to it. For example, Figure 5A This is an example of a watermark pattern W I Please refer to Figure 5A , this watermark pattern W I It is composed of the word "acer".

[0056] The processor 15 converts the watermark pattern W I Convert from a two-dimensional coordinate system to a time-frequency diagram. A two-dimensional coordinate system consists of two dimensions. For example, Figure 5B This is an example of a watermark pattern W in the two-dimensional coordinate system CS. I Please refer to Figure 5B , these two dimensions include the transverse axis X and the longitudinal axis Y. That is, any position on the two-dimensional coordinate system CS can be defined by its distance from the transverse axis X and its distance from the longitudinal axis Y.

[0057] In one embodiment, the processor 15 further extends the watermark pattern W in a dimension corresponding to the two-dimensional coordinate system on the time axis according to the amount of superposition. I This superposition is related to the overlap of adjacent frames. For example, the superposition is 0.5 frames or other time lengths. The superposition of frames will be described in detail later. Figure 5A and Figure 5B For example, assuming that the superposition amount is 0.5 sound frames and the horizontal axis X corresponds to the time axis in the time-frequency diagram, the watermark pattern W I Extend twice along the horizontal axis X. That is, the extended watermark pattern W I The multiple is inversely proportional to the amount of superposition.

[0058] On the other hand, the time-frequency diagram includes a time axis and a frequency axis. Each of the several sound frames on the time axis corresponds to the sine wave signals of different frequencies on the frequency axis. In one embodiment, the processor 15 generates the watermark pattern W according to the watermark pattern W. I A watermark matrix is ​​created in the time-frequency graph. The watermark matrix includes several elements, each of which is a marked element and an unmarked element. The marked element represents the watermark pattern W I The corresponding position in the two-dimensional coordinate system has a value, and the unmarked element represents the watermark pattern W I There is no value at the corresponding position in the two-dimensional coordinate system.

[0059] by Figure 5B For example, the two-dimensional coordinate system CS is divided into 40*8 grids. There is a watermark pattern W at the intersection of any vertical line and horizontal line (which can form a coordinate in the two-dimensional coordinate system CS). I This means that there is a value at this position and there is no watermark pattern W I This means there is no value at this position.

[0060] Figure 5C This is an example Figure 5B Watermark pattern W I Map to time-frequency diagram TFD. Please refer to Figure 5C Similarly, the time-frequency diagram TFD can also be divided into 40*8 grids. The processor 15 compares the two-dimensional coordinate system CS and the time-frequency diagram TFD, and defines the watermark matrix in the time-frequency diagram TFD as having marked elements or unmarked elements accordingly.

[0061] The processor 15 selects one or more sine wave signals in each frame according to the watermark matrix. The one or more selected sine wave signals correspond to the marked elements in those elements. Figure 5C For example, each vertical line on the time axis represents a sound frame. In addition, each horizontal line on the frequency axis represents a sine wave signal of a certain frequency. For example, the bottom horizontal line corresponds to a sine wave signal with a frequency of 16kHz, and the horizontal line above it corresponds to a sine wave signal with a frequency of 16.2kHz, and so on. The processor 15 can record the correspondence between each horizontal line on the frequency axis and the frequency of those sine wave signals. For each sound frame on the time axis, the processor 15 determines whether there is a marker element in the watermark matrix and selects a sine wave signal based on the correspondence.

[0062] The processor 15 superimposes one or more selected sine wave signals on the sound frames in the time-frequency diagram in the time domain to form a watermark sound signal S W . The processor 15 superimposes adjacent sound frames according to the aforementioned superposition amount. For example, Figure 5D This is a diagram illustrating the example of how to stack several sound frames. Figure 5D , the sine wave signal on the first frame overlaps the sine wave signal on the second frame by 0.5 frames, and so on. In addition, compared to Figure 5C , Figure 5D The watermark pattern W in I Zoom out by half in the direction of the time axis.

[0063] Figure 6 This is an example of a watermarked audio signal in a time-frequency diagram. Figure 6 , Figure 5A Watermark pattern W I As if formed on a grid pattern.

[0064] Processor 15 synthesizes speech signal S' in the time domain H With watermark sound signal S W , to generate the embedded watermark signal S H Wed (Step S230). Specifically, the speech signal S HThe transmitting device 10 may receive a voice signal from a speaker through a microphone, or may receive a voice signal from an external device (eg, a conference call server, a voice recorder, or a smartphone).

[0065] In one embodiment, the processor 15 may filter out the original speech signal S H The sine wave signal S is located in the middle f1 ,…,S fN The sound signal of the frequency band is used to generate the speech signal S' H For example, suppose the sine wave signal S f1 ,…,S fN The frequency band is 16kHz~20kHz, and the processor 15 converts the voice signal S H The low-pass filter below 16kHz is used to avoid the voice signal S H Affecting watermark sound signal S W In another embodiment, the processor 15 may convert the original speech signal S H Directly as the speech signal S' H .

[0066] The processor 15 can use methods such as spread spectrum, echo hiding, phase encoding, etc. in the time domain to encode the voice signal S' H Add watermark sound signal S W , to form the embedded watermark signal S H Wed It can be seen that the embodiment of the present invention pre-establishes the watermark sound signal S W , in real time in the time domain with the speech signal S' H synthesis.

[0067] The processor 15 transmits the embedded watermark signal S via the communication transceiver 11 and the network 30. H Wed (Step S240). The processor 55 of the receiving device 50 receives the transmission sound signal S through the communication transceiver 51. A This transmits the sound signal S A is the transmitted watermark signal S H Wed In some cases, the embedded watermark signal S H Wed During the transmission process of the network 30, the audio signal may be distorted (for example, due to interference from other ambient sounds, obstacles, or other noises) to form the transmission audio signal S A(or called the attacked signal). It is worth noting that the transmitting device 10 transmits the watermarked sound signal S W Set to high-frequency sound signal, but high-frequency sound signal may be interfered by pulse signal. For example, Figure 7 is an example of a transmitted sound signal S in a time-frequency diagram A Please refer to Figure 7 , the signal extending vertically from low frequency to high frequency at about 1.05 seconds in the figure is a pulse signal, and the pulse signal will overlap with the watermark sound signal S W , which in turn affects the watermark pattern W I The recognition result of .

[0068] The processor 55 transmits the sound signal S A Map to the time-frequency graph and compare several default watermark signals W1,…,W M (Step S250). Specifically, the processor 55 may use Fast Fourier Transform (FFT) or other time domain to frequency domain conversion to convert the transmitted sound signal S A Each unsuperimposed frame in is switched to the frequency domain, and the overall time-frequency graph composed of all frames is considered.

[0069] On the other hand, the default watermark signals W1,…,W M (M is a positive integer) are used to identify different transmitting devices 10 or different users. The default watermark signal has been stored in the memory 53. The default watermark signal W1, ..., W M Corresponding to several default watermark patterns in the two-dimensional coordinate system. Similarly, each default watermark pattern can be designed according to the needs of the user, and the embodiment of the present invention is not limited thereto.

[0070] The processor 55 transmits S A With the default watermark signal W1,…,W M The correlation between the transmitted sound signal S A With the default watermark signal W1,…,W M The comparison result of the watermark sound signal S is identified W (Step S260). Specifically, the correlation in this article is the transmission of the sound signal S A Compared with those default watermark signals W1,…,W M The default watermark signal with the highest similarity is the watermark sound signal S W .

[0071] Figure 8 This is a flow chart of watermark pattern recognition according to an embodiment of the present invention. Figure 8 , the processor 55 determines to transmit the sound signal SA One or more pulse signals τ x (Step S810). Specifically, the pulse signal τ x The characteristic is that all frequencies have interfered signals in a very short time. In one embodiment, the processor 55 can determine the transmitted sound signal S A The power of each of the several sound frames in the time-frequency diagram at several frequencies, and determine those sound frames with the power of those frequencies greater than the threshold as a pulse signal τ x For example, the processor 55 may determine whether the power of all frequencies of a certain sound frame is greater than a set threshold. If this condition is met (ie, the power of all frequencies is greater than the threshold), the processor 55 may determine that the sound frame is affected by the pulse signal τ x In some embodiments, the processor 55 may select specific frequencies (rather than all frequencies) in the spectrum and determine whether the power at these frequencies is greater than a threshold.

[0072] The processor 55 may be configured to process the signal according to one or more pulse signals τ x Modify the default watermark signals W1,…,W M (Step S830). Specifically, the processor 55 generates the pulse signal τ x The frame position (corresponding to a position in the horizontal axis of the two-dimensional coordinate system) will be the default watermark signal W1,…,W M The pulse interference feature is added or subtracted on the vertical axis (corresponding to the frequency axis) in the two-dimensional coordinate system to generate a modified default watermark signal W'1, ..., W' M .

[0073] For example, Figure 9 This is a diagram illustrating an example of modifying the default watermark signal W1. Figure 9 For a position on the X-axis, the processor 55 fills in a straight line pattern of vertical lines (ie, impulse interference features) at each position on the Y-axis to form a modified default watermark signal W'1.

[0074] In one embodiment, the aforementioned correlation includes a first correlation. The processor 55 may determine the transmitted sound signal S A Compared with the default watermark signals W1,…,W that have not been modified M The first correlation of the default watermark signals W1, ..., W M The processor 55 may only modify the default watermark signals W1, ..., W M The processor 55 can, for example, filter out the candidate watermark signals related to the transmitted sound signal S according to a deep learning-based classifier or cross-correlation. ASome candidate watermark signals with high similarity between them. Taking cross correlation as an example, only those whose cross correlation value is greater than the corresponding threshold can be used as candidate watermark signals.

[0075] In one embodiment, the aforementioned correlation includes a second correlation. The processor 55 may determine the transmitted sound signal S A With the modified default watermark signals W1,…,W M Or the second correlation between the candidate watermark signals, and perform pattern recognition based on it (step S850). Specifically, since the watermark sound signal S W The processor 55 can filter out the original transmission sound signal S A The sine wave signal S is located in the middle f1 ,…,S fN For example, the processor 55 transmits the sound signal S A After the high-pass filter with a frequency of 16kHz or above is passed. In addition, the processor 55 can filter out the signals related to the transmitted sound signal S according to a classifier based on deep learning or cross-correlation. A The candidate watermark signal with the highest similarity between them is selected. Taking cross-correlation as an example, the maximum value of the cross-correlation can be used as the identified watermark sound signal S W For example, the default watermark signal W1 has the highest correlation, so the default watermark signal W1 is the watermark sound signal S W .

[0076] In summary, in the voice communication system and audio watermarking processing method of the embodiments of the present invention, a watermark audio signal composed of a superposition of sine wave signals of different frequencies corresponding to multiple audio frames is predefined at the transmitting end. This watermark audio signal can then be embedded into the voice signal in real time, thereby meeting the requirements of real-time conference calls. Furthermore, at the receiving end, pulse signals are determined and their interference with the default watermark signal is considered, allowing for accurate identification of the watermark audio signal and reducing the impact of pulse signal noise.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing a sound watermark, characterized in that: include: generating a plurality of sine wave signals, wherein the sine wave signals have different frequencies and are high-frequency sound signals; Mapping a watermark pattern to a time-frequency graph to form a watermarked sound signal, wherein two dimensions of the watermark pattern in a two-dimensional coordinate system correspond to a time axis and a frequency axis in the time-frequency graph, respectively, and each of a plurality of sound frames on the time axis corresponds to the sine wave signal of a different frequency on the frequency axis; synthesizing the speech signal and the watermark sound signal in the time domain to generate an embedded watermark signal; receiving a transmission sound signal, wherein the transmission sound signal is the transmitted embedded watermark signal; Mapping the transmitted sound signal to the time-frequency graph and comparing the signal with a plurality of default watermark signals, wherein the default watermark signals correspond to a plurality of default watermark patterns in the two-dimensional coordinate system, and the step of comparing the default watermark signals comprises: determining at least one pulse signal in the transmitted sound signal; Modify the default watermark signal according to the at least one pulse signal; and determining a first correlation between the transmitted sound signal and the modified default watermark signal; The watermark sound signal is identified based on the correlation between the transmitted sound signal and the default watermark signal, wherein the correlation is the degree of similarity between the transmitted sound signal and the default watermark signal, the default watermark signal with the highest degree of similarity is the watermark sound signal, and the correlation includes the first correlation.

2. The method for processing a sound watermark according to claim 1, characterized in that: The step of mapping the watermark pattern to the time-frequency graph to form the watermark sound signal comprises: establishing a watermark matrix in the time-frequency graph according to the watermark pattern, wherein the watermark matrix comprises a plurality of elements, each of the elements being one of a marked element and an unmarked element, the marked element representing that the corresponding position of the watermark pattern in the two-dimensional coordinate system has a value, and the unmarked element representing that the corresponding position of the watermark pattern in the two-dimensional coordinate system has no value; selecting at least one of a plurality of sine wave signals in each of the audio frames according to the watermark matrix, wherein at least one selected sine wave signal corresponds to the marked element among the elements; and At least one selected sine wave signal on the sound frame is superimposed in the time domain to form the watermark sound signal.

3. The method for processing a sound watermark according to claim 2, characterized in that: The step of establishing the watermark matrix in the time-frequency graph according to the watermark pattern comprises: The watermark pattern is extended according to a superposition amount in a dimension on the time axis corresponding to the two-dimensional coordinate system, wherein the superposition amount is related to an overlap amount of adjacent sound frames in superposition.

4. The method for processing a sound watermark according to claim 1, characterized in that: The step of synthesizing the speech signal and the watermark sound signal comprises: The sound signal in the frequency band where the sine wave signal is located is filtered out from the voice signal.

5. The method for processing sound watermark according to claim 1, characterized in that: The step of generating the sine wave signal comprises: Setting the duration of the sine wave signal to the sound frame; and The sine wave signal is windowed.

6. The method for processing sound watermark according to claim 1, characterized in that: The correlation includes a second correlation, and before the step of modifying the default watermark signal according to the at least one pulse signal, the method further includes: determining the second correlation between the transmitted sound signal and the unmodified default watermark signal; and A plurality of candidate watermark signals are selected from the default watermark signal according to the second correlation, wherein only the candidate watermark signals in the default watermark signal are modified.

7. The method for processing sound watermark according to claim 1, characterized in that: The step of determining the at least one pulse signal in the transmission sound signal comprises: determining a power of the transmitted sound signal at a plurality of frequencies for each of a plurality of sound frames in the time-frequency graph; and Determine that the sound frame having the frequency and power greater than a threshold is the at least one pulse signal.

8. A voice communication system, characterized in that: include: A conveyor configured to: generating a plurality of sine wave signals, wherein the sine wave signals have different frequencies and are high-frequency sound signals; Mapping a watermark pattern to a time-frequency graph to form a watermarked sound signal, wherein two dimensions of the watermark pattern in a two-dimensional coordinate system correspond to a time axis and a frequency axis in the time-frequency graph, respectively, and each of a plurality of sound frames on the time axis corresponds to the sine wave signal of a different frequency on the frequency axis; synthesizing the speech signal and the watermark sound signal in the time domain to generate an embedded watermark signal; as well as transmitting the embedded watermark signal; as well as A receiving device configured to: receiving a transmission sound signal, wherein the transmission sound signal is the transmitted embedded watermark signal; mapping the transmitted sound signal to the time-frequency map and comparing the transmitted sound signal to a plurality of default watermark signals, wherein the default watermark signals correspond to a plurality of default watermark patterns in the two-dimensional coordinate system, and the receiving device is further configured to determine at least one pulse signal in the transmitted sound signal, modify the default watermark signal according to the at least one pulse signal, and determine a first correlation between the transmitted sound signal and the modified default watermark signal; as well as The watermark sound signal is identified based on the correlation between the transmitted sound signal and the default watermark signal, wherein the correlation is the degree of similarity between the transmitted sound signal and the default watermark signal, the default watermark signal with the highest degree of similarity is the watermark sound signal, and the correlation includes the first correlation.

Citation Information

Patent Citations

  • Voice communication system encoding and decoding voice and non-voice information

    US20130085751A1

  • Indexing based on time-variant transforms of an audio signal's spectrogram

    US20160148620A1

  • Watermarking in the time-frequency domain

    US6674876B1

  • Additional information embedding method and it's device, and additional information decoding method and its decoding device

    US7299189B1