Sound watermark processing method and sound watermark generating device
By generating reflected sound signals and synthesizing watermark sound signals, the problem of decreased watermark recognition accuracy caused by noise interference in remote conferences is solved, and the effect of combating noise interference and improving watermark recognition accuracy is achieved.
Patent Information
- Application Number
- CN202210043439.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-01-14
AI Technical Summary
In remote conferences, noise interference causes the accuracy of watermark recognition to decrease, affecting the quality of the sound signal on the call transmission path.
A reflected sound signal is generated through virtual reflection conditions, and two watermark sound signals are generated using the watermark identifier and the sound signal spacing value. The watermark sound signal is synthesized and output to reduce the overall watermark sound signal power and improve the recognition accuracy of the watermark identifier.
Effectively combat noise interference, improve call quality and increase the recognition accuracy of watermark identifiers.
Smart Images

Figure CN116486823B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sound signal processing, and in particular to a sound watermark processing method and a sound watermark generating device. Background Art
[0002] Remote conferencing allows people in different locations or spaces to communicate, and conferencing-related equipment, protocols, and applications have become quite mature. It is worth noting that some real-time conferencing programs may synthesize voice signals and watermarked audio signals to identify callers.
[0003] Inevitably, if the sound signal is interfered with by noise, the accuracy of the watermark determination at the receiving end will decrease, thereby affecting the user's voice component in the sound signal on the call transmission path. Summary of the Invention
[0004] The present invention provides a sound watermark processing method and a sound watermark generating device. The generated watermark sound signal can effectively resist noise, thereby improving the call quality.
[0005] The sound watermark processing method of the embodiment of the present invention is applicable to a conference terminal, and the conference terminal includes a microphone. The sound watermark processing method includes (but is not limited to) the following steps: obtaining a call receiving sound signal through the microphone. Generating a reflected sound signal based on a virtual reflection condition and the call receiving sound signal. The virtual reflection condition includes the positional relationship between the microphone, the sound source, and the external object, and the reflected sound signal is a sound signal obtained by simulating the sound emitted by the sound source, which is reflected by the external object and recorded by the microphone. A first watermark sound signal is generated based on the watermark identifier and the reflected sound signal. A second watermark sound signal is generated based on the sound signal spacing value and the first watermark sound signal. The sound signal spacing value is determined based on the high-frequency and low-frequency sound ratio of the reflected sound signal, and the sound signal spacing value is related to the distance difference between the two reflection distances of the sound emitted by the sound source, which is reflected by two external objects under the positional relationship and reaches the microphone. The first watermark sound signal and the second watermark sound signal are synthesized to generate an output watermark sound signal.
[0006] According to an embodiment of the present invention, an audio watermark generating device of the embodiment of the present invention includes (but is not limited to) a memory and a processor. The memory is used to store program code. The processor is coupled to the memory. The processor is configured to load and execute the program code to obtain a call receiving audio signal, and generate a reflected audio signal based on a virtual reflection condition and the call receiving audio signal. The virtual reflection condition includes the positional relationship between the microphone, the sound source, and the external object, and the reflected audio signal is a sound signal obtained by simulating the sound emitted by the sound source reflected by the external object and recorded by the microphone. A first watermark audio signal is generated based on the watermark identifier and the reflected audio signal. A second watermark audio signal is generated based on the audio signal spacing value and the first watermark audio signal. The audio signal spacing value is determined based on the proportion of high and low frequency sounds in the reflected audio signal, and the audio signal spacing value is related to the distance difference between the two reflection distances of the sound emitted by the sound source reflected by two external objects under the positional relationship and reaching the microphone. The first watermark audio signal and the second watermark audio signal are synthesized to generate an output watermark audio signal.
[0007] Based on the foregoing, the audio watermark processing method and audio watermark generation device according to embodiments of the present invention determine the audio signal spacing between two simulated reflected audio signals based on the high- and low-frequency ratios of the received call audio signal, and generate two watermark audio signals accordingly. By outputting the synthesized two watermark audio signals, the power of the overall watermark audio signal can be reduced, and the accuracy of watermark identifier identification can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The accompanying drawings are included to provide a further understanding of the present invention and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present invention and together with the description serve to explain the principles of the present invention.
[0009] Figure 1 is a schematic diagram of a conference call system according to an embodiment of the present invention;
[0010] Figure 2 is a flow chart of a method for processing a sound watermark according to an embodiment of the present invention;
[0011] Figure 3 is a flow chart of a method for generating a sound watermark according to an embodiment of the present invention;
[0012] Figure 4 is a schematic diagram illustrating a virtual reflection condition according to an embodiment of the present invention;
[0013] Figure 5 is a flow chart illustrating watermark recognition according to an embodiment of the present invention;
[0014] Figure 6A This is a simulation diagram illustrating an example of a call receiving audio signal;
[0015] Figure 6B This is a simulation diagram illustrating an example of transmission noise.
[0016] Explanation of Figure Numbers
[0017] 10, 20: conference terminal;
[0018] 50: cloud server;
[0019] 11, 21: radio;
[0020] 13, 23: speakers;
[0021] 15, 25, 55: communication transceiver;
[0022] 17, 27, 57: memory;
[0023] 19, 29, 59: processor;
[0024] 70: sound watermark generating device;
[0025] S210-S290, S310-S330, S510-S595: steps;
[0026] S Rx : Call receiving sound signal;
[0027] S Tx : Call transmission sound signal;
[0028] S WM 、S' WM , S” WM : watermark sound signal;
[0029] S Rx +S WM : Embed watermark signal;
[0030] Δn A : Sound signal spacing value;
[0031] S' Rx , S” Rx 、 Reflect sound signals;
[0032] W1, W2: wall;
[0033] d s d w1 d w2 :distance;
[0034] SS: sound source;
[0035] W E: watermark identifier;
[0036] S A 、 Transmitting sound signals;
[0037] HPF: high-pass filtering;
[0038] LPF: Low-pass filtering. DETAILED DESCRIPTION
[0039] Reference will now be made in detail to exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.
[0040] Figure 1 is a schematic diagram of a conference call system 1 according to an embodiment of the present invention. Figure 1 The voice communication system 1 includes but is not limited to conference terminals 10 , 20 and a cloud server 50 .
[0041] The conference terminals 10 , 20 may be wired telephones, mobile phones, Internet phones, tablet computers, desktop computers, laptop computers, or smart speakers.
[0042] The conference terminal 10 includes (but is not limited to) a sound receiver 11 , a speaker 13 , a communication transceiver 15 , a memory 17 and a processor 19 .
[0043] The microphone 11 can be a dynamic, condenser, or electret condenser microphone. Alternatively, the microphone 11 can be a combination of other electronic components, analog-to-digital converters, filters, and audio processors that receive sound waves (e.g., human voices, ambient sounds, machine operation sounds, etc.) and convert them into sound signals. In one embodiment, the microphone 11 is used to receive / record the caller's voice to obtain a call reception audio signal. In some embodiments, this call reception audio signal may include the caller's voice, the sound emitted by the speaker 13, and / or other ambient sounds.
[0044] Loudspeaker 13 can be a speaker or a loudspeaker. In one embodiment, loudspeaker 13 is used for playing sound.
[0045] The communication transceiver 15 may be, for example, a transceiver that supports wired networks such as Ethernet, fiber optic networks, or cables (which may include, but is not limited to, connection interfaces, signal converters, communication protocol processing chips, and other components). It may also be a transceiver that supports wireless networks such as Wi-Fi, fourth-generation (4G), fifth-generation (5G), or later-generation mobile networks (which may include, but is not limited to, antennas, digital-to-analog / analog-to-digital converters, communication protocol processing chips, and other components). In one embodiment, the communication transceiver 15 is used to transmit or receive data.
[0046] The memory 17 can be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or similar device. In one embodiment, the memory 17 is used to store program code, software modules, configurations, data (e.g., audio signals, watermark identifiers, or watermarked audio signals), or files.
[0047] Processor 19 is coupled to receiver 11, speaker 13, communication transceiver 15, and memory 17. Processor 19 can be a central processing unit (CPU), a graphics processing unit (GPU), or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), programmable logic controller, field programmable gate array (FPGA), application-specific integrated circuit (ASIC), or other similar components or combinations thereof. In one embodiment, processor 19 is used to perform all or part of the operations of conference terminal 10 and can load and execute various software modules, files, and data stored in memory 17.
[0048] Conference terminal 20 includes (but is not limited to) a microphone 21, a speaker 23, a communication transceiver 25, a memory 27, and a processor 29. The implementation and functions of microphone 21, speaker 23, communication transceiver 25, memory 27, and processor 29 can be found in the aforementioned description of microphone 11, speaker 13, communication transceiver 15, memory 17, and processor 19 and are not further elaborated here. Processor 29 is used to perform all or part of the operations of conference terminal 20 and can load and execute various software modules, files, and data stored in memory 27.
[0049] The cloud server 50 is directly or indirectly connected to the conference terminals 10 and 20 via a network. The cloud server 50 can be a computer system, a server, or a signal processing device. In one embodiment, the conference terminals 10 and 20 can also serve as the cloud server 50. In another embodiment, the cloud server 50 can be a separate cloud server from the conference terminals 10 and 20. In some embodiments, the cloud server 50 includes (but is not limited to) the same or similar communication transceiver 55, memory 57, and processor 59, and the implementation and functions of these components will not be detailed again.
[0050] In one embodiment, the sound watermark generating device 70 may be the conference terminal 10, 20 or the cloud server 50. The sound watermark generating device 70 is used to generate a watermark sound signal, which will be described in detail in subsequent embodiments.
[0051] Hereinafter, the method of the embodiment of the present invention will be described with reference to various devices, components and modules in the conference communication system 1. Each process of the method can be adjusted according to the implementation situation and is not limited thereto.
[0052] It should also be noted that for ease of description, identical components may implement identical or similar operations and will not be described in detail. For example, the processor 19 of the conference terminal 10, the processor 19 of the conference terminal 20, and / or the processor 59 of the cloud server 50 may all implement the identical or similar methods of the embodiments of the present invention.
[0053] Figure 2 This is a flow chart of a method for processing a sound watermark according to an embodiment of the present invention. Figure 2 The processor 29 records the call receiving sound signal S through the microphone 21. Rx (Step S210). Specifically, it is assumed that the conference terminals 10 and 20 establish a call conference. For example, the conference is established through video software, voice call software, or by making a phone call, and the speaker can start speaking. After the microphone 21 records / receives the sound, the processor 29 can obtain the call reception sound signal S Rx This call receives the sound signal S RxThe processor 29 of the conference terminal 20 can transmit the call receiving sound signal S through the communication transceiver 25 (ie, via the network interface) to the speech content of the speaker corresponding to the conference terminal 20 (which may also include environmental sounds or other noises). Rx In some embodiments, the call receiving sound signal S Rx The audio signal may be processed by echo cancellation, noise filtering, and / or other sound processing.
[0054] The processor 59 of the cloud server 50 receives the call receiving sound signal S from the conference terminal 20 through the communication transceiver 55. Rx The processor 59 generates a reflected sound signal S' according to the virtual reflection condition and the call receiving sound signal. Rx (Step S230). Specifically, a general echo cancellation algorithm can adaptively eliminate the reference signal component (for example, the call receiving sound signal S in the call receiving path) in the sound signal received by the microphones 11 and 21 from the outside. Rx ). The sound recorded by the microphones 11, 21 includes the shortest path from the speakers 13, 23 to the microphones 11, 21 and different reflection paths of the environment (i.e., the path formed by the sound reflected by external objects). The position of the reflection affects the time delay and attenuation of the sound signal. In addition, the reflected sound signal may also come from different directions, resulting in phase shift. In the embodiment of the present invention, the sound signal S of the known call receiving path is used. Rx To generate a virtual / simulated reflected sound signal that can be eliminated by the echo cancellation mechanism, and to generate a watermark sound signal S WM .
[0055] In one embodiment, the processor 59 may determine the reflected sound signal S' according to the position relationship. Rx Compared to the call receiving sound signal S Rx For example, Figure 4 is a schematic diagram illustrating virtual reflection conditions according to an embodiment of the present invention. Figure 4 Assume that the virtual reflection condition is two walls (ie, two external objects), and the distance between the microphone 21 and the sound source SS is d s (eg, 0.3, 0.5 or 0.8 meters) and the distance between the microphone 21 and the wall W1 is dw1 (eg, 1, 1.5 or 2 meters), the first reflected sound signal S' Rx Receive voice signal S Rx The relationship can be expressed as follows:
[0056] s′ Rx (n) = α1·s Rx (nn w1 )…(1)
[0057] Where α1 is the amplitude attenuation caused by the first reflection (i.e., the reflection of the sound signal blocked by the wall W1), n is the sampling point or time, n w1 is the time delay caused by the first reflection distance (ie, the distance from the sound source SS through the wall W1 to the microphone 21).
[0058] Please refer to Figure 2 , the processor 59 generates a first watermark sound signal based on the watermark identifier and the reflected sound signal (step S250). Specifically, the processor 59 shifts the phase of the reflected sound signal based on the watermark identifier to generate the first watermark sound signal. When a general echo cancellation mechanism is operating, the time delay and amplitude changes of the reflected sound signal have a greater impact on the error of the echo cancellation mechanism than the phase shift of the reflected sound signal. This change is like being in a completely new interference environment, and requires the echo cancellation mechanism to re-adapt. Therefore, the first watermark sound signals corresponding to different values in the watermark identifier of the embodiment of the present invention have only phase differences, but the time delay and amplitude are the same. That is, the first watermark sound signal includes one or more phase-shifted reflected sound signals.
[0059] In one embodiment, the processor 59 may select a filter to generate a filtered reflected sound signal. Specifically, typical echo cancellation mechanisms converge more slowly for low-frequency sound signals (e.g., below 2 kHz or 3 kHz), but converge more quickly (e.g., below 10 milliseconds (ms)) for high-frequency sound signals (e.g., above 3 kHz or 4 kHz). Therefore, the processor 59 may shift the phase of only the reflected sound signal (e.g., the aforementioned first reflected sound signal) that has passed high-pass filtering (e.g., only allowing sound signals with frequencies above 3 kHz or 4 kHz to pass) based on the watermark identifier, thereby making the signal interference less noticeable to humans (i.e., the frequency of the high-frequency sound signal is outside the human hearing range).
[0060] In another embodiment, the processor 59 may not perform frequency-specific filtering on the reflected sound signal.
[0061] In one embodiment, the watermark identifier is encoded in a multi-bit system, and this multi-bit system provides multiple values for each of one or more bits of the watermark identifier. Taking the binary system as an example, the value of each bit in the watermark identifier can be "0" or "1". Taking the hexadecimal system as an example, the value of each bit in the watermark identifier can be "0", "1", "2", ..., "E", "F". In another embodiment, the watermark identifier is encoded using letters, words, and / or symbols. For example, the value of each bit in the watermark identifier can be any of the English letters "A" to "Z".
[0062] In one embodiment, different values of each bit of the watermark identifier correspond to different phase offsets. For example, assuming that the watermark identifier WO is in N-base (N is a positive integer), N values can be provided for each bit. These N different values correspond to different phase offsets. For another example, assuming that the watermark identifier WO is binary, two values (ie, 1 and 0) can be provided for each bit. These two different values correspond to two phase offsets respectively. For example, phase shift is 90° and the phase shift is -90° (i.e., -1).
[0063] The processor 59 can shift the phase of the reflected sound signal (with or without high-pass filtering) according to the value of one or more bits in the watermark identifier. Taking the N-ary system as an example, the processor 59 selects the phase shift according to one or more values in the watermark identifier. One or more of and using the selected phase offset For example, if the value of the first bit of the watermark identifier is 1, the output phase-shifted reflected sound signal Offset relative to reflected sound signal Other reflected sound signals The same can be deduced. The phase shift can be achieved by using Hilbert transform or other phase shift algorithms.
[0064] In one embodiment, if filtering is applied to the reflected sound signal, the processor 59 may further synthesize one or more phase-shifted reflected sound signals and a reflected sound signal (e.g., the first reflected sound signal) that has been low-pass filtered (e.g., only allowing sound signals with frequencies below 4 kHz to pass) to generate the first watermark sound signal. In another embodiment, if filtering is not applied to the reflected sound signal, the processor 59 may use the one or more phase-shifted reflected sound signals as the first watermark sound signal.
[0065] Please refer to Figure 2 The processor 59 generates a second watermark sound signal based on the sound signal spacing value and the first watermark sound signal (step S270). Specifically, the second watermark sound signal is another reflected sound signal (hereinafter referred to as the second reflected sound signal) corresponding to the aforementioned first reflected sound signal, and is related to the difference in time delay between the two reflected sound signals. Figure 4 For example, assuming that the first reflected sound signal S' Rx is a simulated sound signal reflected by the wall W1, then the second reflected sound signal S″ RxIt is a simulated sound signal reflected by the wall W2. Under the condition that the distance between the microphone 21 and the other wall W2 is dw2 (for example, 1, 1.5 or 2 meters), the second reflected sound signal S″ Rx Receive voice signal S Rx The relationship can be expressed as follows:
[0066] S″ Rx (n) = α2·S Rx (nn w2 )…(2)
[0067] Where α2 is the amplitude attenuation caused by the second reflection (i.e., the reflection of the sound signal blocked by the wall W2), n is the sampling point or time, n w2 is the time delay caused by the second reflection distance (ie, the distance from the sound source SS through the wall W2 to the microphone 21). In other words, the two reflected sound signals simulate the sound signals reflected by two external objects respectively.
[0068] It is worth noting that the difference between the time delay caused by the second reflection distance and the time delay caused by the first reflection distance (or the difference between the propagation time of the sound signal reflected by two external objects) (i.e., the sound signal spacing value Δn) can be expressed as follows:
[0069] Δn=n w2 -n w1 …(3)
[0070] The primary cause of sound delay is the distance over which sound signals travel. Therefore, the sound signal spacing value is also related to the difference between the two reflection distances that sound emitted by sound source SS travels after being reflected from two external objects (e.g., walls W1 and W2) and reaching microphone 21, under the positional relationship of the set virtual reflection conditions.
[0071] Assume that the sound signal spacing value Δn is much smaller than the time delay corresponding to any reflected signal (for example, Δn<<n w1 ), the two reflection distances (for example, the first reflection distance and the second reflection distance) are almost equal or completely equal, and the amplitude attenuation of the two reflected sound signals (for example, the first reflected sound signal and the second reflected sound signal) should also be almost equal or completely equal (for example, ). Therefore, the low-frequency parts of the two reflected sound signals after superposition / synthesis cancel each other out, thereby reducing the power of the overall watermark sound signal, making it difficult for users to perceive the added watermark sound signal.
[0072] It is worth noting that the call receiving sound signal S Rx It may change with time. The experiment found that if the sound signal spacing value Δn can be changed with the call receiving sound signal S RxIn the embodiment of the present invention, the sound signal spacing value is determined according to the proportion of high and low frequency sounds of the reflected sound signal (eg, the first reflected sound signal).
[0073] In one embodiment, after generating the reflected sound signal, the processor 59 performs a low-pass filter on the reflected sound signal to generate a low-frequency sound signal. Furthermore, the processor 59 performs a high-pass filter on the reflected sound signal to generate a high-frequency sound signal. The high-low frequency sound ratio is the power ratio between the low-frequency sound signal and the high-frequency sound signal.
[0074] Figure 3 FIG is a flow chart of a method for generating a sound watermark SWM according to an embodiment of the present invention. Figure 3 The processor 59 generates a signal based on the low-frequency sound signal in the reflected sound signal. (For example, sound signals below 2kHz) and high-frequency sound signals (For example, a sound signal above 2kHz) determines the sound signal spacing value Δn (step S310). In one embodiment, if the high frequency sound signal The power is not less than the low-frequency sound signal If the high-frequency sound signal The power is less than the low-frequency sound signal The processor 59 may set the sound signal spacing value to a second value, where the first value is greater than the second value.
[0075] For example, when a call receives a voice signal S Rx High-frequency sound signals The power is not less than its low-frequency sound signal When the voice signal interval value Δn is set to 5 (ie, the first value). Rx High-frequency sound signals The power is less than its low-frequency sound signal When the sound signal spacing value Δn is set to 4 (ie, the second value). and high-frequency sound signals The relationship between can be expressed as follows:
[0076]
[0077] Receive voice signal S for call Rx High-frequency sound signal power, Receive voice signal S for callRx The low-frequency sound signal power. In other words, the proportion of high and low frequency sound is or Furthermore, since the reflected sound signal reflects the incoming call sound signal, changes in the incoming call sound signal also alter the reflected sound signal, and the sound signal spacing value Δn must also be dynamically adjusted. Experiments have shown that dynamic spacing helps improve the accuracy of watermark recognition. It should be noted that the values of the first and second values can still be changed according to actual needs and are not limited by the present embodiment.
[0078] Please refer to Figure 3 The processor 59 is based on the sound signal interval Δn and the first watermark sound signal S′ WM Generate a second watermark sound signal S″ WM (Step S330). Specifically, the second watermark audio signal S″ WM and the first watermark sound signal S′ WM The relationship between the sound signal spacing values Δn with opposite phases and under the above-mentioned virtual reflection conditions can be expressed as follows:
[0079] S″ WM (n) = -S′ WM (n-Δn)…(5)
[0080] That is, the second watermark sound signal S″ WM is the first watermark sound signal S′ with an inverted phase and a time delay of Δn WM .
[0081] Please refer to Figure 2 and Figure 3 , the processor 59 synthesizes the first watermark sound signal S′ WM and the second watermark sound signal S″ WM , to generate the output watermark sound signal S WM (Step S290). In one embodiment, the processor 59 further synthesizes and outputs the watermark audio signal S WM Receive voice signal S Rx , to generate the embedded watermark signal S Rx +S WM and transmits the embedded watermark signal S via the communication transceiver 55 Rx +S WM In another embodiment, the processor 59 transmits the output watermark sound signal S through the communication transceiver 55. WM And call receiving sound signal S Rx .
[0082] The processor 19 of the conference terminal 10 receives the watermarked audio signal S via the network through the communication transceiver 15. WMOr embed watermark signal S Rx +S WM , to obtain the transmitted sound signal S A (ie, the transmitted watermarked sound signal S WM Or embed watermark signal S Rx +S WM ). Since the watermark sound signal S WM The received voice signal (ie, the reflected voice signal) includes a time-delayed and amplitude-attenuated voice signal, so the echo cancellation mechanism of the processor 19 can effectively eliminate the watermarked voice signal S WM In this way, the call transmission voice signal S on the communication transmission path will not be affected. Tx (For example, the conference terminal 10 intends to transmit a call reception audio signal via the network).
[0083] For the watermark sound signal S WM Identification, Figure 5 This is a flow chart illustrating watermark recognition according to an embodiment of the present invention. Figure 5 In one embodiment, the processor 19 may use the same or similar high-pass filter HPF as described above to process the transmitted sound signal S A Perform high-pass filtering (step S510) to output the transmitted audio signal after high-pass filtering. In another embodiment, if the transmitting end does not adopt the filtering process, step S510 (ie, transmitting the sound signal) can be omitted. Equivalent to transmitting sound signal S A In one embodiment, the processor may use the same or similar low-pass filter LPF as described above to process the transmitted sound signal S A Perform low-pass filtering (step S530) to output the transmitted audio signal after the low-pass filtering.
[0084] 6 , the processor 19 shifts the transmitted sound signal S A phase to generate a first offset sound signal (Step S550). It should be noted that this embodiment uses a binary-coded watermark identifier as an example (i.e., only two values are provided), and these two values correspond to phase shifts of 90° and -90°, respectively. However, if other encodings are used, different phase shifts may be possible. Next, the processor 19 processes the transmitted sound signal according to the low-pass filter LPF. Estimate the sound signal spacing value Δn A (Step S570) It should be noted that if the transmitting end adopts filtering processing and only encodes the high-frequency sound signal based on the watermark identifier, it means that the low-frequency sound signal is not affected by the watermark identifier and is helpful to estimate the sound signal spacing value Δn A.
[0085] In one embodiment, the processor 19 may transmit the sound signal Correlation estimation of sound signal spacing Δn at different time delays A For example, the processor 19 measures the transmission sound signal after the low-pass filter LPF by using an auto-cepstrum function (e.g., Mel-Frequency Cepstrum Coefficient (MFCC) or Linear Prediction Cepstrum Coefficient (LPCC)) or other autocorrelation functions. The sound signal spacing value Δn corresponding to the local maximum (Local Maximum) A For example, the sound signal spacing value Δn A 3 or 4.
[0086] The processor 19 generates a first offset sound signal according to the first offset sound signal. And the estimated sound signal spacing value Δn A Generate a second offset sound signal (Step S590) Regarding the second offset sound signal With the first offset sound signal The relationship can be expressed as follows:
[0087]
[0088] That is, the second offset sound signal is the first offset sound signal after a time delay of Δn
[0089] The processor 19 can determine the first offset sound signal and transmit sound signals (S A or ) between them (ie, the first correlation), and determining the second offset sound signal and transmit sound signals (S A or ) to obtain the correlation coefficient. For example, the processor 19 converts the first offset sound signal and transmit sound signals (S A or ) Calculate the cross correlation to get the first correlation And the second offset sound signal and transmit sound signals (S A or ) Calculate the cross correlation to obtain the second correlation The processor 19 converts the first correlation Second correlation Subtract to get the correlation coefficient The correlation coefficient It can be expressed as follows:
[0090]
[0091] The processor 19 can be used according to the correlation coefficient Identify the watermark identifier (step S595). For example, if the processor 19 defines the threshold value Th R (e.g., 0.3, 0.5, or 0.7), the identified watermark identifier W E It can be expressed as:
[0092]
[0093] That is, if the correlation coefficient Above the threshold value Th R , the processor 19 determines that the value of this bit is a value corresponding to a phase shift of 90° (for example, 1); if the correlation coefficient Below the threshold value Th R , the processor 19 determines that the value of this bit is a value corresponding to a phase shift of -90° (for example, 0).
[0094] The following is supplemented by experimental instructions. Figure 6A This is an example of a call receiving audio signal S Rx Please refer to the simulation diagram. Figure 6A , assuming that the call receiving sound signal S Rx The first half of is a white noise sound signal, and the second half is a pink noise sound signal. Figure 6B This is a simulation diagram showing an example of transmission noise NT. Figure 6B , assuming that the sound signal output during the transmission process (for example, the embedded watermark signal S Rx +S WM Or output watermark sound signal S WM ) is attenuated. This attenuation characteristic is 0≤α T ≤1 (e.g., α T = 0.5 or 0.3) and is affected by the transmission noise N T If the transmission noise N T Power P N The larger the value, the more difficult it is for the receiver to determine the watermark identifier. Figure 6BThe transmission noise NT shown is a white noise sound signal with a power of P N Equal to the call receiving sound signal S Rx The power (ie, the same as the call receiving sound signal S Rx Experiments have shown that using a dynamic sound signal spacing value allows for completely accurate watermark identification. For example, the cross-correlation ratio between the watermarked and non-watermarked sound signals is 9.56. A higher ratio indicates a wider range of recognition and more accurate results.
[0095] In summary, in the audio watermark processing method and audio watermark generation device of the embodiments of the present invention, the audio signal spacing between two reflected audio signals to be simulated is dynamically determined based on the power ratio between the high-frequency and low-frequency audio signals in the audio signal. Based on this spacing, two watermark audio signals corresponding to the two reflected audio signals are generated. This reduces the power of the overall watermark audio signal and improves the accuracy of watermark identifier recognition.
[0096] Although the present invention has been disclosed above by way of embodiments, they are not intended to limit the present invention. Any person skilled in the art may make slight changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for processing a sound watermark, applicable to a conference terminal, wherein the conference terminal includes a microphone, characterized in that: The sound watermark processing method includes: Acquiring a call receiving sound signal through the microphone; generating a reflected sound signal based on a virtual reflection condition and the call received sound signal, wherein the virtual reflection condition includes a positional relationship between the microphone, the sound source, and two external objects, and the reflected sound signal is a sound signal simulating a sound emitted by the sound source reflected by one of the external objects and recorded by the microphone; generating a first watermark sound signal according to the watermark identifier and the reflected sound signal; generating a second watermark sound signal based on a sound signal spacing value and the first watermark sound signal, wherein the sound signal spacing value is determined based on a ratio of high-frequency and low-frequency sounds in the reflected sound signal, and the sound signal spacing value is related to a distance difference between two reflection distances of the sound emitted by the sound source at the positional relationship, respectively reflected by the two external objects and reaching the sound receiver; and The first watermarked audio signal and the second watermarked audio signal are synthesized to generate an output watermarked audio signal.
2. The method for processing a sound watermark according to claim 1, characterized in that: After the step of generating the reflected sound signal according to the virtual reflection condition and the call received sound signal, the method further includes: performing low-pass filtering on the reflected sound signal to generate a low-frequency sound signal; and The reflected sound signal is subjected to high-pass filtering to generate a high-frequency sound signal, and the high-low frequency sound ratio is the power ratio between the low-frequency sound signal and the high-frequency sound signal.
3. The method for processing a sound watermark according to claim 2, characterized in that: The step of generating the second watermark sound signal according to the sound signal spacing value and the first watermark sound signal comprises: In response to the power of the high-frequency sound signal being not less than the power of the low-frequency sound signal, setting the sound signal spacing value to a first value; and In response to the power of the high-frequency sound signal being less than the power of the low-frequency sound signal, the sound signal interval value is set to a second value, and the first value is greater than the second value.
4. The method for processing a sound watermark according to claim 2, wherein: The step of generating the first watermark sound signal according to the watermark identifier and the reflected sound signal comprises: shifting the phase of the reflected sound signal processed by the high-pass filter only according to the watermark identifier; and At least one phase-shifted reflected sound signal and the reflected sound signal processed by the low-pass filter are synthesized to generate the first watermarked sound signal.
5. The method for processing a sound watermark according to claim 4, characterized in that: Also includes: receiving a transmission sound signal via a network, the transmission sound signal including the transmitted output watermark sound signal; shifting the phase of the transmitted sound signal to generate a first shifted sound signal; estimating the sound signal interval value based on the transmitted sound signal processed by the low-pass filter; generating a second offset sound signal according to the first offset sound signal and the estimated sound signal interval value; as well as The watermark identifier is identified based on a first correlation and a second correlation, wherein the first correlation is a correlation between the first offset sound signal and the transmission sound signal, and the second correlation is a correlation between the second offset sound signal and the transmission sound signal.
6. The method for processing a sound watermark according to claim 5, characterized in that: Before the step of identifying the watermark identifier, the method further includes: performing high-pass filtering on the transmitted sound signal, The first correlation is a correlation between the first offset sound signal and the transmission sound signal processed by the high-pass filter, and the second correlation is a correlation between the second offset sound signal and the transmission sound signal processed by the high-pass filter.
7. The method for processing sound watermark according to claim 1, characterized in that: The step of generating the reflected sound signal according to the virtual reflection condition and the call receiving sound signal includes: Determining the time delay and amplitude attenuation of the reflected sound signal compared to the call received sound signal based on the positional relationship between the sound source and each of the external objects, The sound signal interval value is the difference between the time delays corresponding to the two external objects.
8. A sound watermark generating device, comprising: a memory for storing program codes; as well as a processor coupled to the memory and configured to load and execute the program code to: Obtain call receiving sound signal through the microphone; generating a reflected sound signal based on a virtual reflection condition and the call received sound signal, wherein the virtual reflection condition includes a positional relationship between the microphone, the sound source, and two external objects, and the reflected sound signal is a sound signal simulating a sound emitted by the sound source reflected by one of the external objects and recorded by the microphone; generating a first watermark sound signal according to the watermark identifier and the reflected sound signal; generating a second watermark sound signal based on a sound signal spacing value and the first watermark sound signal, wherein the sound signal spacing value is determined based on a ratio of high-frequency and low-frequency sounds in the reflected sound signal, and the sound signal spacing value is related to a distance difference between two reflection distances of the sound emitted by the sound source at the positional relationship, respectively reflected by the two external objects and reaching the sound receiver; as well as The first watermarked audio signal and the second watermarked audio signal are synthesized to generate an output watermarked audio signal.
9. The sound watermark generating device according to claim 8, characterized in that: The processor is further configured to: performing low-pass filtering on the reflected sound signal to generate a low-frequency sound signal; as well as The reflected sound signal is subjected to high-pass filtering to generate a high-frequency sound signal, and the high-low frequency sound ratio is the power ratio between the low-frequency sound signal and the high-frequency sound signal.
10. The sound watermark generating device according to claim 9, characterized in that: The processor is further configured to: In response to the power of the high-frequency sound signal being not less than the power of the low-frequency sound signal, setting the sound signal spacing value to a first value; and In response to the power of the high-frequency sound signal being less than the power of the low-frequency sound signal, the sound signal interval value is set to a second value, and the first value is greater than the second value.
11. The sound watermark generating device according to claim 9, characterized in that The processor is further configured to: shifting the phase of the reflected sound signal processed by the high-pass filter only according to the watermark identifier; At least one phase-shifted reflected sound signal and the reflected sound signal processed by the low-pass filter are synthesized to generate the first watermarked sound signal.
12. The sound watermark generating device according to claim 10, characterized in that: The processor is further configured to: receiving a transmission sound signal via a network, the transmission sound signal including the transmitted output watermark sound signal; shifting the phase of the transmitted sound signal to generate a first shifted sound signal; estimating the sound signal interval value based on the transmitted sound signal processed by the low-pass filter; generating a second offset sound signal according to the first offset sound signal and the estimated sound signal interval value; The watermark identifier is identified based on a first correlation and a second correlation, wherein the first correlation is a correlation between the first offset sound signal and the transmission sound signal, and the second correlation is a correlation between the second offset sound signal and the transmission sound signal.
13. The sound watermark generating device according to claim 12, characterized in that: The processor is further configured to: performing high-pass filtering on the transmitted sound signal, The first correlation is a correlation between the first offset sound signal and the transmission sound signal processed by the high-pass filter, and the second correlation is a correlation between the second offset sound signal and the transmission sound signal processed by the high-pass filter.
14. The sound watermark generating device according to claim 8, characterized in that: The processor is further configured to: Determining the time delay and amplitude attenuation of the reflected sound signal compared to the call received sound signal based on the positional relationship between the sound source and each of the external objects, The sound signal interval value is the difference between the time delays corresponding to the two external objects.
Citation Information
Patent Citations
Audio watermark embedding and extracting method and device
CN103413552A
Method and apparatus for separating sound-source signal and method and device for detecting pitch
CN1658283A