Method for processing an audio watermark and audio watermark generating device
By generating a sound watermark that can be eliminated by an echo cancellation mechanism, and by using virtual reflection conditions and phase shift technology, the problem that the sound watermark signal cannot be eliminated on the call transmission path is solved, thus improving call quality.
Patent Information
- Application Number
- CN202110914948.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-08-10
AI Technical Summary
In existing technologies, voice watermark signals cannot be effectively eliminated during call transmission, affecting call quality.
By generating a sound watermark that can be eliminated by an echo cancellation mechanism, and using virtual reflection conditions and phase shift technology, the reflected sound signal is simulated and the watermark signal is encoded so that the call reception signal and the watermark signal are retained at the speaker end at the same time.
This method achieves simultaneous preservation of the received call signal and the watermark signal under the echo cancellation algorithm, thereby improving call quality and avoiding the impact of the watermark signal on the call transmission path.
Smart Images

Figure CN115705847B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a sound signal processing technique, and more particularly to a method for processing a sound watermark and a sound watermark generating apparatus. BACKGROUND
[0002] Teleconference allows people in different locations or spaces to have a conversation, and the conference related equipments, protocols and applications are quite mature. It is worth noting that some real-time conference programs can synthesize a voice signal and a sound watermark signal, and use it to identify the caller.
[0003] For example, Figure 1 is a schematic diagram illustrating a mobile device M used for conference call. Please refer to Figure 1 The mobile device M can receive a sound signal S1 via a network. The sound signal S1 includes a call receiving signal obtained by recording a speaker and a sound watermark signal. The sound watermark signal can be used to identify another device transmitting the sound signal S1. The call receiving signal can be further played through a speaker S so that a user sp of the mobile device M can listen to the voice of the other party. On the other hand, a microphone R (e.g., a microphone) records the user sp to obtain a sound signal S2.
[0004] Generally, the main function of echo cancellation C on the call transmission path is to eliminate the component of the call receiving signal in the sound signal S2 received by the microphone R, thereby obtaining a sound signal S3 without echo. However, the generation path of the sound watermark signal can be different from the path of the general call receiving signal. When the microphone R receives the sound signal of the speaker S via the feedback path fp, the component of the sound watermark signal in the sound signal S1 can not be eliminated and further transmitted via the network, thereby affecting the voice component of the user sp in the sound signal S3 on the call transmission path. SUMMARY
[0005] The present invention is directed to a method for processing a sound watermark and a sound watermark generating apparatus, which generates a sound watermark that can be eliminated by an echo cancellation mechanism, thereby improving the call quality.
[0006] According to an embodiment of the present application, a method for processing a sound watermark is applied to a conference terminal, and the conference terminal includes a microphone. The method for processing a sound watermark includes (but not limited to) the following steps: obtaining a call-received sound signal via the microphone. A reflected sound signal is generated according to a virtual reflection condition and the call-received sound signal. The virtual reflection condition includes a positional relationship among the microphone, a sound source and an external object, and the reflected sound signal is a sound signal simulated from a sound emitted by the sound source, reflected by the external object and recorded by the microphone. A phase of the reflected sound signal is shifted according to a watermark identifier, so as to generate a watermark sound signal. The watermark sound signal includes the phase-shifted reflected sound signal.
[0007] According to an embodiment of the present application, a sound watermark generation device includes (but not limited to) a memory and a processor. The memory is used to store program codes. The processor is coupled to the memory. The processor is configured to load and execute the program codes to obtain a call-received sound signal, generate a reflected sound signal according to a virtual reflection condition and the call-received sound signal, and shift a phase of the reflected sound signal according to a watermark identifier, so as to generate a watermark sound signal. The call-received sound signal is obtained via a microphone. The virtual reflection condition includes a positional relationship among the microphone, a sound source and an external object, and the reflected sound signal is a sound signal simulated from a sound emitted by the sound source, reflected by the external object and recorded by the microphone. The watermark sound signal includes the phase-shifted reflected sound signal.
[0008] Based on the above, according to the method for processing a sound watermark and the sound watermark generation device of an embodiment of the present application, a sound signal reflected by an external object is simulated, and the simulated sound signal is encoded by shifting a phase, so as to generate a watermark sound signal. In this way, a general call-received signal and a sound watermark signal can be simultaneously preserved at a speaker end. Moreover, both of the two signals can be eliminated by an existing echo cancellation algorithm, so that a voice signal on a call transmission path is not affected. BRIEF DESCRIPTION OF DRAWINGS
[0009] The accompanying drawings are included to provide a further understanding of the present application, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application.
[0010] Figure 1 is a schematic diagram of a mobile device for a conference call;
[0011] Figure 2 is a schematic diagram of a conference call system according to an embodiment of the present application;
[0012] Figure 3 is a flowchart of a method for processing a sound watermark according to an embodiment of the present application;
[0013] Figure 4 is a flowchart of a method of generating a sound watermark according to an embodiment of the present invention;
[0014] Figure 5 is a schematic diagram illustrating a virtual reflection condition according to an embodiment of the present invention;
[0015] Figure 6 is a schematic diagram illustrating a filtering process according to an embodiment of the present invention;
[0016] Figure 7 is a schematic diagram illustrating a multi-phase offset according to an embodiment of the present invention;
[0017] Figure 8 is a schematic diagram illustrating a two-phase offset according to an embodiment of the present invention;
[0018] Figure 9A is a simulation diagram illustrating a call receiving sound signal according to an embodiment of the present invention;
[0019] Figure 9B is a simulation diagram illustrating an embedded watermark signal according to an embodiment of the present invention;
[0020] Figure 10 is a flowchart of a method of watermark identification according to an embodiment of the present invention.
[0021] BRIEF DESCRIPTION OF DRAWINGS
[0022] M: mobile device;
[0023] S1-S3: sound signal;
[0024] S: speaker;
[0025] R: receiver;
[0026] sp: user;
[0027] C: echo cancellation;
[0028] fp: feedback path;
[0029] 1: voice communication system;
[0030] 10, 20: conference terminal;
[0031] 50: cloud server;
[0032] 11, 21: receiver;
[0033] 13, 21: speaker;
[0034] 15, 25, 55: communication transceiver;
[0035] 17, 27, 57: memory;
[0036] 19, 29, 59: processor
[0037] 70: sound watermark generating device
[0038] S310-S350, S410-S450, S910-S950: step
[0039] S Rx : call receiving sound signal
[0040] S Tx : call transmitting sound signal
[0041] S WM , S WM1 : watermark sound signal
[0042] S Rx + S WM : embedded watermark signal
[0043] S' Rx , S" Rx , S 90° , S WO : reflected sound signal
[0044] W: wall
[0045] γ w : reflection coefficient
[0046] d s , d w : distance
[0047] SS: sound source
[0048] W O , W E : watermark identifier
[0049] phase shift
[0050] S A , transmitting sound signal DETAILED DESCRIPTION
[0051] Reference will now be made in detail to exemplary embodiments of the application, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers will be used in the drawings and the description to refer to the same or like parts.
[0052] Figure 2 is a schematic diagram of a conference call system 1 according to an embodiment of the application. Please refer to Figure 2The voice communication system 1 includes, but is not limited to, conference terminals 10, 20 and a cloud server 50.
[0053] The conference terminals 10, 20 can be wired phones, mobile phones, web phones, tablet computers, desktop computers, notebook computers or smart speakers.
[0054] The conference terminal 10 includes, but is not limited to, a microphone 11, a speaker 13, a communication transceiver 15, a memory 17 and a processor 19.
[0055] The microphone 11 can be a dynamic, condenser or electret condenser microphone, or a combination of other electronic components, analog-to-digital converters, filters and audio processors that can receive sound waves (e.g., human voice, environmental sound, machine operation sound, etc.) and convert them into sound signals. In an embodiment, the microphone 11 is used to pick up / record the sound of a speaker to obtain a call-receiving sound signal. In some embodiments, the call-receiving sound signal can include the sound of the speaker, the sound emitted by the speaker 13 and / or other environmental sounds.
[0056] The speaker 13 can be a loudspeaker or a loudspeaker. In an embodiment, the speaker 13 is used to play sound.
[0057] The communication transceiver 15 is, for example, a transceiver that supports wired networks such as Ethernet, fiber optic networks or cables (which can include, but are not limited to, connection interfaces, signal converters, communication protocol processing chips and other components), and can also be a transceiver that supports wireless networks such as Wi-Fi, fourth generation (4G), fifth generation (5G) or later generation mobile networks (which can include, but are not limited to, antennas, digital-to-analog / analog-to-digital converters, communication protocol processing chips and other components). In an embodiment, the communication transceiver 15 is used to transmit or receive data.
[0058] The memory 17 can be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, traditional hard disk drive (HDD), solid-state drive (SSD) or similar component. In an embodiment, the memory 17 is used to store program code, software modules, configuration settings, data (e.g., sound signals, watermark identifiers or watermark sound signals) or files.
[0059] The processor 19 is coupled to the microphone 11, the speaker 13, the communication transceiver 15, and the memory 17. The processor 19 can be a central processing unit (CPU), a graphic processing unit (GPU), or other programmable general purpose or special purpose microprocessors, digital signal processors (DSPs), programmable controllers, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or other similar components or a combination thereof. In an embodiment, the processor 19 is configured to perform all or part of the operations of the conference terminal 10 and can load and execute various software modules, files, and data stored in the memory 17.
[0060] The conference terminal 20 includes (but not limited to) a microphone 21, a speaker 23, a communication transceiver 25, a memory 27, and a processor 29. The embodiments and functions of the microphone 21, the speaker 23, the communication transceiver 25, the memory 27, and the processor 29 can be referred to the aforementioned descriptions of the microphone 11, the speaker 13, the communication transceiver 15, the memory 17, and the processor 19, and will not be repeated here. The processor 29 is configured to perform all or part of the operations of the conference terminal 20 and can load and execute various software modules, files, and data stored in the memory 27.
[0061] The cloud server 50 is connected to the conference terminals 10, 20 directly or indirectly via a network. The cloud server 50 can be a computer system, a server, or a signal processing device. In an embodiment, the conference terminals 10, 20 can also serve as the cloud server 50. In another embodiment, the cloud server 50 can be a separate cloud server different from the conference terminals 10, 20. In some embodiments, the cloud server 50 includes (but not limited to) the same or similar communication transceiver 55, memory 57, and processor 59, and the embodiments and functions of the components will not be repeated here.
[0062] In an embodiment, the sound watermark generation device 70 can be the conference terminal 10, 20, or the cloud server 50. The sound watermark generation device 70 is configured to generate a sound watermark signal, and will be described in detail in subsequent embodiments.
[0063] In the following, the method described in the embodiments of the present disclosure will be described in conjunction with the devices, components, and modules in the conference communication system 1. The various processes of the method can be adjusted according to the implementation situation, and are not limited thereto.
[0064] It is also noted that, for the sake of convenience, the same components can implement the same or similar operations, and will not be repeated. For example, the processor 19 of the conference terminal 10, the processor 19 of the conference terminal 20, and / or the processor 59 of the cloud server 50 can all implement the same or similar methods of the embodiments of the present application.
[0065] Figure 3 is a flowchart of a processing method of a sound watermark according to an embodiment of the present application. Referring to Figure 3 , the processor 29 obtains a call receiving sound signal S Rx (step S310) by recording through the microphone 21. Specifically, assume that the conference terminals 10, 20 establish a call conference. For example, the conference is established by means of a video software, a voice call software, or a telephone call, and a speaker can start to speak. After recording / sounding through the microphone 21, the processor 29 can obtain the call receiving sound signal S Rx . This call receiving sound signal S Rx relates to the voice content of the speaker corresponding to the conference terminal 20 (which can also include environmental sound or other noise). The processor 29 of the conference terminal 20 can transmit the call receiving sound signal S Rx through the communication transceiver 25 (i.e., via a network interface). In some embodiments, the call receiving sound signal S Rx may be subjected to echo cancellation, noise filtering, and / or other sound signal processing.
[0066] The processor 59 of the cloud server 50 receives the call receiving sound signal S Rx from the conference terminal 20 through the communication transceiver 55. The processor 59 generates a reflected sound signal S' Rx (step S330) according to a virtual reflection condition and the call receiving sound signal. Specifically, a general echo cancellation algorithm can adaptively cancel the component belonging to a reference signal in the sound signal received by the microphone 11, 21 from the outside (e.g., the call receiving sound signal S Rx of the call receiving path). The sound recorded by the microphone 11, 21 includes the shortest path from the speaker 13, 23 to the microphone 11, 21 and different reflection paths of the environment (i.e., the paths formed by the reflection of sound from external objects). The reflected sound signal is affected according to the reflection coefficient of the reflected object, and the position of the reflection affects the time delay and the attenuation of the sound signal. In addition, the reflected sound signal can also come from different directions, thereby causing phase shift. In the embodiments of the present application, the known sound signal S Rx of the call receiving path is used to generate a virtual / simulated reflected sound signal that can be cancelled by the echo cancellation mechanism, and a sound watermark signal S WM is generated accordingly.
[0067] Figure 4 is a sound watermark S according to an embodiment of the present application WM The flowchart of the generating method of the sound watermark S Figure 4 The processor 59 can set a virtual reflection condition and generate the reflection sound signal S' Rx (step S410) according to the virtual reflection condition. Specifically, the virtual reflection condition includes a positional relationship among the microphone 11, 21, a sound source (e.g., a speaker, the microphone 13, 23), and an external object (e.g., a wall, a ceiling, furniture, or a person). For example, a distance between the microphone 11 and the external object, a distance between the microphone 11 and the sound source, and / or a distance between the sound source and the external object. The reflection sound signal S' Rx is a sound signal obtained by simulating a sound emitted from the sound source, reflected by the external object, and recorded by the microphone 11, 21.
[0068] In an embodiment, the processor 59 can determine the time delay and the amplitude attenuation of the reflection sound signal S' Rx compared to the conversation receiving sound signal S Rx For example, the reflection sound signal S' Figure 5 is a schematic diagram illustrating the virtual reflection condition according to an embodiment of the present application. Referring to Figure 5 , it is assumed that the virtual reflection condition is a single wall (i.e., the external object), and a reflection coefficient of the wall W is γ w (e.g., 0.7, 0.3, or 1). A distance between the microphone 21 and the sound source SS is d s (e.g., 0.3, 0.5, or 0.8 meters), and a distance between the microphone 21 and the wall W is d w (e.g., 1, 1.5, or 2 meters). The relationship between the reflection sound signal S' Rx and the conversation receiving sound signal S Rx can be expressed as follows:
[0069]
[0070] where T s is a sampling time, v s is a speed of sound, and n is a sampling point or time.
[0071] If the reflection sound signal S' Rx is set to have a time delay γ w and an amplitude attenuation α w compared to the conversation receiving sound signal S Rx , the relationship between the reflection sound signal S' Rx and the conversation receiving sound signal S Rx can be expressed as follows:
[0072] s′ Rx (n) = a w ·s Rx (n-n w )…(2)
[0073] And according to equations (1), (2), we have:
[0074]
[0075]
[0076] wherein n f is a time delay caused by the filter (optionally, and to be detailed in subsequent embodiments), is a time delay caused by the phase shift (optionally, and to be detailed in subsequent embodiments).
[0077] It is to be noted that the variables in the virtual reflection condition can be further adjusted according to different design requirements. For example, there are not only one external object or relative position.
[0078] Referring to Figure 3 , the processor 59 phase shifts the reflected sound signal s' O according to the watermark identifier W Rx to generate a watermark sound signal s WM (step S350). Specifically, when a general echo cancellation mechanism is in operation, the time delay and amplitude change of the reflected sound signal have a greater error impact on the echo cancellation mechanism than the phase shift of the reflected sound signal. This change is like being in a completely new interference environment and makes the echo cancellation mechanism need to adapt again. Therefore, the sound watermark signals s WM corresponding to different values in the watermark identifier W O of the embodiment of the present application only have phase differences but have the same time delay and amplitude. That is, the watermark sound signal s WM includes one or more phase-shifted reflected sound signals s' Rx .
[0079] Referring to Figure 4 , in an embodiment, the processor 59 can select a filter to generate a filtered reflected sound signal s" Rx (step S430). Specifically, a general echo cancellation mechanism has a slower convergence speed for processing low-frequency (for example, 3 kilohertz (kHz) or below 4 kHz) sound signals but has a faster convergence speed (for example, below 10 milliseconds (ms)) for processing high-frequency sound signals (for example, above 3 kHz or 4 kHz). Therefore, the processor 59 can only filter the reflected sound signal s' RxPhase shifting is performed, and the interference of the signal is not easily perceived by a person (i.e., the frequency of the high-frequency sound signal is outside the range of human hearing).
[0080] For example, Figure 6 is a schematic diagram illustrating the filtering process according to an embodiment of the present application. Please refer to Figure 6 The processor 59 can perform low-pass filtering on the reflected sound signal S' Rx to output a low-pass filtered reflected sound signal S For example, the low-pass filter LPF blocks signals above 4 kHz and only allows signals below 4 kHz to pass. On the other hand, the processor 59 can perform high-pass filtering on the reflected sound signal S' Rx to output a high-pass filtered reflected sound signal S For example, the high-pass filter HPF blocks signals below 4 kHz and only allows signals above 4 kHz to pass.
[0081] In another embodiment, the processor 59 can also not perform filtering on the reflected sound signal S' Rx to output a filtered reflected sound signal S Rx which is equivalent to the reflected sound signal S' Rx .
[0082] Please refer to Figure 4 The processor 59 can perform phase shifting on the reflected sound signal S" O according to the watermark identifier W Rx In an embodiment, the watermark identifier W O is encoded in a multi-radix, and the multi-radix provides multiple values in each of one or more bits of the watermark identifier W O For example, in binary, the value of each bit in the watermark identifier W O may be "0" or "1". For example, in hexadecimal, the value of each bit in the watermark identifier W O may be "0", "1", "2", …, "E", "F". In another embodiment, the watermark identifier is encoded in letters, characters, and / or symbols. For example, the value of each bit in the watermark identifier W O may be any one of English "A" to "Z".
[0083] In an embodiment, the different values on the bits of the watermark identifier W O correspond to different phase shifts. For example, Figure 7 is a schematic diagram illustrating multiple phase shifts according to an embodiment of the present application. Please refer to Figure 7Assuming the watermark identifier W O If it's an N-ary number system (where N is a positive integer), then each bit can provide N values. These N different values correspond to different phase offsets.
[0084] Figure 8 This is a schematic diagram illustrating two phase shifts according to an embodiment of the present invention. Please refer to... Figure 7 Assuming the watermark identifier W O If it's a binary system, then each bit can have two values (i.e., 1 and 0). These two different values correspond to two phase offsets. For example, phase shift It is 90°, and the phase shift is 90°. It is -90° (i.e., -1).
[0085] Processor 59 can determine the watermark identifier W O One or more bits of the value are offset from the reflected sound signal S” Rx The phase. Figure 7 For example, processor 59 uses the watermark identifier W O One or more values in the selection phase offset One or more of them, and using the selected phase offset Phase shifting is performed. For example, the watermark identifier W O If the value of the first bit is 1, then the output is a phase-shifted reflected sound signal. Relative to the reflected sound signal S” Rx Offset Other reflected sound signals And so on. Phase shifting can be achieved using the Hilbert transform or other phase shifting algorithms.
[0086] In one embodiment, the watermark identifier includes multiple bits. This watermark sound signal S WM It includes multiple phase-shifted reflected sound signals, and each phase-shifted reflected sound signal occupies the watermark sound signal S. WM The duration of each element. Assume the duration of each element is L. b (For example, 0.1, 0.5, or 1 second, and greater than the time delay n) w This indicates that, similar to the concept of time-sharing multitasking, processor 59 will process the watermarked audio signal S. WM The time period (i.e., the main time unit) is based on the watermark identifier W. O The included bits are divided into sub-time units of the same or different time lengths, and each sub-time unit carries a phase-shifted reflected sound signal corresponding to a different bit.
[0087] In one embodiment, if the filter processing of Figure 6 is employed, the processor 59 can synthesize one or more phase-shifted reflected sound signals and the reflected sound signal filtered by the low-pass filter processing For example, the reflected sound signal filtered by the high-pass filter processing Figure 8 is phase-shifted by 90° (creating the phase-shifted reflected sound signal S 90° ), and the phase-shifted reflected sound signal S WO is outputted. The processor 59 further synthesizes the reflected sound signal filtered by the low-pass filter processing and the phase-shifted reflected sound signal S WO to generate the watermark sound signal S WM1 .
[0088] In some embodiments, the processor 59 can generate a plurality of identical watermark sound signals. These watermark sound signals respectively correspond to different primary time units. That is, the watermark sound signal is outputted cyclically. In order to distinguish adjacent watermark sound signals, the processor 59 can add a gap between adjacent watermark sound signals. For example, a mute signal or other known high-frequency sound signal is added at the gap.
[0089] In one embodiment, the processor 59 can transmit the call-received sound signal S Rx and the watermark sound signal S WM through the communication transceiver 55 respectively. In another embodiment, the processor 59 can synthesize the call-received sound signal S Rx and the watermark sound signal S WM to generate the watermark-embedded signal S Rx +S WM . Then, the processor 59 can transmit the watermark-embedded signal S Rx +S WM through the communication transceiver 55.
[0090] Figure 9A is an example simulation diagram of the call-received sound signal S Rx , and Figure 9B is an example simulation diagram of the watermark-embedded signal S Rx +S WM . Please refer to Figure 9A and Figure 9B , the two sounds are very close, and it is difficult or impossible for a person to distinguish them.
[0091] The processor 19 of the conference terminal 10 receives the watermark sound signal S WM or the watermark-embedded signal S Rx +S WM through the communication transceiver 15 via the network.To obtain the transmitted sound signal S A (That is, the transmitted watermarked sound signal S) WM Or embed watermark signal S Rx +S WM Due to the watermark sound signal S WM Including the received audio signal after time delay and amplitude attenuation (i.e., the reflected audio signal), the echo cancellation mechanism of processor 19 can effectively eliminate the watermarked audio signal S. WM This ensures that the voice signal S transmitted during a call is not affected on the communication transmission path. Tx (For example, the conference terminal 10 receives the audio signal of a call that it wants to transmit via the network).
[0092] For the watermark sound signal S WM The identification, Figure 10 This is a flowchart illustrating watermark recognition according to an embodiment of the present invention. Please refer to... Figure 10 In one embodiment, if using Figure 6 If the filtering process is performed, the processor 19 can use the same or similar high-pass filter HPF to filter the transmitted audio signal S. A Perform high-pass filtering (step S910) to output the transmitted audio signal that has passed the high-pass filtering process. In another embodiment, if not adopted Figure 6 If the filtering process is performed, then step S910 (i.e., transmitting the audio signal) can be ignored. Equivalent to transmitting sound signal S A ).
[0093] Processor 19 can shift the transmitted sound signal according to the correspondence between the value and phase shift described in step S450. The phase (i.e., step S930, phase shifting). With Figure 8 For example, processor 19 generates a transmitted sound signal with a 90° phase shift. Processor 19 can transmit sound signals and the transmission of sound signals via phase shift The correlation between the watermark identifier W E (Step S950). For example, processor 19 will transmit an audio signal. With transmitting sound signals Due to time delay n w Calculate the orthogonal cross correlation R at the location xy (n w And -1 ≤ R xy (n w )≤1. Processor 19 defines a threshold Th R Then the watermark identifier W E It can be represented as:
[0094]
[0095] That is, if the correlation is higher than the threshold Th R , the processor 19 determines the value of this bit is the value corresponding to the phase shift 90° (for example, 1); if the correlation is lower than the threshold Th R , the processor 19 determines the value of this bit is the value corresponding to the phase shift -90° (for example, 0). In another embodiment, the processor 19 can identify the transmitted sound signal corresponding values on different time units.
[0096] In summary, in the sound watermark processing method and the sound watermark generation device according to the embodiments of the present application, the reflected sound signal is simulated according to the principle of the echo cancellation mechanism, and the sound watermark signal is encoded by phase shifting the reflected sound signal. In this way, at the receiving end, the sound watermark signal obtained through the feedback path can be eliminated by the echo cancellation mechanism, and the sound watermark signal will not affect the communication transmission signal on the communication transmission path.
[0097] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method of processing an acoustic watermark, suitable for a conference terminal, the conference terminal comprising a sound receiver, characterized in that, The method for processing the sound watermark comprises: obtaining a call-receiving sound signal through the sound receiver; generating a reflected sound signal according to a virtual reflection condition and the call-received sound signal, wherein the virtual reflection condition comprises a positional relationship among the microphone, a sound source and an external object, the positional relationship comprises a distance between the microphone and the external object, a distance between the microphone and the sound source and a distance between the sound source and the external object, the reflected sound signal is a sound signal simulating a sound signal of a sound emitted by the sound source, reflected by the external object and recorded by the microphone, the reflected sound signal has a time delay and an amplitude attenuation compared with the call-received sound signal, the time delay is the amplitude attenuation is γ w a reflection coefficient of the external object, d s the distance between the microphone and the sound source, d w the distance between the microphone and the external object, T s a sampling time, and v s a speed of sound; and phase-shifting the reflected sound signal according to a watermark identifier to generate a watermark sound signal, wherein the watermark sound signal comprises at least one phase-shifted reflected sound signal.
2. The method for processing the sound watermark according to claim 1, wherein the watermark identifier is encoded in a multi-radix, the multi-radix providing a plurality of values in each of at least one bit of the watermark identifier, and the step of phase-shifting the reflected sound signal according to the watermark identifier comprises: phase-shifting the reflected sound signal according to a value of the bit in the watermark identifier, wherein different values correspond to different phase shifts.
3. The method for processing the sound watermark according to claim 2, wherein the bit of the watermark identifier comprises a plurality of bits, the watermark sound signal comprises a plurality of phase-shifted reflected sound signals, and each of the phase-shifted reflected sound signals occupies a length of time in the watermark sound signal.
4. The processing method of sound watermarking according to claim 1, characterized in that, before the step of phase-shifting the reflected sound signal according to the watermark identifier, further comprising: low-pass filtering the reflected sound signal; and high-pass filtering the reflected sound signal, wherein only the phase of the reflected sound signal passing through the high-pass filtering is shifted, and the step of generating the watermark sound signal further comprises: synthesizing the phase-shifted reflected sound signal and the reflected sound signal passing through the low-pass filtering.
5. The processing method of sound watermarking according to claim 1, characterized in that, The method for processing the sound watermark further comprises: receiving a transmitted sound signal via a network, wherein the transmitted sound signal comprises a transmitted watermark sound signal; phase-shifting the transmitted sound signal; and identifying the watermark identifier according to a correlation between the transmitted sound signal and the phase-shifted transmitted sound signal.
6. A sound watermark generation apparatus, comprising: a memory for storing program code; and a processor coupled to the memory, wherein the processor is configured to load and execute the program code to: obtain a call-receiving sound signal, wherein the call-receiving sound signal is obtained by sound recording through a sound receiver; phase-shift the reflected sound signal according to a watermark identifier to generate a watermark sound signal, wherein the watermark sound signal comprises at least one phase-shifted reflected sound signal. generating a reflected sound signal according to a virtual reflection condition and the call-received sound signal, wherein the virtual reflection condition comprises a positional relationship among the microphone, a sound source and an external object, the positional relationship comprises a distance between the microphone and the external object, a distance between the microphone and the sound source and a distance between the sound source and the external object, the reflected sound signal is a sound signal simulating a sound signal of a sound emitted by the sound source, reflected by the external object and recorded by the microphone, the reflected sound signal has a time delay and an amplitude attenuation compared with the call-received sound signal, the time delay is the amplitude attenuation is γ w a reflection coefficient of the external object, d s the distance between the microphone and the sound source, d w the distance between the microphone and the external object, T s a sampling time, and v s a speed of sound; and The watermark identifier is encoded in a multi-radix, the multi-radix providing a plurality of values in each of at least one bit of the watermark identifier, and the processor is further configured to:
7. The sound watermark generation apparatus of claim 6, wherein phase-shift the reflected sound signal according to a value of the bit in the watermark identifier, wherein different values correspond to different phase shifts. The bit of the watermark identifier comprises a plurality of bits, the watermark sound signal comprises a plurality of phase-shifted reflected sound signals, and each of the phase-shifted reflected sound signals occupies a length of time in the watermark sound signal.
8. The sound watermark generation apparatus of claim 7, wherein The processor is further configured to:
9. The sound watermark generation apparatus of claim 6, wherein low-pass filter the reflected sound signal; and high-pass filter the reflected sound signal, wherein only the phase of the reflected sound signal passing through the high-pass filtering is shifted, and the step of generating the watermark sound signal further comprises: synthesize the phase-shifted reflected sound signal and the reflected sound signal passing through the low-pass filtering. high pass filtering the reflected sound signal, wherein only a phase of the reflected sound signal that passes through the high pass filtering is shifted; and combining the phase shifted reflected sound signal and the reflected sound signal that passes through the low pass filtering.
10. The sound watermark generation apparatus of claim 6, wherein The watermark identifier is identified from a correlation between the transmitted watermark sound signal and the phase shifted watermark sound signal. The watermark identifier is identified from a correlation between the transmitted watermark sound signal and the phase shifted watermark sound signal.
Citation Information
Patent Citations
Electronic watermark embedding device and electronic watermark detecting device, and electronic watermark embedding method and electronic watermark detection method
JP2009210828A