Voice watermark encoding and decoding method, device, equipment and medium
By down-tuning the audio and embedding the watermark signal in the high-frequency hole area, the problem of high computing power consumption in the existing technology is solved, and low-cost and efficient voice watermark authentication is achieved. It is suitable for intelligent assistants or customer service systems in medical and financial scenarios.
Patent Information
- Application Number
- CN202511059542.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing voice watermarking technology consumes a lot of computing power and is difficult to effectively identify the authenticity of generated speech, resulting in the abuse of intelligent speech generation technology in illegal activities.
By down-tuning the audio, obtaining the high-frequency hole area, detecting the bandwidth, generating a watermark signal, and embedding the watermark signal in the high-frequency hole area, decoding is performed using shared spectrum analysis to reduce frequency domain conversion, and constructing an integrated process of pitch shifting and watermark encoding.
It reduces computing costs and power consumption, improves processing efficiency, and achieves rapid authentication. It is suitable for intelligent assistants or customer service systems in medical and financial scenarios.
Smart Images

Figure CN120708628A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and speech processing technology, and in particular to a speech watermark encoding and decoding method, device, equipment and medium. Background Art
[0002] With the rapid development of deepfake technology, the authenticity, naturalness, and timbre of generated speech have been greatly improved, reaching levels of authenticity that are indistinguishable from the real thing. While intelligent speech generation technology facilitates intelligent interactive applications and devices, it also poses threats to information cognition and social security. In recent years, intelligent speech generation software, primarily based on speech synthesis and voice conversion, has become widely available online, lowering the technical barriers and costs of speech production. This has led to its widespread use by criminals for various illegal or fraudulent activities. For example, generated speech is used to impersonate intelligent customer service representatives for medical insurance applications to conduct fraudulent health insurance verification, or to mimic intelligent assistants in financial scenarios to obtain users' online banking information. Therefore, the speech generated by intelligent customer service representatives or assistants needs to be watermarked to facilitate authentication. However, existing speech watermarking technologies mostly rely on feature extraction and adaptive modulation, which consumes a lot of computing power. Summary of the Invention
[0003] The present invention provides a voice watermark encoding and decoding method, device, equipment and medium to solve the technical problem of high computing power consumption in existing voice watermark adding technology.
[0004] In a first aspect, a speech watermark encoding and decoding method is provided, comprising:
[0005] Down-convert the audio to obtain the high-frequency hole area of the down-converted audio;
[0006] Detect the bandwidth of the high-frequency hole area of the down-tuned audio;
[0007] Generate a watermark signal according to a preset encoding character string;
[0008] According to the preset watermark embedding rules, the watermark signal is embedded in the high-frequency hole area of the down-tuned audio to obtain the watermarked audio and output it;
[0009] receiving the watermarked audio and extracting the high-frequency audio segment of the watermarked audio;
[0010] The high-frequency audio segment of the extracted watermarked audio is divided according to the number of frames corresponding to a single watermark character set of a preset watermark embedding rule to obtain a character set unit frame group, and the watermark signal is decoded for each character set unit frame group.
[0011] In a second aspect, a speech watermark encoding and decoding device is provided, comprising:
[0012] An audio down-conversion module is used to down-convert the audio and obtain the high-frequency hole area of the down-converted audio;
[0013] A bandwidth detection module is used to detect the bandwidth of the high-frequency hole area of the down-tuned audio;
[0014] A watermark signal generating module, configured to generate a watermark signal according to a preset coding string;
[0015] A watermark embedding module is used to embed a watermark signal in the high-frequency hole area of the down-tuned audio according to a preset watermark embedding rule, obtain the watermarked audio and output it;
[0016] A high-frequency audio segment extraction module is used to receive the watermarked audio and extract the high-frequency audio segment of the watermarked audio;
[0017] The watermark signal decoding module is used to divide the high-frequency audio segment of the extracted watermarked audio according to the number of frames corresponding to a single watermark character set of the preset watermark embedding rule, obtain character set unit frame groups, and decode the watermark signal of each character set unit frame group.
[0018] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned speech watermark encoding and decoding method when executing the computer program.
[0019] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned voice watermark encoding and decoding method are implemented.
[0020] In the scheme implemented by the above-mentioned voice watermark encoding and decoding method, device, equipment and medium, audio can be received through the client, the audio can be down-tuned to obtain the high-frequency hole area of the down-tuned audio; the bandwidth of the high-frequency hole area of the down-tuned audio can be detected; a watermark signal can be generated according to a preset coding string; according to a preset watermark embedding rule, a watermark signal can be embedded in the high-frequency hole area of the down-tuned audio to obtain and output the watermarked audio; the watermarked audio can be received and the high-frequency audio segment of the watermarked audio can be extracted; the high-frequency audio segment of the watermarked audio can be divided according to the number of frames corresponding to a single watermark character set of the preset watermark embedding rule to obtain a character set unit frame group, and the watermark signal can be decoded for each character set unit frame group. In the present invention, for an intelligent assistant for medical insurance authentication procedures under medical services, or for an intelligent customer service for a voice-protected bank account under financial services, a voice watermark encoding and decoding scheme can be used to obtain the down-tuned audio by down-tuning the audio. The high-frequency hole area of the frequency is detected, the bandwidth of the high-frequency hole area of the down-tuned audio is detected, a watermark signal is generated according to a preset coding string, and a watermark signal is embedded in the high-frequency hole area of the down-tuned audio according to a preset watermark embedding rule to obtain the watermarked audio and output it for audio authentication. The high-frequency hole area after the speech down-tuning is used as the embedding carrier, and no additional feature extraction and adaptive adjustment process is required. The computing power consumption is low, the computing cost can be reduced, and the timeliness response is fast. Moreover, an integrated processing of pitch shifting and watermark encoding is constructed, and the down-tuning processing and watermark embedding are deeply integrated; by receiving the watermarked audio, the high-frequency audio segment of the watermarked audio is extracted, and the high-frequency audio segment of the extracted watermarked audio is divided according to the number of frames corresponding to the single watermark character set of the preset watermark embedding rule to obtain the character set unit frame group, and the watermark signal is decoded for each character set unit frame group. Through shared spectrum analysis and signal processing, repeated frequency domain conversion is reduced, and there is no need for complex time-frequency transformation, thereby improving processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0022] Figure 1 This is a schematic diagram of an application environment of a voice watermark encoding and decoding method according to an embodiment of the present invention;
[0023] Figure 2 This is a flow chart of a voice watermark encoding and decoding method according to an embodiment of the present invention;
[0024] Figure 3yes Figure 2 A schematic flow chart of a specific implementation of step S40;
[0025] Figure 4 It is a structural diagram of a speech watermark encoding and decoding device according to an embodiment of the present invention;
[0026] Figure 5 is a structural diagram of a computer device in one embodiment of the present invention;
[0027] Figure 6 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0029] The speech watermark encoding and decoding method provided by the embodiment of the present invention can be applied in the following fields: Figure 1In the application environment, it is used in intelligent assistants or intelligent customer service in application scenarios such as medical care and finance, and is usually implemented through a server, wherein the client communicates with the server through a network. The server can receive audio through the client, down-tune the audio, and obtain the high-frequency hole area of the down-tuned audio; detect the bandwidth of the high-frequency hole area of the down-tuned audio; generate a watermark signal according to a preset coding string; embed the watermark signal in the high-frequency hole area of the down-tuned audio according to the preset watermark embedding rule, obtain the watermarked audio and output it; receive the watermarked audio, extract the high-frequency audio segment of the watermarked audio; divide the extracted high-frequency audio segment of the watermarked audio according to the number of frames corresponding to the single watermark character set of the preset watermark embedding rule, obtain the character set unit frame group, and decode the watermark signal for each character set unit frame group. In the present invention, for the intelligent assistant of the medical insurance authentication procedure under the medical business, or for the intelligent customer service of the voice-protected bank account under the financial business, the voice watermark encoding and decoding scheme can be used to down-tune the audio, obtain the high-frequency hole area of the down-tuned audio, detect the down-tuned audio, and decode the watermark signal. The method uses the bandwidth of the high-frequency hole region of the down-tuned audio to generate a watermark signal based on a preset encoding string. Based on a preset watermark embedding rule, the watermark signal is embedded in the high-frequency hole region of the down-tuned audio to obtain and output the watermarked audio for audio authentication. Using the high-frequency hole region of the down-tuned speech as an embedding carrier eliminates the need for additional feature extraction and adaptive adjustment processes, resulting in low computing power consumption, reduced computational costs, and fast time-to-response. Furthermore, an integrated process for pitch shifting and watermark encoding is constructed, deeply integrating the down-tuning processing and watermark embedding. The method receives watermarked audio, extracts high-frequency audio segments of the watermarked audio, divides the extracted high-frequency audio segments of the watermarked audio based on the number of frames corresponding to a single watermark character set according to the preset watermark embedding rule, obtains character set unit frame groups, and decodes the watermark signal for each character set unit frame group. Through shared spectrum analysis and signal processing, repeated frequency domain conversions are reduced, eliminating the need for complex time-frequency transformations and improving processing efficiency. The client can include, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server side can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0030] See also Figure 2 As shown, Figure 2 A flow chart of a voice watermark encoding and decoding method provided in an embodiment of the present invention includes the following steps:
[0031] S10: Down-tuning the audio to obtain a high-frequency hole region of the down-tuned audio.
[0032] The voice watermark encoding and decoding method provided by the present invention can be applied to intelligent customer service or intelligent assistants in various application scenarios such as medical care, finance, and insurance, and is usually implemented through the server. For example, in an intelligent assistant under a medical insurance authentication program in the medical application field, a user can verify the applicant's identity, confirm their application intention, or obtain a verbal description of their health status through voice communication methods such as telephone authentication or remote consultation. The intelligent assistant generates a response audio based on the user's audio. Alternatively, in an intelligent customer service of a voice-protected bank account in the financial application field, a user can perform information setting operations and transaction operations on the bank account through voice. The intelligent customer service can generate corresponding operation prompt audio based on the user's operation. After generating the response audio or operation prompt audio, the intelligent assistant or intelligent customer service can down-tune the audio to obtain the high-frequency hole area of the down-tune audio, so that the high-frequency hole area can be subsequently used as a carrier for watermark embedding, without the need for additional feature extraction and adaptive modulation processes, thereby reducing computing power consumption and computing costs.
[0033] In step S10 , the audio is down-tuned, which may be performed by down-tuning the audio by 2 semitones.
[0034] S20: Detecting the bandwidth of the high-frequency hole region of the down-tuned audio.
[0035] Preferably, step S20, i.e., detecting the bandwidth of the high-frequency hole region of the down-tuned audio, is specifically as follows:
[0036] The bandwidth of the high-frequency hole region of the down-converted audio is determined according to the sampling rate of the audio and the parameters of the down-conversion processing.
[0037] The audio sampling rate can be 16kHz, and the down-tuning parameter can be a down-tuning interval. In this embodiment, the down-tuning parameter is 2 semitones. After down-tuning the audio, a high-frequency hole bandwidth of approximately 1kHz is obtained. The spectrum of the down-tuned audio has holes between approximately 7kHz and 8kHz, representing the high-frequency hole region. This region has almost no energy, and adding a watermark to the high-frequency hole region can effectively decouple it from other regions. Detecting the bandwidth of the high-frequency hole region can ensure the reliability of the high-frequency hole region.
[0038] S30: Generate a watermark signal according to a preset encoding character string.
[0039] The preset encoding string can be set according to actual needs, for example, "PINGAN".
[0040] S40: According to a preset watermark embedding rule, a watermark signal is embedded in the high-frequency hole area of the down-tuned audio to obtain and output the watermarked audio.
[0041] In an embodiment of the invention, in step S40, the preset watermark embedding rules include a watermark character set, a short-time Fourier transform frame length, a short-time Fourier transform frame shift, and the number of frames corresponding to a single watermark character set. The number of frames corresponding to a single watermark character set represents the audio frame interval corresponding to a single watermark character in the watermark signal. The short-time Fourier transform frame length can be 32ms, and the short-time Fourier transform frame shift can be 16ms. By outputting watermarked audio, users can easily determine that the output of the operation prompt audio or response audio is genuine, such as in the case of intelligent customer service for voice-protected bank accounts in financial applications or intelligent assistants for medical insurance authentication procedures in medical applications.
[0042] Preferably, the watermark character set includes 26 uppercase English characters, 10 Arabic numerals, and four special symbols, and the number of frames corresponding to a single watermark character set is four. The four special symbols may be slash, underscore, asterisk, and pound sign, and the watermark character set includes 40 characters, namely A-Z, 0-9, / , _, *, and #.
[0043] Among them, such as Figure 3 As shown, in step S40, that is, according to the preset watermark embedding rule, the watermark signal is embedded in the high-frequency hole area of the down-tuned audio, including the following steps:
[0044] S41: Mapping the watermark character set to frames according to the character sequence in the watermark character set and the number of frames corresponding to the single watermark character set to obtain the watermark character range represented by each frame in the number of frames corresponding to the single watermark character set. If the number of frames corresponding to the single watermark character set is four, then in the four frames corresponding to the single watermark character set, the watermark character range represented by the first frame is ABCDEFGHIJ, the watermark character range represented by the second frame is KLMNOPQRST, the watermark character range represented by the third frame is UVWXYZ0123, and the watermark character range represented by the fourth frame is 456789 / _*#.
[0045] S42: Based on the short-time Fourier transform frame length, obtain the number of frequency bands for each audio frame, and sort each audio frame from high to low frequency, so that the frequency band with the highest frequency has the largest sequence number. The short-time Fourier transform frame length is 32 ms, and the number of frequency bands for each audio frame is 257.
[0046] S43: Assign characters to each audio frame in descending order of frequency bands based on the watermark character range represented by each frame in the number of frames corresponding to the single watermark character set, and obtain a mapping relationship between the frequency band position and the watermark character of each frame in the number of frames corresponding to the single watermark character set. Then, in the four frames corresponding to the single watermark character set, the watermark characters corresponding to the 257th to 248th frequency bands of the first audio frame are A to J, respectively; the watermark characters corresponding to the 257th to 248th frequency bands of the second audio frame are K to T, respectively; the watermark characters corresponding to the 257th to 248th frequency bands of the third audio frame are U to Z and 0 to 3, respectively; and the watermark characters corresponding to the 257th to 248th frequency bands of the fourth audio frame are 4 to 9, / , _, *, and #, respectively.
[0047] S44: Embed the watermark signal in the high-frequency hole region of the down-tuned audio based on the mapping relationship between the frequency band position of each frame in the number of frames corresponding to the single watermark character set and the watermark character, in combination with a preset encoding string corresponding to the watermark signal. Embedding the watermark signal in the high-frequency hole region of the down-tuned audio involves embedding the corresponding watermark character in the frequency band position of the audio frame corresponding to the high-frequency hole region of the down-tuned audio based on the mapping relationship between the frequency band position of each frame in the number of frames corresponding to the single watermark character set and the watermark character, in combination with a preset encoding string corresponding to the watermark signal. For example, when the preset encoding string is "PINGAN", according to the mapping relationship between the frequency band position of each frame in the number of frames corresponding to the single watermark character set and the watermark character, it can be determined that the frequency band positions where the watermark character needs to be embedded are the 252nd frequency band (P) of the second frame of the number of frames corresponding to the first single watermark character set, the 249th frequency band (I) of the first frame of the number of frames corresponding to the second single watermark character set, the 254th frequency band (N) of the second frame of the number of frames corresponding to the third single watermark character set, and the 255th frequency band (N) of the second frame of the number of frames corresponding to the fourth single watermark character set. The 251st frequency band (G) of the first frame, the 257th frequency band (A) of the first frame corresponding to the fifth single watermark character set, and the 254th frequency band (N) of the second frame corresponding to the sixth single watermark character set, that is, the frequency band positions where the watermark characters need to be embedded are the 252nd frequency band (P) of the second frame, the 249th frequency band (I) of the fifth frame, the 254th frequency band (N) of the tenth frame, the 251st frequency band (G) of the thirteenth frame, the 257th frequency band (A) of the seventeenth frame, and the 254th frequency band (N) of the twenty-second frame of the total audio.
[0048] Preferably, when a clear watermark is required, that is, the watermark needs to be audible from the audio, the embedded energy can be increased; when a dark watermark is required, that is, the watermark needs to be inaudible from the audio, the embedded energy can be reduced.
[0049] S50: Receive the watermarked audio, and extract the high-frequency audio segment of the watermarked audio.
[0050] In some embodiments of the present invention, step S50, i.e., extracting the high-frequency audio segment of the watermarked audio, comprises:
[0051] According to the preset watermark embedding rule of short-time Fourier transform frame length and short-time Fourier transform frame shift, the audio with watermark is short-time Fourier transformed to obtain the audio amplitude spectrum;
[0052] The amplitude spectra corresponding to a preset number of frequency bands are extracted from the audio amplitude spectrum for each frame of audio in descending order of frequency to obtain the high-frequency audio segment of the watermarked audio, wherein the preset number is equal to the number of characters in the watermark character range represented by each frame. Since the watermark character set may include forty characters, the number of frames corresponding to a single watermark character set is four, and the number of characters in the watermark character range represented by each frame is ten, the preset number is ten. The amplitude spectra corresponding to a preset number of frequency bands are extracted from the audio amplitude spectrum for each frame of audio in descending order of frequency, that is, the amplitude spectra corresponding to the ten highest frequency bands are extracted from the audio amplitude spectrum for each frame of audio in descending order of frequency to obtain the high-frequency audio segment of the watermarked audio.
[0053] S60: Divide the high-frequency audio segment of the extracted watermarked audio according to the number of frames corresponding to a single watermark character set of a preset watermark embedding rule to obtain character set unit frame groups, and decode the watermark signal for each character set unit frame group.
[0054] The watermark encoding and decoding share spectrum analysis, which can reduce repeated frequency domain conversion operations and improve processing efficiency. In step S60, the watermark signal is decoded for each character set unit frame group, including:
[0055] Calculate the frequency point with the maximum energy in each character set unit frame group, and obtain the corresponding character according to the preset decoding mapping rules;
[0056] The characters corresponding to all the obtained character set unit frame groups are combined in sequence into a character string.
[0057] It can be seen that in the above scheme, the intelligent assistant for the medical insurance authentication procedure under the medical business, or the intelligent customer service for the voice-protected bank account under the financial business, can use the voice watermark encoding and decoding scheme to down-tune the audio, obtain the high-frequency hole area of the down-tuned audio, detect the bandwidth of the high-frequency hole area of the down-tuned audio, generate a watermark signal according to the preset coding string, embed the watermark signal in the high-frequency hole area of the down-tuned audio according to the preset watermark embedding rule, obtain the watermarked audio and output it to facilitate audio authentication, and use the high-frequency hole area after the voice down-tune as the embedding carrier without additional The feature extraction and adaptive adjustment process has low computing power consumption, can reduce computing costs, and has a fast time response. Moreover, an integrated processing of pitch shifting and watermark encoding is constructed, and the pitch-down processing is deeply integrated with watermark embedding. By receiving the watermarked audio, the high-frequency audio segment of the watermarked audio is extracted, and the high-frequency audio segment of the extracted watermarked audio is divided according to the number of frames corresponding to the single watermark character set of the preset watermark embedding rule to obtain the character set unit frame group, and the watermark signal is decoded for each character set unit frame group. Through shared spectrum analysis and signal processing, repeated frequency domain conversion is reduced, and there is no need for complex time-frequency transformation, thereby improving processing efficiency.
[0058] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0059] In one embodiment, a speech watermark encoding and decoding device is provided, which corresponds to the speech watermark encoding and decoding method in the above embodiment. Figure 4 As shown, the speech watermark encoding and decoding device includes an audio down-conversion module 101, a bandwidth detection module 102, a watermark signal generation module 103, a watermark embedding module 104, a high-frequency audio segment extraction module 105, and a watermark signal decoding module 106. The functional modules are described in detail as follows:
[0060] The audio down-tuning module 101 is used to down-tune the audio and obtain the high-frequency hole area of the down-tuned audio;
[0061] A bandwidth detection module 102 is configured to detect the bandwidth of a high-frequency hole region of the down-tuned audio;
[0062] The watermark signal generating module 103 is configured to generate a watermark signal according to a preset encoding string;
[0063] The watermark embedding module 104 is configured to embed a watermark signal in the high-frequency hole region of the down-tuned audio according to a preset watermark embedding rule, obtain the watermarked audio, and output the obtained audio;
[0064] A high-frequency audio segment extraction module 105 is configured to receive the watermarked audio and extract the high-frequency audio segment of the watermarked audio;
[0065] The watermark signal decoding module 106 is used to divide the high-frequency audio segment of the extracted watermarked audio according to the number of frames corresponding to a single watermark character set of the preset watermark embedding rule, obtain character set unit frame groups, and decode the watermark signal for each character set unit frame group.
[0066] In one embodiment, the bandwidth detection module 102 is specifically configured to:
[0067] The bandwidth of the high-frequency hole region of the down-converted audio is determined according to the sampling rate of the audio and the parameters of the down-conversion processing.
[0068] In one embodiment, the watermark embedding module 104 is specifically configured to:
[0069] Mapping the watermark character set according to the character sequence in the watermark character set and the number of frames corresponding to the single watermark character set to obtain the watermark character range represented by each frame in the number of frames corresponding to the single watermark character set;
[0070] According to the frame length of the short-time Fourier transform, the number of frequency bands of each frame of audio is obtained, and each frame of audio is sorted from high to low according to frequency, so that the frequency band with the highest frequency has the largest sequence number;
[0071] According to the watermark character range represented by each frame in the number of frames corresponding to the single watermark character set, characters are assigned to each frame of audio in descending order of frequency band, and a mapping relationship between the frequency band position of each frame in the number of frames corresponding to the single watermark character set and the watermark character is obtained;
[0072] According to the mapping relationship between the frequency band position of each frame in the number of frames corresponding to a single watermark character set and the watermark character, combined with the preset encoding string corresponding to the watermark signal, the watermark signal is embedded in the high-frequency hole area of the down-tuned audio.
[0073] In one embodiment, the high frequency audio segment extraction module 105 is specifically configured to:
[0074] According to the preset watermark embedding rule of short-time Fourier transform frame length and short-time Fourier transform frame shift, the audio with watermark is short-time Fourier transformed to obtain the audio amplitude spectrum;
[0075] The amplitude spectrum corresponding to a preset number of frequency bands is extracted from each frame of audio in descending order of frequency to obtain the high-frequency audio segment of the watermarked audio, wherein the preset number is equal to the number of characters in the watermark character range represented by each frame.
[0076] In one embodiment, the watermark decoding module 102 is specifically configured to:
[0077] Calculate the frequency point with the maximum energy in each character set unit frame group, and obtain the corresponding character according to the preset decoding mapping rules;
[0078] The characters corresponding to all the obtained character set unit frame groups are combined in sequence into a character string.
[0079] The present invention provides a speech watermark encoding and decoding device, which performs down-tuning processing on audio to obtain the high-frequency hole area of the down-tuned audio, detects the bandwidth of the high-frequency hole area of the down-tuned audio, generates a watermark signal according to a preset coding string, embeds the watermark signal in the high-frequency hole area of the down-tuned audio according to a preset watermark embedding rule, obtains and outputs watermarked audio, so as to facilitate audio authentication. The high-frequency hole area after the speech down-tuning is used as an embedding carrier, and no additional feature extraction and adaptive adjustment process is required. The computing power consumption is low, the computing cost can be reduced, and the timeliness response is fast. Moreover, an integrated processing of pitch change and watermark encoding is constructed, and the down-tuning processing and watermark embedding are deeply integrated. By receiving the watermarked audio, the high-frequency audio segment of the watermarked audio is extracted, and the high-frequency audio segment of the extracted watermarked audio is divided according to the number of frames corresponding to the single watermark character set of the preset watermark embedding rule to obtain a character set unit frame group, and the watermark signal is decoded for each character set unit frame group. Through shared spectrum analysis and signal processing, repeated frequency domain conversion is reduced, and complex time-frequency conversion is unnecessary, thereby improving processing efficiency.
[0080] The specific limitations of the speech watermark encoding and decoding device can be found in the limitations of the speech watermark encoding and decoding method described above and will not be further elaborated here. Each module in the speech watermark encoding and decoding device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the modules described above can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0081] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a voice watermark encoding and decoding method.
[0082] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, a network interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps on the client side of a voice watermark encoding and decoding method.
[0083] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0084] Down-convert the audio to obtain the high-frequency hole area of the down-converted audio;
[0085] Detect the bandwidth of the high-frequency hole area of the down-tuned audio;
[0086] Generate a watermark signal according to a preset encoding character string;
[0087] According to the preset watermark embedding rules, the watermark signal is embedded in the high-frequency hole area of the down-tuned audio to obtain the watermarked audio and output it;
[0088] receiving the watermarked audio and extracting the high-frequency audio segment of the watermarked audio;
[0089] The high-frequency audio segment of the extracted watermarked audio is divided according to the number of frames corresponding to a single watermark character set of a preset watermark embedding rule to obtain a character set unit frame group, and the watermark signal is decoded for each character set unit frame group.
[0090] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0091] Down-convert the audio to obtain the high-frequency hole area of the down-converted audio;
[0092] Detect the bandwidth of the high-frequency hole area of the down-tuned audio;
[0093] Generate a watermark signal according to a preset encoding character string;
[0094] According to the preset watermark embedding rules, the watermark signal is embedded in the high-frequency hole area of the down-tuned audio to obtain the watermarked audio and output it;
[0095] receiving the watermarked audio and extracting the high-frequency audio segment of the watermarked audio;
[0096] The high-frequency audio segment of the extracted watermarked audio is divided according to the number of frames corresponding to a single watermark character set of a preset watermark embedding rule to obtain a character set unit frame group, and the watermark signal is decoded for each character set unit frame group.
[0097] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0098] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0099] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0100] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A speech watermark encoding and decoding method, characterized in that: include: Down-convert the audio to obtain the high-frequency hole area of the down-converted audio; Detect the bandwidth of the high-frequency hole area of the down-tuned audio; Generate a watermark signal according to a preset encoding character string; According to the preset watermark embedding rules, the watermark signal is embedded in the high-frequency hole area of the down-tuned audio to obtain the watermarked audio and output it; receiving the watermarked audio and extracting the high-frequency audio segment of the watermarked audio; The high-frequency audio segment of the extracted watermarked audio is divided according to the number of frames corresponding to a single watermark character set of a preset watermark embedding rule to obtain a character set unit frame group, and the watermark signal is decoded for each character set unit frame group.
2. The speech watermark encoding and decoding method according to claim 1, wherein: The bandwidth of the high-frequency hole region of the down-tuned audio is specifically: The bandwidth of the high-frequency hole region of the down-converted audio is determined according to the sampling rate of the audio and the parameters of the down-conversion processing.
3. The speech watermark encoding and decoding method according to claim 1, wherein: The preset watermark embedding rule includes a watermark character set, a short-time Fourier transform frame length, a short-time Fourier transform frame shift, and the number of frames corresponding to a single watermark character set.
4. The speech watermark encoding and decoding method according to claim 3, wherein: The watermark character set includes 26 uppercase English characters, 10 Arabic numeral characters and 4 special symbol characters, and the number of frames corresponding to a single watermark character set is four.
5. The speech watermark encoding and decoding method according to claim 3, wherein: The method of embedding a watermark signal in a high-frequency hole region of the down-tuned audio according to a preset watermark embedding rule includes: Mapping the watermark character set according to the character sequence in the watermark character set and the number of frames corresponding to the single watermark character set to obtain the watermark character range represented by each frame in the number of frames corresponding to the single watermark character set; According to the frame length of the short-time Fourier transform, the number of frequency bands of each frame of audio is obtained, and each frame of audio is sorted from high to low according to frequency, so that the frequency band with the highest frequency has the largest sequence number; According to the watermark character range represented by each frame in the number of frames corresponding to the single watermark character set, characters are assigned to each frame of audio in descending order of frequency band, and a mapping relationship between the frequency band position of each frame in the number of frames corresponding to the single watermark character set and the watermark character is obtained; According to the mapping relationship between the frequency band position of each frame in the number of frames corresponding to a single watermark character set and the watermark character, combined with the preset encoding string corresponding to the watermark signal, the watermark signal is embedded in the high-frequency hole area of the down-tuned audio.
6. The speech watermark encoding and decoding method according to claim 5, wherein: The extracting of the high-frequency audio segment of the watermarked audio comprises: According to the preset watermark embedding rule of short-time Fourier transform frame length and short-time Fourier transform frame shift, the audio with watermark is short-time Fourier transformed to obtain the audio amplitude spectrum; The amplitude spectrum corresponding to a preset number of frequency bands is extracted from each frame of audio in descending order of frequency to obtain the high-frequency audio segment of the watermarked audio, wherein the preset number is equal to the number of characters in the watermark character range represented by each frame.
7. The speech watermark encoding and decoding method according to claim 1, wherein: The step of decoding the watermark signal for each character set unit frame group includes: Calculate the frequency point with the maximum energy in each character set unit frame group, and obtain the corresponding character according to the preset decoding mapping rules; The characters corresponding to all the obtained character set unit frame groups are combined in sequence into a character string.
8. A speech watermark encoding and decoding device, characterized in that: include: An audio down-conversion module is used to down-convert the audio and obtain the high-frequency hole area of the down-converted audio; A bandwidth detection module is used to detect the bandwidth of the high-frequency hole area of the down-tuned audio; A watermark signal generating module, configured to generate a watermark signal according to a preset coding string; A watermark embedding module is used to embed a watermark signal in the high-frequency hole area of the down-tuned audio according to a preset watermark embedding rule, obtain the watermarked audio and output it; A high-frequency audio segment extraction module is used to receive the watermarked audio and extract the high-frequency audio segment of the watermarked audio; The watermark signal decoding module is used to divide the high-frequency audio segment of the extracted watermarked audio according to the number of frames corresponding to a single watermark character set of the preset watermark embedding rule, obtain character set unit frame groups, and decode the watermark signal of each character set unit frame group.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the speech watermark encoding and decoding method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the speech watermark encoding and decoding method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Audio watermark embedding method and device, electronic equipment and storage medium
CN119152861A
Audio watermark generation method and device and computer storage medium
CN120236594A
Creating spectral wells for inserting watermarks in audio signals
EP3326079A1