Audio identification method, system and storage medium

By adjusting the random frequency point offset and amplitude of the audio clips, the watermark information is embedded, which solves the problem of easy tampering with plain text watermarks, and achieves high security and copyright protection of audio content.

CN120299464BActive Publication Date: 2025-08-22SHANGHAI CANGUANG VIDEO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510777875.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-22
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

In existing audio identification technology, plain text watermarks are easily tampered with, deleted and attacked, resulting in the owner of the audio being unable to prove the source or track the propagation path, and cannot effectively protect the copyright and integrity of the audio content.

Method used

By dividing the audio to be identified into audio clips by frame, a random frequency point offset sequence and modulation parameters are generated, the amplitude of the audio clip at the identified frequency point is adjusted, and watermark information is embedded to ensure the concealment of the watermark information and the ability to resist attacks.

Benefits of technology

It improves the security and copyright protection effect of audio content, enhances the difficulty of watermark information to be removed and tampered, and improves the ability of audio content to resist attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299464B_ABST
    Figure CN120299464B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an audio identification method, system and storage medium, wherein the audio identification method includes: obtaining audio to be identified, watermark information and a master key; dividing the audio to be identified into a plurality of audio segments by frame, each of the audio segments having its own sequence information; determining a first subkey based on the master key; generating a frequency offset sequence based on the first subkey; the frequency offsets included in the frequency offset sequence correspond one-to-one to the audio segments; determining the modulation parameters of each of the audio segments; determining the identification frequency of each of the audio segments based on the sequence information of each of the audio segments, the frequency offset corresponding to each of the audio segments and the modulation parameters of each of the audio segments; and adjusting the amplitude of each of the audio segments at the corresponding identification frequency based on the watermark information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio identification methods, and in particular to an audio identification method, system and storage medium. Background Art

[0002] With the rapid development of digital audio technology, the dissemination and sharing of audio content has become extremely convenient. However, this convenience also brings with it problems such as misuse, piracy, and unauthorized modification of audio content. To protect the integrity and copyright of audio content, as well as to prevent illegal modification, existing technologies often embed specific information into audio clips. This information can be used to identify the audio source, copyright information, or track the transmission path without significantly affecting audio quality.

[0003] Existing audio identification technology usually embeds plaintext watermarks directly into audio clips. However, when plaintext watermarks are directly embedded into audio clips, they are easily tampered with, deleted, and attacked, resulting in the owner or creator of the audio still being unable to prove the audio's source, copyright information, or track the transmission path. Summary of the Invention

[0004] The embodiments of the present invention provide an audio identification method, system, and storage medium, which are intended to improve the anti-attack capability of the identified audio, making the watermark information more difficult to remove, thereby enhancing the security and copyright protection of the audio content.

[0005] In order to solve the above technical problems, the embodiments of the present invention disclose the following technical solutions:

[0006] In a first aspect, a method for audio identification is provided, comprising:

[0007] Obtain the audio to be identified, watermark information and master key;

[0008] Dividing the audio to be identified into a plurality of audio segments by frame, each of the audio segments having its own sequence information;

[0009] determining a first subkey based on the master key;

[0010] inputting the first subkey into a pseudo-random number generation function to generate a frequency offset sequence; wherein the frequency offsets included in the frequency offset sequence correspond one-to-one with the audio segments, the frequency offsets being random natural numbers, and the frequency offsets being used to determine the identification frequency of the corresponding audio segment; and determining the modulation parameters of each audio segment;

[0011] determining, based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment, an identification frequency of each audio segment;

[0012] Based on the watermark information, the amplitude of each of the audio segments at the corresponding identified frequency point is adjusted.

[0013] In addition to or as an alternative to one or more features disclosed above, dividing the audio to be identified into a plurality of audio segments by frame includes:

[0014] Determining an initial frame length and an initial overlap amount, and dividing the audio to be identified into a plurality of initial audio segments according to the initial frame length and the initial overlap amount;

[0015] Extracting audio features of each of the initial audio segments; the audio features at least include spectral centroid and loudness features;

[0016] Dividing the plurality of initial audio segments into a high-dynamic segment set and a low-dynamic segment set based on the spectral centroid and loudness characteristics of each of the initial audio segments;

[0017] The frame lengths of the initial audio segments belonging to the high-dynamic segment set are shortened and the overlap amount is reduced, while the frame lengths of the initial audio segments belonging to the low-dynamic segment set are lengthened and the overlap amount is increased, so as to obtain the plurality of audio segments.

[0018] In addition to or as an alternative to one or more features disclosed above, determining the modulation parameters of each of the audio segments includes:

[0019] If the audio segment belongs to the high dynamic segment set, determining the first modulation parameter as the modulation parameter of the audio segment; or;

[0020] If the audio segment belongs to the low-dynamic segment set, the second modulation parameter is determined to be the modulation parameter of the audio segment.

[0021] In addition to or as an alternative to one or more features disclosed above, the modulation parameters include a preset initial frequency, a frequency step, and a frequency threshold;

[0022] The preset initial frequency point in the first modulation parameter is lower than the preset initial frequency point in the second modulation parameter, the frequency step in the first modulation parameter is smaller than the frequency step in the second modulation parameter, and the frequency threshold in the first modulation parameter is smaller than the frequency threshold in the second modulation parameter.

[0023] In addition to or as an alternative to one or more features disclosed above, determining the identified frequency of each audio segment based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment includes:

[0024] Get the sequence information corresponding to the audio clip;

[0025] Generate a first frequency point based on a product of the sequence information and the frequency step length and a sum of the preset initial frequency point and the frequency point offset;

[0026] When the first frequency point is greater than the frequency point threshold, performing a modulo operation on the first frequency point to obtain an identified frequency point corresponding to the audio segment; or;

[0027] When the first frequency point is less than the frequency point threshold, the first frequency point is an identification frequency point.

[0028] In addition to or as an alternative to one or more features disclosed above, adjusting the amplitude of each audio segment at the corresponding identified frequency point based on the watermark information includes:

[0029] Converting the watermark information into a binary sequence and extending the binary sequence, wherein bits included in the extended binary sequence correspond one-to-one to the plurality of audio clips;

[0030] For each of the audio segments, the amplitude of the audio segment at the identified frequency point is adjusted according to the bit corresponding to the audio segment.

[0031] In addition to or as an alternative to one or more features disclosed above, converting the watermark information into a binary sequence and expanding the binary sequence includes:

[0032] generating a second subkey based on the master key;

[0033] The binary sequence is expanded using the second subkey.

[0034] In addition to or as an alternative to one or more features disclosed above, adjusting, for each audio segment, the amplitude of the audio segment at the identified frequency point according to a bit corresponding to the audio segment includes:

[0035] When the audio segment belongs to the high-dynamic segment set, increasing or decreasing the amplitude of the audio segment at the identified frequency point according to the bit corresponding to the audio segment; the intensity of the amplitude increase or decrease is a first intensity;

[0036] When the audio segment belongs to the low-dynamic segment set, increasing or decreasing the amplitude of the audio segment at the identified frequency point according to the bit corresponding to the audio segment; the intensity of the amplitude increase or decrease is a second intensity;

[0037] The first intensity is smaller than the second intensity.

[0038] In a second aspect, an audio identification system is provided, comprising:

[0039] An acquisition module is used to obtain the audio to be identified, watermark information and master key;

[0040] a framing module, configured to divide the audio to be identified into a plurality of audio segments by frame, each of the audio segments having its own sequence information;

[0041] a key module, configured to determine a first subkey based on the master key;

[0042] a frequency offset generation module, configured to input the first subkey into a pseudo-random number generation function to generate a frequency offset sequence; wherein the frequency offsets included in the frequency offset sequence correspond one-to-one with the audio segments, the frequency offsets being random natural numbers and used to determine the identification frequency of the corresponding audio segment; and a parameter determination module, configured to determine the modulation parameters of each audio segment;

[0043] a frequency point generation module, configured to determine an identification frequency point of each audio segment based on sequence information of each audio segment, a frequency point offset corresponding to each audio segment, and a modulation parameter of each audio segment;

[0044] The identification module is configured to adjust the amplitude of each of the audio segments at the corresponding identified frequency points based on the watermark information.

[0045] In a third aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement any of the above-mentioned audio identification methods.

[0046] One of the above-mentioned technical solutions has the following advantages or beneficial effects: This application generates a frequency offset sequence using a key, and the frequency offsets included in the frequency offset sequence are random numbers. This ensures that the identification frequency of each audio segment, determined by its own sequence information, corresponding frequency offset, and modulation parameters, is also randomly distributed. Furthermore, by adjusting the amplitude of each audio segment at the corresponding identification frequency based on the watermark information, the watermark information can be inserted at different positions in audio segments with adjacent sequence information, making the watermark information more subtle and enhancing the identified audio's resistance to attacks, thereby strengthening the security and copyright protection of the audio content. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The technical solutions and other beneficial effects of the present invention will be made apparent by describing in detail the specific embodiments of the present invention in conjunction with the accompanying drawings.

[0048] Figure 1 This is a flow diagram of an audio identification method provided by an embodiment of the present application. Figure 1 ;

[0049] Figure 2 This is a flow diagram of an audio identification method provided by an embodiment of the present application. Figure 2 ;

[0050] Figure 3 This is a flow diagram of an audio identification method provided by an embodiment of the present application. Figure 3 ;

[0051] Figure 4 This is a flow diagram of an audio identification method provided by an embodiment of the present application. Figure 4 . DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and beneficial effects of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described in this specification are only for the purpose of explaining the present invention and are not intended to limit the present invention.

[0053] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate the orientation or position relationship based on the orientation or position relationship shown in the accompanying drawings, which is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" refers to two or more, unless otherwise clearly and specifically defined.

[0054] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections, direct connections, or indirect connections through an intermediate medium; they may refer to internal communication between two components or the interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0055] In the present invention, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may include the first and second features being in direct contact, or may also include the first and second features not being in direct contact but being in contact via another feature between them. Furthermore, a first feature being "above," "above," and "above" a second feature may include the first feature being directly above or diagonally above the second feature, or may simply mean that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may include the first feature being directly above or diagonally above the second feature, or may simply mean that the first feature is lower in level than the second feature.

[0056] Figure 1 It is a flowchart of an audio identification method provided by an embodiment of the present application. This specification provides method operation steps such as the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative work. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in the order shown in the embodiment or the figure or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, Figure 1 As shown, an audio identification method may include:

[0057] S100: Obtain the audio to be identified, watermark information and master key.

[0058] S200: Divide the audio to be identified into several audio segments by frame, and each audio segment has its own sequence information.

[0059] S300: Determine the first subkey based on the master key.

[0060] S400: Input the first subkey into a pseudo-random number generation function to generate a frequency offset sequence; the frequency offsets included in the frequency offset sequence correspond one-to-one with the audio clips, and the frequency offsets are random natural numbers. The frequency offsets are used to determine the identification frequency of the corresponding audio clip.

[0061] S500: Determine the modulation parameters of each audio segment.

[0062] S600: Determine the identification frequency of each audio segment based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment.

[0063] S700: Based on the watermark information, adjust the amplitude of each audio clip at the corresponding marked frequency point.

[0064] The present application generates a frequency offset sequence through a key, and the frequency offset contained in the frequency offset sequence is a random natural number, so that the identification frequency determined by each audio segment through its own sequence information, its corresponding frequency offset and modulation parameters is also randomly distributed. By confirming the identification frequency in this way, on the time axis of the entire audio to be identified, the identification frequency changes with time and can appear at different positions in different time periods. Moreover, the identification frequency is spread over multiple frequencies and is not concentrated at predictable frequencies, making it difficult for attackers to retrieve the specific location of the watermark information through traditional filtering or frequency analysis and launch attacks on the watermark information. Furthermore, by adjusting the amplitude of each audio segment at the corresponding identification frequency based on the watermark information, the watermark information of the audio segments of adjacent sequence information can be inserted at different positions, and the amplitude adjustment strategy can be changed at different positions, making the embedding of the watermark information more covert, improving the anti-attack capability of the identified audio, and thus enhancing the security and copyright protection of the audio content.

[0065] In S100 , the audio to be identified, watermark information, and a master key are obtained.

[0066] In some embodiments, the audio to be identified is obtained from a storage medium or transmitted over a network. The watermark information can be a copyright identifier, a user ID, or other authentication information. The master key can be a high-entropy key. The watermark information and master key are typically generated by the owner of the audio to be identified and are sufficiently random and unpredictable.

[0067] In step S200 , the audio to be identified is divided into a number of audio segments by frame, and each audio segment has its own sequence information.

[0068] In some embodiments, multiple audio segments are divided into multiple segments with a fixed frame length and a fixed overlap, for example, 20 milliseconds per frame with a 50% overlap. After segmentation, each audio segment is assigned unique sequence information, which typically represents the sequential position of the audio segment in the audio to be identified.

[0069] Reference Figure 2 In some embodiments, step S200 includes:

[0070] S201: Determine an initial frame length and an initial overlap amount, and divide the audio to be identified into a plurality of initial audio segments according to the initial frame length and the initial overlap amount;

[0071] In step S201, the audio to be identified is divided into a number of initial audio segments according to the initial frame length and initial overlap amount. The number of audio segments is determined based on the total duration of the audio to be identified, the initial frame length, and the initial overlap amount.

[0072] S202: Extracting audio features of each initial audio segment; the audio features include at least spectral centroid and loudness features;

[0073] In step S202, before extracting audio features from each initial audio segment, a short-time Fourier transform is performed on each initial audio segment to obtain its spectral information. Based on this spectral information, the spectral centroid and loudness feature are calculated. The spectral centroid reflects the spectral distribution density of each audio segment, while the loudness feature measures the audio energy level of each initial audio segment.

[0074] S203: Divide the initial audio segments into a high-dynamic segment set and a low-dynamic segment set based on the spectral centroid and loudness characteristics of each initial audio segment;

[0075] In step S203, each initial audio segment is divided into a set of high-dynamic segments and a set of low-dynamic segments. Specifically, a centroid threshold and a loudness threshold are set. When at least one of the spectral centroid or loudness feature of each initial audio segment is above the corresponding threshold, the segment is classified as a high-dynamic segment. When both the spectral centroid or loudness feature are below the corresponding threshold, the segment is classified as a low-dynamic segment.

[0076] In some other embodiments, step S203 divides the initial audio segments into a high-dynamic segment set, a medium-dynamic segment set, and a low-dynamic segment set. The initial audio segments in the high-dynamic segment set have loudness characteristics that are higher than a loudness threshold and a spectral centroid that is higher than a centroid threshold; the initial audio segments in the medium-dynamic segment set have loudness characteristics that are higher than the loudness threshold and a spectral centroid that is lower than the centroid threshold; or have loudness characteristics that are lower than the loudness threshold and a spectral centroid that is higher than the centroid threshold; and the initial audio segments in the low-dynamic segment set have loudness characteristics that are lower than the loudness threshold and a spectral centroid that is lower than the centroid threshold.

[0077] S204: shortening the frame length of the initial audio segments belonging to the high-dynamic segment set and reducing the overlap amount, and lengthening the frame length of the initial audio segments belonging to the low-dynamic segment set and increasing the overlap amount, to obtain a plurality of audio segments.

[0078] The frame length of the initial audio segments in the high-dynamic segment set is shortened, and the overlap is reduced to generate corresponding audio segments. For example, the frame length is shortened from 20 milliseconds to 10 milliseconds, and the overlap is reduced from 50% to 25%. This improves temporal resolution, more accurately captures rapid changes in the audio segments, and reduces inter-frame interference. The frame length of the initial audio segments in the low-dynamic segment set is lengthened, and the overlap is increased to generate corresponding audio segments. For example, the frame length is shortened from 20 milliseconds to 40 milliseconds, and the overlap is reduced from 50% to 75%. This improves frequency resolution, better analyzes the spectral characteristics of the audio segments, enhances processing efficiency, and improves the concealment of the watermark.

[0079] After adjusting the frame length and overlap of the initial audio segments, the audio to be identified can be re-divided using the sliding window technology to generate several audio segments for subsequent identification. Among the generated audio segments, some belong to a high-dynamic segment set and some belong to a low-dynamic segment set.

[0080] In some other embodiments, the initial audio segment is further divided into a set of medium dynamic segments, and the frame lengths and overlap amounts of the initial audio segments in the set of medium dynamic segments remain unchanged.

[0081] In step S300 , a first subkey is determined based on the master key.

[0082] In some embodiments, the first subkey can be determined by inputting the master key into a key derivation function, including but not limited to PBKDF2 (Password-Based Key Derivation Function 2) and HKDF (HMAC-based Extract-and-Expand Key Derivation Function). PBKDF2 is a password-based key derivation function that converts the original password or master key into a high-entropy subkey through multiple hashing iterations. PBKDF2 combines a salt value and a number of iterations to enhance security. When generating the first subkey using PBKDF2, the master key, salt value, number of iterations, and the length of the first subkey are input to generate a unique first subkey.

[0083] It is worth mentioning that after the audio to be identified is identified, the watermark information, master key and relevant parameters of the key derivation function used to generate the first subkey associated with the audio to be identified can all be stored in relevant storage media, databases or cloud storage, so that the required data can be obtained accurately and quickly when the watermark is extracted subsequently.

[0084] In S400, the first subkey is input into a pseudo-random number generation function to generate a frequency offset sequence; the frequency offsets included in the frequency offset sequence correspond one-to-one with the audio clips, the frequency offsets are random natural numbers, and the frequency offsets are used to determine the identification frequency of the corresponding audio clip.

[0085] In some embodiments, a frequency offset sequence can be generated by inputting the first subkey into a pseudo-random number generation function. Pseudo-random number generation functions include, but are not limited to, those based on hash functions (e.g., using the SHA-256 hash algorithm, generating a hash value as a random number using the first subkey) and those based on encryption algorithms (e.g., in AES-CTR mode, generating a continuous sequence of random numbers using the first subkey as the AES key). The number of frequency offsets included in the frequency offset sequence generated by the pseudo-random number generation function can be set.

[0086] In the embodiment disclosed in the present application, the number of frequency offsets included in the frequency offset sequence corresponds one-to-one to the number of audio clips. For example, if the audio clip is divided into 50, the frequency offset sequence contains 50 frequency offsets, and the frequency offsets in the frequency offset sequence correspond one-to-one to the audio clips, that is, the frequency offset in the first position corresponds to the audio clip with sequence information of 1, and the frequency offset in the second position corresponds to the audio clip with sequence information of 2.

[0087] In step S500 , the modulation parameters of each audio segment are determined.

[0088] In some embodiments, the modulation parameters include a first modulation parameter and a second modulation parameter. The first modulation parameter and the second modulation parameter both include a preset initial frequency, a frequency step, and a frequency threshold. The preset initial frequency in the first modulation parameter is lower than the preset initial frequency in the second modulation parameter, the frequency step in the first modulation parameter is smaller than the frequency step in the second modulation parameter, and the frequency threshold in the first modulation parameter is smaller than the frequency threshold in the second modulation parameter.

[0089] If the audio segment belongs to the high-dynamic segment set, the first modulation parameter is determined as the modulation parameter of the audio segment; if the audio segment belongs to the low-dynamic segment set, the second modulation parameter is determined as the modulation parameter of the audio segment.

[0090] Audio clips in the high-dynamic segment set typically exhibit large volume variations (from very quiet to very loud) and are highly complex. This large amount of audio information makes even small amplitude adjustments more difficult to detect. Therefore, the first modulation parameter is selected for audio clips in the high-dynamic segment set, enabling the watermark to be embedded in the low-frequency region. The low-frequency region (for example, 60 Hz to 1000 Hz) is less sensitive to the human ear, so embedding the watermark in this range can reduce the impact on sound quality and make the watermark less noticeable. Furthermore, a smaller frequency step size (for example, 10 Hz) is used for audio clips in the high-dynamic segment set to ensure a more even distribution of the marker frequencies within the low-frequency region, which helps ensure comprehensive watermark coverage and improves its anti-attack capabilities.

[0091] The audio segments in the low-dynamic segment set have smaller volume changes, making larger amplitude adjustments less noticeable. At the same time, it is necessary to ensure that the watermark signal is strong enough to facilitate extraction and recovery. Therefore, the audio segments in the low-dynamic segment set select the second modulation parameter, which can select an identification frequency point in the high-frequency region to embed the watermark information. The high-frequency region (for example, 5000Hz–22050Hz) can more easily accommodate larger amplitude adjustments, ensuring that the watermark signal is sufficiently significant within this range and enhancing its extractability. The audio segments in the low-dynamic segment set use a larger frequency step (for example, 500Hz) to ensure that the frequency hopping points are widely distributed in the high-frequency band, improving the dispersion of the watermark signal and its ability to resist various audio processing attacks (such as filtering and compression).

[0092] In step S600 , the identification frequency of each audio segment is determined based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment.

[0093] Reference Figure 3 In some embodiments, determining the identified frequency of each audio segment based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment includes:

[0094] S601: Obtain sequence information corresponding to the audio clip;

[0095] S602: Generate a first frequency based on the product of the sequence information and the frequency step length and the sum of the preset initial frequency and the frequency offset;

[0096] S603: When the first frequency point is greater than the frequency point threshold, perform a modulo operation on the first frequency point to obtain an identification frequency point corresponding to the audio segment; or when the first frequency point is less than the frequency point threshold, the frequency point is an identification frequency point.

[0097] Specifically, from the above content, it can be seen that the confirmation formula for the first frequency point is:

[0098]

[0099] The sequence information is The first frequency point of the audio clip, To preset the initial frequency, is the sequence information of the audio segment, is the frequency step size, is the frequency offset corresponding to the audio segment with sequence information k, is a natural number, for example, if 50 audio clips are divided, then That is, a natural number between 1 and 50, including the two endpoints.

[0100] When the first frequency point is greater than the frequency point threshold corresponding to the audio clip, a modulo operation adjustment needs to be performed on the first frequency point. The specific formula is:

[0101]

[0102] To mark the frequency point, is the frequency threshold, which ensures that the identified frequency is within or near the preset frequency band. The preset frequency band is , avoid frequency overflow and ensure the accuracy and stability of watermark embedding.

[0103] In some embodiments, in the first modulation parameters corresponding to the audio segments in the high-dynamic segment set, the preset initial frequency can be around 60 Hz, which can reduce the sensitivity of the frequency to the human ear and reduce the significant impact of the watermark information on the sound quality; the frequency step can be 10 Hz, which can increase the density of the identification frequency, ensure that the watermark information is evenly distributed in the low-frequency band, and improve the concealment; the frequency threshold can be around 1000 Hz, which can limit the identification frequency of the audio segments in the high-dynamic set to the low-frequency band, and prevent frequency overflow from affecting important audio content in the high-frequency area.

[0104] For example, if the audio clip with sequence information 32 belongs to the high-dynamic segment set, its corresponding frequency offset is 15Hz, and its first frequency is 395Hz. This first frequency is less than the frequency threshold, so the identified frequency of the audio clip with sequence information 32 is 395Hz. If the audio clip with sequence information 100 belongs to the high-dynamic segment set, its corresponding frequency offset is 25Hz, and its first frequency is 1085Hz. This first frequency is greater than the frequency threshold. After performing a modulo operation, the identified frequency of the audio clip with sequence information 100 is 85Hz.

[0105] In some embodiments, in the second modulation parameters corresponding to the audio segments in the low-dynamic segment set, the preset initial frequency can be around 3000 Hz, so that the high frequency band can be used to carry a stronger watermark signal, ensuring the watermark extractability of the audio segments in the low-dynamic segment set; the frequency step can be 500 Hz, so that the setting can reduce the density of the identification frequency points, increase the coverage range of the watermark signal, and improve the robustness; the frequency threshold can be around 22050 Hz, so that the setting can allow the identification frequency points to spread throughout the high frequency band, thereby enhancing the watermark's anti-attack capability.

[0106] For example, when the audio clip with sequence information 10 belongs to the low-dynamic clip set, its corresponding frequency offset is 250 Hz, and its first frequency is 10250 Hz. The first frequency is less than the frequency threshold, so the identified frequency of the audio clip with sequence information 10 is 10250 Hz.

[0107] It is worth mentioning that in order to expand the frequency range of the identification frequency points in the entire audio to be identified, the preset initial frequency points and frequency thresholds of the first modulation parameter and the second modulation parameter can be reasonably set so that the two modulation parameters cover a wider frequency range, thereby making the identification frequency points more dispersed, increasing the complexity of the identified audio and the imperceptibility of the watermark information.

[0108] In step S700 , the amplitude of each audio segment at the corresponding marked frequency point is adjusted based on the watermark information.

[0109] Reference Figure 4 In some embodiments, step S700 includes:

[0110] S701: Convert the watermark information into a binary sequence and expand the binary sequence, where the bits of the expanded binary sequence correspond one-to-one to the plurality of audio clips;

[0111] The watermark information is converted into binary code, for example, the character A can be converted into the ASCII code 01000001. In order to enhance the robustness of the watermark, the spread spectrum technology binary coding can be used for expansion, and each original bit can be expanded into multiple redundant bits.

[0112] In some embodiments, the binary code is expanded by: generating a second subkey based on the master key; and expanding the binary sequence using the second subkey.

[0113] Specifically, the method of using the master key to generate the second subkey can refer to the above-disclosed method of using the master key to generate the first subkey, and the second subkey is used to expand the binary sequence of the watermark information, thereby improving the security and unpredictability of the expansion process.

[0114] The method for using the second subkey to expand the binary sequence can be to use the second subkey to expand or encrypt the original binary sequence through an encryption algorithm (such as AES encryption or pseudo-random number generation). Using spread spectrum technology, a single bit is expanded into multiple redundant bits. For example, each bit can be expanded to three bits. This increases the redundancy and improves the robustness of the watermark.

[0115] It is worth mentioning that when expanding the binary encoding of the watermark information, it is necessary to consider how many audio segments are divided as mentioned above. The number of audio segments is the same as the number of bits of the expanded binary encoding, so that the bits contained in the expanded binary encoding correspond to the sequence information of the audio segments according to the sequence of the binary encoding, for example, the first bit corresponds to the audio segment with sequence information of 1.

[0116] S702: For each audio segment, adjust the amplitude of the audio segment at the identified frequency point according to the bit corresponding to the audio segment.

[0117] In some embodiments, when the bit is 1, the amplitude at the identified frequency point is increased, and when the bit is 0, the amplitude at the identified frequency point is decreased.

[0118] Specifically, in some embodiments, when an audio clip belongs to a high-dynamic clip set, the amplitude of the audio clip at the identified frequency is increased or decreased based on the bit position corresponding to the audio clip; the intensity of the amplitude increase or decrease is a first intensity. When an audio clip belongs to a low-dynamic clip set, the amplitude of the audio clip at the identified frequency is increased or decreased based on the bit position corresponding to the audio clip; the intensity of the amplitude increase or decrease is a second intensity. The first intensity is less than the second intensity.

[0119] Setting the amplitude adjustment strength (i.e., the first strength) for audio clips in the high-dynamic segment set to a lower value can reduce the amplitude adjustment value, thereby minimizing the impact on sound quality and ensuring the concealment of the watermark information. Setting the amplitude adjustment strength (i.e., the second strength) for audio clips in the low-dynamic segment set to a higher value can ensure sufficient watermark strength for subsequent extraction. In some embodiments, the first strength can be 0.02, and the second strength can be 0.05.

[0120] In some embodiments, in order to facilitate the extraction of complete watermark information when extracting watermark information, synchronization identification information will also be embedded when identifying audio. Specifically, a group of specific frequency points are first selected as synchronization identification frequency points, such as the specific frequency points are 500Hz and 3000Hz. Then, at each preset frame length, the amplitude is adjusted at the above-mentioned synchronization identification frequency point. If the preset frame length is in the middle of the audio segment in the high-dynamic segment set, the amplitude at 500Hz is adjusted. If the preset frame length is in the middle of the audio segment in the low-dynamic segment set, the amplitude at 1000Hz is adjusted. When extracting the watermark subsequently, the synchronization identification signal is first identified, and the audio frame position of the detected synchronization identification frequency point is aligned with the audio segment according to the position before extracting the watermark information.

[0121] Furthermore, after the watermark information is added to the audio to be identified by the above method, the watermark information can be extracted by the following method.

[0122] First, the identified audio, watermark information, and master key are obtained, along with the correlation functions involved in deriving the first and second subkeys from the master key. The first and second subkeys are then determined, along with a frequency offset sequence based on the first subkey and an expanded binary code of the watermark information based on the second subkey. The frame sequence of the identified audio is then adjusted, and the audio is divided into a predetermined number of audio segments based on the frame length and overlap of the high-dynamic segment set and the frame length and overlap of the low-dynamic segment set. Modulation parameters are then used to determine the identified frequencies of each audio segment. At each identified frequency, the amplitude change before and after the watermark is calculated. A positive amplitude change corresponds to a binary code of 1, while a negative amplitude change corresponds to a binary code of 0. This expanded binary code is then restored to the watermark information using the second subkey.

[0123] The present application also provides an audio identification system, including:

[0124] An acquisition module is used to obtain the audio to be identified, watermark information and master key;

[0125] A framing module is used to divide the audio to be identified into several audio segments by frame, each audio segment has its own sequence information;

[0126] A key module, configured to determine a first subkey based on a master key;

[0127] a frequency offset generation module, configured to input the first subkey into a pseudo-random number generation function to generate a frequency offset sequence; wherein the frequency offsets included in the frequency offset sequence correspond one-to-one with the audio segments, the frequency offsets being random natural numbers and used to determine the identification frequency of the corresponding audio segment; and a parameter determination module, configured to determine the modulation parameters of each audio segment;

[0128] A frequency point generation module is used to determine the identification frequency point of each audio segment based on the sequence information of each audio segment, the frequency point offset corresponding to each audio segment, and the modulation parameter of each audio segment;

[0129] The identification module is used to adjust the amplitude of each audio segment at the corresponding identified frequency point based on the watermark information.

[0130] An embodiment of the present application further provides a computer-readable storage medium, in which at least one instruction or at least one program is stored. The at least one instruction or at least one program is loaded and executed by a processor to implement any of the above-mentioned audio identification methods.

[0131] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0132] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. An audio identification method, characterized in that: include: Obtain the audio to be identified, watermark information and master key; Determining an initial frame length and an initial overlap amount, and dividing the audio to be identified into a plurality of initial audio segments according to the initial frame length and the initial overlap amount; Extracting audio features of each of the initial audio segments, wherein the audio features include at least a spectral centroid and a loudness feature; and dividing the plurality of initial audio segments into a high-dynamic segment set and a low-dynamic segment set based on the spectral centroid and the loudness feature of each of the initial audio segments; shortening the frame lengths of the initial audio segments belonging to the high-dynamic segment set and reducing the amount of overlap, and lengthening the frame lengths of the initial audio segments belonging to the low-dynamic segment set and increasing the amount of overlap, to obtain the plurality of audio segments; each of the audio segments having its own sequence information; determining a first subkey based on the master key; Inputting the first subkey into a pseudorandom number generation function to generate a frequency offset sequence; the frequency offsets included in the frequency offset sequence correspond one-to-one with the audio clips, the frequency offsets being random natural numbers, and the frequency offsets being used to determine the identification frequency of the corresponding audio clip; determining a modulation parameter for each of the audio segments; determining, based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment, an identification frequency of each audio segment; Based on the watermark information, the amplitude of each of the audio segments at the corresponding identified frequency point is adjusted.

2. The audio identification method according to claim 1, characterized in that The determining of the modulation parameters of each of the audio segments includes: If the audio segment belongs to the high dynamic segment set, determining the first modulation parameter as the modulation parameter of the audio segment; or; If the audio segment belongs to the low-dynamic segment set, the second modulation parameter is determined to be the modulation parameter of the audio segment.

3. The audio identification method according to claim 2, characterized in that: The modulation parameters include a preset initial frequency, a frequency step, and a frequency threshold; The preset initial frequency point in the first modulation parameter is lower than the preset initial frequency point in the second modulation parameter, the frequency step in the first modulation parameter is smaller than the frequency step in the second modulation parameter, and the frequency threshold in the first modulation parameter is smaller than the frequency threshold in the second modulation parameter.

4. The audio identification method according to claim 3, characterized in that: The determining, based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment, of the identified frequency of each audio segment includes: Get the sequence information corresponding to the audio clip; Generate a first frequency point based on a product of the sequence information and the frequency step length and a sum of the preset initial frequency point and the frequency point offset; When the first frequency point is greater than the frequency point threshold, performing a modulo operation on the first frequency point to obtain an identified frequency point corresponding to the audio segment; or; When the first frequency point is less than the frequency point threshold, the first frequency point is an identification frequency point.

5. The audio identification method according to claim 1, characterized in that: The adjusting the amplitude of each audio segment at the corresponding identified frequency point based on the watermark information includes: Converting the watermark information into a binary sequence and extending the binary sequence, wherein bits included in the extended binary sequence correspond one-to-one to the plurality of audio clips; For each of the audio segments, the amplitude of the audio segment at the identified frequency point is adjusted according to the bit corresponding to the audio segment.

6. The audio identification method according to claim 5, characterized in that: The converting the watermark information into a binary sequence and expanding the binary sequence includes: generating a second subkey based on the master key; The binary sequence is expanded using the second subkey.

7. The audio identification method according to claim 5, characterized in that: The adjusting, for each of the audio segments, the amplitude of the audio segment at the identified frequency point according to the bit corresponding to the audio segment includes: When the audio segment belongs to the high-dynamic segment set, increasing or decreasing the amplitude of the audio segment at the identified frequency point according to the bit corresponding to the audio segment; the intensity of the amplitude increase or decrease is a first intensity; When the audio segment belongs to the low-dynamic segment set, increasing or decreasing the amplitude of the audio segment at the identified frequency point according to the bit corresponding to the audio segment; the intensity of the amplitude increase or decrease is a second intensity; The first intensity is smaller than the second intensity.

8. An audio identification system, characterized in that: The audio identification method according to any one of claims 1 to 7 is applied, comprising: An acquisition module is used to obtain the audio to be identified, watermark information and master key; a framing module, configured to divide the audio to be identified into a plurality of audio segments by frame, each of the audio segments having its own sequence information; a key module, configured to determine a first subkey based on the master key; a frequency offset generation module, configured to input the first subkey into a pseudo-random number generation function to generate a frequency offset sequence; wherein the frequency offsets included in the frequency offset sequence correspond one-to-one with the audio segments, the frequency offsets being random natural numbers, and the frequency offsets being used to determine the identification frequency of the corresponding audio segment; a parameter determination module, configured to determine a modulation parameter of each of the audio segments; a frequency point generation module, configured to determine an identification frequency point of each audio segment based on sequence information of each audio segment, a frequency point offset corresponding to each audio segment, and a modulation parameter of each audio segment; The identification module is configured to adjust the amplitude of each of the audio segments at the corresponding identified frequency points based on the watermark information.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the audio identification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audio watermark adding method, audio watermark extracting method and related products

    CN118841019A

  • System and method to prevent audio watermark detection

    US20110243327A1