Audio identification method and system and storage medium
Through the audio identification method of random frequency offset and amplitude adjustment of audio clips, the problem of plain text watermarks is easily tampered with, and the security and copyright protection of audio content are improved.
Patent Information
- Application Number
- CN202510777875.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In the existing audio identification technology, plain text watermarks are easily tampered with, deleted and attacked, making it difficult to track the source and copyright information of the audio.
By dividing the audio to be identified into audio segments by frames, a frequency point offset sequence is generated using a pseudo-random number generation function, and the amplitude of the audio segment at the identified frequency points is adjusted in combination with modulation parameters and watermark information to ensure the embedded position and amplitude of the watermark information are randomly distributed.
It improves the anti-attack ability of audio content, enhances the security and copyright protection effect of audio content, and it is difficult to detect the location of watermark information through traditional methods.
Smart Images

Figure CN120299464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio identification methods, and particularly to an audio identification method, system, and storage medium. Background Art
[0002] With the rapid development of digital audio technology, the dissemination and sharing of audio content have become extremely convenient. However, this convenience has also brought problems such as the abuse, piracy, and unauthorized tampering of audio content. To protect the integrity, copyright attribution of audio content, and prevent illegal tampering, existing technologies usually embed specific information in audio segments, which can be used to identify the source, copyright information of the audio, or trace the dissemination path, without significantly affecting the audio quality.
[0003] Existing audio identification technologies usually directly embed plaintext watermarks into audio segments. However, when directly embedding plaintext watermarks into audio segments, the plaintext watermarks are easily tampered with, deleted, and attacked, resulting in the inability of the audio owner or creator to still prove the source, copyright information of the audio, or trace the dissemination path. Summary of the Invention
[0004] Embodiments of the present invention provide an audio identification method, system, and storage medium, aiming to enhance the anti-attack ability of the identified audio, make the watermark information more difficult to remove, thereby enhancing the security and copyright protection effect of audio content.
[0005] To solve the above technical problems, embodiments of the present invention disclose the following technical solutions: In a first aspect, an audio identification method is provided, including: Obtain an audio to be identified, watermark information, and a master key; Divide the audio to be identified into a plurality of audio segments by frames, and each of the audio segments has its own sequence information; Determine a first sub-key based on the master key; Input the first sub-key into a pseudo-random number generation function to generate a frequency offset sequence; the frequency offset sequence contains frequency offsets corresponding to the audio segments one by one, the frequency offset is a random natural number, and the frequency offset is used to determine the identification frequency of the corresponding audio segment; determine the modulation parameters of each audio segment; Determine the identification frequency of each audio segment based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameters of each audio segment; Adjust the amplitude of each audio segment at the corresponding identification frequency based on the watermark information.
[0006] In addition to, or as an alternative to, one or more of the features disclosed above, the step of dividing the audio to be identified into a plurality of audio segments by frames includes: Determine an initial frame length and an initial overlap amount, and divide the audio to be identified into a plurality of initial audio segments according to the initial frame length and the initial overlap amount; Extract the audio features of each of the initial audio segments; the audio features at least include spectral centroid and loudness features; Based on the spectral centroid and loudness features of each of the initial audio segments, divide the plurality of initial audio segments into a high-dynamic segment set and a low-dynamic segment set; Shorten the frame length of the initial audio segments belonging to the high-dynamic segment set and reduce the overlap amount, and lengthen the frame length of the initial audio segments belonging to the low-dynamic segment set and increase the overlap amount to obtain the plurality of audio segments.
[0007] In addition to, or as an alternative to, one or more of the features disclosed above, the step of determining the modulation parameter of each of the audio segments includes: If the audio segment belongs to the high-dynamic segment set, determine the first modulation parameter as the modulation parameter of the audio segment; or; If the audio segment belongs to the low-dynamic segment set, determine the second modulation parameter as the modulation parameter of the audio segment.
[0008] In addition to, or as an alternative to, one or more of the features disclosed above, the modulation parameter includes a preset initial frequency point, a frequency step size, and a frequency point threshold; The preset initial frequency point in the first modulation parameter is lower than the preset initial frequency point in the second modulation parameter, the frequency step size in the first modulation parameter is smaller than the frequency step size in the second modulation parameter, and the frequency point threshold in the first modulation parameter is smaller than the frequency point threshold in the second modulation parameter.
[0009] In addition to, or as an alternative to, one or more of the features disclosed above, the step of determining the identification frequency point of each of the audio segments based on the sequence information of each of the audio segments, the frequency offset corresponding to each of the audio segments, and the modulation parameter of each of the audio segments includes: Obtain the sequence information corresponding to the audio segment; Generate a first frequency point based on the product of the sequence information and the frequency step size and the sum of the preset initial frequency point and the frequency offset; When the first frequency point is greater than the frequency point threshold, perform a modulo operation on the first frequency point to obtain the identification frequency point corresponding to the audio segment; or; When the first frequency point is less than the frequency point threshold, the first frequency point is the identification frequency point.
[0010] In addition to one or more of the features disclosed above, or alternatively, adjusting the amplitude of each of the audio segments at the corresponding identification frequency points based on the watermark information includes: Converting the watermark information into a binary sequence, and expanding the binary sequence, where the number of bits included in the expanded binary sequence corresponds one-to-one to the several audio segments; For each of the audio segments, adjusting the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment.
[0011] In addition to one or more of the features disclosed above, or alternatively, converting the watermark information into a binary sequence and expanding the binary sequence includes: Generating a second sub-key based on the master key; Using the second sub-key to expand the binary sequence.
[0012] In addition to one or more of the features disclosed above, or alternatively, for each of the audio segments, adjusting the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment includes: When the audio segment belongs to the high-dynamic segment set, increasing or decreasing the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment; the intensity of the increase or decrease in the amplitude is the first intensity; When the audio segment belongs to the low-dynamic segment set, increasing or decreasing the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment; the intensity of the increase or decrease in the amplitude is the second intensity; Wherein, the first intensity is less than the second intensity.
[0013] In a second aspect, an audio identification system is provided, including: An acquisition module, configured to acquire an audio to be identified, watermark information, and a master key; A frame division module, configured to divide the audio to be identified into several audio segments by frames, and each of the audio segments has its own sequence information; A key module, configured to determine a first sub-key based on the master key; A frequency point offset generation module, configured to input the first sub-key into a pseudo-random number generation function to generate a frequency point offset sequence; the frequency point offsets included in the frequency point offset sequence correspond one-to-one to the audio segments, the frequency point offset is a random natural number, and the frequency point offset is used to determine the identification frequency point of the corresponding audio segment; a parameter determination module, configured to determine the modulation parameters of each of the audio segments; A frequency point generation module, configured to determine the identification frequency points of the audio segments based on the sequence information of the audio segments, the frequency point offsets corresponding to the audio segments, and the modulation parameters of the audio segments. An identification module, configured to adjust the amplitude of each audio segment at the corresponding identification frequency point based on the watermark information.
[0014] In a third aspect, a computer-readable storage medium is provided. At least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement any one of the above audio identification methods.
[0015] One of the above technical solutions has the following advantages or beneficial effects: In this application, a frequency point offset sequence is generated through a key, and the frequency point offsets included in the frequency point offset sequence are random numbers. In this way, it can be ensured that the identification frequency points determined by each audio segment through its own sequence information, the frequency point offset corresponding to itself, and the modulation parameters are also randomly distributed. Further, adjusting the amplitude of each audio segment at the corresponding identification frequency point based on the watermark information can make the insertion positions of the watermark information of adjacent audio segments with sequence information different, and the embedding of the watermark information is more concealed, improving the anti-attack ability of the marked audio, thereby enhancing the security of the audio content and the copyright protection effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The following, in conjunction with the drawings, through a detailed description of the specific embodiments of the present invention, will make the technical solutions and other beneficial effects of the present invention obvious.
[0017] Figure 1 is a flowchart of an audio identification method provided by an embodiment of the present application Figure 1 ; Figure 2 is a flowchart of an audio identification method provided by an embodiment of the present application Figure 2 ; Figure 3 is a flowchart of an audio identification method provided by an embodiment of the present application Figure 3 ; Figure 4 is a flowchart of an audio identification method provided by an embodiment of the present application Figure 4 . DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to make the objectives, technical solutions, and beneficial effects of the present invention clearer, the following further details the present invention in conjunction with the drawings and specific embodiments. It should be understood that the specific embodiments described in this specification are only for explaining the present invention and not for limiting the present invention.
[0019] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0020] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "mounted", "connected" and "coupled" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0021] In the present invention, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through additional features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature.
[0022] Figure 1It is a schematic flowchart of an audio identification method provided by an embodiment of the present application. This specification provides method operation steps such as in the embodiment or flowchart, but based on routine or non-creative labor, it may include more or fewer operation steps. The step order listed in the embodiment is only one of the ways of the execution order of numerous steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the method order shown in the embodiment or the accompanying drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 1 shown, an audio identification method may include: S100: Obtain the audio to be identified, watermark information, and the master key.
[0023] S200: Divide the audio to be identified into a plurality of audio segments by frames, and each audio segment has its own sequence information.
[0024] S300: Determine the first sub-key based on the master key.
[0025] S400: Input the first sub-key into a pseudo-random number generation function to generate a frequency point offset sequence; the frequency point offsets included in the frequency point offset sequence correspond to the audio segments one by one, the frequency point offset is a random natural number, and the frequency point offset is used to determine the identification frequency point of the corresponding audio segment.
[0026] S500: Determine the modulation parameters of each audio segment.
[0027] S600: Determine the identification frequency point of each audio segment based on the sequence information of each audio segment, the frequency point offset corresponding to each audio segment, and the modulation parameters of each audio segment.
[0028] S700: Adjust the amplitude of each audio segment at the corresponding identification frequency point based on the watermark information.
[0029] In this application, a frequency offset sequence is generated through a secret key. The frequency offsets included in the frequency offset sequence are random natural numbers. In this way, it can be ensured that the identification frequencies determined by each audio segment through its own sequence information, the corresponding frequency offset, and the modulation parameters are also randomly distributed. By confirming the identification frequencies in this way, on the time axis of the entire audio to be identified, the identification frequencies change with time, can appear at different positions in different time periods, and the identification frequencies are distributed over multiple frequencies rather than concentrated at predictable frequencies, making it difficult for attackers to retrieve the specific position of the watermark information through traditional filtering or frequency analysis and launch an attack on the watermark information. Further, by adjusting the amplitude of each audio segment at the corresponding identification frequency based on the watermark information, it is possible to make the insertion positions of the watermark information of adjacent audio segments with sequence information different and at the same time change the amplitude adjustment strategy at different positions, making the embedding of the watermark information more concealed, improving the anti-attack ability of the identified audio, and thus enhancing the security and copyright protection effect of the audio content.
[0030] In S100, the audio to be identified, the watermark information, and the master key are obtained.
[0031] In some embodiments, the audio to be identified is obtained from a storage medium or through network transmission. The watermark information can be a copyright identifier, a user ID, or other verification information. The master key can be a high-entropy key. The watermark information and the master key are usually generated by the owner of the audio to be identified and have sufficient randomness and unpredictability.
[0032] In step S200, the audio to be identified is divided into a plurality of audio segments by frames, and each audio segment has its own sequence information.
[0033] In some embodiments, a plurality of audio segments are divided into several according to a fixed frame length and a fixed overlap amount. For example, each frame is 20 milliseconds and the overlap amount is 50%. After being divided, each audio segment is assigned a unique sequence information, and the sequence information is usually the order position of the audio segment in the audio to be identified.
[0034] Refer to Figure 2 , in some embodiments, step S200 includes: S201: Determine the initial frame length and the initial overlap amount, and divide the audio to be identified into a plurality of initial audio segments according to the initial frame length and the initial overlap amount; In step S201, the audio to be identified is divided into a plurality of initial audio segments according to the initial frame length and the initial overlap amount. The number of audio segments is determined based on the total duration of the audio to be identified, the initial frame length, and the initial overlap amount.
[0035] S202: Extract the audio features of each initial audio segment; the audio features at least include the spectral centroid and the loudness feature; In step S202, before extracting the audio features of each initial audio segment, it is necessary to first perform a short-time Fourier transform on each initial audio segment to obtain the spectral information of each initial audio segment, and calculate its spectral centroid and loudness feature based on the spectral information of each initial audio segment. The spectral centroid reflects the spectral distribution density of each audio segment, and the loudness feature is used to measure the audio energy level of each initial audio segment.
[0036] S203: Based on the spectral centroid and loudness feature of each initial audio segment, divide a number of initial audio segments into a high-dynamic segment set and a low-dynamic segment set; In step S203, each initial audio segment is divided into a high-dynamic segment set and a low-dynamic segment set. Specifically, set a centroid threshold and a loudness threshold. When at least one of the spectral centroid or the loudness feature in each initial audio segment is higher than the corresponding threshold, it is divided into a high-dynamic segment. When both the spectral centroid and the loudness feature in each initial audio segment are lower than the corresponding threshold, it is divided into a low-dynamic segment.
[0037] In some other embodiments, step S203 divides each initial audio segment into a high-dynamic segment set, a medium-dynamic segment set, and a low-dynamic segment set. The initial audio segments in the high-dynamic segment set have a loudness feature higher than the loudness threshold and a spectral centroid higher than the centroid threshold; the initial audio segments in the medium-dynamic segment set have a loudness feature higher than the loudness threshold and a spectral centroid lower than the centroid threshold; or a loudness feature lower than the loudness threshold and a spectral centroid higher than the centroid threshold; the initial audio segments in the low-dynamic segment set have a loudness feature lower than the loudness threshold and a spectral centroid lower than the centroid threshold.
[0038] S204: Shorten the frame length and reduce the overlap amount of the initial audio segments belonging to the high-dynamic segment set, and extend the frame length and increase the overlap amount of the initial audio segments belonging to the low-dynamic segment set to obtain a number of audio segments.
[0039] The frame length of the initial audio segments in the high-dynamic segment set is shortened and the overlap amount is reduced to generate the corresponding audio segments. For example, the frame length is shortened from 20 milliseconds to 10 milliseconds, and the overlap amount is reduced from 50% to 25%. This can improve the time resolution, more accurately capture the rapid changes in the audio segments, and reduce the inter-frame interference. The frame length of the initial audio segments in the low-dynamic segment set is extended and the overlap amount is increased to generate the corresponding audio segments. For example, the frame length is extended from 20 milliseconds to 40 milliseconds, and the overlap amount is increased from 50% to 75%. This can improve the frequency resolution, better analyze the spectral characteristics of the audio segments, and enhance the processing efficiency and the concealment of the watermark.
[0040] After adjusting the frame length and overlap amount of the initial audio segment, the audio to be identified can be re-partitioned by means of a sliding window technique to generate several subsequent audio segments for identification, and among the generated several audio segments, some belong to the high-dynamic segment set and some belong to the low-dynamic segment set.
[0041] In some other embodiments, the initial audio segment is also partitioned into a medium-dynamic segment set, and the frame length and overlap amount of the initial audio segment in the medium-dynamic segment set remain unchanged.
[0042] In step S300, a first sub-key is determined based on the master key.
[0043] In some embodiments, the first sub-key can be determined by inputting the master key into a key derivation function, and the key derivation function includes but is not limited to PBKDF2 (Password-Based Key Derivation Function 2), HKDF (HMAC-based Extract-and-Expand Key Derivation Function), etc. PBKDF2 is a password-based key derivation function that mainly converts the original password or master key into a high-entropy sub-key through multiple hash iterations. PBKDF2 combines a salt value and the number of iterations to enhance security. When generating the first sub-key through PBKDF2, by inputting the master key, the salt value, the number of iterations, and the length of the first sub-key, a dedicated first sub-key can be generated.
[0044] It is worth mentioning that after the audio to be identified is identified, the watermark information associated with the audio to be identified, the master key, and the relevant parameters of the key derivation function used to generate the first sub-key can all be stored in a relevant storage medium, database, or cloud storage so that the required data can be accurately and quickly obtained when extracting the watermark subsequently.
[0045] In S400, the first sub-key is input into a pseudo-random number generation function to generate a frequency offset sequence; the frequency offsets included in the frequency offset sequence correspond one-to-one to the audio segments, the frequency offset is a random natural number, and the frequency offset is used to determine the identification frequency of the corresponding audio segment.
[0046] In some embodiments, the frequency offset sequence can be generated by inputting the first sub - key into a pseudo - random number generation function. The pseudo - random number generation function includes, but is not limited to, a pseudo - random number generation function based on a hash function (such as using the SHA - 256 hash algorithm to generate a hash value as a random number through the first sub - key), and a PRNG based on an encryption algorithm (such as the AES - CTR mode, using the first sub - key as the AES key to generate a continuous random number sequence). The number of frequency offsets included in the frequency offset sequence generated by the pseudo - random number generation function can be set.
[0047] In the embodiments disclosed in the present application, the number of frequency offsets included in the frequency offset sequence corresponds one - to - one with the number of audio segments. For example, if the audio segments are divided into 50, then the frequency offset sequence contains 50 frequency offsets, and the frequency offsets in the frequency offset sequence correspond one - to - one with the audio segments, that is, the frequency offset in the first position corresponds to the audio segment with sequence information 1, and the frequency offset in the second position corresponds to the audio segment with sequence information 2.
[0048] Determine the modulation parameters of each audio segment in step S500.
[0049] In some embodiments, the modulation parameters include a first modulation parameter and a second modulation parameter. Both the first modulation parameter and the second modulation parameter include a preset initial frequency point, a frequency step size, and a frequency point threshold. The preset initial frequency point in the first modulation parameter is lower than the preset initial frequency point in the second modulation parameter, the frequency step size in the first modulation parameter is smaller than the frequency step size in the second modulation parameter, and the frequency point threshold in the first modulation parameter is smaller than the frequency point threshold in the second modulation parameter.
[0050] If the audio segment belongs to the high - dynamic segment set, determine the first modulation parameter as the modulation parameter of the audio segment; if the audio segment belongs to the low - dynamic segment set, determine the second modulation parameter as the modulation parameter of the audio segment.
[0051] The audio segments belonging to the high - dynamic segment set usually have a large volume change (from very quiet to very loud), with a high complexity. A large amount of audio information makes it more difficult for small amplitude adjustments to be perceptible by the human ear. Therefore, for the audio segments in the high - dynamic segment set, selecting the first modulation parameter can select the identification frequency points in the low - frequency region to embed the watermark information. The low - frequency region (such as 60 Hz - 1000 Hz) is less sensitive to the human ear. Embedding the watermark in this interval can reduce the impact on the sound quality and make the watermark less perceptible. And for the audio segments belonging to the high - dynamic segment set, choosing a smaller frequency step size (such as 10 Hz) can ensure that the identification frequency points are more evenly distributed in the low - frequency region, which helps to comprehensively cover the watermark information and improve its anti - attack ability.
[0052] The volume change of the audio segments belonging to the set of low-dynamic segments is small, and larger amplitude adjustments are less likely to be perceived. At the same time, it is necessary to ensure that the strength of the watermark signal is sufficient for easy extraction and recovery. Therefore, the audio segments in the set of low-dynamic segments select the second modulation parameter, which can select the identification frequency points in the high-frequency region and embed the watermark information. The high-frequency region (such as 5000 Hz - 22050 Hz) is more likely to accommodate larger amplitude adjustments, ensuring that the watermark signal is sufficiently prominent in this interval and enhancing its extractability. And the audio segments belonging to the set of low-dynamic segments select a larger frequency step (such as 500 Hz), which can ensure that the frequency hopping points are widely distributed in the high-frequency band, improving the dispersion of the watermark signal and the ability to resist various audio processing attacks (such as filtering, compression).
[0053] In step S600, based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment, determine the identification frequency point of each audio segment.
[0054] Refer to Figure 3 , in some embodiments, based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameter of each audio segment, determining the identification frequency point of each audio segment includes: S601: Obtain the sequence information corresponding to the audio segment; S602: Generate a first frequency point based on the product of the sequence information and the frequency step, and the sum of the preset initial frequency point and the frequency offset; S603: When the first frequency point is greater than the frequency threshold, perform a modulo operation on the first frequency point to obtain the identification frequency point corresponding to the audio segment; or; when the first frequency point is less than the frequency threshold, the frequency point is the identification frequency point.
[0055] Specifically, from the above content, the confirmation formula for the first frequency point is:
[0056] is the sequence information is of the audio segment's first frequency point, is the preset initial frequency point, is the sequence information of the audio segment, is the frequency step, is the frequency offset corresponding to the audio segment with sequence information k, is a natural number. For example, if 50 audio segments are divided, then is a natural number between 1 and 50, including both endpoints.
[0057] When the first frequency point is greater than the frequency threshold corresponding to the audio segment, it is necessary to perform a modulo operation adjustment on the first frequency point. The specific formula is:
[0058] To identify the frequency points, a frequency point threshold is used, which can ensure that the obtained identified frequency points are within or near a preset frequency band. The preset frequency band is , avoiding the occurrence of frequency overflow and ensuring the accuracy and stability of watermark embedding.
[0059] In some embodiments, among the first modulation parameters corresponding to the audio segments in the high-dynamic segment set, the preset initial frequency point can be around 60 Hz, which can reduce the sensitivity of the frequency to the human ear and reduce the significant impact of the watermark information on the sound quality; the frequency step can be 10 Hz. Such a setting can increase the density of the identified frequency points, ensure the uniform distribution of the watermark information in the low-frequency band, and improve the concealment; the frequency point threshold can be around 1000 Hz. Such a setting can limit the identified frequency points of the audio segments belonging to the high-dynamic set within the low-frequency band and prevent frequency overflow from affecting the important audio content in the high-frequency region.
[0060] For example, when the audio segment with sequence information 32 belongs to the high-dynamic segment set, its corresponding frequency point offset is 15 Hz, and its first frequency point is 395 Hz. Since the first frequency point is less than the frequency point threshold, the identified frequency point of the audio segment with sequence information 32 is 395 Hz. When the audio segment with sequence information 100 belongs to the high-dynamic segment set and its corresponding frequency point offset is 25 Hz, its first frequency point is 1085 Hz. Since the first frequency point is greater than the frequency point threshold, after modulo operation, the identified frequency point of the audio segment with sequence information 100 is 85 Hz.
[0061] In some embodiments, among the second modulation parameters corresponding to the audio segments in the low-dynamic segment set, the preset initial frequency point can be around 3000 Hz, which can utilize the high-frequency band to carry a stronger watermark signal and ensure the extractability of the watermark in the audio segments of the low-dynamic segment set; the frequency step can be 500 Hz. Such a setting can reduce the density of the identified frequency points, increase the coverage range of the watermark signal, and enhance the robustness; the frequency point threshold can be around 22050 Hz. Such a setting can allow the identified frequency points to cover the high-frequency band and enhance the anti-attack ability of the watermark.
[0062] For example, when the audio segment with sequence information 10 belongs to the low-dynamic segment set, its corresponding frequency point offset is 250 Hz, and its first frequency point is 10250 Hz. Since the first frequency point is less than the frequency point threshold, the identified frequency point of the audio segment with sequence information 10 is 10250 Hz.
[0063] It is worth mentioning that, in order to make the frequency range where the identification frequency points are located in the entire audio to be identified larger, by reasonably setting the preset initial frequency points and frequency thresholds of the first modulation parameter and the second modulation parameter, the coverage of the two modulation parameters can be made to cover a wider frequency range, so that the identification frequency points are more dispersed, increasing the complexity of the audio to be identified and the imperceptibility of the watermark information.
[0064] In step S700, based on the watermark information, adjust the amplitude of each audio segment at the corresponding identification frequency point.
[0065] Refer to Figure 4 , in some embodiments, step S700 includes: S701: Convert the watermark information into a binary sequence, and expand the binary sequence. The number of bits contained in the expanded binary sequence corresponds one by one to a number of audio segments; Convert the watermark information into a binary code. For example, the character A can be converted into the ASCII code 01000001. In order to enhance the robustness of the watermark, spread spectrum technology can be used to expand the binary code, and each original bit is expanded into multiple redundant bits.
[0066] In some embodiments, the way to expand the binary code is: generate a second sub-key based on the master key; use the second sub-key to expand the binary sequence.
[0067] Specifically, the way to generate the second sub-key using the master key can refer to the way to generate the first sub-key using the master key disclosed above. Use the second sub-key to expand the binary sequence of the watermark information, improving the security and unpredictability of the expansion process.
[0068] The method of using the second sub-key to expand the binary sequence can be to use the second sub-key to expand or encrypt the original binary sequence through an encryption algorithm (such as AES encryption, pseudo-random number generation, etc.). Through spread spectrum technology, a single bit is expanded into multiple redundant bits. For example, each bit is expanded into 3 bits, and by increasing the redundancy, the robustness of the watermark is improved.
[0069] It is worth mentioning that when expanding the binary code of the watermark information, it is necessary to consider how many audio segments are divided as described above. The number of audio segments is the same as the number of bits of the expanded binary code, so that the bits contained in the expanded binary code correspond to the sequence information of the audio segments according to the binary code sequence. For example, the first bit corresponds to the audio segment with sequence information 1.
[0070] S702: For each audio segment, adjust the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment.
[0071] In some embodiments, when the bit is 1, the amplitude at the identification frequency point is increased, and when the bit is 0, the amplitude at the identification frequency point is decreased.
[0072] Specifically, in some embodiments, when the audio segment belongs to the high-dynamic segment set, the amplitude of the audio segment at the identification frequency point is increased or decreased according to the bit corresponding to the audio segment; the intensity of the amplitude increase or decrease is the first intensity. When the audio segment belongs to the low-dynamic segment set, the amplitude of the audio segment at the identification frequency point is increased or decreased according to the bit corresponding to the audio segment; the intensity of the amplitude increase or decrease is the second intensity. The first intensity is less than the second intensity.
[0073] Setting the amplitude adjustment intensity (i.e., the first intensity) of the audio segments in the high-dynamic segment set to a lower value can reduce the amplitude adjustment value to reduce the impact on the sound quality and ensure the concealment of the watermark information. Setting the amplitude adjustment intensity (i.e., the second intensity) of the audio segments in the low-dynamic segment set to a higher value can ensure that the strength of the watermark information is sufficient for subsequent extraction. In some embodiments, the first intensity can be 0.02 and the second intensity can be 0.05.
[0074] In some embodiments, in order to facilitate the extraction of complete watermark information when extracting watermark information, synchronization identification information is also embedded when identifying the audio. Specifically, a group of specific frequency points are first selected as the synchronization identification frequency points. For example, the specific frequency points are 500 Hz and 3000 Hz. Then, at every preset frame length, the amplitude is adjusted at the above synchronization identification frequency points. If the preset frame length is for an audio segment in the high-dynamic segment set, the amplitude at 500 Hz is adjusted. If the preset frame length is for an audio segment in the low-dynamic segment set, the amplitude at 1000 Hz is adjusted. When extracting the watermark subsequently, the synchronization identification signal is first identified, and the audio frame position where the detected synchronization identification frequency point is located. After aligning the audio segments according to this position, the watermark information is extracted.
[0075] Further, when the watermark information is added to the audio to be identified by the above method, the watermark information can be extracted by the following method.
[0076] First, it is necessary to obtain the identified audio, watermark information, and master key, as well as the relevant functions involved in deriving the first sub-key and the second sub-key from the master key, to determine the first sub-key and the second sub-key, and to determine the frequency offset sequence based on the first sub-key, and to determine the extended binary code of the watermark information based on the second sub-key. Then, adjust the frame order of the identified audio, and divide it into a reserved number of audio segments according to the frame length and overlap amount of the high-dynamic segment set and the frame length and overlap amount of the low-dynamic segment set. And determine the identification frequency points of each audio segment according to the modulation parameters, calculate the change amount of the amplitude before and after the watermark at each identification frequency point. If the amplitude change is a positive change, the corresponding binary code is 1, and if the amplitude change is a negative change, the corresponding binary code is 0, so as to obtain the extended binary code, and restore it to the watermark information through the second sub-key.
[0077] An embodiment of the present application also provides an audio identification system, including: An acquisition module, configured to acquire the audio to be identified, watermark information, and the master key; A frame division module, configured to divide the audio to be identified into a plurality of audio segments by frames, and each audio segment has its own sequence information; A key module, configured to determine the first sub-key based on the master key; A frequency offset generation module, configured to input the first sub-key into a pseudo-random number generation function to generate a frequency offset sequence; the frequency offsets included in the frequency offset sequence correspond to the audio segments one by one, the frequency offset is a random natural number, and the frequency offset is used to determine the identification frequency point of the corresponding audio segment; a parameter determination module, configured to determine the modulation parameters of each audio segment; A frequency point generation module, configured to determine the identification frequency point of each audio segment based on the sequence information of each audio segment, the frequency offset corresponding to each audio segment, and the modulation parameters of each audio segment; An identification module, configured to adjust the amplitude of each audio segment at the corresponding identification frequency point based on the watermark information.
[0078] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction or at least one program is stored, and at least one instruction or at least one program is loaded and executed by a processor to implement any one of the above audio identification methods.
[0079] The technical features of the above embodiments can be combined arbitrarily. For the sake of brief description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0080] The above embodiments merely illustrate several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An audio identification method, characterized in that, Including: Obtain the audio to be marked, watermark information, and master key; Divide the audio to be marked into several audio segments by frames, and each of the audio segments has its own sequence information; Determine a first sub-key based on the master key; Input the first sub-key into a pseudo-random number generation function to generate a frequency offset sequence; the frequency offsets included in the frequency offset sequence correspond one-to-one to the audio segments, the frequency offset is a random natural number, and the frequency offset is used to determine the marked frequency of the corresponding audio segment; Determine the modulation parameters of each of the audio segments; Based on the sequence information of each of the audio segments, the frequency offset corresponding to each of the audio segments, and the modulation parameters of each of the audio segments, determine the marked frequency of each of the audio segments; Based on the watermark information, adjust the amplitude of each of the audio segments at the corresponding marked frequency.
2. The audio identification method according to claim 1, wherein The step of dividing the audio to be marked into several audio segments by frames includes: Determine an initial frame length and an initial overlap amount, and divide the audio to be marked into several initial audio segments according to the initial frame length and the initial overlap amount; Extract the audio features of each of the initial audio segments; the audio features at least include spectral centroid and loudness features; Based on the spectral centroid and loudness features of each of the initial audio segments, divide the several initial audio segments into a high-dynamic segment set and a low-dynamic segment set; Shorten the frame length of the initial audio segments belonging to the high-dynamic segment set and reduce the overlap amount, and extend the frame length of the initial audio segments belonging to the low-dynamic segment set and increase the overlap amount to obtain the several audio segments.
3. The audio identification method according to claim 2, wherein The step of determining the modulation parameters of each of the audio segments includes: If the audio segment belongs to the high-dynamic segment set, determine the first modulation parameter as the modulation parameter of the audio segment; or; If the audio segment belongs to the low-dynamic segment set, determine the second modulation parameter as the modulation parameter of the audio segment.
4. The audio identification method according to claim 3, wherein The modulation parameters include a preset initial frequency, a frequency step, and a frequency threshold; The preset initial frequency in the first modulation parameter is lower than the preset initial frequency in the second modulation parameter, the frequency step in the first modulation parameter is less than the frequency step in the second modulation parameter, and the frequency threshold in the first modulation parameter is less than the frequency threshold in the second modulation parameter.
5. The audio identification method according to claim 4, wherein The step of determining the marked frequency of each of the audio segments based on the sequence information of each of the audio segments, the frequency offset corresponding to each of the audio segments, and the modulation parameters of each of the audio segments includes: Obtain the sequence information corresponding to the audio segment; Generate a first frequency based on the product of the sequence information and the frequency step, and the sum of the preset initial frequency and the frequency offset; When the first frequency is greater than the frequency threshold, perform a modulo operation on the first frequency to obtain the marked frequency corresponding to the audio segment; or; When the first frequency is less than the frequency threshold, the first frequency is the marked frequency.
6. The audio identification method according to claim 2, wherein The step of adjusting the amplitude of each of the audio segments at the corresponding marked frequency based on the watermark information includes: Convert the watermark information into a binary sequence, and expand the binary sequence. The number of bits included in the expanded binary sequence corresponds to the several audio segments one by one; For each of the audio segments, adjust the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment.
7. The audio identification method according to claim 6, characterized in that The step of converting the watermark information into a binary sequence and expanding the binary sequence includes: Generate a second sub-key based on the master key; Use the second sub-key to expand the binary sequence.
8. The audio identification method according to claim 6, wherein The step of, for each of the audio segments, adjusting the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment includes: When the audio segment belongs to the high-dynamic segment set, increase or decrease the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment; the intensity of the increase or decrease in the amplitude is the first intensity; When the audio segment belongs to the low-dynamic segment set, increase or decrease the amplitude of the audio segment at the identification frequency point according to the bit corresponding to the audio segment; the intensity of the increase or decrease in the amplitude is the second intensity; Wherein, the first intensity is less than the second intensity.
9. An audio identification system, characterized in that, It includes: An acquisition module, configured to acquire the audio to be marked, watermark information, and a master key; A framing module, configured to divide the audio to be marked into several audio segments by frames, and each of the audio segments has its own sequence information; A key module, configured to determine a first sub-key based on the master key; A frequency offset generation module, configured to input the first sub-key into a pseudo-random number generation function to generate a frequency offset sequence; the frequency offsets included in the frequency offset sequence correspond to the audio segments one by one, the frequency offset is a random natural number, and the frequency offset is used to determine the identification frequency point of the corresponding audio segment; A parameter determination module, configured to determine the modulation parameters of each of the audio segments; A frequency point generation module, configured to determine the identification frequency point of each of the audio segments based on the sequence information of each of the audio segments, the frequency offset corresponding to each of the audio segments, and the modulation parameters of each of the audio segments; An identification module, configured to adjust the amplitude of each of the audio segments at the corresponding identification frequency point based on the watermark information.
10. A computer-readable storage medium, characterized in that, At least one instruction or at least one program is stored in the computer-readable storage medium, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the audio marking method according to any one of claims 1-8.
Citation Information
Patent Citations
Deep-learning-based automatic audio annotation method
CN108053836A
Audio signal processing method and device, and storage medium
CN113362837A
Audio processing method and device, storage medium and electronic equipment
CN114242111A
Watermark batch embedding method for audio library
CN117935820A
Audio watermark adding method, audio watermark extracting method and related products
CN118841019A