Audio-based real-time power grid frequency following and tamper discrimination method and system

CN117558294BActive Publication Date: 2026-09-18HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311656023.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2026-09-18
Estimated Expiration
2043-12-05

AI Technical Summary

Technical Problem

文献([21]WANG T T. The segmented chirp Z-transformand its application in spectrum analysis [J]. 1990, 39(2): 318-323.)探讨了CZT的优缺点,提出一种分段Chirp Z算法([22]MAX, YANG Q, CHEN C, et al. Harmonic andInterharmonic Analysis of Mixed Dense Frequency Signals [J]. IEEETransactions on Industrial Electronics, 2021, 68(10): 10142-10153.),它能够有效实现部分频谱高分辨率计算,克服CZT无法处理较大输入数据的问题

Benefits of technology

1、本发明方法通过针对每一帧音频信号作短时线性调频Z变换,并查找指定窗口内最大幅值对应的频率kl,将所有窗口内获得的频率拼接,得到频率估计值序列F,对频率估计值序列F进行平滑滤波得到滤波后的频率估计值序列fz,能够实现对电网频率信号的准确提取,而且即使在低信噪比条件下,仍然能够保持较高的提取精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117558294B_ABST
    Figure CN117558294B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for real-time tracking and tamper detection of power grid frequencies based on audio. The method includes noise reduction, framing, and short-time linear frequency modulation of the audio signal containing the power grid frequency. Z Transform to find the frequency corresponding to the maximum amplitude within a specified window. k l By concatenating the frequencies obtained within all windows, a frequency estimation sequence is obtained. F The filtered frequency estimate sequence is obtained by performing a smoothing filter. f z Then it is compared with the reference frequency sequence. r Perform point-by-point calculation of Pearson similarity CC and European distance D The optimal matching point is obtained, and finally, the optimal matching point is used as the reference frequency sequence. r The starting point and f z A comparison is made, and tampering is judged based on whether there is a sudden change. The present invention aims to achieve accurate extraction of power grid frequency signals, especially accurate extraction of power grid frequency signals under low signal-to-noise ratio conditions, and to achieve real-time tracking and tampering identification of audio signals containing power grid frequencies based on power grid frequency signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital audio tampering detection technology, specifically to a method and system for real-time tracking and tampering detection of power grid frequencies based on audio. Background Technology

[0002] With the rapid development of digital technology, secondary editing and tampering of audio have become more covert, and the unavoidable noise interference in actual audio increases the difficulty of identifying genuine audio. Power grid frequency technology is one of the typical audio forensics methods proposed in recent years. Frequency fluctuations are caused by the imbalance between power supply and demand. Because these fluctuations are closely related to changes in power grid load, they are highly random. Furthermore, interconnected power grids share the same power network frequency (ENF) variation trend. Based on this characteristic, Catalin Grigoras proposed a method for authenticating digital audio based on power grid frequency (GRIGORAS C. Applications of ENF criterion in forensic audio, video, computer and telecommunication analysis [J]. Forensic Sci Int, 2007, 167(2-3): 136-145.). The literature (GRIGORAS C. Applications of ENF Analysis in Forensic Authentication of Digital Audio and Video Recordings [J]. J Audio Eng Soc, 2009, 57(9): 643-661. and Bao Yongqiang, Liang Ruiyu, Cong Yun, et al. Research progress on several key technologies of audio forensics [J]. Data Acquisition and Processing, 2016, 31(2):252-259.) elucidates the principles and methods of ENF, explores the classification of audio forensics, and constructs a framework for audio forensics. The analysis method of ENF mainly involves the following three steps: 1) Extracting the ENF signal from the audio signal to be tested; 2) Collecting reference data of ENF; 3) Matching the two to obtain the audio recording time.

[0003] The Short-Time Fourier Transform (STFT) is the most commonly used algorithm for ENF extraction. However, STFT suffers from insufficient spectral resolution due to its fixed step size. The literature (HUA G, GOH J, THING VL L. A Dynamic Matching Algorithm for Audio Timestamp Identification Using the ENF Criterion [J]. IEEE Transactions on Information Forensics and Security, 2014, 9(7): 1045-1055.) uses a threshold-based dynamic matching algorithm, where the threshold is selected based on the frequency resolution determined by the size of the STFT window. The literature (HAJJ-AHMAD A, GARG R, WU M, et al. Instantaneous frequency estimation and localization for ENF signals;proceedings of the Proceedings of The 2012 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, Hollywood, CA, F Dec 03-06, 2012 [C]. 2012.) uses an ENF signal parameter estimation method to estimate the instantaneous frequency with as few samples as possible, thereby improving the frequency estimation accuracy. The literature (OJOWU O, KARLSSON J, LI J, et al. ENFExtraction From Digital Recordings Using Adaptive Techniques and FrequencyTracking [J]. IEEE Transactions on Information Forensics and Security, 2012,7(4): 1330-1338.) proposes a frequency tracking method based on dynamic programming based on the characteristic that ENF changes slowly over time, thereby improving the estimation accuracy of ENF. The literature (Liu Yuming. Frequency measurement of power system and its application in the identification of the authenticity of digital voice [D]; Chongqing University, 2012.) used an external sound card to record voice signals and a computer's built-in sound card to record power grid signals, thus completing the acquisition of digital audio signals containing ENF.The paper (AIELLO M, CATALIOTTI A, NUCCIO S. Achirp-z transform-based synchronizer for power system measurements [J]. 2005,54(3): 1025-1032.) proposes a fundamental frequency detection algorithm based on linear frequency modulated Z-transform (CZT). It can make the CZT synchronizer perform better than the interpolated Fourier transform synchronizer without increasing the observation window length and computational workload. The literature (

[21] WANG T T. The segmented chirp Z-transform and its application in spectrum analysis [J]. 1990, 39(2): 318-323.) discusses the advantages and disadvantages of CZT and proposes a segmented Chirp Z algorithm (

[22] MAX, YANG Q, CHEN C, et al. Harmonic and Interharmonic Analysis of Mixed Dense Frequency Signals [J]. IEEE Transactions on Industrial Electronics, 2021, 68(10): 10142-10153.), which can effectively realize high-resolution calculation of part of the spectrum and overcome the problem that CZT cannot handle large input data. There is a lot of background noise in actual audio signals, and the existing power grid frequency estimation methods do not consider the extraction performance of digital audio under low signal-to-noise ratio conditions. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and system for real-time tracking and tamper identification of power grid frequency based on audio, which addresses the above-mentioned problems in the prior art. The present invention aims to achieve accurate extraction of power grid frequency signals, especially accurate extraction of power grid frequency signals under low signal-to-noise ratio conditions, and to achieve real-time tracking and tamper identification of audio signals containing power grid frequency signals based on power grid frequency signals.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: An audio-based method for real-time tracking and tamper detection of power grid frequencies includes: S101, performs noise reduction on audio signals containing power grid frequency signals; S102, divides the noise-reduced audio signal into frames according to the set frame length and movement step size, and performs short-time linear frequency modulation on each frame of audio signal.Z Transform to find the frequency corresponding to the maximum amplitude within a specified window. k l , all frequencies obtained within the window k l The frequency estimate sequence is obtained by splicing. F ; S103, for the frequency estimate sequence F Smoothing filtering is performed to obtain the filtered frequency estimate sequence. f z ; S104, the filtered frequency estimate sequence f z With reference frequency sequence composed of power grid frequency signals r Perform point-by-point calculation of Pearson similarity sequence CC and Euclidean distance sequence D In the Pearson similarity sequence CC and Euclidean distance sequence D Find the best matching point between them. If the best matching point is found, proceed to step S105; otherwise, proceed to step S102. S105, using the optimal matching point as the reference frequency sequence. r The starting point, from the reference frequency sequence r The starting point and the filtered frequency estimate sequence fz A point-by-point comparison is performed. If a sudden change occurs at a certain point, it is determined that the audio containing the power grid frequency signal has been tampered with and the location of the sudden change is the location of the tampering; otherwise, it is determined that the audio containing the power grid frequency signal has not been tampered with.

[0006] Optionally, the function expression for the short-time linear frequency modulated Z-transform of each frame of audio in step S102 is: , In the above formula, For the audio of this frame at frequency k place Z Transformation value, For along Z A segment of a spiral on a plane is used as sampling points to bisect the angle. For the frame window size, This is the audio signal after bandpass filtering. For sampling points, As the starting point of the spiral profile, This is the ratio of adjacent points on the spiral profile. For window functions, Centered at the window, and having: , , , In the above formula, The window number. For the movement step size, The starting sampling point for the spiral profile. The imaginary unit, The phase angle of the initial sampling point z0, The angle difference between adjacent sampling points. denoted as the elongation of the spiral.

[0007] Optionally, in step S102, all frequencies obtained within the window are... k l The frequency estimate sequence is obtained by splicing. F The function expression is: , In the above formula, For frequency estimate sequence F The Middle l One frequency estimate, For the first l The frequency corresponding to the maximum amplitude within each window The frequency corresponding to the maximum amplitude. The window number. For window length, The number of frames.

[0008] Optionally, the reference frequency sequence composed of the power grid frequency signal in step S104 r The reference frequency sequence was obtained by synchronously detecting power grid signals using power grid potential sensing equipment. r It is a signal sequence ordered by time from the power grid frequency signal, and its time covers the acquisition time of the audio signal containing the power grid frequency signal.

[0009] Optionally, in step S104, the Pearson similarity sequence CC and Euclidean distance sequence D Finding the best matching point between them includes: S201, Search to determine Pearson similarity sequence CC The position corresponding to the maximum value ICC max Search to determine the Euclidean distance sequence D The position corresponding to the minimum value ID min ; S202, judgment ICC max ≈ ID minIf the condition is met, the best matching point is found; otherwise, the best matching point is not found.

[0010] Optionally, the judgment ICC max ≈ ID min Whether it holds true includes: calculating the Pearson similarity sequence. CC The position corresponding to the maximum value ICC max Euclidean distance sequence D The position corresponding to the minimum value ID min The difference between them, if the difference is less than the set value, is then determined. ICC max ≈ ID min If true, otherwise judge ICC max ≈ ID min This is not true.

[0011] Optionally, a mutation at a certain point in step S104 means that the point in the reference frequency sequence... r The power grid frequency signal and the frequency estimation sequence after filtering. fz The difference between the frequency estimates and the set value exceeds the set value.

[0012] Optionally, the noise reduction in step S101 refers to filtering and noise reduction using a bandpass filter. Before step S101, the audio signal containing the power grid frequency signal is also downsampled.

[0013] Furthermore, the present invention also provides an audio-based real-time power grid frequency tracking and tamper detection system, comprising a microprocessor and a memory interconnected thereto, wherein the microprocessor is programmed or configured to execute the audio-based real-time power grid frequency tracking and tamper detection method.

[0014] Furthermore, the present invention also provides a computer-readable storage medium storing a computer program that is programmed or configured by a microprocessor to execute the audio-based real-time power grid frequency tracking and tamper detection method.

[0015] Compared with the prior art, the present invention has the following main advantages: 1. The method of the present invention performs short-time linear frequency modulation on each frame of audio signal. Z Transform and find the frequency corresponding to the maximum amplitude value within the specified window. k l By concatenating the frequencies obtained within all windows, a frequency estimation sequence is obtained. FFor the frequency estimate sequence F Smoothing filtering is performed to obtain the filtered frequency estimate sequence. f z It can accurately extract power grid frequency signals, and maintain high extraction accuracy even under low signal-to-noise ratio conditions.

[0016] 2. This invention includes filtering the frequency estimation sequence. f z With reference frequency sequence r Perform point-by-point calculation of Pearson similarity sequence CC and Euclidean distance sequence D Then search for the position corresponding to the maximum Pearson similarity. ICCmax The position corresponding to the minimum value of the Euclidean distance sequence IDmax If index point ICC max ≈ ID min If the index point is the best match, then the index point is the optimal match point; otherwise, the frequency estimate sequence is extracted again, and the optimal match point is used as the reference frequency sequence. r The starting point and the filtered frequency estimate sequence fz By comparison, if a sudden change is found, it is determined that the audio containing the power grid frequency signal has been tampered with and the change position is the tampered position; otherwise, it is determined that the audio containing the power grid frequency signal has not been tampered with. This method can easily and quickly achieve real-time tracking and tamper identification of audio signals containing the power grid frequency signal based on the power grid frequency signal. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the basic process of the method in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the system structure of the method in an embodiment of the present invention.

[0019] Figure 3 The grid frequency signal (a) generated in this embodiment of the invention and its original frequency (b) are shown.

[0020] Figure 4 This refers to the audio signal synthesized in this embodiment of the invention, which includes the power grid frequency signal.

[0021] Figure 5 These are the results of power grid frequency signals extracted by different methods in the embodiments of the present invention.

[0022] Figure 6 The power grid frequency sequence extracted by the STFT method for comparison f z1 With reference frequency sequence rPearson similarity and Euclidean distance between different points.

[0023] Figure 7 The power grid frequency sequence extracted by the STFT method for comparison f z1 With reference frequency sequence r The matching results.

[0024] Figure 8 The filtered frequency estimate sequence extracted by the method in this embodiment of the invention. f z With reference frequency sequence r Pearson similarity and Euclidean distance between different points.

[0025] Figure 9 The filtered frequency estimate sequence extracted by the method in this embodiment of the invention. f z With reference frequency sequence r The matching results.

[0026] Figure 10 The results of the digital audio authenticity detection experiment are shown in the embodiments of the present invention. Detailed Implementation

[0027] like Figure 1 As shown, the audio-based real-time power grid frequency tracking and tamper detection method in this embodiment includes: S101, performs noise reduction on audio signals containing power grid frequency signals; S102, divides the noise-reduced audio signal into frames according to the set frame length and movement step size, and performs short-time linear frequency modulation on each frame of audio signal. Z Transform to find the frequency corresponding to the maximum amplitude within a specified window. k l , all frequencies obtained within the window k l The frequency estimate sequence is obtained by splicing. F ; S103, for the frequency estimate sequence F Smoothing filtering is performed to obtain the filtered frequency estimate sequence. f z ; S104, the filtered frequency estimate sequence f z With reference frequency sequence composed of power grid frequency signals r Perform point-by-point calculation of Pearson similarity sequence CC and Euclidean distance sequence D In the Pearson similarity sequence CC and Euclidean distance sequence DFind the best matching point between them. If the best matching point is found, proceed to step S105; otherwise, proceed to step S102. S105, using the optimal matching point as the reference frequency sequence. r The starting point, from the reference frequency sequence r The starting point and the filtered frequency estimate sequence fz A point-by-point comparison is performed. If a sudden change occurs at a certain point, it is determined that the audio containing the power grid frequency signal has been tampered with and the location of the sudden change is the location of the tampering; otherwise, it is determined that the audio containing the power grid frequency signal has not been tampered with.

[0028] Since the audio signal sampling rate is 48000Hz, the data volume will become very large when the audio recording time is long. In this embodiment, before denoising the audio signal containing the power grid frequency signal in step S101, the audio containing the power grid frequency signal is further downsampled to improve the system response speed.

[0029] Since audio signals contain various noises, unwanted sounds, and harmonics, bandpass filtering is required for the original signal. In this embodiment, noise reduction in step S101 refers to using a bandpass filter for filtering and noise reduction.

[0030] The method in this embodiment includes performing short-time linear frequency modulation on each frame of audio signal. Z Transformation, linear frequency modulation Z The Chirp Z-Transform (CZT) can sample along a spiral trajectory and can be rapidly computed using the FFT algorithm. Furthermore, its frequency resolution can be adaptively adjusted, enabling high-resolution analysis of narrowband signals and offering great flexibility, making it particularly suitable for speech signal analysis.

[0031] In step S102 of this embodiment, short-time linear frequency modulation is performed on each frame of audio. Z The functional expression for the transformation is: (1) In the above formula, For the audio of this frame at frequency k place Z Transformation value, For along Z A segment of a spiral on a plane is used as sampling points to bisect the angle. For the frame window size, This is the audio signal after bandpass filtering. For sampling points, As the starting point of the spiral profile, This is the ratio of adjacent points on the spiral profile. For window functions, Centered at the window, and having: (2) (3) (4) In the above formula, The window number. For the movement step size, The starting sampling point for the spiral profile. The imaginary unit, The phase angle of the initial sampling point z0, The angle difference between adjacent sampling points. The extension rate of the spiral. In step S102, short-time linear frequency modulation is performed on each frame of audio. Z The derivation of the functional expression for the transformation is as follows: Considering that a real audio signal contains both speech and electrical signal, it can be represented as: (5) In the above formula, s ( t (This refers to the actual audio signal.) c ( t () is a power grid signal; ɛ ( t () represents the speech signal.

[0032] Power grid signals c ( t ) can be seen as being caused by M Composed of a cosine component and a noise signal, it can be represented as: (6) In the above formula, A i Let be the amplitude of the i-th cosine component; α i Let be the attenuation factor of the i-th cosine component; f i Let be the frequency of the i-th cosine component; θ i Let i be the phase of the i-th cosine component; n ( t The signal is noise. Euclidean transformation of equation (6) and simplification yield the discretized power grid signal. c ( n ): (7) In the above formula, 0 ≤n≤N- 1; N For signal c ( t The number of sampling points; p i =A i e jθ Residue information representing the sampled signal; z i =e (α i +jω i ) Represents pole information, where ω i = 2 πf i It is the angular frequency. Therefore, the discretized audio signal... s ( n )for: (8) In the above formula, ɛ ( n ) represents the discretized speech signal.

[0033] Because the audio signal sampling rate is 48000Hz, the data volume becomes extremely large when the audio recording time is long, requiring signal processing. Y Downsampling reduces the amount of data and improves the system's response speed. The rated frequency of the Chinese power grid is 50Hz, with a permissible frequency deviation of ±0.5Hz. Since audio signals contain various noises, interferences, and harmonics, bandpass filtering is required on the original signal. The functional expression for this is: (9) In the above formula, x ( n () represents the output signal of the bandpass filter. s ( nk ) is the first nk A discretized sampled audio signal h ( k ) represents the coefficients of the filter; Q This refers to the order of the filter. For the output signal of a bandpass filter... x ( n )conduct Z Transformation: (10) In the above formula, X ( z () is the output signal of the bandpass filter. x ( n )of Z Transformation, Nfor x ( n The length of z -n Represents a complex number. Along Z Sampling is performed on a spiral segment in a plane, dividing the angle into equal parts. The sampling points are... z k for: (11) In the above formula, , M The total number of sampling points. A It is the starting point of the spiral profile. W It is the ratio of adjacent points on the spiral profile, which can be expressed as: (12) Substituting formula (12) into formula (11) yields: (13) In the above formula, A 0 is the starting sampling point. When hour, z 0 represents the length of the vector radius. When hour, z 0 is located on the unit circle Inside, otherwise, z 0 will be on the unit circle The outside. Indicates the starting sampling point z The phase angle is 0. ϕ 0 represents the angle difference between adjacent sampling points. When ϕ When 0 is positive, it means z k The path rotates counterclockwise; when ϕ When 0 is negative, it means z k The path rotates clockwise. W 0 represents the elongation of the spiral. When W When 0 < 1, follow k Increase the spiral extension; when W When 0>1, then follow k The increase of spiral inward; when W 0 = When 1, it means that the radius is A An arc segment of a circle. At this point, if... A 0 = 1. Then this arc is part of the unit circle.

[0034] Z The values ​​at these sampling points are transformed as follows: (14) Substituting formula (11) into formula (14) yields: (15) Let signal x ( n The length of ) is N Take a window with a length of M The step size between each two adjacent windows is H Then the signal can be decomposed into L windows, and we have: (16) For the l A window, with its center defined as: (17) Define the window function as ω [ k If the processing result in the window is the function expression of the short-time linear frequency modulation Z-transform for each frame of audio in step S102 of this embodiment, that is, formula (1). X [ z k ] is the frequency of this window. k place Z Transformation value.

[0035] Search for the maximum amplitude within this window. X [ z k ] max Corresponding position Index peak The frequency estimate is obtained. f max Finally, the frequencies obtained from all windows are concatenated to form the frequency estimate sequence. F In step S102 of this embodiment, all frequencies obtained within the window are... k l The frequency estimate sequence is obtained by splicing. F The function expression is: (18) In the above formula, For frequency estimate sequence F The Middle l One frequency estimate, For the first l The frequency corresponding to the maximum amplitude within each window The frequency corresponding to the maximum amplitude. The window number. For window length, The number of frames.

[0036] In step S104 of this embodiment, the reference frequency sequence is composed of the power grid frequency signal. r The reference frequency sequence was obtained by synchronously detecting power grid signals using power grid potential sensing equipment. r This is a signal sequence ordered by time from the power grid frequency signal, and its time span covers the acquisition time of the audio signal containing the power grid frequency signal. See details below. Figure 2 In this embodiment, a grid situational awareness device (GSAD) is used to synchronously acquire the grid frequency, and a frequency database is established as a sequence of filtered frequency estimates. f z Provide reference frequency sequence r The system in this embodiment includes a laptop computer, an external sound card, a microphone, a smartphone, a GPS navigation antenna, a step-down device, a power system, and a satellite. The external sound card is a Sound Blaster X-Fi Surround 5.1 Pro v3. The step-down device consists of a transformer and voltage divider resistors; its main function is to reduce the 220V / 50Hz AC voltage to a small voltage signal with an amplitude of 0.1V through the transformer and resistors. Because the dynamic range of human voice is very large, and the low frequencies are significantly higher than normal, the phone speaker cannot be placed too close to the microphone during recording. Therefore, the smartphone is placed 5cm away from the microphone. The recorded audio is saved as an m4a file.

[0037] In step S104 of this embodiment, the filtered frequency estimation value sequence is... f z With reference frequency sequence composed of power grid frequency signals r Perform point-by-point calculation of Pearson similarity sequence CC and Euclidean distance sequence D At that time, Pearson similarity sequence CC The expression for calculating the Pearson similarity of each element is as follows: (19) In the above formula, The filtered frequency estimate sequence f z The frequency estimate of the i-th element in the equation. The filtered frequency estimate sequence f z The average of all frequency estimates, Reference frequency sequence r The i-th reference frequency in Reference frequency sequence r The reference frequency mean is N, where N is the filtered frequency estimate sequence.f z The sequence length.

[0038] In step S104 of this embodiment, the filtered frequency estimation value sequence is... f z With reference frequency sequence composed of power grid frequency signals r Perform point-by-point calculation of Pearson similarity sequence CC and Euclidean distance sequence D At that time, Euclidean distance sequence D The function expression for calculating each Euclidean distance in the equation is: (20) In the above formula, The filtered frequency estimate sequence f z The frequency estimate of the i-th element in the equation. Reference frequency sequence r The i-th reference frequency in the sequence, where N is the filtered frequency estimate sequence. f z The sequence length.

[0039] In step S104 of this embodiment, the Pearson similarity sequence... CC and Euclidean distance sequence D Finding the best matching point between them includes: S201, Search to determine Pearson similarity sequence CC The position corresponding to the maximum value ICC max Search to determine the Euclidean distance sequence D The position corresponding to the minimum value ID min ; S202, judgment ICC max ≈ ID min If the condition is met, the best matching point is found; otherwise, the best matching point is not found.

[0040] Specifically, this embodiment determines ICC max ≈ ID min Whether it holds true includes: calculating the Pearson similarity sequence. CC The position corresponding to the maximum value ICC max Euclidean distance sequence D The position corresponding to the minimum value ID min The difference between them, if the difference is less than the set value, is then determined. ICCmax ≈ ID min If true, otherwise judge ICC max ≈ ID min This is not true. Alternatively, other approximate implementations can be used to achieve a similar judgment effect, depending on the requirements.

[0041] In this embodiment, a sudden change at a certain point in step S104 means that the point in the reference frequency sequence... r The power grid frequency signal and the frequency estimation sequence after filtering. fz The difference between the frequency estimates and the set value exceeds the set value. Alternatively, other mathematical methods can be used to calculate the abrupt change value, such as the proportion of the difference, as needed.

[0042] To verify and analyze the method in this embodiment, the Chirp function is used to generate a power grid signal, and the signal expression is as follows: ,(twenty one) In the above formula, f 0 is the starting frequency. b For frequency modulation. In this embodiment, the lower frequency limit is set to 49.9Hz, the upper frequency limit is set to 50.1Hz, the chirp duration is set to 60s, the sampling rate is set to 48000Hz, and the amplitude is set to 0.05V. This generates a complex power grid signal containing a standard power grid signal with an initial frequency of 50Hz, a 25Hz simple harmonic wave signal with an amplitude of 0.005V, and a 100Hz second harmonic wave signal, and saves it as an m4a file. Figure 3 As shown, Figure 3 In the diagram, (a) represents the mains frequency signal generated by the Chirp function. Figure 3 (b) shows the change of the power grid frequency signal over time. The audio signal is synthesized from the speech signal and the power grid frequency signal to obtain an audio signal containing the power grid frequency signal, as shown in the image. Figure 4 As shown.

[0043] In this embodiment, the downsampling factor will be... Y Set to 100 to reduce the sampling rate of the audio signal from 48000Hz to 480Hz. Perform a bandpass filter [49.95Hz, 50.05Hz] on the downsampled signal to eliminate noise while retaining frequency components. Divide the signal into sub-signals of 1 second per frame, with a frame length of 480, a movement step size of 48, a transformation point count of 2048, a refinement start frequency of 49.5Hz, and a refinement cutoff frequency of 51Hz. Calculate the spiral profile start point using formula (12). The ratio of adjacent points on the spiral profile is 0.7973 + j0.6036.W The value is 1-j0.00001. At a signal-to-noise ratio of 15 dB, the comparison between the power grid frequency signal (ENF signal) extracted from the audio signal by the method of this embodiment (SCZT), the existing RDFT method, and the STFT method, and the real power grid frequency signal is shown below. Figure 5 As shown. From Figure 5 As can be seen, when the signal-to-noise ratio (SNR) is 15 dB, compared with the RDFT method and the STFT method, the power grid frequency signal extracted by the method in this embodiment (SCZT) is significantly closer to the real power grid frequency signal and has higher detection accuracy.

[0044] To further quantify the detection accuracy of the method in this embodiment (SCZT), the existing RDFT method, and the STFT method, the signal sequences of the power grid frequency signals extracted by the three methods under different signal-to-noise ratio (SNR) conditions were matched point by point with the real frequency sequences to obtain the corresponding Pearson similarity and Euclidean distance, as shown in Tables 1 and 2.

[0045] Table 1. Pearson similarity of the three methods under different SNR conditions.

[0046] As shown in Table 1, with the increase of signal-to-noise ratio (SNR), the Pearson similarity between the power grid frequency signal extracted by the method in this embodiment (SCZT) and the real power grid frequency signal remains almost unchanged and remains above 0.999. However, under the condition of a SNR of 1 dB, the Pearson similarity between the power grid frequency signal extracted by the RDFT method and the STFT method and the real power grid frequency signal is 0.137 and 0.8639, respectively, which is significantly lower than that of the method in this embodiment (SCZT).

[0047] Table 2. Euclidean distances for the three algorithms under different SNR conditions.

[0048] As shown in Table 2, with the increase of signal-to-noise ratio (SNR), the Euclidean distance between the grid frequency signal extracted by the method in this embodiment (SCZT) and the actual grid frequency signal remains almost unchanged and stays below 0.258. However, under the condition of a SNR of 1 dB, the Euclidean distances between the grid frequency signal extracted by the RDFT method and the actual grid frequency signal are 10.2446 and 1.0044, respectively, which are much higher than the 0.258 of the method in this embodiment (SCZT).

[0049] This embodiment uses a 15-minute digital audio segment as an example, with a signal-to-noise ratio of 20.59 dB and a sampling rate of 48000 Hz. The STFT algorithm and the algorithm proposed in this paper are used to extract the ENF signal from the digital audio. The first cutoff frequency of the bandpass filter is set to 49.5 Hz, the second cutoff frequency is set to 50.5 Hz, and the filter order is... Q The value is set to 1501, and the window function is a Hamming window. Then, the STFT method is used to extract the power grid frequency sequence from the digital audio. f z1 And calculate its frequency sequence with the reference frequency sequence recorded by GSAD for 30 minutes. r Pearson similarity and Euclidean distance between different points, such as Figure 6 As shown. From Figure 6 It can be seen that the STFT method extracts the power grid frequency sequence. f z1 With reference frequency sequence r The maximum similarity (Pearson similarity) between different points is 0.961, corresponding to sampling point 5043; the minimum Euclidean distance is 0.8273, also corresponding to sampling point 5043. Therefore, the 5043rd point can be determined as the optimal matching sampling point. The reference frequency sequence is then processed starting from the 5043rd point. r (The curve is labeled GSAD) and the power grid frequency sequence f z1 (The curve is labeled STFT) Matching was performed, and the result is as follows: Figure 7 As shown, the extracted power grid frequency sequence f z1 It can reflect the trend of frequency changes.

[0050] Figure 8 It is a filtered frequency estimate sequence extracted from a digital audio signal using the method (SCZT) of this embodiment. f z And calculate its frequency sequence with the reference frequency sequence recorded by GSAD for 30 minutes. r Pearson similarity sequence, composed of similarity and Euclidean distance between different points. CC and Euclidean distance sequence D .from Figure 8 As can be seen, the filtered frequency estimate sequence extracted by the method in this embodiment... f z (The curve is labeled SCZT) and the reference frequency sequence r(The curve is labeled GSAD) The maximum Pearson similarity between different points is 0.9932, corresponding to sampling point 4561; the minimum Euclidean distance is 0.2263, also corresponding to sampling point 4561. Therefore, the 4561st point can be determined as the optimal matching sampling point, and the time corresponding to the optimal matching sampling point is the creation time of the digital recording. Starting from the 4561st point, the reference frequency sequence... r (The curve is labeled GSAD) and the filtered frequency estimate sequence f z (The curve is labeled SCZT) Matching is performed, and the matching results are as follows: Figure 9 As shown. From Figure 9 As can be seen from the above, the filtered frequency estimate sequence obtained using the method of this embodiment... f z (The curve is marked SCZT) not only accurately reflects the frequency change trend, but also provides detailed frequency information. (Comparison) Figure 6 and Figure 7 It can be seen that the filtered frequency estimate sequence obtained using the method of this embodiment (SCZT) f z The Pearson similarity to the true frequency sequence was improved by 3.22 × 10⁻⁶ compared to the STFT method. -2 The Euclidean distance decreased by 0.6009. Therefore, the method of this embodiment (SCZT) is superior to the STFT method.

[0051] The method (SCZT) in this embodiment is used to extract filtered frequency estimation sequences of digital audio signals with different signal-to-noise ratios. f z The Pearson similarity and Euclidean distance between the frequency sequence and the measured frequency sequence of GSAD were calculated, and the results are shown in Table 3.

[0052] Table 3. Pearson similarity and Euclidean distance of the method (SCZT) in this embodiment under different SNR conditions.

[0053] As shown in Table 3, under a signal-to-noise ratio of 1.22 dB, the Pearson similarity between the ENF estimate sequence extracted by the method (SCZT) in this embodiment and the true frequency sequence is 0.9584, and the Euclidean distance is less than 0.57. This demonstrates that the method (SCZT) in this embodiment can maintain high extraction accuracy even under low signal-to-noise ratio conditions.

[0054] To test the integrity detection of digital audio and determine its tampering type, this embodiment uses the voice editing software GoldWave to perform deletion and insertion tampering on 15 minutes of digital audio. Deletion tampering involves removing the content between the 5th minute and the 2nd second of the 5th minute; insertion tampering involves inserting a 2-second audio segment at the 10th minute. Figure 10 As shown, the power grid frequency signal with the tampered audio removed is shifted downwards by 0.02 Hz, and the power grid frequency signal with the tampered audio inserted is shifted upwards by 0.02 Hz. The filtered frequency estimation sequence of the audio signal after removal and tampering is obtained using the method of this embodiment. f z Compared to the power grid frequency signal value acquired by GSAD (GSAD data), there is a significant abrupt change at sampling point 2988. Therefore, it can be determined that the audio signal has been deleted and tampered with, and the deletion / tampering location is at sampling point 2988, which is consistent with the actual deletion / tampering location. The filtered frequency estimation sequence of the inserted / tampered audio signal obtained using the method of this embodiment is shown below. f z Compared with the power grid frequency signal value collected by GSAD, there is a significant abrupt change at sampling points 5934 and 5985. Therefore, it can be determined that the audio is inserted and tampered with. Moreover, the inserted and tampered audio is between sampling points 5934 and 5985, which is consistent with the actual insertion and tampering location.

[0055] In summary, this embodiment addresses the problems of insufficient extraction accuracy of existing ENF methods under low signal-to-noise ratio conditions and insufficient spectral resolution caused by the fixed moving step size of STFT. The method described in this embodiment first establishes a mathematical model of the audio signal containing ENF, then extracts the grid frequency component through low-pass filtering, downsampling, band-pass filtering, and short-time linear frequency modulation Z-transform, uses a smoothing filter to eliminate residual noise in the signal, and finally evaluates the matching degree between the ENF estimate and the grid frequency using Pearson similarity and Euclidean distance. Furthermore, the method is validated using the system involved in this embodiment, confirming its practicality. Experimental results show that this embodiment's method can achieve ENF signal extraction and database matching, demonstrating good theoretical research value and engineering application prospects in the field of digital audio ENF extraction. Experimental results also show that, under low signal-to-noise ratio conditions, the filtered frequency estimate sequence obtained by this embodiment's method (SCZT method) is... f z The Pearson similarity and Euclidean distance with the true frequency sequence were improved and reduced to 0.9932 and 0.2264 respectively compared with the traditional STFT algorithm, verifying the superiority of the method in this embodiment. Under the conditions of low signal-to-noise ratio and audio signal tampering, the method in this embodiment can effectively determine whether the audio is complete and locate the tampering position of the audio, verifying the effectiveness of the method in this embodiment.

[0056] Furthermore, this embodiment also provides an audio-based real-time power grid frequency tracking and tamper detection system, including a microprocessor and a memory interconnected, wherein the microprocessor is programmed or configured to execute the audio-based real-time power grid frequency tracking and tamper detection method. This embodiment also provides a computer-readable storage medium storing a computer program for being programmed or configured by the microprocessor to execute the audio-based real-time power grid frequency tracking and tamper detection method.

[0057] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0058] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principle of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for real-time tracking and tamper detection of power grid frequencies based on audio, characterized in that, include: S101, performs noise reduction on audio signals containing power grid frequency signals; S102, divides the noise-reduced audio signal into frames according to the set frame length and movement step size, and performs short-time linear frequency modulation on each frame of audio signal. Z Transform to find the frequency corresponding to the maximum amplitude within a specified window. k l , all frequencies obtained within the window k l The frequency estimate sequence is obtained by splicing. F ; S103, for the frequency estimate sequence F Smoothing filtering is performed to obtain the filtered frequency estimate sequence. f z ; S104, the filtered frequency estimate sequence f z With reference frequency sequence composed of power grid frequency signals r Perform point-by-point calculation of Pearson similarity sequence CC and Euclidean distance sequence D In the Pearson similarity sequence CC and Euclidean distance sequence D Find the best matching point between them. If the best matching point is found, proceed to step S105. Otherwise, proceed to step S102; S105, using the optimal matching point as the reference frequency sequence. r The starting point, from the reference frequency sequence r The starting point and the filtered frequency estimate sequence fz A point-by-point comparison is performed. If a sudden change occurs at a certain point, it is determined that the audio containing the power grid frequency signal has been tampered with and the location of the sudden change is the location of the tampering; otherwise, it is determined that the audio containing the power grid frequency signal has not been tampered with.

2. The method for real-time tracking and tamper detection of power grid frequency based on audio according to claim 1, characterized in that, The function expression for performing a short-time linear frequency modulated Z-transform on each frame of audio in step S102 is as follows: , In the above formula, For the audio of this frame at frequency k place Z Transformation value, For along Z A segment of a spiral on a plane is used as sampling points to bisect the angle. For the frame window size, This is the audio signal after bandpass filtering. For sampling points, As the starting point of the spiral profile, This is the ratio of adjacent points on the spiral profile. For window functions, Centered at the window, and having: , , , In the above formula, The window number. For the movement step size, The starting sampling point for the spiral profile. The imaginary unit, The phase angle of the initial sampling point z0, The angle difference between adjacent sampling points. denoted as the elongation of the spiral.

3. The method for real-time tracking and tamper detection of power grid frequency based on audio according to claim 2, characterized in that, In step S102, all frequencies obtained within the window are... k l The frequency estimate sequence is obtained by splicing. F The function expression is: , In the above formula, For frequency estimate sequence F The Middle l One frequency estimate, For the first l The frequency corresponding to the maximum amplitude within each window The frequency corresponding to the maximum amplitude. The window number. For window length, The number of frames.

4. The method for real-time tracking and tamper detection of power grid frequency based on audio according to claim 1, characterized in that, The reference frequency sequence composed of the power grid frequency signal in step S104 r The reference frequency sequence was obtained by synchronously detecting power grid signals using power grid potential sensing equipment. r It is a signal sequence ordered by time from the power grid frequency signal, and its time covers the acquisition time of the audio signal containing the power grid frequency signal.

5. The method for real-time tracking and tamper detection of power grid frequency based on audio according to claim 1, characterized in that, In step S104, in the Pearson similarity sequence CC and Euclidean distance sequence D Finding the best matching point between them includes: S201, Search to determine Pearson similarity sequence CC The position corresponding to the maximum value ICC max Search to determine the Euclidean distance sequence D The position corresponding to the minimum value ID min ; S202, judgment ICC max ≈ ID min If the condition is met, the best matching point is found; otherwise, the best matching point is not found.

6. The method for real-time tracking and tamper detection of power grid frequency based on audio according to claim 1, characterized in that, The judgment ICC max ≈ ID min Whether it holds true includes: calculating the Pearson similarity sequence. CC The position corresponding to the maximum value ICC max Euclidean distance sequence D The position corresponding to the minimum value ID min The difference between them, if the difference is less than the set value, is then determined. ICC max ≈ ID min If true, otherwise judge ICC max ≈ ID min This is not true.

7. The method for real-time tracking and tamper detection of power grid frequency based on audio according to claim 1, characterized in that, In step S104, a mutation at a certain point means that the point in the reference frequency sequence... r The power grid frequency signal and the frequency estimation sequence after filtering. fz The difference between the frequency estimates and the set value exceeds the set value.

8. The method for real-time tracking and tamper detection of power grid frequency based on audio according to claim 1, characterized in that, The noise reduction in step S101 refers to filtering and noise reduction using a bandpass filter. Before step S101, the audio signal containing the power grid frequency signal is also downsampled.

9. A real-time power grid frequency tracking and tamper detection system based on audio, comprising a microprocessor and a memory interconnected, characterized in that, The microprocessor is programmed or configured to execute the audio-based real-time power grid frequency tracking and tamper detection method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program is used to be programmed or configured by a microprocessor to execute the audio-based real-time power grid frequency tracking and tamper detection method according to any one of claims 1 to 8.