Deep power grid frequency enhancement method, device and equipment for audio forensics

Through deep learning technology and the bidirectional condition least squares generation adversarial network framework, the generator structure and loss function are optimized, and the problem of difficulty in extracting frequency signals and insufficient quality of power grid is solved, efficient and real-time power grid frequency signal enhancement is achieved, and the application scope of audio forensics technology is expanded.

CN119152884BActive Publication Date: 2025-08-29WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410990708.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-08-29
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

In the prior art, the extraction of power grid frequency signals is difficult, insufficient quality, high computational complexity, limited application range, and insufficient standardization and interoperability, especially in complex multimedia environments, which are difficult to effectively enhance and apply.

Method used

Deep learning technology, especially the bidirectional condition least squares generation adversarial network framework, is adopted to optimize the generator structure and loss function, and process large-scale grid frequency data sets to enhance the grid frequency signal quality, including preprocessing, generative adversarial network training and target evidence for.

Benefits of technology

It significantly enhances the quality of the power grid frequency signal, simplifies the calculation process, realizes real-time processing, ensures signal integrity, and expands the application boundaries of audio forensic technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152884B_ABST
    Figure CN119152884B_ABST
Patent Text Reader

Abstract

The present application relates to the field of digital signal detection and processing technology, and in particular to a method, device and equipment for deep power grid frequency enhancement for audio forensics, wherein the method comprises: obtaining audio data required for forensic applications; preprocessing the audio data and estimating the power grid frequency signal of the preprocessed audio data; inputting the power grid frequency signal into a generative adversarial network, and the generative adversarial network outputting an enhanced power grid frequency signal, wherein the generative adversarial network comprises a generator and a discriminator, the generator and the discriminator have different numbers of input channels, the encoder part of the generator uses multiple one-dimensional convolutional layers with the same structure, and the encoder part and the decoder part of the generator are mirror images of each other; target forensics are performed based on the enhanced power grid frequency signal and a reference signal database. Thus, the problems of signal extraction difficulty, insufficient quality, high computational complexity, limited application scope, and insufficient standardization and interoperability in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of digital signal detection and processing, and in particular to a method, device and equipment for deep power grid frequency enhancement for audio forensics. Background Art

[0002] With the rapid advancement of digital technology, the recording, dissemination, and utilization of audio have become more convenient and efficient than ever before, greatly expanding the application areas of digital audio files and leading to a surge in their number. However, this convenience also comes with risks. The widespread use of audio and video editing tools has empowered users with powerful processing capabilities, making audio file modification easy. This has exacerbated the risks to audio information security and provided an opportunity for criminals to forge and disseminate false information. In key areas such as military intelligence analysis, case investigation, and judicial evidence collection, the authenticity and integrity of audio files are extremely demanding. Therefore, in-depth research on digital audio forensics technology is of vital importance to maintaining national information security and promoting social stability.

[0003] Grid frequency, a fundamental characteristic of power distribution networks, is determined by turbine speed and generally maintains a nominal value of 50Hz or 60Hz, exhibiting slight random fluctuations. This characteristic makes grid frequency signals ubiquitous in multimedia content such as audio, video, and images, making them a key passive indicator in digital multimedia forensics. Grid frequency forensics, with its unique fluctuation patterns and geographical consistency, demonstrates strong potential for geolocation tracing, timestamp verification, and tamper detection.

[0004] Currently, research on power grid frequency-based forensics primarily focuses on accurately estimating power grid frequency and exploring its subsequent application in forensic practice. However, insufficient research exists on whether the quality of power grid frequency signals extracted from complex multimedia environments is sufficient to support high-standard forensic requirements. Power grid frequency signal enhancement, particularly noise suppression techniques, has become a research hotspot in recent years. Existing enhancement methods primarily include RFA (Robust Filtering Algorithm), HRFA ​​(Harmonic Robust Filtering Algorithm), and robust media timestamps. These methods, all rooted in traditional signal processing techniques, effectively mitigate additive noise interference and improve power grid frequency signal quality in low signal-to-noise ratio environments. However, these methods generally suffer from high computational complexity and long processing times. Defragmentation algorithms, in particular, while improving signal quality by removing noise fragments, sacrifice signal integrity and tamper detection capabilities, limiting their widespread application in forensic applications and prone to producing misleading results when processing short-duration signals. Summary of the Invention

[0005] This application provides a deep power grid frequency enhancement method, device and equipment for audio forensics to solve the problems of difficult signal extraction, insufficient quality, high computational complexity, limited application scope, and lack of standardization and interoperability in the existing technology.

[0006] The first aspect of the present application provides a deep power grid frequency enhancement method for audio forensics, comprising the following steps: obtaining audio data required for forensic applications; preprocessing the audio data and estimating the power grid frequency signal of the preprocessed audio data; inputting the power grid frequency signal into a generative adversarial network, and the generative adversarial network outputs an enhanced power grid frequency signal, wherein the generative adversarial network includes a generator and a discriminator, the generator and the discriminator have different numbers of input channels, the encoder part of the generator uses multiple one-dimensional convolutional layers with the same structure, and the encoder part and decoder part of the generator are mirror images of each other; and performing target forensics based on the enhanced power grid frequency signal and a reference signal database.

[0007] Optionally, before inputting the grid frequency signal into the generative adversarial network, it also includes: obtaining a training data set, wherein the training data set includes a reference grid frequency signal and a noisy grid frequency signal extracted from actual recorded data; inputting the reference grid frequency signal and the noisy grid frequency signal into a discriminator, the discriminator learning a first data distribution of the reference grid frequency signal and the noisy grid frequency signal, performing binary classification based on the first data distribution to obtain real samples, and updating the parameters of the discriminator through back propagation; inputting the noisy grid frequency signal and the potential sample vector into a generator, the generator outputs a denoised grid frequency signal, inputting the noisy grid frequency signal and the denoised grid frequency signal into the discriminator, the discriminator learning a second data distribution of the noisy grid frequency signal and the denoised grid frequency signal, performing binary classification based on the second data distribution to obtain false samples, and updating the parameters of the discriminator through back propagation; freezing the parameters of the discriminator, calculating a training loss value based on the discriminator output and the real sample, and allowing the generator to back propagate and learn the data distribution of the reference grid frequency signal.

[0008] Optionally, the target forensics based on the enhanced grid frequency signal and the reference signal database includes: calculating the mean square error or correlation coefficient between the grid frequency signal and the reference signal in the reference signal database; and determining a reference signal that matches the grid frequency signal based on the minimum mean square error in the mean square error or the maximum correlation coefficient in the correlation coefficient.

[0009] Optionally, the calculation formula of the mean square error is:

[0010]

[0011] Wherein, N is the length, x(n) is the grid frequency signal extracted from the recording file to be tested, n is the nth frequency point of the signal, r(n) is the grid frequency reference signal within the search range, and k is the kth segment reference signal; the calculation formula of the correlation coefficient is:

[0012]

[0013] in, is the mean value of the grid frequency reference signal within the search time range, is the mean value of the power grid frequency extracted from the audio file to be tested.

[0014] Optionally, before performing target forensics based on the enhanced grid frequency signal and reference signal database, it also includes: synchronously collecting the grid frequency signal and corresponding time information; storing the grid frequency signal and corresponding time information in different storage locations; and generating the reference signal database based on the data stored in the different storage locations.

[0015] Optionally, estimating the grid frequency signal of the preprocessed audio data includes: performing short-time Fourier transform on the preprocessed audio data; and using a quadratic interpolation method to perform grid frequency signal estimation on the short-time Fourier transformed data to obtain the grid frequency signal.

[0016] The second aspect of the present application provides a deep power grid frequency enhancement device for audio forensics, including: an acquisition module for acquiring audio data required for forensic applications; a processing module for preprocessing the audio data and estimating the power grid frequency signal of the preprocessed audio data; an output module for inputting the power grid frequency signal into a generative adversarial network, and the generative adversarial network outputs an enhanced power grid frequency signal, wherein the generative adversarial network includes a generator and a discriminator, the generator and the discriminator have different numbers of input channels, the encoder part of the generator uses multiple one-dimensional convolutional layers with the same structure, and the encoder part and decoder part of the generator are mirror images of each other; an forensics module for performing target forensics based on the enhanced power grid frequency signal and a reference signal database.

[0017] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the deep power grid frequency enhancement method for audio forensics as described in the above embodiment.

[0018] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the deep power grid frequency enhancement method for audio forensics as described in the above embodiment.

[0019] The fifth aspect of the present application provides a computer program product, including a computer program or instructions, which, when executed, are used to implement the deep power grid frequency enhancement method for audio forensics as described in the above embodiment.

[0020] Therefore, this application has the following beneficial effects:

[0021] The embodiments of this application utilize deep learning technology, specifically a bidirectional conditional least squares generative adversarial network framework, to effectively process large-scale power grid frequency datasets by optimizing the generator structure and loss function. This significantly enhances the quality of power grid frequency signals, particularly in complex environments with low signal-to-noise ratios. This not only simplifies the computational process and enables real-time processing, but also ensures the integrity of power grid frequency signals, further expanding the application boundaries of audio forensics technology. This addresses technical issues in existing technologies, such as difficulty in signal extraction, insufficient quality, high computational complexity, limited application scope, and insufficient standardization and interoperability.

[0022] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0024] Figure 1 A flowchart of a deep power grid frequency enhancement method for audio forensics provided according to an embodiment of the present application;

[0025] Figure 2 A schematic diagram of a generator structure of a generative adversarial network for grid frequency enhancement provided according to an embodiment of the present application;

[0026] Figure 3 A training flow chart of a generative adversarial network for grid frequency enhancement according to an embodiment of the present application;

[0027] Figure 4 This is a schematic diagram of the results of grid frequency signal enhancement provided according to one embodiment of the present application;

[0028] Figure 5 This is a schematic diagram of timestamp matching results after the grid frequency signal is enhanced according to one embodiment of the present application;

[0029] Figure 6 A flowchart of a deep power grid frequency enhancement method for audio forensics provided according to one embodiment of the present application;

[0030] Figure 7 A schematic diagram of a deep power grid frequency enhancement system for audio forensics provided according to an embodiment of the present application;

[0031] Figure 8 This is an example diagram of a deep power grid frequency enhancement device for audio forensics according to an embodiment of the present application;

[0032] Figure 9 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0033] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0034] The following describes the deep power grid frequency enhancement method, device and equipment for audio forensics of the embodiments of the present application with reference to the accompanying drawings. In response to the problem of inability to accurately obtain evidence mentioned in the above background technology, the present application provides a deep power grid frequency enhancement method for audio forensics. In this method, deep learning technology, especially the bidirectional conditional least squares generative adversarial network framework, is used to effectively process large-scale power grid frequency data sets by optimizing the generator structure and loss function, significantly enhancing the quality of the power grid frequency signal, especially in complex environments with low signal-to-noise ratio. It not only simplifies the calculation process and realizes real-time processing, but also ensures the integrity of the power grid frequency signal, further expanding the application boundary of audio forensics technology. As a result, the problems of signal extraction difficulty, insufficient quality, high computational complexity, limited application scope, and lack of standardization and interoperability in the prior art are solved.

[0035] Specifically, Figure 1 A flowchart of a deep power grid frequency enhancement method for audio forensics provided in an embodiment of the present application.

[0036] like Figure 1 As shown in FIG, the deep power grid frequency enhancement method for audio forensics includes the following steps:

[0037] In step S101 , audio data required for a forensic application is acquired.

[0038] Among them, forensic application can be the process of obtaining and analyzing evidence using technical means in fields such as law, security or investigation, and audio data can be sound information recorded by recording equipment.

[0039] It can be understood that the embodiments of the present application not only provide key evidence for case investigation by obtaining the audio data required for forensic applications, but also serve as the basis for deep learning model processing, facilitating subsequent support for the estimation and enhancement of power grid frequency signals.

[0040] Specifically, the recorded audio file is read to obtain the read audio data and sampling frequency. In general, the audio sampling frequency is 8000Hz or 44100Hz. In order to reduce the computing resources required, the audio recording is downsampled. The downsampled audio data is S and the downsampled sampling frequency is f. s , f s Can be set to 400Hz.

[0041] In step S102 , the audio data is preprocessed, and a power grid frequency signal of the preprocessed audio data is estimated.

[0042] The grid frequency signal may be the power supply frequency of the power distribution network in the grid.

[0043] It can be understood that the embodiments of the present application can reduce noise and interference in audio data, remove some unnecessary redundant information, improve data quality, and facilitate more accurate estimation of power grid frequency signals by preprocessing audio data. Through preprocessing and power grid frequency signal estimation, the authenticity and integrity of the audio data can be further verified.

[0044] Specifically, since the power grid frequency is a narrowband signal, the out-of-band audio content and noise will affect the power grid frequency estimation, so a bandpass filter is selected to filter the down-sampled result, with the center frequency set to 100 Hz and the passband width to 4 Hz, to obtain the audio data S after bandpass filtering and preprocessing. down .

[0045] In an embodiment of the present application, estimating the grid frequency signal of the preprocessed audio data includes: performing a short-time Fourier transform on the preprocessed audio data; and using a quadratic interpolation method to perform a grid frequency signal estimation on the short-time Fourier transformed data to obtain a grid frequency signal.

[0046] Among them, the short-time Fourier transform can be a mathematical transformation related to the Fourier transform, which is used to determine the frequency and phase of the sine wave in the local area of ​​the time-varying signal; the quadratic interpolation method can be a polynomial approximation method established by utilizing the property that the function has a quadratic function near the extreme point.

[0047] It can be understood that the embodiment of the present application applies short-time Fourier transform to the preprocessed audio data to extract spectral features, and uses quadratic interpolation to accurately estimate the grid frequency signal, thereby improving the efficiency and accuracy of audio signal processing.

[0048] Specifically, the preprocessed data is subjected to a short-time Fourier transform to obtain a time-frequency domain sequence X(k,n) corresponding to the audio data, where k represents the kth frame of the short-time Fourier transform, n represents the nth frequency point of the short-time Fourier transform, and the total number of short-time Fourier transform frequency points is N;

[0049] Use the quadratic interpolation method to estimate the grid frequency signal and obtain the grid frequency signal; find the point m corresponding to the maximum value β of the fast Fourier transform within a window length, and then find the points m-1 and m+1 adjacent to m within the window length respectively, and calculate the value of their fast Fourier transform. The number of points selected is N FFT Assuming that the fast Fourier transform corresponding to point m-1 is a and the fast Fourier transform corresponding to point m+1 is λ, the relative position δ of the top of the quadratic model interpolation is:

[0050]

[0051] The best estimated position of the instantaneous frequency of the grid frequency signal obtained by the interpolation estimation method is m+δ, so the best estimated value of the instantaneous frequency of the grid frequency signal is:

[0052]

[0053] The frequencies within a certain fluctuation range are retained to obtain a preliminary estimate. Since the actual fluctuation range of the power grid frequency signal is very limited, in order to reduce the impact of outliers during estimation, the amplitudes of fluctuations exceeding the specified threshold are truncated and replaced with corresponding boundary values.

[0054] It should be noted that, considering the characteristics of frequency signal fluctuations in the central China power grid, the specified threshold can be 50 Hz ± 0.05 Hz.

[0055] In step S103, the grid frequency signal is input into a generative adversarial network, and the generative adversarial network outputs a strong grid frequency signal.

[0056] Among them, the generative adversarial network can include a generator and a discriminator. The generator and the discriminator have different numbers of input channels. The encoder part of the generator uses multiple one-dimensional convolutional layers with the same structure, and the encoder part and the decoder part of the generator are mirror images of each other.

[0057] It is understood that the embodiments of the present application enhance the grid frequency signal through a generative adversarial network. The generative adversarial network includes a generator and a discriminator. The generator uses an encoder and decoder structure composed of multiple one-dimensional convolutional layers to generate pseudo signals, while the discriminator is responsible for distinguishing between real signals and pseudo signals. Through adversarial training between the two, the generative adversarial network can output a clearer and more realistic grid frequency signal, thereby achieving purposes such as signal enhancement, noise suppression, data amplification, and anomaly detection.

[0058] Specifically, a bidirectional conditional least squares generative adversarial network is used, and an improved generator structure and loss function are adopted for training. When improving the generator, the architectural idea of ​​the autoencoder is borrowed to achieve an efficient mapping from the latent space to the data space. Figure 2 As shown in the figure, the encoder part of the generator cleverly uses four one-dimensional convolutional layers, each of which is configured with a convolution kernel of size 4 and a stride of 2. This not only effectively reduces the spatial dimension of the data, but also increases the number of channels of the feature map, thereby achieving effective feature extraction and compression.

[0059] For example, if m is the number of data samples, l is the length of a single sample, and m×l is used to represent the size of each convolutional layer, then the resulting sizes of all 8 layers are: 250×64, 125×128, 63×256, and 32×512. Therefore, in the encoder stage of the generator, the intermediate thought vector c guided by the latent vector z has a size of 32×512.

[0060] The decoder portion of the generator mirrors the encoder structure, gradually restoring the spatial dimensions of the data through a series of deconvolution operations until an output close to the original data is generated. This symmetrical encoder-decoder structure not only ensures efficient information transfer but also enhances the network's robustness to complex data transformations.

[0061] Unlike the generator, the discriminator is designed to process both the grid frequency reference signal and the noisy grid frequency signal. Therefore, it has two input channels to receive these two signals separately. Structurally, the discriminator uses a similar stack of convolutional layers as the generator encoder, but with appropriate adjustments based on the task requirements to ensure accurate distinction between real and generated samples, providing strong feedback to the generator.

[0062] In an embodiment of the present application, before the grid frequency signal is input into the generative adversarial network, it also includes: obtaining a training data set, wherein the training data set includes a reference grid frequency signal and a noisy grid frequency signal extracted from the actual recorded data; inputting the reference grid frequency signal and the noisy grid frequency signal into the discriminator, the discriminator learns the first data distribution of the reference grid frequency signal and the noisy grid frequency signal, performs binary classification based on the first data distribution to obtain real samples, and updates the parameters of the discriminator through back propagation; inputting the noisy grid frequency signal and the potential sample vector into the generator, the generator outputs the denoised grid frequency signal, inputting the noisy grid frequency signal and the denoised grid frequency signal into the discriminator, the discriminator learns the second data distribution of the noisy grid frequency signal and the denoised grid frequency signal, performs binary classification based on the second data distribution to obtain false samples, and updates the parameters of the discriminator through back propagation; freezing the parameters of the discriminator, calculating the training loss value based on the discriminator output and the real sample, and allowing the generator to back propagate and learn the data distribution of the reference grid frequency signal.

[0063] The reference grid frequency signal may be a clean grid frequency signal sample that is not interfered with by noise.

[0064] It can be understood that in this embodiment, by alternately training the generator and discriminator, the generator gradually learns to produce more realistic denoised signals through their mutual competition, while the discriminator continuously improves its ability to distinguish between real and false signals. By continuously optimizing the generator and discriminator, effective denoising of the power grid frequency signal is achieved.

[0065] For example, if Figure 3 As shown in Figure 1, the training process can be divided into two steps. In the first step, the training data consists of a reference grid frequency signal x and a noisy grid frequency signal y extracted from actual recordings. These two types of signals are fed into the discriminator for preliminary learning. The discriminator distinguishes between real samples and noisy samples using a binary classification mechanism and continuously optimizes its parameters using a backpropagation algorithm to enhance its discrimination ability.

[0066] The second step can be broken down into two phases to fine-tune the generator's denoising capabilities. In the first phase, the discriminator faces a new challenge: it again applies its learned knowledge of the data distribution to attempt to classify pairs of noisy and denoised signals as fake, and continues to update its parameters through backpropagation to counter the generator's ever-improving deceptive capabilities.

[0067] In the second phase, the discriminator's parameters are frozen to stabilize the training process and ensure the generator focuses on learning the essential characteristics of the reference signal. This allows the generator to independently optimize its parameters through backpropagation, producing a denoised version that more closely resembles the reference grid frequency signal. As the number of iterations increases, the generator gradually grasps the data distribution of the reference signal, significantly improving its output quality.

[0068] like Figure 4 The results are shown in Figure 2, which compares the signal before enhancement, the original reference signal, and the results of three existing advanced enhancement algorithms. To more clearly observe the results, the amplitudes of the different signals are shifted up and down to highlight the significant improvements in clarity and signal-to-noise ratio of the enhanced signal.

[0069] In step S104, target evidence is collected based on the enhanced grid frequency signal and the reference signal database.

[0070] The reference signal database may include known, standard or verified grid frequency signals.

[0071] It can be understood that the embodiments of the present application utilize the enhanced grid frequency signal and reference signal database to perform specific forensic analysis or verification work, such as determining the timestamp of the audio recording, verifying the authenticity of the audio, etc., thereby improving the accuracy, capability and application scope of the forensic process and effectively ensuring the authenticity of the evidence.

[0072] In an embodiment of the present application, target forensics is performed based on the enhanced grid frequency signal and the reference signal database, including: calculating the mean square error or correlation coefficient between the grid frequency signal and the reference signal in the reference signal database; determining the reference signal that matches the grid frequency signal based on the minimum mean square error in the mean square error or the maximum correlation coefficient in the correlation coefficient.

[0073] It can be understood that the embodiments of the present application improve the accuracy of signal recognition by calculating the mean square error or maximum correlation coefficient between the grid frequency signal and the reference signal database to determine the most matching reference signal, thereby supporting grid monitoring and analysis, helping to optimize grid operation, predict and prevent faults, and promote the development of smart grids.

[0074] In the embodiment of the present application, the calculation formula of the mean square error is:

[0075]

[0076] Where N is the length, x(n) is the grid frequency signal extracted from the recording file to be tested, n is the nth frequency point of the signal, r(n) is the grid frequency reference signal within the search range, and k is the kth segment reference signal;

[0077] The calculation formula for the correlation coefficient is as follows:

[0078]

[0079] Where, is the mean value of the grid frequency reference signal within the search time range, is the mean value of the grid frequency extracted from the audio file to be measured.

[0080] Specifically, data matching is performed between the enhanced grid frequency signal and the reference signal. It is required that the time-frequency sequence of the obtained reference signal is longer than the time-frequency sequence extracted from the audio file to be measured, so as to enable searching within the possible audio recording time range. Therefore, in order to determine the part of the reference signal corresponding to the extracted grid frequency signal, the minimum mean square error or the maximum correlation coefficient can be found between all possible reference signals and the extracted grid frequency signal. If the grid frequency reference signal within the search range is r(n) with a length of L, and the grid frequency signal extracted from the audio file to be measured is x(n) with a length of N, and N < L, the mean square error formula is:

[0081]

[0082] Where, r k (n) = r(n + k), k = 0, 1, 2…L - N.

[0083] The corresponding optimal matching position at this time is:

[0084]

[0085] The calculation formula for the correlation coefficient is as follows:

[0086]

[0087] Where, is the mean value of the grid frequency reference signal within the search time range, is the mean value of the grid frequency extracted from the audio file to be measured.

[0088] The corresponding optimal matching position at this time is:

[0089]

[0090] In the embodiment of the present application, the matching criterion of the maximum correlation coefficient is selected, and the timestamp matching result of a certain enhanced grid frequency signal is obtained. As Figure 5 shown, it shows the comparison of the grid frequency signal before and after enhancement, as well as the effects of three advanced enhancement algorithms. Each signal is shown by a dotted line representing the reference signal and the upper and lower offsets, clearly comparing the improvement of the signal quality, which provides strong support for evaluating the enhancement algorithm.

[0091] In an embodiment of the present application, before performing target evidence collection based on the strengthened grid frequency signal and reference signal database, it also includes: synchronously collecting the grid frequency signal and the corresponding time information; storing the grid frequency signal and the corresponding time information in different storage locations; and generating a reference signal database based on the data stored in different storage locations.

[0092] It can be understood that the embodiment of the present application ensures the accuracy and completeness of the data, optimizes data management, supports real-time monitoring and early warning, and improves the stability and reliability of the power grid by synchronously collecting the power grid frequency signal and its time information and storing them separately, thereby generating a reference signal database.

[0093] Specifically, a reference signal acquisition and storage system was constructed to process grid frequency signals in real time and archive the grid frequency reference signals. During grid frequency reference signal acquisition, a step-down transformer captures grid frequency fluctuations and transmits them to a computer. The computer's sound card and control program read the signals and save them to a buffer. To facilitate customizing short-time Fourier transform parameters and accuracy, the time-domain reference signal was directly stored.

[0094] The Python program based on the Windows operating system was carefully written to ensure that the system can run uninterruptedly 24 hours a day, and at the same time, the collected grid frequency signal values ​​and their corresponding time information are regularly stored in two different storage locations.

[0095] In addition, the embodiment of the present application arranges a power grid frequency collector for synchronous collection, which can reduce the impact of local power grid activity instability, obtain a more accurate and stable power grid frequency reference signal, effectively avoid interference with data collection caused by special circumstances such as local power outages and equipment failures, and ensure data continuity and reliability.

[0096] The grid frequency reference signal used in this embodiment is stored using the following parameters: a sampling rate of 400 Hz, a quantization precision of 16 bits, a mono channel, and a storage interval of one hour. All audio files are saved in pulse code modulation waveform format, which helps ensure the quality and accuracy of the audio data. The audio file naming convention is "year_month_day_week_hour_minute_second.wav," for example, 2024_06_26_Wed_07_00_00.wav, to facilitate intuitive organization and retrieval of audio files.

[0097] The deep grid frequency enhancement method for audio forensics proposed in the embodiments of this application utilizes deep learning technology, specifically the bidirectional conditional least squares generative adversarial network framework, to effectively process large-scale grid frequency datasets by optimizing the generator structure and loss function. This significantly enhances the quality of grid frequency signals, particularly in complex environments with low signal-to-noise ratios. This not only simplifies the computational process and enables real-time processing, but also ensures the integrity of grid frequency signals, further expanding the application boundaries of audio forensics technology. This solves the problems of existing technologies, such as difficulty in signal extraction, insufficient quality, high computational complexity, limited application scope, and insufficient standardization and interoperability.

[0098] The following will be combined Figure 6 The deep power grid frequency enhancement method for audio forensics is described in detail as follows:

[0099] Step 1: Downsample the recorded audio file and pass it through a bandpass filter to obtain preprocessed data;

[0100] Step 2: Perform short-time Fourier transform on the preprocessed data and use quadratic interpolation to estimate the grid frequency signal. Obtain the grid frequency signal and retain the frequency within a certain fluctuation range to obtain a preliminary estimate.

[0101] Step 3: Build and train a grid frequency-generative adversarial network for grid frequency enhancement, input the estimated grid frequency signal into the neural network, and obtain the enhanced grid frequency signal;

[0102] Step 4: Build a grid frequency reference signal database and use the enhanced grid frequency signal for forensic applications such as timestamp verification, geolocation estimation, and tamper detection.

[0103] In addition, if Figure 7 As shown, the deep power grid frequency enhancement system 10 for audio forensics provided in the embodiment of the present application includes: an audio recording module 100, a PC (Personal Computer) data processing module 200 and a power grid frequency reference signal database 300.

[0104] Among them, the audio recording module 100 is used to record the audio content to be collected for evidence, save it and transmit it to the PC data processing module; the PC processing module 200 is mainly used to save and process the audio files, and use the grid frequency generation adversarial network to enhance the extracted grid frequency signal for forensic application; the grid frequency reference signal database 300 mainly stores reference signals to provide reference values ​​for audio evidence collection.

[0105] In summary, the embodiments of the present application utilize deep learning technology to process large-scale grid frequency data sets, achieving efficient real-time enhancement of grid frequency signals of different qualities, especially improving the quality of low signal-to-noise ratio signals, while ensuring signal integrity and broadening the scope of audio analysis and application.

[0106] Next, a deep power grid frequency enhancement device for audio forensics proposed according to an embodiment of the present application will be described with reference to the accompanying drawings.

[0107] Figure 8 It is a block diagram of a deep power grid frequency enhancement device for audio forensics according to an embodiment of the present application.

[0108] like Figure 8 As shown, the deep power grid frequency enhancement device 20 for audio forensics includes: an acquisition module 210, a processing module 220, an output module 230 and an evidence collection module 240.

[0109] Among them, the acquisition module 210 is used to obtain the audio data required for the forensic application; the processing module 220 is used to preprocess the audio data and estimate the power grid frequency signal of the preprocessed audio data; the output module 230 is used to input the power grid frequency signal into the generative adversarial network, and the generative adversarial network outputs the enhanced power grid frequency signal, wherein the generative adversarial network includes a generator and a discriminator, the generator and the discriminator have different numbers of input channels, the encoder part of the generator uses multiple one-dimensional convolutional layers with the same structure, and the encoder part and the decoder part of the generator are mirror images of each other; the forensics module 240 is used to perform target forensics based on the enhanced power grid frequency signal and the reference signal database.

[0110] It should be noted that the aforementioned explanation of the embodiment of the deep power grid frequency enhancement method for audio forensics is also applicable to the deep power grid frequency enhancement device for audio forensics in this embodiment, and will not be repeated here.

[0111] The deep grid frequency enhancement device for audio forensics proposed in the embodiment of this application utilizes deep learning technology, particularly the bidirectional conditional least squares generative adversarial network framework, to effectively process large-scale grid frequency data sets by optimizing the generator structure and loss function. This significantly enhances the quality of grid frequency signals, particularly in complex environments with low signal-to-noise ratios. This not only simplifies the calculation process and enables real-time processing, but also ensures the integrity of grid frequency signals, further expanding the application boundaries of audio forensics technology. This solves the problems of existing technologies such as difficulty in signal extraction, insufficient quality, high computational complexity, limited application scope, and insufficient standardization and interoperability.

[0112] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0113] A memory 901 , a processor 902 , and a computer program stored in the memory 901 and executable on the processor 902 .

[0114] When the processor 902 executes the program, the deep power grid frequency enhancement method for audio forensics provided in the above embodiment is implemented.

[0115] Furthermore, the electronic device further includes:

[0116] The communication interface 903 is used for communication between the memory 901 and the processor 902 .

[0117] The memory 901 is used to store computer programs that can be run on the processor 902 .

[0118] The memory 901 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.

[0119] If the memory 901, the processor 902, and the communication interface 903 are implemented independently, the communication interface 903, the memory 901, and the processor 902 can be connected to each other via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0120] Optionally, in a specific implementation, if the memory 901, the processor 902 and the communication interface 903 are integrated on a chip, the memory 901, the processor 902 and the communication interface 903 can communicate with each other through an internal interface.

[0121] The processor 902 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0122] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned deep power grid frequency enhancement method for audio forensics.

[0123] The present application also provides a computer program product, including a computer program or instructions, which, when executed, is used to implement the deep power grid frequency enhancement method for audio forensics as described in the above embodiment.

[0124] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0125] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0126] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0127] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, the steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement the method: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.

[0128] A person skilled in the art may understand that all or part of the steps carried out in the method for implementing the above-mentioned embodiment may be completed by instructing the relevant hardware through a program, and the above-mentioned program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.

[0129] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A deep power grid frequency enhancement method for audio forensics, characterized by: The following steps are involved: Acquire audio data required for forensic applications; preprocessing the audio data and estimating a power grid frequency signal of the preprocessed audio data; Inputting the power grid frequency signal into a generative adversarial network, the generative adversarial network outputting an enhanced power grid frequency signal, wherein the generative adversarial network includes a generator and a discriminator, the generator and the discriminator having different numbers of input channels, the encoder portion of the generator using multiple one-dimensional convolutional layers with the same structure, and the encoder portion and decoder portion of the generator being mirror images of each other; Target forensics is performed based on the enhanced grid frequency signal and the reference signal database.

2. The deep power grid frequency enhancement method for audio forensics according to claim 1 is characterized in that: Before inputting the grid frequency signal into a generative adversarial network, the method further includes: Acquire a training data set, wherein the training data set includes a reference grid frequency signal and a noisy grid frequency signal extracted from actual recorded data; inputting the reference grid frequency signal and the noisy grid frequency signal into a discriminator, wherein the discriminator learns a first data distribution of the reference grid frequency signal and the noisy grid frequency signal, performs binary classification based on the first data distribution to obtain real samples, and updates the parameters of the discriminator through back propagation; Inputting the noisy grid frequency signal and the potential sample vector into a generator, the generator outputting a de-noised grid frequency signal, inputting the noisy grid frequency signal and the de-noised grid frequency signal into a discriminator, the discriminator learning a second data distribution of the noisy grid frequency signal and the de-noised grid frequency signal, performing binary classification based on the second data distribution to obtain false samples, and updating the parameters of the discriminator through back propagation; The parameters of the discriminator are frozen, a training loss value is calculated based on the discriminator output and the real sample, and the generator is allowed to back-propagate and learn the data distribution of the reference grid frequency signal.

3. The deep power grid frequency enhancement method for audio forensics according to claim 1 is characterized in that: The target forensics based on the enhanced grid frequency signal and the reference signal database includes: Calculating a mean square error or a correlation coefficient between the grid frequency signal and a reference signal in a reference signal database; A reference signal matching the grid frequency signal is determined according to a minimum mean square error among the mean square errors or a maximum correlation coefficient among the correlation coefficients.

4. The deep power grid frequency enhancement method for audio forensics according to claim 3 is characterized in that: The calculation formula of the mean square error is: Where N is the length, x(n) is the grid frequency signal extracted from the recording file to be tested, n is the nth frequency point of the signal, r(n) is the grid frequency reference signal within the search range, and k is the kth segment reference signal; The calculation formula of the correlation coefficient is: in, is the mean value of the grid frequency reference signal within the search time range, is the mean value of the power grid frequency extracted from the audio file to be tested.

5. The deep power grid frequency enhancement method for audio forensics according to claim 1 or 3, characterized in that: Before performing target forensics based on the enhanced grid frequency signal and reference signal database, the method further includes: Synchronously collect grid frequency signals and corresponding time information; Storing the grid frequency signal and the corresponding time information in different storage locations; The reference signal database is generated according to the data stored in the different storage locations.

6. The deep power grid frequency enhancement method for audio forensics according to claim 1 is characterized in that: The estimating the grid frequency signal of the preprocessed audio data includes: Perform short-time Fourier transform on the preprocessed audio data; The grid frequency signal is estimated by using the quadratic interpolation method on the data after short-time Fourier transform to obtain the grid frequency signal.

7. A deep power grid frequency enhancement device for audio forensics, characterized in that: include: An acquisition module, used to obtain audio data required for forensic applications; a processing module, configured to pre-process the audio data and estimate a power grid frequency signal of the pre-processed audio data; an output module, configured to input the grid frequency signal into a generative adversarial network, wherein the generative adversarial network outputs an enhanced grid frequency signal, wherein the generative adversarial network includes a generator and a discriminator, the generator and the discriminator have different numbers of input channels, the encoder portion of the generator uses multiple one-dimensional convolutional layers with the same structure, and the encoder portion and decoder portion of the generator are mirror images of each other; A forensics module is used to perform target forensics based on the enhanced grid frequency signal and the reference signal database.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the deep power grid frequency enhancement method for audio forensics according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed, the deep power grid frequency enhancement method for audio forensics according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instructions are executed, the deep power grid frequency enhancement method for audio forensics according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Audio identification method based on analysis on ENF phase spectrum and instantaneous frequency spectrum

    CN108806718A

  • Automatic detection method and system for digital audio deletion and insertion tampering operation

    CN114048770A