Grid frequency reference signal network construction and audio positioning method and device
By building a power grid frequency reference signal network through online real-time multimedia data, the high cost and low accuracy problems of traditional audio positioning methods are solved, and efficient and extensive audio positioning is achieved.
Patent Information
- Application Number
- CN202411066311.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-08-05
AI Technical Summary
Traditional audio positioning methods require the deployment of a large number of frequency interference recorders, which are expensive and complex in maintenance, have poor positioning accuracy and low accuracy, and are not widely used enough to meet the actual application needs.
The power grid frequency reference signal network is built using online real-time multimedia data, and the power grid frequency reference signal network is built through extraction, downsampling, grid frequency signal enhancement and multi-harmonic processing, and the correlation coefficient is calculated for audio positioning.
It effectively reduces construction and maintenance costs, improves positioning accuracy and accuracy, expands the application scope, and facilitates practical application.
Smart Images

Figure CN119170048B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing technology, and in particular to a method and device for constructing a power grid frequency reference signal network and audio positioning. Background Art
[0002] With the rapid development of science and technology, audio storage and dissemination technologies are constantly improving. As a result, the amount of audio data in our daily lives is experiencing exponential growth. However, when audio is recorded, the original recording location is often not recorded. In applications that rely strictly on recording location information, such as security surveillance, accurately determining the recording location of audio is particularly critical.
[0003] Traditional audio positioning methods that use power grid frequency for audio positioning rely on building a power grid frequency reference signal network, which typically requires deploying a large number of frequency interference recorders at different locations. With the continuous advancement of network transmission technology, a large amount of multimedia content can now be transmitted in real time over the network, providing a new way to acquire data. This real-time multimedia data not only contains rich geographic location information but also naturally embeds power grid frequency signals, providing a new type of power grid frequency reference signal network data source for audio positioning.
[0004] However, traditional audio positioning methods in related technologies require the deployment of a large number of frequency interference recorders, which are costly and complex to maintain. They also face harsh usage conditions, poor positioning accuracy, and low positioning accuracy. Their application scope is not wide enough and cannot meet the needs of audio positioning in practical applications, which needs to be solved urgently. Summary of the Invention
[0005] The present application provides a method and device for constructing a power grid frequency reference signal network and audio positioning, so as to solve the problems in the traditional audio positioning method in the related art, such as the need to deploy a large number of frequency interference recorders, high cost and complex maintenance, and also facing the problems of harsh usage conditions, poor positioning precision, low positioning accuracy, and a limited application scope, which cannot meet the needs of audio positioning in actual applications.
[0006] The first aspect of the present application provides a method for constructing a grid frequency reference signal network and positioning audio, comprising the following steps: extracting original audio from online real-time multimedia data in multiple regions, and downsampling the original audio to obtain downsampled audio; detecting whether the audio has a grid frequency, and processing the audio to obtain a grid frequency signal of a single audio when it is detected that the audio has the grid frequency; processing the grid frequency signals of multiple audios in the same region to obtain a single-region grid frequency reference signal, and processing the single-region grid frequency reference signals in multiple regions to obtain a grid frequency reference signal network; processing the audio to be located to obtain a grid frequency signal of the audio to be located, and calculating the correlation coefficient between the grid frequency signal of the audio to be located and each node of the grid frequency reference signal network to obtain a final positioning result of the audio based on the correlation coefficient.
[0007] Optionally, in one embodiment of the present application, detecting whether the audio has a grid frequency includes: obtaining the signal length of the grid frequency signal of the audio; calculating the noise segment rate based on the signal length, so as to determine whether the audio has a grid frequency based on the noise segment rate.
[0008] Optionally, in one embodiment of the present application, the processing of the audio to obtain a single audio grid frequency signal includes: enhancing the grid frequency signal of the audio to obtain an enhanced audio signal; extracting a multi-harmonic grid frequency signal of the enhanced audio signal; and selecting and fusing the multi-harmonic grid frequency signal to obtain the single audio grid frequency signal.
[0009] Optionally, in one embodiment of the present application, the grid frequency signals of multiple audios in the same area are processed to obtain a single-area grid frequency reference signal, including: constructing a target spatial domain selection matrix based on the signal length of the grid frequency signal; converting the target spatial domain selection matrix into a target spatial domain selection undirected graph, and fusing the spatial domain signals of the target spatial domain selection undirected graph to obtain the single-area grid frequency reference signal.
[0010] Optionally, in one embodiment of the present application, the processing of the single-region power grid frequency reference signal of multiple regions to obtain a power grid frequency reference signal network includes: dividing the total area to be located in the multiple regions into multiple grids, and obtaining the grid point coordinates of the multiple grids; collecting the center point coordinates of the single-region power grid frequency reference signal of the multiple regions; calculating the distance between the grid point coordinates of the multiple grids and the center point coordinates, to obtain the power grid frequency reference signal network based on the distance.
[0011] The second aspect of the present application provides a power grid frequency reference signal network construction and audio positioning device, including: an extraction module for extracting original audio from online real-time multimedia data in multiple regions, and downsampling the original audio to obtain the downsampled audio; a first processing module for detecting whether the audio has a power grid frequency, and when detecting that the audio has the power grid frequency, processing the audio to obtain a power grid frequency signal of a single audio; a second processing module for processing the power grid frequency signals of multiple audios in the same region to obtain a single-region power grid frequency reference signal, and processing the single-region power grid frequency reference signals of multiple regions to obtain a power grid frequency reference signal network; a positioning module for processing the audio to be positioned to obtain the power grid frequency signal of the audio to be positioned, and calculating the correlation coefficient between the power grid frequency signal of the audio to be positioned and each network point of the power grid frequency reference signal network to obtain the final positioning result of the audio based on the correlation coefficient.
[0012] Optionally, in one embodiment of the present application, the first processing module includes: a first acquisition unit, used to obtain the signal length of the power grid frequency signal of the audio; and a determination unit, used to calculate the noise segment rate based on the signal length, so as to determine whether the audio has a power grid frequency based on the noise segment rate.
[0013] Optionally, in one embodiment of the present application, the first processing module includes: an enhancement unit for enhancing the grid frequency signal of the audio to obtain an enhanced audio signal; an extraction unit for extracting the multi-harmonic grid frequency signal of the enhanced audio signal; and a first processing unit for selecting and fusing the multi-harmonic grid frequency signal to obtain the single audio grid frequency signal.
[0014] Optionally, in one embodiment of the present application, the second processing module includes: a construction unit for constructing a target spatial domain selection matrix based on the signal length of the grid frequency signal; a second processing unit for converting the target spatial domain selection matrix into a target spatial domain selection undirected graph, and fusing the spatial domain signals of the target spatial domain selection undirected graph to obtain the single regional grid frequency reference signal.
[0015] Optionally, in one embodiment of the present application, the second processing module includes: a second acquisition unit, used to divide the total area to be located in the multiple regions into multiple grids, and obtain the grid point coordinates of the multiple grids; a collection unit, used to collect the center point coordinates of the single-region power grid frequency reference signal of the multiple regions; a calculation unit, used to calculate the distance between the grid point coordinates of the multiple grids and the center point coordinates, so as to obtain the power grid frequency reference signal network based on the distance.
[0016] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the grid frequency reference signal network construction and audio positioning method as described in the above embodiment.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned grid frequency reference signal network construction and audio positioning method.
[0018] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned grid frequency reference signal network construction and audio positioning method.
[0019] The embodiment of the present application can use online real-time multimedia data to construct a power grid frequency reference signal network, and use the reference network to locate the audio data, thereby realizing the use of online real-time multimedia data as the source of the power grid frequency reference signal, avoiding the construction cost and maintenance cost of the power grid frequency reference network, while taking advantage of the large amount of online multimedia data, which can be collected at multiple points, effectively ensuring the geographical resolution of power grid frequency signal collection, and improving the accuracy of the power grid frequency reference signal network. Furthermore, by using the power grid frequency reference network to locate audio, the efficiency of audio positioning in this application is effectively improved, the scope of positioning application is expanded, and it is convenient for actual implementation. Thus, it solves the problem that the traditional audio positioning method in the related art needs to deploy a large number of frequency interference recorders, which is costly and complex to maintain, and also faces the problems of harsh use conditions, poor positioning accuracy, low positioning accuracy, insufficient application scope, and inability to meet the needs of audio positioning in actual applications.
[0020] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0022] Figure 1 A flowchart of a method for constructing a power grid frequency reference signal network and audio positioning according to an embodiment of the present application;
[0023] Figure 2 This is a schematic diagram of the result of determining the existence of an online audio power grid frequency signal according to one embodiment of the present application;
[0024] Figure 3 This is a schematic diagram of online audio grid frequency signal enhancement and multi-harmonic extraction results according to one embodiment of the present application;
[0025] Figure 4 This is a schematic diagram of online audio grid frequency harmonic selection and fusion results according to one embodiment of the present application;
[0026] Figure 5 This is a schematic diagram of grid frequency spatial domain selection and fusion results according to one embodiment of the present application;
[0027] Figure 6 This is a flowchart of a method for constructing a grid frequency reference signal network and audio positioning according to an embodiment of the present application;
[0028] Figure 7 A schematic diagram of the structure of a grid frequency reference signal network construction and an audio positioning device according to an embodiment of the present application;
[0029] Figure 8 Schematic diagram of the structure of an electronic device according to an embodiment of the present application.
[0030] Reference numerals:
[0031] 10-Grid frequency reference signal network construction and audio positioning device: 100-extraction module, 200-first processing module, 300-second processing module and 400-positioning module; 801-memory, 802-processor and 803-communication interface. DETAILED DESCRIPTION
[0032] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0033] The following describes a method and apparatus for constructing a grid frequency reference signal network and audio positioning according to an embodiment of the present application with reference to the accompanying drawings. The conventional audio positioning methods in the related art mentioned in the background art require the deployment of a large number of frequency interference recorders, which are costly and complex to maintain. Furthermore, they face the problems of harsh operating conditions, poor positioning accuracy, and low positioning accuracy. Their application scope is limited and they cannot meet the needs of audio positioning in practical applications. The present application provides a method for constructing a grid frequency reference signal network and audio positioning. In this method, a grid frequency reference signal network can be constructed using online real-time multimedia data, and the reference network can be used to locate audio data. This method utilizes online real-time multimedia data as a source of grid frequency reference signals, avoiding the construction and maintenance costs of the grid frequency reference network. While utilizing the large amount of online multimedia data and the ability to collect data at multiple points, it effectively ensures the geographical resolution of grid frequency signal collection and improves the accuracy of the grid frequency reference signal network. Furthermore, by using the grid frequency reference network to locate audio, the efficiency of audio positioning in the present application is effectively improved, the scope of positioning applications is expanded, and practical applications are facilitated. This solves the problem that traditional audio positioning methods in related technologies require the deployment of a large number of frequency interference recorders, which are costly and complex to maintain. At the same time, they also face harsh usage conditions, poor positioning accuracy, low positioning accuracy, a limited application scope, and cannot meet the needs of audio positioning in actual applications.
[0034] Specifically, Figure 1 This is a flowchart of a grid frequency reference signal network construction and audio positioning method provided in an embodiment of the present application.
[0035] like Figure 1 As shown, the grid frequency reference signal network construction and audio positioning method includes the following steps:
[0036] In step S101, original audio is extracted from multi-region online real-time multimedia data, and down-sampled to obtain down-sampled audio.
[0037] It is understood that the present application can implement the construction of a power grid frequency reference signal network and audio localization based on multi-region online real-time multimedia data. In some embodiments, before implementing this task, it is first necessary to extract the original audio from the multi-region online real-time multimedia data. Furthermore, the original audio can be downsampled to facilitate the use of the downsampled audio to implement the construction of the power grid frequency reference signal network and audio localization in the subsequent process.
[0038] For example, the embodiment of the present application may define the relevant total area to be located as , you can choose different regions As an online real-time multimedia data collection area, , Indicates the total number of regions where online real-time multimedia data is collected.
[0039] By searching the online real-time multimedia data platform, you can find areas with online real-time multimedia data collection The identified data source is downloaded and saved in real time to obtain several online real-time original audio data for each region ,in, Indicates from the region The collected Audio, , Indicates the total amount of audio collected in this area. Then, each original audio is downsampled to obtain the available downsampled audio data. .
[0040] The embodiment of the present application can extract original audio from online real-time multimedia data in multiple regions and downsample the original audio. By using the audio data after downsampling, the storage space and transmission bandwidth requirements of the audio data can be reduced, and the speed of audio processing may be accelerated.
[0041] Step S102 : detecting whether the audio has a grid frequency, and if it is detected that the audio has a grid frequency, processing the audio to obtain a grid frequency signal of a single audio.
[0042] It is understandable that the nominal value of the Electric Network Frequency (ENF), that is, the transmission frequency of alternating current, is usually determined by the region. Due to the influence of various factors such as power generation adjustment and changes in consumer demand, the actual grid frequency is not constant, but needs to be monitored and adjusted in real time by the power control departments in each region to maintain stability. This frequency adjustment change is not completed instantaneously for the same power grid, but propagates within the power grid at a certain speed. Therefore, for the same power grid, the grid frequencies in different locations are slightly different. Based on this, the embodiment of the present application can first process the audio to obtain a single audio grid frequency signal.
[0043] In some embodiments, after extracting the original audio from the online real-time multimedia data and downsampling the original audio to obtain usable audio, the collected audio can be detected, that is, whether the power grid frequency is present in the audio, so that audio carrying the power grid frequency signal can be obtained.
[0044] Furthermore, the embodiment of the present application can also perform grid frequency signal enhancement, multi-harmonic extraction, selection and fusion on the audio carrying the grid frequency signal to obtain a single audio grid frequency signal.
[0045] Next, how to detect whether the audio contains the grid frequency and how to process the audio to obtain a single audio grid frequency signal will be further explained.
[0046] Optionally, in one embodiment of the present application, detecting whether the audio has a grid frequency includes: obtaining the signal length of the grid frequency signal of the audio; and calculating the noise segment rate based on the signal length to determine whether the audio has a grid frequency based on the noise segment rate.
[0047] In some embodiments, when detecting whether there is a grid frequency in the audio, the present application can make a judgment based on the signal length of the grid frequency signal in the audio, such as calculating the noise fragment rate of the grid frequency signal in the audio by the signal length of the grid frequency signal in the audio. If the noise fragment rate meets certain conditions, it can be determined that there is a grid frequency in the audio.
[0048] For example, each downsampled audio data can be Perform short-time Fourier transform to obtain the time-frequency domain sequence corresponding to the audio data ,in express The frequency domain sequence, Represents the first step of the Short-Time Fourier Transform (STFT) frame, represents the STFT transform frequency points, the total number of STFT frequency points can be set to .
[0049] Then, we use quadratic interpolation to obtain ENF multiple harmonic estimates ,in, Represents audio data of The ENF prediction value estimated by the subharmonics is .
[0050] Then, yes Perform grid frequency existence judgment. Since the grid frequency fluctuation range and fluctuation change rate are usually kept in a fixed range, we can set the following grid frequency signal judgment criteria for each harmonic, which can be expressed as follows:
[0051]
[0052]
[0053] in, 、 are the upper and lower limits of the grid frequency change, Indicates the maximum rate of change of the grid frequency. For example, if the ENF value at a certain point does not meet the grid frequency signal criterion, then the point and the left and right sides of the point Each point will be judged as a noise signal.
[0054] Therefore, based on the above criteria, if a certain grid frequency signal , the signal length is , the length of the signal judged as noise is , then the noise fragment rate It can be defined as:
[0055]
[0056] Generally speaking, for audio data with ENF , there is at least one harmonic whose noise fragment rate satisfies the condition and can be expressed as follows:
[0057]
[0058] in, is the noise control factor. In the embodiment of the present application, the noise control factor can be 50%. According to this judgment method, There is audio frequency of grid frequency signal in , Indicates from the region The collected ENF signal Audio, , Indicates the total number of audios containing ENF signals collected in the area.
[0059] Optionally, in one embodiment of the present application, the audio is processed to obtain a single audio grid frequency signal, including: enhancing the grid frequency signal of the audio to obtain an enhanced audio signal; extracting a multi-harmonic grid frequency signal of the enhanced audio signal; and selecting and fusing the multi-harmonic grid frequency signal to obtain a single audio grid frequency signal.
[0060] In actual implementation, to facilitate accurate audio positioning, embodiments of the present application may process the audio to obtain a single audio grid frequency signal. The audio processing may include, but is not limited to, performing grid frequency signal enhancement, multi-harmonic extraction, selection, and fusion processing on the audio carrying the grid frequency signal to obtain a single audio grid frequency signal.
[0061] For example, the grid frequency signal of the audio is first enhanced to obtain the enhanced audio signal, and then the multi-harmonic grid frequency signal of the enhanced audio signal is extracted. Then, the multi-harmonic grid frequency signal is selected and fused to obtain a single audio grid frequency signal.
[0062] For example, for each audio signal with an ENF (Electric Network Frequency) signal For example, we can first use, but are not limited to, the Harmonic Robust Filtering Algorithm (HRFA) to enhance the ENF signal. After obtaining the enhanced audio signal, we use the STFT and quadratic interpolation method to extract the multi-harmonic ENF signal and obtain the extracted multi-harmonic grid frequency signal. ,in, Represents audio data of The ENF prediction signal estimated by the subharmonic has a signal length of .
[0063] Then, the extracted To select harmonics, first construct a The harmonic selection matrix can be expressed as follows:
[0064]
[0065] in, For audio data The harmonic selection matrix, Represents the matrix The element at position, Indicates calculating the Pearson correlation coefficient between the two. 、 Represents audio data of Subharmonics and ENF prediction signal estimated from sub-harmonics. is the frequency domain correlation coefficient control factor, which can be calculated by the following formula:
[0066]
[0067] in, It means taking the minimum value between the two. is the frequency domain minimum correlation coefficient control parameter, which can be 0.88 in the embodiment of the present application; is the frequency domain coefficient magnification factor. In the embodiment of the present application, the frequency domain coefficient magnification factor can be 4; is the pseudo-random signal control factor, which can be calculated using the following formula:
[0068]
[0069] in, Indicates that the generated length is Gaussian white noise sequence, Indicates calculating the Pearson correlation coefficient between the two. Indicates that the operation is repeated Second-rate, is the number of Gaussian white noise frequency domain generation times. In the embodiment of the present application, the number of Gaussian white noise frequency domain generation times can be taken as , Indicates taking the brackets The maximum number of.
[0070] Then, the harmonic selection matrix Convert to harmonic selection undirected graph , select the available harmonic number through the maximum weight group algorithm in graph theory Finally, the maximum weighted likelihood estimation method is used for fusion to obtain a single audio grid frequency signal. ,in, Represents audio data The final grid frequency estimation result.
[0071] Figure 3 This is a graph showing the results of online audio grid frequency signal enhancement and multi-harmonic extraction according to an embodiment of the present application. Figure 4 for Figure 3 The harmonic selection and fusion results of the multi-harmonic grid frequency signal are shown in the figure, where the harmonic selection result is The local reference signal is the power grid frequency signal directly recorded by a frequency interference recorder in the area. The results show that the power grid frequency signal extracted from the online audio is The correlation coefficient with the local reference signal is 99.089%.
[0072] Step S103 , processing the grid frequency signals of multiple audio frequencies in the same region to obtain a single-region grid frequency reference signal, and processing the single-region grid frequency reference signals of multiple regions to obtain a grid frequency reference signal network.
[0073] In other embodiments, after obtaining a single-frequency grid frequency signal, the present application may process grid frequency signals of multiple frequencies in the same region based on the single-frequency grid frequency signal to obtain a single-region grid frequency reference signal.
[0074] Furthermore, the embodiment of the present application may also use a single-region power grid frequency reference signal from multiple regions as an anchor point, process it, and thereby obtain a power grid frequency reference signal network.
[0075] Next, we will further explain how to process the grid frequency signals of multiple audio frequencies in the same area to obtain a single-area grid frequency reference signal and how to use the single-area grid frequency reference signal of multiple regions as an anchor point to obtain a grid frequency reference signal network.
[0076] Optionally, in one embodiment of the present application, the grid frequency signals of multiple audio frequencies in the same region are processed to obtain a single-region grid frequency reference signal, including: constructing a target spatial domain selection matrix based on the signal length of the grid frequency signal; converting the target spatial domain selection matrix into a target spatial domain selection undirected graph, and fusing the spatial domain signals of the target spatial domain selection undirected graph to obtain a single-region grid frequency reference signal.
[0077] During actual implementation, when processing power grid frequency signals of multiple audio frequencies in the same region, the present application may, but is not limited to, perform spatial selection and fusion processing on the power grid frequency signals of multiple audio frequencies in the same region, thereby obtaining a single regional power grid frequency reference signal. Specifically, a target spatial selection matrix can be first constructed based on the signal length of the power grid frequency signal, and then the target spatial selection matrix can be converted into a target spatial selection undirected graph, and the spatial signals of the target spatial selection undirected graph can be fused to obtain a single regional power grid frequency reference signal.
[0078] For example, for the same region Extracted grid frequency signal , the length of each signal is First, we can construct a Order space selection matrix:
[0079]
[0080] in, For the region The spatial domain selection matrix, Represents the matrix The element at position, Indicates calculating the Pearson correlation coefficient between the two. 、 Indicates area Mid-range audio data and The ENF estimation results are: is the spatial correlation coefficient control factor, which can be calculated by the following formula:
[0081]
[0082] in, It means taking the minimum value between the two. is the spatial minimum correlation coefficient control parameter, which can be 0.9 in the embodiment of the present application; is the spatial coefficient magnification factor. In the embodiment of the present application, the spatial coefficient magnification factor can be 4; is the spatial pseudo-random signal control factor, which can be calculated using the following formula:
[0083]
[0084] in, Indicates that the generated length is Gaussian white noise sequence, Indicates calculating the Pearson correlation coefficient between the two. Indicates that the operation is repeated Second-rate, is the number of Gaussian white noise spatial domain generation times. In the embodiment of the present application, the number of Gaussian white noise spatial domain generation times can be taken as , Indicates taking the brackets The maximum number of.
[0085] Then, the spatial domain selection matrix Converted into a spatial selection undirected graph , select the available spatial signal set through the maximum weight group algorithm in graph theory Finally, the maximum weighted likelihood estimation method is used for fusion to obtain the frequency signals of the power grids in various places. ,in Indicates area The final grid frequency reference signal estimation result.
[0086] Figure 5 This is a diagram showing the grid frequency space selection and fusion results of an embodiment of the present application. Figure 5 As shown in the figure, the local reference signal is the power grid frequency signal directly recorded by a frequency interference recorder in the area. The results show that the correlation coefficient between the ENF signal after spatial domain fusion and the local reference signal is 99.473%.
[0087] Optionally, in one embodiment of the present application, the single-region power grid frequency reference signals of multiple regions are processed to obtain a power grid frequency reference signal network, including: dividing the total area to be located in the multiple regions into multiple grids, and obtaining the grid point coordinates of the multiple grids; collecting the center point coordinates of the single-region power grid frequency reference signals of the multiple regions; calculating the distance between the grid point coordinates of the multiple grids and the center point coordinates to obtain the power grid frequency reference signal network based on the distance.
[0088] In certain embodiments, when processing a single-region grid frequency reference signal from multiple regions to obtain a grid frequency reference signal network, it is possible but not limited to using the multi-region grid frequency reference signal as an anchor point and using the inverse distance weighted method for geographic interpolation to obtain the grid frequency reference signal network.
[0089] Specifically, the total area to be located in multiple regions can be first divided into multiple grids, and the grid point coordinates of these grids can be obtained; then the center point coordinates of the power grid frequency reference signal of a single region in multiple regions can be collected; finally, the power grid frequency reference signal network can be obtained by calculating the distance between the grid point coordinates of multiple grids and the center point coordinates.
[0090] For example, for the total area to be located For example, a size of The interpolation grid, the coordinates of each grid point can be expressed as:
[0091]
[0092] At the same time, the grid frequency signal area is collected The center point coordinates can be expressed as:
[0093]
[0094] Then, each point of the interpolation network can be interpolated using, but not limited to, the following formula, which is expressed as follows:
[0095]
[0096] in, Indicates that the grid point coordinates are The grid frequency signal at is the inverse distance weight coefficient, which can be calculated using the following formula:
[0097]
[0098] in, is the inverse distance weighting coefficient. In the embodiment of the present application, the inverse distance weighting coefficient can be 2; For each area where the grid frequency signal is collected The distance of the center point coordinates can be calculated using the following formula:
[0099]
[0100] Perform inverse distance weighted interpolation on each grid point in the interpolation grid to obtain the grid frequency reference signal network .
[0101] It should be noted that the parameter values in the embodiments of this application are for illustrative purposes only. Professionals and technicians in this field can make adjustments and changes according to specific circumstances in actual applications, and this application does not impose any specific restrictions.
[0102] Step S104: Process the audio to be located to obtain a grid frequency signal of the audio to be located, and calculate the correlation coefficient between the grid frequency signal of the audio to be located and each grid point of the grid frequency reference signal network to obtain a final positioning result of the audio according to the correlation coefficient.
[0103] As a possible implementation method, after obtaining the grid frequency reference signal network, the embodiment of the present application can perform grid frequency signal enhancement, multi-harmonic extraction and fusion on the audio to be located to obtain the grid frequency signal of the audio to be located, and then calculate the correlation coefficient between the grid frequency signal of the audio to be located and each point of the grid frequency reference signal network respectively. Finally, the final positioning result of the audio can be obtained according to the size of the correlation coefficient.
[0104] Specifically, for an audio signal to be located , you can first downsample it to obtain the downsampled audio signal Then, The HRFA algorithm is used to enhance the ENF signal. After obtaining the enhanced audio signal, the multi-harmonic ENF signal is extracted using the STFT and quadratic interpolation method to obtain the extracted multi-harmonic grid frequency signal. ,in Represents audio data of The ENF prediction signal estimated by the sub-harmonic is then used to obtain the audio data using the multi-harmonic selection and fusion algorithm. Final grid frequency signal .
[0105] To obtain the audio ENF signal to be located , can be calculated using but not limited to the following formula with the grid frequency reference signal network The correlation coefficient between each point:
[0106]
[0107] in, The grid coordinates are The grid frequency reference signal at Calculate the Pearson correlation coefficient between the two.
[0108] After obtaining the correlation coefficient, the final positioning result can be obtained according to the following criterion:
[0109]
[0110] in, The area containing all the points that meet this determination condition is the final audio positioning result.
[0111] In this example, the positioning correlation coefficient determination factor If 97.3% is taken, the positioning result area contains the audio recording location to be located.
[0112] The present application is described in detail below using a specific embodiment.
[0113] Figure 6 This is a flow chart of a method for constructing a grid frequency reference signal network and audio positioning according to an embodiment of the present application. Figure 6 As shown:
[0114] Step 1: Collect online real-time multimedia data from multiple regions, extract its audio, and downsample the audio;
[0115] Step 2: Determine the presence of grid frequency on the collected audio to obtain audio with grid frequency signal;
[0116] Step 3: Perform grid frequency signal enhancement, multi-harmonic extraction, selection, and fusion on the audio carrying the grid frequency signal to obtain a single audio grid frequency signal;
[0117] Step 4: Perform spatial domain selection and fusion on the grid frequency signals of multiple audio frequencies in the same region to obtain a single-region grid frequency reference signal;
[0118] Step 5: Using the multi-regional grid frequency reference signals as anchor points, perform geographic interpolation using the inverse distance weighted method to obtain the grid frequency reference signal network;
[0119] Step 6: Perform grid frequency signal enhancement, multi-harmonic extraction and fusion on the audio to be located to obtain the grid frequency signal of the audio to be located;
[0120] Step 7: Calculate the correlation coefficient between the audio grid frequency signal to be located and each point of the grid frequency reference signal network, and obtain the final positioning result based on the size of the correlation coefficient.
[0121] According to the grid frequency reference signal network construction and audio positioning method proposed in the embodiment of the present application, online real-time multimedia data can be used to construct a grid frequency reference signal network, and the reference network can be used to locate the audio data. This realizes the use of online real-time multimedia data as the source of the grid frequency reference signal, avoiding the construction cost and maintenance cost of the grid frequency reference network. While taking advantage of the large amount of online multimedia data, the characteristics of being able to collect at multiple points can effectively ensure the geographical resolution of grid frequency signal collection and improve the accuracy of the grid frequency reference signal network. Furthermore, by using the grid frequency reference network to locate the audio, the efficiency of audio positioning in the present application is effectively improved, the scope of positioning application is expanded, and it is convenient for actual implementation. Thus, the problems of the traditional audio positioning method in the related art requiring the deployment of a large number of frequency interference recorders, high cost and complex maintenance, as well as harsh use conditions, poor positioning accuracy, low positioning accuracy, insufficient application scope, and inability to meet the needs of audio positioning in actual applications are solved.
[0122] Next, the grid frequency reference signal network construction and audio positioning device proposed in accordance with the embodiments of the present application will be described with reference to the accompanying drawings.
[0123] Figure 7 It is a structural diagram of the power grid frequency reference signal network construction and audio positioning device according to an embodiment of the present application.
[0124] like Figure 7 As shown, the grid frequency reference signal network construction and audio positioning device 10 includes: an extraction module 100 , a first processing module 200 , a second processing module 300 and a positioning module 400 .
[0125] The extraction module 100 is used to extract original audio from multi-region online real-time multimedia data and downsample the original audio to obtain downsampled audio.
[0126] The first processing module 200 is configured to detect whether the audio has a power grid frequency, and if the power grid frequency is detected, process the audio to obtain a single audio power grid frequency signal.
[0127] The second processing module 300 is used to process the grid frequency signals of multiple audio frequencies in the same area to obtain a single-area grid frequency reference signal, and to process the single-area grid frequency reference signals of multiple areas to obtain a grid frequency reference signal network.
[0128] The positioning module 400 is used to process the audio to be located to obtain the grid frequency signal of the audio to be located, and calculate the correlation coefficient between the grid frequency signal of the audio to be located and each grid point of the grid frequency reference signal network to obtain the final positioning result of the audio based on the correlation coefficient.
[0129] Optionally, in one embodiment of the present application, the first processing module 200 includes: a first acquiring unit and a determining unit.
[0130] The first acquisition unit is configured to acquire the signal length of the audio power grid frequency signal.
[0131] The determining unit is configured to calculate a noise segment rate based on the signal length, so as to determine whether the audio has a power grid frequency based on the noise segment rate.
[0132] Optionally, in one embodiment of the present application, the first processing module 200 includes: an enhancement unit, an extraction unit and a first processing unit.
[0133] The enhancement unit is used to enhance the power grid frequency signal of the audio to obtain an enhanced audio signal.
[0134] The extraction unit is used to extract the multi-harmonic power grid frequency signal of the enhanced audio signal.
[0135] The first processing unit is used to select and fuse the multi-harmonic grid frequency signals to obtain a single audio grid frequency signal.
[0136] Optionally, in one embodiment of the present application, the second processing module 300 includes: a construction unit and a second processing unit.
[0137] The construction unit is used to construct a target spatial domain selection matrix based on the signal length of the power grid frequency signal.
[0138] The second processing unit is used to convert the target airspace selection matrix into a target airspace selection undirected graph, and fuse the airspace signals of the target airspace selection undirected graph to obtain a single regional power grid frequency reference signal.
[0139] Optionally, in one embodiment of the present application, the second processing module 300 includes: a second acquisition unit, a collection unit, and a calculation unit.
[0140] The second acquisition unit is configured to divide the total area to be located in the multiple regions into a plurality of grids, and acquire the coordinates of the points of the plurality of grids.
[0141] The acquisition unit is used to acquire the center point coordinates of the frequency reference signal of the power grid of a single region in multiple regions.
[0142] The calculation unit is used to calculate the distance between the coordinates of the grid points of multiple grids and the coordinates of the center point, so as to obtain the grid frequency reference signal network according to the distance.
[0143] It should be noted that the above explanations of the embodiment of the grid frequency reference signal network construction and audio positioning method are also applicable to the grid frequency reference signal network construction and audio positioning device of this embodiment, and will not be repeated here.
[0144] According to the power grid frequency reference signal network construction and audio positioning device proposed in the embodiment of the present application, online real-time multimedia data can be used to construct a power grid frequency reference signal network, and the reference network can be used to locate the position of audio data, thereby realizing the use of online real-time multimedia data as the source of the power grid frequency reference signal, avoiding the construction cost and maintenance cost of the power grid frequency reference network, while taking advantage of the large amount of online multimedia data, which can be collected at multiple points, effectively ensuring the geographical resolution of power grid frequency signal collection, and improving the accuracy of the power grid frequency reference signal network. Furthermore, by using the power grid frequency reference network to locate audio, the efficiency of audio positioning in this application is effectively improved, the scope of positioning application is expanded, and it is convenient for actual implementation. Thus, the problems of the traditional audio positioning method in the related art requiring the deployment of a large number of frequency interference recorders, high cost and complex maintenance, as well as harsh use conditions, poor positioning accuracy, low positioning accuracy, insufficient application scope, and inability to meet the needs of audio positioning in actual applications are solved.
[0145] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0146] A memory 801 , a processor 802 , and a computer program stored in the memory 801 and executable on the processor 802 .
[0147] When the processor 802 executes the program, the grid frequency reference signal network construction and audio positioning method provided in the above embodiment is implemented.
[0148] Furthermore, the electronic device further includes:
[0149] The communication interface 803 is used for communication between the memory 801 and the processor 802 .
[0150] The memory 801 is used to store computer programs that can be run on the processor 802.
[0151] The memory 801 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0152] If the memory 801, processor 802, and communication interface 803 are implemented independently, the communication interface 803, memory 801, and processor 802 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0153] Optionally, in a specific implementation, if the memory 801, the processor 802 and the communication interface 803 are integrated on a chip, the memory 801, the processor 802 and the communication interface 803 can communicate with each other through an internal interface.
[0154] The processor 802 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0155] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned grid frequency reference signal network construction and audio positioning method.
[0156] An embodiment of the present application also provides a computer program product, including a computer program, which can run computer instructions. When the computer instructions are executed by a processor, the power grid frequency reference signal network construction and audio positioning method provided in the embodiment of the present application is implemented.
[0157] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0158] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0159] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0160] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" is any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable media include: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0161] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, it can be implemented using any one or a combination of the following technologies known in the art: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0162] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0163] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0164] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for constructing a power grid frequency reference signal network and audio positioning, characterized in that: The following steps are involved: Extracting original audio from multi-region online real-time multimedia data, and downsampling the original audio to obtain downsampled audio; detecting whether the audio has a power grid frequency, and if it is detected that the audio has the power grid frequency, processing the audio to obtain a power grid frequency signal of a single audio; Processing the grid frequency signals of multiple audio frequencies in the same region to obtain a single-region grid frequency reference signal, and processing the single-region grid frequency reference signals in multiple regions to obtain a grid frequency reference signal network; The audio to be located is processed to obtain a grid frequency signal of the audio to be located, and a correlation coefficient between the grid frequency signal of the audio to be located and each grid point of the grid frequency reference signal network is calculated to obtain a final positioning result of the audio according to the correlation coefficient.
2. The method according to claim 1, characterized in that The detecting whether the audio has a power grid frequency includes: Obtaining a signal length of a power grid frequency signal of the audio; A noise segment rate is calculated based on the signal length to determine whether the audio has a power grid frequency based on the noise segment rate.
3. The method according to claim 1, characterized in that The processing of the audio to obtain a single audio grid frequency signal includes: enhancing the power grid frequency signal of the audio to obtain an enhanced audio signal; extracting a multi-harmonic power grid frequency signal from the enhanced audio signal; The multi-harmonic grid frequency signal is selected and fused to obtain the single audio grid frequency signal.
4. The method according to claim 1, wherein The processing of the power grid frequency signals of multiple audio frequencies in the same region to obtain a single region power grid frequency reference signal includes: constructing a target spatial domain selection matrix based on the signal length of the power grid frequency signal; The target spatial domain selection matrix is converted into a target spatial domain selection undirected graph, and the spatial domain signals of the target spatial domain selection undirected graph are fused to obtain the single-region power grid frequency reference signal.
5. The method according to claim 1, characterized in that The processing of the single-region power grid frequency reference signal of multiple regions to obtain a power grid frequency reference signal network includes: Dividing the total area to be located in the multiple regions into a plurality of grids, and obtaining the coordinates of the points of the plurality of grids; Collecting the center point coordinates of the power grid frequency reference signal of a single region in the multiple regions; The distances between the grid point coordinates of the plurality of grids and the center point coordinates are calculated to obtain the grid frequency reference signal network according to the distances.
6. A power grid frequency reference signal network construction and audio positioning device, characterized in that: include: An extraction module is used to extract original audio from multi-region online real-time multimedia data and downsample the original audio to obtain downsampled audio; a first processing module, configured to detect whether the audio has a power grid frequency, and, if the power grid frequency is detected in the audio, process the audio to obtain a power grid frequency signal of a single audio; a second processing module, configured to process the grid frequency signals of multiple audio frequencies in the same region to obtain a single-region grid frequency reference signal, and to process the single-region grid frequency reference signals of multiple regions to obtain a grid frequency reference signal network; The positioning module is used to process the audio to be located to obtain the grid frequency signal of the audio to be located, and calculate the correlation coefficient between the grid frequency signal of the audio to be located and each grid point of the grid frequency reference signal network to obtain the final positioning result of the audio according to the correlation coefficient.
7. The device according to claim 6, characterized in that The first processing module includes: an acquiring unit, configured to acquire a signal length of a power grid frequency signal of the audio; A determining unit is configured to calculate a noise segment rate based on the signal length, so as to determine whether the audio has a power grid frequency based on the noise segment rate.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the grid frequency reference signal network construction and audio positioning method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the grid frequency reference signal network construction and audio positioning method as described in any one of claims 1 to 5.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed, it is used to implement the grid frequency reference signal network construction and audio positioning method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Digital audio tamper automatic detection method based on feature fusion
CN107274915A
Digital audio encryption method and system based on power grid frequency characteristic embedding
CN116155623A