Denoising method and device based on big data statistical analysis

By using big data statistical analysis to remove flying point data from the spectral data and using principal component analysis to fit curves, the problem of significant noise impact in time-frequency electromagnetic exploration was solved, improving the signal-to-noise ratio and quality of electromagnetic exploration data.

CN122063685APending Publication Date: 2026-05-19CHINA NAT PETROLEUM CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA NAT PETROLEUM CORP
Filing Date
2024-11-19
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In time-frequency electromagnetic exploration in areas with active human activity, conventional denoising methods are complex and prone to signal loss, especially at high frequencies where superposition denoising is ineffective, resulting in poor data quality.

Method used

By employing a big data statistical analysis approach, we remove stray data from the spectral data, calculate the coefficient of variation, and use principal component analysis to fit the spectral data curve, thereby reducing the impact of noise and improving denoising capabilities.

Benefits of technology

It effectively removes noise from the spectral data, reduces the impact of flying point data, and improves the signal-to-noise ratio and data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122063685A_ABST
    Figure CN122063685A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a noise reduction method based on big data statistical analysis. The noise reduction method comprises the following steps: removing flying spot data from frequency spectrum data (amplitude and phase); calculating a discrete coefficient of the frequency spectrum data of which the flying spot data is removed; and determining high-dispersion data of the spectrum data according to the dispersion coefficient, and performing regression fitting on a frequency domain amplitude curve and a frequency domain phase curve of the frequency by using a statistical method based on principal component analysis. According to the noise reduction method based on big data statistical analysis provided by the embodiment of the invention, the flying spot data is removed, the influence of the flying spot data on the spectrum data is reduced, and the noise of the spectrum data is reduced; the high-dispersion data is processed by adopting a statistical method based on principal component analysis, redundant denoising can be performed on the high-dispersion data, and the denoising capability of the denoising method based on big data statistical analysis provided by the embodiment of the invention is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electromagnetic exploration, and in particular to a noise reduction method and apparatus based on big data statistical analysis. Background Technology

[0002] In areas with high human activity, time-frequency electromagnetic exploration data can be affected by strong interference sources. Conventional noise suppression methods involve first removing invalid data periods, then filtering each period using multiple notch filtering or approximate spline curve smoothing, followed by Fourier transform and multi-period superposition for noise reduction in the frequency domain. This noise suppression method is complex and involves many human factors. Especially in the mid-to-high frequency band (above 1Hz), when the number of acquisition periods exceeds 100 pairs, superposition denoising is too simplistic. When non-random noise is present, the frequency domain data exhibits flying-point phenomena, resulting in poor data quality in areas with strong interference. Furthermore, multi-period superposition can easily lead to signal loss. Summary of the Invention

[0003] In view of the above problems, the present invention is proposed to provide a noise reduction method and apparatus based on big data statistical analysis that overcomes or at least partially solves the above problems.

[0004] In a first aspect, embodiments of the present invention provide a noise reduction method based on big data statistical analysis, comprising:

[0005] For each frequency in the pre-acquired spectrum data of multiple frequencies, the flying point data of the frequency spectrum data is removed from the spectrum data of the frequency to obtain the filtered spectrum data of the frequency; the flying point data includes: spectrum data of at least one period in the spectrum data of each frequency.

[0006] Based on the spectral data of each filtered frequency, the discrete coefficients of the spectral data of the filtered frequency are obtained.

[0007] For each frequency in the filtered frequency spectrum data, the high dispersion data is used to regress and fit the frequency spectrum data curve using a statistical method based on principal component analysis; the high dispersion data is the spectrum data obtained after filtering out the frequency whose dispersion coefficient is greater than a preset first threshold.

[0008] In one embodiment, the pre-acquired spectrum data of multiple frequencies is obtained in the following manner:

[0009] The analog-to-digital converter code value data of time-domain analog signal data of multiple frequencies is converted into digital signal data potential difference data of multiple frequencies;

[0010] The digital signal data potential difference data of all frequencies are normalized to obtain normalized digital signal data potential difference data of multiple frequencies.

[0011] Eliminate DC drift in the normalized digital signal data potential difference data at all frequencies;

[0012] The normalized digital signal data potential difference data of each frequency, after eliminating DC drift, is converted into the spectrum data of the frequency to obtain spectrum data of multiple frequencies.

[0013] In one embodiment, the step of removing the flying point data of the frequency's spectrum data from the spectrum data of each of the pre-acquired multiple frequencies to obtain the filtered spectrum data of the frequency includes:

[0014] For the spectral data at each frequency, calculate the average value of the spectral data across all periods;

[0015] For each frequency of the spectrum data, the root mean square error between the spectrum data of each period and the average value of the spectrum data is calculated based on the spectrum data of each period and the average value of the spectrum data.

[0016] The periodic spectral data in which the root mean square error of all the spectral data and the average value of the spectral data is greater than a preset error threshold is used as the flying point data.

[0017] The flying point data is removed from the frequency spectrum data to obtain the filtered frequency spectrum data.

[0018] In one embodiment, obtaining the discrete coefficients of the spectrum data of the filtered frequencies based on the spectrum data of each filtered frequency includes:

[0019] Based on the spectrum data after filtering for each frequency, calculate the standard deviation and mean of the spectrum data after filtering for each frequency;

[0020] The ratio of the standard deviation to the mean of the spectrum data after each frequency filtering is used as the coefficient of variation.

[0021] In one embodiment, the method further includes: classifying the spectral data of the filtered frequencies according to a first threshold and a second threshold to obtain high-dispersion data, medium-dispersion data, and low-dispersion data.

[0022] In one embodiment, the first threshold is 0.5; the second threshold is 0.2;

[0023] Correspondingly, the step of classifying the spectral data of the filtered frequencies according to the first threshold and the second threshold to obtain high-dispersion data, medium-dispersion data, and low-dispersion data includes:

[0024] When the coefficient of dispersion of the frequency spectrum data after screening is less than 0.2, the frequency spectrum data after screening is classified as low dispersion data.

[0025] When the coefficient of dispersion of the frequency spectrum data after screening is greater than or equal to 0.2 and less than or equal to 0.5, the frequency spectrum data after screening is classified as medium dispersion data.

[0026] When the coefficient of dispersion of the selected frequency spectrum data is greater than 0.5, the selected frequency spectrum data is classified as high-dispersion data.

[0027] In one embodiment, the method further includes:

[0028] For the low-dispersion data, frequency domain spectrum data curves are directly plotted based on the average value of the spectrum data of the filtered frequencies;

[0029] For the medium-discreteness data, the least squares method is used to fit the spectrum data curve based on the filtered frequency spectrum data.

[0030] In one embodiment, the spectral data of the plurality of frequencies includes: spectral data of the dominant frequency and harmonics;

[0031] The step of using a statistical method based on principal component analysis to regress and fit the high dispersion data of each frequency in the filtered frequency spectrum data to the frequency spectrum data curve includes:

[0032] Acquire spectrum data at multiple frequencies from multiple measurement points; the number of measurement points is equal to the number of frequencies.

[0033] The spectral data of each frequency at each measurement point is dimensionality reduced to obtain the dimensionality-reduced spectral data;

[0034] The least squares method is used to fit the dimensionality-reduced spectral data, and the resulting fitted curve is the spectral data curve.

[0035] In one embodiment, the dimensionality reduction processing of the spectral data for each frequency at each measurement point includes:

[0036] Based on the pre-acquired spectral data of each frequency at each measurement point, calculate the average value of the spectral data of each frequency at each measurement point;

[0037] A spectrum matrix is ​​constructed based on the average value of the spectrum data for each frequency at each measurement point; the columns of the spectrum matrix are the average values ​​of each frequency at each measurement point;

[0038] The spectrum matrix is ​​standardized to obtain a standardized spectrum matrix;

[0039] Calculate the covariance matrix of the normalized spectrum matrix;

[0040] The covariance matrix is ​​subjected to eigenvalue decomposition to obtain eigenvalues ​​and eigenvectors;

[0041] Calculate the number of eigenvalues ​​whose proportion of the sum of all eigenvalues ​​is greater than or equal to a preset threshold to obtain the principal component frequencies; retain the rows of the spectrum matrix of the principal component frequencies to obtain the dimension-reduced spectrum matrix; retain the eigenvectors of the eigenvectors of the principal component frequencies to obtain the dimension-reduced eigenvectors.

[0042] Multiplying the dimension-reduced spectrum matrix with the dimension-reduced eigenvector yields the dimension-reduced spectrum data.

[0043] Thirdly, embodiments of the present invention provide a computing device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the program executed by the processor implements a noise reduction method based on big data statistical analysis.

[0044] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a noise reduction method based on big data statistical analysis.

[0045] Fifthly, embodiments of the present invention provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements a noise reduction method based on big data statistical analysis.

[0046] The beneficial effects of the above-described technical solutions provided in the embodiments of the present invention include at least the following:

[0047] This invention provides a noise reduction method based on big data statistical analysis, comprising: removing fly-spot data from spectral data; calculating the coefficient of variation of the spectral data after removing fly-spot data; determining the highly discrete data of the spectral data based on the coefficient of variation; and using a statistical method based on principal component analysis to regress and fit the frequency domain amplitude curve and frequency domain phase curve. The noise reduction method based on big data statistical analysis provided by this invention removes fly-spot data, reducing its impact on spectral data and lowering the noise level of the spectral data; by using a statistical method based on principal component analysis to process the highly discrete data, it can perform redundant noise reduction on the highly discrete data, thus improving the noise reduction capability of the noise reduction method based on big data statistical analysis provided by this invention.

[0048] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings.

[0049] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0050] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0051] Figure 1 A flowchart illustrating a noise reduction method based on big data statistical analysis provided in an embodiment of the present invention;

[0052] Figure 2 A pre-acquired amplitude and phase dot matrix diagram of multiple frequencies provided for embodiments of the present invention;

[0053] Figure 3 A graph of the frequency domain phase curve provided in the embodiments of the present invention;

[0054] Figure 4 This is a structural block diagram of a noise reduction device based on big data statistical analysis provided in an embodiment of the present invention. Detailed Implementation

[0055] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0056] To address the aforementioned "flying point" problem, "flying point" data refers to spectral data among the multiple frequency spectral data processed by this invention whose average phase or average frequency root mean square error with the spectral data exceeds a preset error threshold. This invention provides a noise reduction method based on big data statistical analysis, the flowchart of which is shown below. Figure 1 As shown, it includes the following steps:

[0057] S11. For the spectrum data of each frequency in the pre-acquired spectrum data of multiple frequencies, remove the flying point data of the frequency spectrum data from the frequency spectrum data to obtain the filtered frequency spectrum data; the flying point data includes: the spectrum data of at least one cycle in the spectrum data of each frequency.

[0058] S12. Based on the spectrum data of each frequency after screening, obtain the discrete coefficients of the spectrum data of the screened frequencies.

[0059] S13. For each frequency in the filtered frequency spectrum data, use a statistical method based on principal component analysis to regress and fit the frequency spectrum data curve; the high dispersion data is the spectrum data obtained after filtering out the dispersion coefficients that are greater than the preset first threshold.

[0060] The noise reduction method based on big data statistical analysis provided in this embodiment of the invention removes flying point data, reduces the impact of flying point data on spectrum data, and lowers the noise of spectrum data; it uses a statistical method based on principal component analysis to process highly discrete data, which can perform redundant noise reduction on highly discrete data, thereby improving the noise reduction capability of the noise reduction method based on big data statistical analysis provided in this embodiment of the invention.

[0061] Acquiring spectral data at multiple frequencies can be achieved, for example, through the time-frequency electromagnetic method (TFM), a novel approach in petroleum exploration. This method involves supplying a strong current to the ground to excite oil and gas exploration targets, then measuring the secondary electromagnetic field and its spectrum generated by discharge in the porous media of the oil and gas reservoir. This method simultaneously obtains signals in both the time and frequency domains. Through joint processing of these signals, a subsurface physical property model can be accurately reconstructed, revealing resistivity and polarizability anomalies of the oil and gas exploration targets. The excited signals can be acquired, for example, through measurement points.

[0062] In step S11 above, the pre-acquired spectrum data of multiple frequencies can be processed, for example, in the following manner:

[0063] The analog-to-digital converter (ADC) code values ​​of time-domain analog signal data at multiple frequencies collected from the measurement points are converted into digital signal data (potential difference data) at multiple frequencies. Since adjacent measurement points have high correlation, the optimal positional relationship between these points is adjacent; for example, this can be calculated using the following formula:

[0064] ΔU = ADC × factor / Gain;

[0065] Where: ΔU is the digital signal data potential difference data; ADC is the ADC code value of time-domain analog signal data at multiple frequencies; factor is the digital-to-analog conversion coefficient; Gain is the gain, which can be set to 1, 10 or 100.

[0066] Normalize the digital signal data potential difference data for all frequencies to obtain normalized digital signal data potential difference data for multiple frequencies; for example, it can be calculated using the following formula:

[0067] ΔU N =ΔU / Distance;

[0068] Where: ΔU N ΔU is the normalized digital signal data potential difference data; Distance is the electrode distance of the electric field that generates the aforementioned multiple frequencies of time-domain analog signal data; electrode distance refers to the distance between the cathode and anode of the electric field.

[0069] Eliminate DC drift in normalized digital signal data potential difference data of all frequencies; for example, half-cycle superposition processing can be performed on normalized digital signal data potential difference data to eliminate DC drift; half-cycle superposition processing is to subtract the data of the second half-cycle from the data of the first half-cycle of normalized digital signal data potential difference data of each cycle and then divide by 2, and set the data of the second half-cycle to zero to obtain the data after eliminating DC drift in that cycle.

[0070] The normalized digital signal data potential difference data of each frequency to eliminate DC drift is converted into frequency spectrum data to obtain spectrum data of multiple frequencies. The normalized digital signal data potential difference data of each frequency in the time domain is converted into frequency spectrum data, for example, by using Fourier transform.

[0071] After processing the spectral data of multiple frequencies, before step S11, the noise reduction method based on big data statistical analysis may also include the following two steps:

[0072] Step 1: Since there is a multiple difference between the harmonic energy and the fundamental energy, it is necessary to compensate for the harmonic amplitude: compensate for the harmonic amplitude of all frequencies; for example, the time-frequency electromagnetic method can be used to process the odd harmonics, and the 3rd and 5th harmonics are usually used to compensate for the harmonic amplitude.

[0073] Step 2: Combine the spectrum data of all frequencies together and sort them according to frequency size to obtain spectrum data of multiple frequencies.

[0074] Specifically, step S11 can be implemented in the following manner, for example:

[0075] For each frequency's spectral data, calculate the average of the spectral data across all periods;

[0076] For each frequency's spectral data, the root mean square error between the spectral data of each period and the average value is calculated based on the average value of the spectral data of each period.

[0077] In the frequency spectrum data, the period of spectrum data where the root mean square error between all spectrum data and the average of the spectrum data is greater than a preset error threshold is used as flying point data.

[0078] By removing flying point data from the frequency spectrum data, the filtered frequency spectrum data is obtained.

[0079] For example, if the aforementioned spectral data could be amplitude data, then the aforementioned steps could be implemented in the following manner:

[0080] For each frequency of amplitude data, calculate the average amplitude data for all cycles;

[0081] For the spectral data of each frequency, the root mean square error between the amplitude data of each period and the average value of the amplitude data is calculated based on the amplitude data of each period and the average value of the amplitude data.

[0082] The amplitude data of the period in which the root mean square error of all amplitude data and the average amplitude data is greater than a preset error threshold is taken as flying point data.

[0083] By removing the flying point data from the frequency amplitude data, the filtered frequency amplitude data is obtained.

[0084] The aforementioned spectral data can also be phase data, for example. Since the implementation of phase data is similar to that of the aforementioned amplitude data, it will not be described in detail here.

[0085] In step S12 above, the discrete coefficients can be calculated, for example, in the following manner:

[0086] Based on the spectrum data after filtering for each frequency, calculate the standard deviation and mean of the spectrum data after filtering for each frequency;

[0087] The ratio of the standard deviation to the mean of the spectrum data after each frequency screening is used as the coefficient of variation.

[0088] In one embodiment, the aforementioned noise reduction method based on big data statistical analysis may further include: classifying the spectral data of the filtered frequencies according to a first threshold and a second threshold to obtain high-dispersion data, medium-dispersion data, and low-dispersion data; the aforementioned first threshold may be, for example, 0.5; the aforementioned second threshold may be, for example, 0.2.

[0089] Correspondingly, based on the first threshold and the second threshold, the filtered frequency spectrum data may include, for example, the following categories:

[0090] When the coefficient of dispersion of the spectral data of the selected frequencies is less than 0.2, the spectral data of the selected frequencies are classified as low-dispersion data.

[0091] When the coefficient of dispersion of the selected frequency spectrum data is greater than or equal to 0.2 and less than or equal to 0.5, the selected frequency spectrum data is classified as medium dispersion data.

[0092] When the coefficient of dispersion of the spectral data of the selected frequencies is greater than 0.5, the spectral data of the selected frequencies are classified as high-dispersion data.

[0093] For the aforementioned low-dispersion data, for example, spectral data curves can be directly plotted based on the spectral data of the filtered frequencies.

[0094] For the aforementioned medium-discreteness data, for example, the least squares method can be used to fit the spectrum data curve based on the filtered frequency spectrum data.

[0095] In one embodiment, the spectral data of multiple frequencies may include, for example, the number of frequencies; then the aforementioned step S13 may include, for example, the following steps:

[0096] Acquire spectrum data at multiple frequencies from multiple measurement points; the number of measurement points equals the number of frequencies.

[0097] The spectral data of each frequency at each measurement point is dimensionality reduced to obtain the dimensionality-reduced spectral data;

[0098] The least squares method is used to fit the dimensionality-reduced spectral data, and the resulting fitted curve is the spectral data curve.

[0099] Taking spectral data as phase data as an example, the aforementioned step S13 includes, for example, the following steps: acquiring phase data of spectral data at multiple frequencies from multiple measurement points; the number of measurement points being equal to the number of frequencies; performing dimensionality reduction processing on the phase data of the spectral data at each frequency from each measurement point to obtain dimensionality-reduced phase data; fitting the dimensionality-reduced phase data using the least squares method, and the resulting fitted curve is the frequency domain phase curve; the frequency domain phase curve obtained using the aforementioned three curve plotting methods can be, for example, as shown in the figure. Figure 3 As shown.

[0100] Specifically, dimensionality reduction processing can be performed on the spectral data of each measurement point at each frequency, for example, in the following manner:

[0101] Based on the pre-acquired spectral data of each frequency at each measurement point, calculate the average value of the spectral data of each frequency at each measurement point;

[0102] A spectrum matrix is ​​constructed based on the average spectral data of each frequency at each measurement point; the columns of the spectrum matrix are the average spectral data of each frequency at each measurement point.

[0103] The spectrum matrix is ​​standardized to obtain a standardized spectrum matrix; the covariance matrix of the standardized spectrum matrix is ​​calculated; the covariance matrix of the spectrum matrix is ​​eigenvalued and eigenvectors are obtained by eigenvalue decomposition.

[0104] Calculate the number of eigenvalues ​​whose proportion of the sum of all eigenvalues ​​is greater than or equal to a preset threshold to obtain the number of principal component frequencies; retain the rows of the spectrum matrix containing the principal component frequencies to obtain the dimension-reduced spectrum matrix; retain the eigenvectors of the eigenvectors containing the principal component frequencies to obtain the dimension-reduced eigenvectors.

[0105] Multiplying the dimension-reduced spectrum matrix by the dimension-reduced eigenvector yields the dimension-reduced spectrum data.

[0106] Taking spectral data as amplitude data as an example, dimensionality reduction processing includes the following steps:

[0107] Based on the pre-acquired amplitude data of each frequency at each measuring point, calculate the average value of the amplitude data of each frequency at each measuring point;

[0108] An amplitude matrix is ​​constructed based on the average amplitude data for each frequency at each measuring point; the columns of the amplitude matrix are the average amplitude values ​​for each frequency at each measuring point.

[0109] The amplitude matrix is ​​standardized to obtain a standardized amplitude matrix; the amplitude covariance matrix of the standardized amplitude matrix is ​​calculated; the amplitude covariance matrix is ​​decomposed into eigenvalues ​​and amplitude eigenvectors.

[0110] Calculate the number of amplitude eigenvalues ​​whose proportion of the sum of all amplitude eigenvalues ​​is greater than or equal to a preset threshold to obtain the number of amplitude principal component frequencies; retain the rows of the amplitude matrix with the number of amplitude principal component frequencies to obtain the dimension-reduced amplitude matrix; retain the amplitude eigenvectors of the amplitude eigenvectors with the number of amplitude principal component frequencies to obtain the dimension-reduced amplitude eigenvectors.

[0111] Multiplying the dimension-reduced amplitude matrix with the dimension-reduced amplitude eigenvector yields the dimension-reduced amplitude data.

[0112] The above-mentioned noise reduction method based on big data statistical analysis was used to process the spectrum data of multiple frequencies of an oilfield. The signal-to-noise ratio of the spectrum data increased by an average of 12.8%, and the rate of high-quality products increased by 3.4%.

[0113] Based on the same inventive concept, embodiments of the present invention provide a noise reduction device based on big data statistical analysis, the structural block diagram of which is shown below. Figure 4 As shown, it includes:

[0114] The filtering module 41 is used to remove the flypoint data of the frequency spectrum data from the spectrum data of each frequency in the spectrum data of multiple pre-acquired frequencies, so as to obtain the filtered frequency spectrum data; the flypoint data includes: spectrum data of at least one cycle in the spectrum data of each frequency.

[0115] The coefficient calculation module 42 is used to obtain the discrete coefficients of the spectrum data of the selected frequencies based on the spectrum data of each selected frequency.

[0116] The curve fitting module 43 is used to use a statistical method based on principal component analysis to regress and fit the high dispersion data of each frequency in the filtered frequency spectrum data to the frequency spectrum data curve; the high dispersion data is the spectrum data of the filtered frequencies with a dispersion coefficient greater than a preset first threshold.

[0117] Based on the same inventive concept, embodiments of the present invention also provide a computing device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the program executed by the processor is a noise reduction method based on big data statistical analysis.

[0118] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a noise reduction method based on big data statistical analysis.

[0119] Based on the same inventive concept, embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements a noise reduction method based on big data statistical analysis.

[0120] Since the principle behind these devices is similar to the aforementioned noise reduction method based on big data statistical analysis, the implementation of these devices can be found in the implementation of the aforementioned methods, and the repetitions will not be repeated.

[0121] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0122] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0123] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0124] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0125] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A noise reduction method based on big data statistical analysis, characterized in that, include: For each frequency in the pre-acquired spectrum data of multiple frequencies, the flying point data of the frequency spectrum data is removed from the spectrum data of the frequency to obtain the filtered spectrum data of the frequency. The flying point data includes: at least one period of spectrum data in the spectrum data of each frequency; Based on the spectral data of each filtered frequency, the discrete coefficients of the spectral data of the filtered frequency are obtained. For each frequency in the filtered frequency spectrum data, the high dispersion data is used to regress and fit the frequency spectrum data curve using a statistical method based on principal component analysis; the high dispersion data is the spectrum data obtained after filtering out the frequency whose dispersion coefficient is greater than a preset first threshold.

2. The method as described in claim 1, characterized in that, The pre-acquired spectral data for multiple frequencies are obtained in the following manner: The analog-to-digital converter code value data of time-domain analog signal data of multiple frequencies is converted into digital signal data potential difference data of multiple frequencies; The digital signal data potential difference data of all frequencies are normalized to obtain normalized digital signal data potential difference data of multiple frequencies. Eliminate DC drift in the normalized digital signal data potential difference data at all frequencies; The normalized digital signal data potential difference data of each frequency, after eliminating DC drift, is converted into the spectrum data of the frequency to obtain spectrum data of multiple frequencies.

3. The method as described in claim 1, characterized in that, The step of removing the flying point data of the frequency's spectrum data from the spectrum data of each of the pre-acquired multiple frequencies to obtain the filtered spectrum data of the frequency includes: For the spectral data at each frequency, calculate the average value of the spectral data across all periods; For each frequency of the spectrum data, the root mean square error between the spectrum data of each period and the average value of the spectrum data is calculated based on the spectrum data of each period and the average value of the spectrum data. The periodic spectral data in which the root mean square error of all the spectral data and the average value of the spectral data is greater than a preset error threshold is used as the flying point data. The flying point data is removed from the frequency spectrum data to obtain the filtered frequency spectrum data.

4. The method as described in claim 1, characterized in that, The step of obtaining the discrete coefficients of the spectrum data of the filtered frequencies based on the spectrum data of each filtered frequency includes: Based on the spectrum data after filtering for each frequency, calculate the standard deviation and mean of the spectrum data after filtering for each frequency; The ratio of the standard deviation to the mean of the spectrum data after each frequency filtering is used as the coefficient of variation.

5. The method as described in claim 1, characterized in that, The method further includes classifying the spectral data of the filtered frequencies according to a first threshold and a second threshold to obtain high-dispersion data, medium-dispersion data and low-dispersion data.

6. The method as described in claim 5, characterized in that, The first threshold is 0.5; the second threshold is 0.2; Correspondingly, the step of classifying the spectral data of the filtered frequencies according to the first threshold and the second threshold to obtain high-dispersion data, medium-dispersion data, and low-dispersion data includes: When the coefficient of dispersion of the frequency spectrum data after screening is less than 0.2, the frequency spectrum data after screening is classified as low dispersion data. When the coefficient of dispersion of the frequency spectrum data after screening is greater than or equal to 0.2 and less than or equal to 0.5, the frequency spectrum data after screening is classified as medium dispersion data. When the coefficient of dispersion of the selected frequency spectrum data is greater than 0.5, the selected frequency spectrum data is classified as high-dispersion data.

7. The method as described in claim 5, characterized in that, The method further includes: For the low-dispersion data, frequency domain spectrum data curves are directly plotted based on the filtered frequency spectrum data; For the medium-discreteness data, the least squares method is used to fit the spectrum data curve based on the filtered frequency spectrum data.

8. The method as described in claim 1, characterized in that, The spectral data for the multiple frequencies includes: the number of frequencies; The step of using a statistical method based on principal component analysis to regress and fit the high dispersion data of each frequency in the filtered frequency spectrum data to the frequency spectrum data curve includes: Acquire spectrum data at multiple frequencies from multiple measurement points; the number of measurement points is equal to the number of frequencies. The spectral data of each frequency at each measurement point is dimensionality reduced to obtain the dimensionality-reduced spectral data; The least squares method is used to fit the dimensionality-reduced spectral data, and the resulting fitted curve is the spectral data curve.

9. The method as described in claim 8, characterized in that, The dimensionality reduction processing of the spectral data at each frequency for each measurement point includes: Based on the pre-acquired spectral data of each frequency at each measurement point, calculate the average value of the spectral data of each frequency at each measurement point; A spectrum matrix is ​​constructed based on the average value of the spectrum data for each frequency at each measurement point; the columns of the spectrum matrix are the average values ​​of each frequency at each measurement point; The spectrum matrix is ​​standardized to obtain a standardized spectrum matrix; Calculate the covariance matrix of the normalized spectrum matrix; The covariance matrix is ​​subjected to eigenvalue decomposition to obtain eigenvalues ​​and eigenvectors; Calculate the number of eigenvalues ​​whose proportion of the sum of all eigenvalues ​​is greater than or equal to a preset threshold to obtain the principal component frequencies; retain the rows of the spectrum matrix of the principal component frequencies to obtain the dimension-reduced spectrum matrix; retain the eigenvectors of the eigenvectors of the principal component frequencies to obtain the dimension-reduced eigenvectors. Multiplying the dimension-reduced spectrum matrix with the dimension-reduced eigenvector yields the dimension-reduced spectrum data.

10. A noise reduction device based on big data statistical analysis, characterized in that, include: The filtering module is used to remove the flying point data of the frequency spectrum data from the spectrum data of each frequency in a pre-acquired spectrum data set, so as to obtain the filtered spectrum data of the frequency. The flying point data includes: at least one period of spectrum data in the spectrum data of each frequency; The coefficient calculation module is used to obtain the discrete coefficients of the spectrum data of the filtered frequencies based on the spectrum data of each frequency. The curve fitting module is used to regress and fit the high dispersion data of each frequency in the filtered frequency spectrum data using a statistical method based on principal component analysis to the spectrum data curve of the frequency; the high dispersion data refers to the spectrum data of the filtered frequencies whose dispersion coefficient is greater than a preset first threshold.

11. A computing device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the program executed by the processor implements the noise reduction method based on big data statistical analysis as described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the noise reduction method based on big data statistical analysis as described in any one of claims 1-9.

13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the noise reduction method based on big data statistical analysis as described in any one of claims 1-9.