Human noise removal method and system based on combination of neural network and wavelet transform
By converting the audio earth electromagnetic signal into a two-dimensional timing chart and combining wavelet transformation and convolutional neural network methods, the problem of traditional methods being difficult to denoise with high precision under strong human interference is solved, and efficient audio earth electromagnetic signal denoising is achieved.
Patent Information
- Application Number
- CN202510962963.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing traditional denoising methods are difficult to achieve high-precision electromagnetic denoising of audio ground under strong human interference, and neural networks are difficult to effectively deal with strong human interference due to neglecting human noise characteristics.
The method of combining neural network and wavelet transformation is adopted to convert the one-dimensional audio earth electromagnetic timing signal into a two-dimensional timing diagram through Gram angle and field transformation. The two-dimensional wavelet positive transformation and convolutional neural network are used to decompose and extract subband features of different frequency bands, and combine the two-dimensional wavelet inverse transformation fusion characteristics to achieve the removal of humanistic noise.
It improves the denoising accuracy under strong human noise interference, can effectively capture and remove human noise, and improves the high-precision performance of the denoising model.
Smart Images

Figure CN120452470A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of human noise removal, and in particular to a method and system for removing human noise by combining a neural network and a wavelet transform. Background Art
[0002] Audio magnetotelluric (AMT) plays a vital role in the study of the Earth's internal electrical structure and is widely used in the exploration of various metal ores. However, the energy of AMT signals from natural sources is weak and susceptible to interference from human-induced noise.
[0003] Traditional denoising methods (such as remote referencing, impedance estimation, inversion, time period selection, mathematical decomposition, sparse representation, etc.) are difficult to achieve high-precision audio magnetotelluric denoising under strong human interference due to limitations in usage conditions and denoising performance. However, neural networks, when a rich sample set and sufficient network depth are established, use optimization algorithms to train the network and find a mapping relationship from samples to labels. This is expected to break through the various limitations of traditional methods and remove strong human interference from a deeper level. However, at this stage, neural networks still have difficulty dealing with strong human interference because they ignore the characteristics of human noise. Summary of the Invention
[0004] The main purpose of the embodiments of the present disclosure is to propose a method and system for removing human noise by combining a neural network and a wavelet transform, which can improve the denoising accuracy of the denoising model under strong human interference.
[0005] A first aspect of the embodiments of the present application provides a method for removing human noise by combining a neural network and a wavelet transform, the method comprising: Obtain the target audio magnetotelluric time series signal from which human noise is to be removed; Converting the target audio magnetotelluric time series signal into a target two-dimensional time series diagram based on Gram angle and field transformation; The target two-dimensional time series graph is input into a denoising model, so as to decompose a plurality of sub-band features of different frequency bands from the target two-dimensional time series graph based on a two-dimensional wavelet forward transform according to the denoising model, extract a plurality of corresponding sub-band intermediate features from the plurality of sub-band features based on a two-dimensional convolutional neural network, and fuse the plurality of sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human noise.
[0006] In some embodiments, the denoising model is composed of the two-dimensional convolutional neural network, n times the two-dimensional wavelet forward transform, and n times the two-dimensional wavelet inverse transform; n is a positive integer greater than 1; The method comprises: decomposing a plurality of sub-band features of different frequency bands from a target two-dimensional time series graph based on a two-dimensional wavelet forward transform, extracting a plurality of corresponding sub-band intermediate features from the plurality of sub-band features based on a two-dimensional convolutional neural network, and fusing the plurality of sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out human noise, including: Decomposing the target two-dimensional time series graph into m sub-band features of different frequency bands according to the first two-dimensional wavelet forward transform; m is 4; Extracting m corresponding first sub-band intermediate features from the m sub-band features according to the two-dimensional convolutional neural network; Decomposing the m first sub-band intermediate features into corresponding m second sub-band features according to the second two-dimensional wavelet forward transform; According to the two-dimensional convolutional neural network From the second sub-band features, the corresponding The second sub-band intermediate features; And so on, until we get the two-dimensional convolutional neural network from From the nth sub-band features, extract the corresponding nth sub-band intermediate features; According to the first two-dimensional wavelet inverse transform, Each m corresponding n-th sub-band intermediate features in the n-th sub-band intermediate features are fused to obtain n+1th sub-band features; According to the two-dimensional convolutional neural network From the n+1th sub-band features, extract the corresponding The n+1th sub-band intermediate features; And so on, until the m 2n-th sub-band intermediate features are extracted from the m 2n-th sub-band features according to the two-dimensional convolutional neural network, and the m 2n-th sub-band intermediate features are fused according to the n-th two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human noise.
[0007] In some embodiments, the denoising model is trained in the following manner: Acquire a time series sample; the time series sample is an audio magnetotelluric time series signal obtained by superimposing human noise on a noise-free audio magnetotelluric time series, and the label of the time series sample is the audio magnetotelluric time series signal without superimposed human noise; converting the time series samples into two-dimensional time series graph samples based on Gram angle and field transformation; The two-dimensional time series graph samples and their labels are input into the denoising model for training until the trained denoising model is obtained.
[0008] In some embodiments, obtaining time series samples includes: Acquire real audio magnetotelluric time series containing human noise; Generating a noise-free simulated audio magnetotelluric time series based on forward simulation; wherein the simulated audio magnetotelluric time series and the actual audio magnetotelluric time series have the same number of electromagnetic channels; Sampling the actual audio magnetotelluric time series using a sliding time window to obtain an audio magnetotelluric time series containing human noise; wherein the step size of the sampling sliding time window is smaller than the length of the sliding time window; Sampling the simulated audio magnetotelluric time series using the sliding time window and the step size of the sliding time window to obtain a noise-free audio magnetotelluric time series; Human noise is extracted from the audio magnetotelluric time series containing human noise, and the extracted human noise is superimposed on the noise-free audio magnetotelluric time series to obtain a time series sample.
[0009] In some embodiments, converting the time series samples into two-dimensional time series samples based on Gram angle and field transformation includes: Normalizing and scaling the time series samples; Mapping the normalized and scaled time series samples onto polar coordinates; Converting the time series samples mapped onto polar coordinates into two-dimensional time series graph samples based on Gram angle and field transformation; The converting of the target audio magnetotelluric time sequence signal into a target two-dimensional time sequence diagram based on Gram angle and field transformation includes: Normalizing and scaling the target audio frequency magnetotelluric time series signal; Mapping the normalized and scaled target audio magnetotelluric time series signal to polar coordinates; Based on Gram angle and field transformation, the target audio frequency magnetotelluric time series signal mapped onto polar coordinates is converted into a target two-dimensional time series graph.
[0010] In some embodiments, converting the target audio frequency magnetotelluric time sequence signal into a target two-dimensional time sequence diagram includes: Normalizing and scaling the target audio frequency magnetotelluric time series signal; Mapping the normalized and scaled target audio magnetotelluric time series signal to polar coordinates; Based on Gram angle and field transformation, the target audio frequency magnetotelluric time series signal mapped onto polar coordinates is converted into a target two-dimensional time series graph.
[0011] In some embodiments, the two-dimensional convolutional neural network includes at least one convolution block, normalization, and an ELU activation function.
[0012] In some embodiments, after obtaining the target audio magnetotelluric time series signal after filtering out human noise, the method further includes: Extract all frequency points from the target audio magnetotelluric time series signal after filtering out human noise; Calculating the degree of distortion of all the frequency points based on the Nyquist diagram; Evaluate the quality of the denoising model according to the distortion levels of all the frequency points; Wherein, the real part of the impedance in the Nyquist plot is and the imaginary part The expression is: ; ; ; ; in, 、 represents different frequency variables, represents a constant matrix, Represents a constant matrix The real coefficients of the matrix, Represents a constant matrix The imaginary coefficient of the matrix, represents an imaginary number, express A constant coefficient matrix, represents the Cauchy principal value operation, express The corresponding constant parameter, express The corresponding constant parameter, Indicates the An arbitrary real constant.
[0013] A second aspect of the embodiments of the present application provides a system for removing human noise by combining a neural network and a wavelet transform, the system comprising: A signal acquisition module is used to obtain the target audio magnetotelluric time series signal from which human noise is to be removed; a two-dimensional graph generating module, configured to convert the target audio frequency magnetotelluric time series signal into a target two-dimensional time series graph based on Gram angle and field transformation; The noise removal module is used to input the target two-dimensional time series graph into a denoising model, so as to decompose a plurality of sub-band features of different frequency bands from the target two-dimensional time series graph based on a two-dimensional wavelet forward transform according to the denoising model, extract a plurality of corresponding sub-band intermediate features from the plurality of sub-band features based on a two-dimensional convolutional neural network, and fuse the plurality of sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out human noise.
[0014] A third aspect of an embodiment of the present application proposes an electronic device, comprising at least one controller and a memory for communicating with the controller; the memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the at least one controller to enable the at least one controller to execute a method for removing human noise that combines a neural network and a wavelet transform as described above.
[0015] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed, the method for removing human noise by combining a neural network and a wavelet transform as described above is implemented.
[0016] The method provided in this embodiment has the following advantages: Based on the Gram angle and field transform, this method converts the one-dimensional target audio magnetotelluric time series signal into a two-dimensional target two-dimensional time series diagram without destroying the temporal relationship of the time series, adding two-dimensional spatial information and making the noise shape more prominent, which is conducive to the denoising model capturing these human-induced noises. Then, the denoising model decomposes multiple sub-band features of different frequency bands from the target two-dimensional time series diagram based on the two-dimensional wavelet forward transform, and extracts the corresponding multiple sub-band intermediate features from the multiple sub-band features based on the two-dimensional convolutional neural network. The multiple sub-band intermediate features are fused based on the two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human-induced noise. Here, a two-dimensional wavelet transform is combined with a neural network to decompose the input target two-dimensional time series diagram into wavelet sparse sub-band features of different frequencies. In the denoising model, the neural network can analyze the spectral information of the human-induced noise in different frequency bands in the form of wavelet sparse coefficients, so that the denoising model can explore the sparse characteristics of the human-induced noise and improve the high-precision denoising accuracy of the denoising model under strong human-induced noise interference.
[0017] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0019] Figure 1 1 is a flow chart of a method for removing human noise by combining a neural network and a wavelet transform according to an embodiment of the present application; Figure 2 Schematic diagram of the structure of the denoising model provided in the embodiment of the present application; Figure 3 for Figure 2 Schematic diagram of the structure of the two-dimensional convolutional neural network in; Figure 4 This is a flowchart of the sample data collection process provided by the embodiment of the present application; Figure 5 This is a flowchart of the method for removing human noise provided in an embodiment of the present application; Figure 6 Schematic diagram of a noisy high-frequency magnetotelluric time series provided in an embodiment of the present application; Figure 7 for Figure 6 Schematic diagram of the audio magnetotelluric time series fragment sampled in the medium time window; Figure 8 Schematic diagram of a noise-free or low-noise frequency magnetotelluric time series provided by an embodiment of the present application; Figure 9 for Figure 8 Schematic diagram of the audio magnetotelluric time series fragment sampled in the medium time window; Figure 10 Schematic diagram of audio magnetotelluric time series signals of different forms in the samples provided in the embodiments of the present application; Figure 11 2 is a schematic diagram of a comparative experiment on denoising a simulated audio magnetotelluric time series provided in an embodiment of the present application; Figure 12 2 is a schematic diagram of a comparative experiment on denoising actual audio magnetotelluric time series provided in an embodiment of the present application; Figure 13 The apparent resistivity phase curve and Nyquist diagram of the noisy high-frequency magnetotelluric time series provided in the embodiment of the present application are as follows; Figure 14 for Figure 13 Apparent resistivity phase curve and Nyquist plot under wavelet transform; Figure 15 for Figure 13Apparent resistivity phase curve and Nyquist plot under DDTF; Figure 16 for Figure 13 Apparent resistivity phase curve and Nyquist plot under CNN; Figure 17 for Figure 13 Apparent resistivity phase curve and Nyquist diagram under the method of this embodiment; Figure 18 Schematic diagram of the structure of the human noise removal system combining neural network and wavelet transform provided in an embodiment of the present application; Figure 19 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0021] In the description of this application, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0022] In the description of this application, it should be understood that descriptions involving orientation, such as the orientation or positional relationship indicated by up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and function in a specific orientation, and therefore cannot be understood as a limitation on this application.
[0023] Before introducing the embodiments, the basic concepts and current research in this field are introduced: Audio magnetotellurics (AMT) plays a vital role in studying the electrical structure of the Earth's interior and is widely used in the exploration of various metal minerals. However, AMT signals based on natural sources are weak and susceptible to interference from anthropogenic noise. As with other geophysical exploration methods, the reliability of observational data is fundamental to accurate results. Due to the development of industry, the rapid expansion of various industrial power grids, and urban agglomerations, AMT practitioners are encountering an increasing number of artificial sources during field exploration, making obtaining reliable AMT data increasingly challenging. This is particularly true in deep exploration, where anthropogenic noise has been a persistent problem for AMT practitioners for over a decade. With the advancement of three-dimensional electromagnetic inversion over the past decade, the reliability of observational data has become a key bottleneck in the overall AMT exploration method. Therefore, removing anthropogenic noise and ensuring the reliability of AMT data is crucial for the precise exploration of various mineral resources.
[0024] 1. In recent decades, scholars have proposed a variety of denoising methods to ensure the reliability of observation data. One type of method does not directly remove the human noise in the time series, but suppresses the impact of noise on impedance through other methods, such as using remote referencing, impedance estimation, inversion and other methods to suppress human noise.
[0025] Remote referencing is one of the earliest methods used to suppress human-induced noise. It utilizes the different distribution characteristics of measurement point noise and reference station noise, and suppresses human-induced noise by introducing the magnetic track of the reference station to participate in impedance estimation.
[0026] Robust estimation is also commonly used in AMT denoising. This method assigns weights to different data segments through residual analysis, reducing the impact of outliers and improving noise immunity. Examples include least squares estimation, M-regression estimation, bounded influence estimation, repeated median estimation, multi-station robust estimation, and tensor impedance estimation. Currently, a time-domain impedance estimation method using Bayesian estimation can effectively reduce error bars at lower frequencies, improving the effectiveness of suppressing human-induced noise.
[0027] The inversion-based Dplus and Rhoplus can also effectively suppress the influence of human noise on the apparent resistivity and phase curves.
[0028] Although such methods have been developed for decades in AMT denoising, they often have various limitations in their use. For example, the denoising effect of far-reference denoising is related to the selection of reference points. However, with the trend of industrialization, good reference points are becoming increasingly difficult to find. Robust estimation has high requirements for the length of high-quality time series, and inversion requires that the number of noisy frequency points should not be too large.
[0029] 2. Directly removing human noise from time series with the help of various signal processing methods is also a development trend of AMT denoising, such as using time period selection, mathematical decomposition, sparse representation and other methods to remove human noise.
[0030] Time period selection can effectively remove time series segments with obvious noise interference, and improve data quality by manually removing noisy segments through experience.
[0031] Mathematical decomposition extracts human noise from time series based on shape structure, time scale, and frequency characteristics through preset decomposition steps, such as mathematical morphological filtering, empirical mode decomposition, ensemble empirical mode decomposition, and variational mode decomposition (VMD).
[0032] Sparse representation utilizes the different distribution characteristics of human noise and signals in the sparse domain to remove human noise through thresholds, such as Fourier transform, wavelet transform, K-th singular value decomposition, data-driven tight framework (DDTF), etc.
[0033] Although these digital signal processing methods acting on time series are quick to implement and have no usage restrictions, they are essentially filtering of time series, which can damage the AMT signal when removing human interference. Therefore, it is difficult to obtain reliable denoising results under strong human interference.
[0034] These traditional AMT denoising methods are limited in their usage conditions and denoising performance, making them incapable of dealing with strong human-induced interference. However, neural networks, when equipped with a rich sample set and sufficient network depth, use optimization algorithms to train the network and find a mapping relationship from samples to labels. This approach is expected to overcome the limitations of traditional methods and remove strong human-induced interference at a deeper level. However, the neural networks currently used for AMT denoising largely draw on denoising processes from other signal processing fields and do not consider the characteristics of human-induced noise during the denoising process. This makes it difficult for neural networks to obtain reliable denoising results even with strong human-induced interference.
[0035] In summary, traditional denoising methods (such as remote referencing, impedance estimation, inversion, time period selection, mathematical decomposition, sparse representation, etc.) are difficult to achieve high-precision AMT denoising under strong human interference due to limitations in usage conditions and denoising performance. Neural networks also have difficulty dealing with strong human interference because they ignore the characteristics of human noise.
[0036] Human-induced noise not only exhibits temporal characteristics in time series but also often affects both the electrical and magnetic tracks of the AMT, exhibiting correlation characteristics. From the perspective of sparse representation, human-induced noise is concentrated in large sparse coefficients, also exhibiting sparsity characteristics. Therefore, designing a neural network denoising algorithm based on the correlation, temporal characteristics, and sparsity of human-induced noise is key to achieving high-precision AMT denoising under strong human-induced interference.
[0037] So, if Figure 1 and Figure 2 One embodiment of the present application provides a method for removing human noise by combining a neural network and a wavelet transform, the method comprising steps S100 to S300: Step S100: obtaining a target audio magnetotelluric time series signal from which human-induced noise is to be removed.
[0038] Step S200 : converting the target audio magnetotelluric time series signal into a target two-dimensional time series graph based on Gram angle and field transformation.
[0039] In step S300, the target two-dimensional time series graph is input into the denoising model, so as to decompose the sub-band features of multiple frequency bands from the target two-dimensional time series graph based on the two-dimensional wavelet forward transform according to the denoising model, extract the corresponding multiple sub-band intermediate features from the multiple sub-band features based on the two-dimensional convolutional neural network, and fuse the multiple sub-band intermediate features based on the two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human noise.
[0040] In this step S100, the audio magnetotelluric method plays an important role in the study of the electrical structure of the earth's interior and is widely used in the exploration of various metal minerals. In the exploration of metal minerals, the target audio magnetotelluric time series signal in step S100 can be collected by a relevant collection device, and then the human noise in the target audio magnetotelluric time series signal is removed according to the subsequent denoising model.
[0041] The time series signal is a one-dimensional signal. Expanding the one-dimensional signal into a two-dimensional time series graph in step S200 can increase spatial information (the second dimension), making the shape of the noise more prominent, which is conducive to the denoising model capturing this noise information during the training process.
[0042] In some embodiments, methods for expanding a one-dimensional timing signal into a two-dimensional timing graph include but are not limited to: Gramian Angular Summation Field (GASF), Short Time Fourier Transform (STFT), and Hilbert-Huang Transform (HHT).
[0043] In step S300, a denoising model is constructed based on a neural network, which has: (1) Two-dimensional wavelet forward transform is mainly used to extract subbands of different frequency bands from the input data, such as approximate subbands, horizontal detail subbands, vertical detail subbands, and diagonal detail subbands. Since the input data is a two-dimensional time series graph, a two-dimensional wavelet transform is used here.
[0044] (2) Two-dimensional convolutional neural network.
[0045] (3) Two-dimensional wavelet inverse transform is mainly used to fuse sub-bands of different frequency bands. There is a corresponding relationship between the two-dimensional wavelet inverse transform and the two-dimensional wavelet forward transform.
[0046] After the original signal is mapped to a two-dimensional time-series image, a multi-level wavelet decomposition is used to separate high-frequency noise from low-frequency valid signals layer by layer. The sub-band features generated at each decomposition level are then processed through a two-dimensional convolutional neural network to extract key features and suppress noise-related components. During the reconstruction phase, the processed features of each frequency band are fused layer by layer through an inverse transform, ultimately restoring the target audio magnetotelluric time-series signal while removing interfering artifacts.
[0047] This method converts the one-dimensional target audio magnetotelluric time series signal into a two-dimensional target two-dimensional time series graph, adds two-dimensional spatial information to the target audio magnetotelluric time series signal, makes the shape of the human noise more prominent, and is conducive to the denoising model to capture these human noises; then, the denoising model decomposes multiple sub-band features of different frequency bands from the target two-dimensional time series graph based on the two-dimensional wavelet forward transform, and extracts the corresponding multiple sub-band intermediate features from the multiple sub-band features based on the two-dimensional convolutional neural network, and fuses the multiple sub-band intermediate features based on the two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human noise. Here, a two-dimensional wavelet transform and a neural network are combined to decompose the input target two-dimensional time series graph into wavelet sparse sub-band features of different frequencies. In the denoising model, the neural network can analyze the spectral information of the human noise in different frequency bands in the form of wavelet sparse coefficients, so that the denoising model can mine the sparse characteristics of the human noise, thereby improving the high-precision denoising accuracy of the denoising model under strong human noise interference.
[0048] like Figure 2 and Figure 3 ,Furthermore, the denoising model consists of a two-dimensional convolutional neural network, n-times two-dimensional wavelet forward transform and n-times two-dimensional wavelet inverse transform;,n,is a positive integer greater than 1; Step S300 decomposes a plurality of sub-band features of different frequency bands from the target two-dimensional time series graph based on a two-dimensional wavelet forward transform, extracts a plurality of corresponding sub-band intermediate features from the plurality of sub-band features based on a two-dimensional convolutional neural network, and fuses the plurality of sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out human-induced noise. The method includes the following steps S310 to S370: Step S310 , decomposing the target two-dimensional time series graph into m sub-band features of different frequency bands according to the first two-dimensional wavelet forward transform; m is 4.
[0049] Step S320: extracting corresponding m first sub-band intermediate features from the m sub-band features according to a two-dimensional convolutional neural network.
[0050] Step S330 : Decompose the m first sub-band intermediate features into corresponding m second sub-band features according to the second two-dimensional wavelet forward transform.
[0051] Step S330, according to the two-dimensional convolutional neural network from From the second sub-band features, the corresponding The second sub-band intermediate features.
[0052] Step S340, and so on, until the two-dimensional convolutional neural network is obtained. From the nth sub-band features, extract the corresponding The nth sub-band intermediate features.
[0053] Step S350, according to the first two-dimensional wavelet inverse transform Each m corresponding n-th sub-band intermediate features in the n-th sub-band intermediate features are fused to obtain The n+1th sub-band features.
[0054] Step S360, according to the two-dimensional convolutional neural network from From the n+1th sub-band features, extract the corresponding The intermediate features of the n+1th sub-band.
[0055] Step S370, and so on, until obtaining m 2n-th sub-band intermediate features extracted from m 2n-th sub-band features according to the two-dimensional convolutional neural network, and fusing the m 2n-th sub-band intermediate features according to the n-th two-dimensional wavelet inverse transform, to obtain the target audio magnetotelluric time series signal after filtering out the human noise.
[0056] The model of this embodiment utilizes the sparsity characteristics of human noise, regards the wavelet forward transform function as a downsampling function to replace the pooling layer, regards the wavelet inverse transform function as an upsampling function to replace the linear interpolation, and uses the wavelet forward and inverse transform functions to connect the two-dimensional convolutional neural network, so that the neural network can complete the efficient screening of noise information in the form of wavelet sparse coefficients during training.
[0057] like Figure 2 The model executes in the following order: forward wavelet transform - 2D convolutional neural network (rectangular box in the figure) - forward wavelet transform - 2D convolutional neural network - forward wavelet transform - 2D convolutional neural network - ... - 2D convolutional neural network - inverse wavelet transform - ... - 2D convolutional neural network - inverse wavelet transform. Each subband requires a corresponding 2D convolutional neural network to extract the characteristics of human noise.
[0058] After the first wavelet forward transform is performed on the target two-dimensional time series graph, the target two-dimensional time series graph is decomposed (downsampled) into approximate sub-bands, horizontal detail sub-bands, vertical detail sub-bands and diagonal detail sub-bands of different frequency bands. Each sub-band is then input into a corresponding two-dimensional convolutional neural network to extract the corresponding features (here named sub-band intermediate features). Then, a second wavelet forward transform is performed to downsample the sub-bands of each frequency band. The wavelet forward transform function is performed n times, and the wavelet inverse transform function is performed n times. After each wavelet forward transform function, it is divided into 4 sub-band features, and so on. The wavelet inverse transform function here is the inverse process of the wavelet forward transform function, which will not be described in detail here.
[0059] For example, the input target two-dimensional time series graph is decomposed into four frequency sub-bands by two-dimensional wavelet forward transform , and then each sub-band is input into the two-dimensional convolutional neural network to obtain the corresponding sub-band features. Then each sub-band feature is decomposed into four frequency sub-bands by two-dimensional wavelet forward transform, for example The corresponding sub-band features are decomposed into: , and so on, Figure 2 The process has been fully demonstrated and will not be described in detail here.
[0060] Through continuous wavelet decomposition, the denoising model analyzes noise information across different frequency bands at multiple scales. Compared to the existing pooling layer of two-dimensional neural networks, which simply retains the maximum value in a region and discards smaller values in downsampling, this model uses a forward wavelet transform to better capture the spectral information of human-induced noise. Similarly, compared to conventional neural network upsampling methods such as linear interpolation, the use of an inverse wavelet transform for upsampling subbands in each frequency band allows network output to be obtained without losing noise information, effectively transferring human-induced noise information within the two-dimensional neural network, thereby achieving high-precision denoising.
[0061] Furthermore, the denoising model in step S300 is trained in the following manner: Step S510, obtaining a time series sample; the time series sample is an audio magnetotelluric time series signal obtained by superimposing human noise on a noise-free audio magnetotelluric time series, and the label of the time series sample is an audio magnetotelluric time series signal without superimposed human noise.
[0062] Step S520: convert the time series samples into two-dimensional time series graph samples.
[0063] Step S530 : Input the two-dimensional time series graph samples and their labels into the denoising model for training until a trained denoising model is obtained.
[0064] In order to enable the denoising model to obtain the noise-free target audio magnetotelluric time series signal, in the actual process of collecting time series samples, by superimposing human noise on the noise-free audio magnetotelluric time series, and setting the label of the time series sample to the audio magnetotelluric time series signal without superimposed human noise, through denoising model mapping training, the trained denoising model can obtain noise-free output.
[0065] While traditional methods rely on collecting noisy data in real environments, this method generates a large number of controllable training samples by superimposing noise. This method enhances the expressive power of temporal features through two-dimensional image transformation, while combining the multi-scale decomposition of wavelet transforms allows the model to more effectively separate noise and valid signal frequency bands.
[0066] This method constructs a training dataset with clear noise distribution characteristics, improving the model's ability to identify strong artifacts. The combination of two-dimensional image transformation and multi-level wavelet decomposition enables the two-dimensional neural network to gradually extract and fuse features from different frequency bands, thereby accurately removing complex noise patterns. This significantly improves the stability of the denoising results, especially in low signal-to-noise ratio environments.
[0067] Furthermore, obtaining time series samples in step S510 includes the following steps S5110 to S5150: Step S5110: collecting actual audio magnetotelluric time series containing human noise.
[0068] Step S5120: Generate a noise-free simulated audio magnetotelluric time series based on the forward simulation; wherein the simulated audio magnetotelluric time series and the actual audio magnetotelluric time series have the same number of electromagnetic channels.
[0069] Step S5130: Sampling the actual audio magnetotelluric time series using a sliding window to obtain an audio magnetotelluric time series containing human noise; wherein the step length of the sampling sliding window is smaller than the length of the sliding window.
[0070] Step S5140 : Sampling from the simulated audio magnetotelluric time series using a sliding window with a step size of the sliding window to obtain a noise-free audio magnetotelluric time series.
[0071] Step S5150 , extracting human-induced noise from the audio magnetotelluric time series containing human-induced noise, and superimposing the extracted human-induced noise onto the noise-free audio magnetotelluric time series to obtain a time series sample.
[0072] In steps S5130 and S5140, a sliding window is used to sample the actual audio magnetotelluric time series and the simulated audio magnetotelluric time series to obtain an audio magnetotelluric time series containing human-induced noise and a noise-free audio magnetotelluric time series. The sliding window slides from left to right within the actual audio magnetotelluric time series or the simulated audio magnetotelluric time series with a step size less than the window length, simultaneously sampling the electrical and magnetic tracks from both sequences. This ensures that the samples contain information about the electrical and magnetic tracks collected within the same time period, thereby maintaining the correlation characteristics of the human-induced noise and ensuring that the previous and next data segments overlap, thus preserving the temporal characteristics of the human-induced noise.
[0073] In step S5140, after the sliding window cutting is completed, the adaptive sparse transform method (Data-Driven Tight Frame, DDTF) is used to extract the human noise from the audio magnetotelluric time series containing human noise, and then superimpose it on the noise-free audio magnetotelluric time series to preliminarily form a sample label pair.
[0074] The method provided in this embodiment simultaneously scans the electrical and magnetic tracks of an audio magnetotelluric time series to produce multi-component samples, extracts the correlation characteristics of human-induced noise, and maintains the temporal characteristics while ensuring overlap between the previous and next data segments, thereby improving the denoising capability of the subsequent model.
[0075] Furthermore, the step S520 of converting the time series samples into two-dimensional time series graph samples based on the Gram angle and field transformation includes the following steps S5210 to S5230: Step S5210: normalize and scale the time series samples.
[0076] Step S5220: Map the normalized and scaled time series samples to polar coordinates.
[0077] Step S5230: Based on the Gram angle and field transformation, the time series samples mapped to the polar coordinates are converted into two-dimensional time series graph samples.
[0078] In this embodiment, the GASF transformation expands one-dimensional time series samples into two dimensions, adding spatial information and making the noise more prominent, which helps the two-dimensional neural network capture this noise information during training. Furthermore, this transformation is performed through polar coordinate conversion based on time relationships, without destroying the original temporal relationship of the time series samples.
[0079] Furthermore, the conversion of the target audio magnetotelluric time sequence signal into a target two-dimensional time sequence diagram in step S200 includes steps S210 to S230: Step S210 , normalizing and scaling the target audio magnetotelluric time series signal.
[0080] Step S220 , mapping the normalized and scaled target audio magnetotelluric time series signal to polar coordinates.
[0081] Step S230 : Based on the Gram angle and field transformation, the target audio magnetotelluric time series signal mapped to the polar coordinates is converted into a target two-dimensional time series graph.
[0082] Similar to the above embodiment, this embodiment expands the one-dimensional target audio magnetotelluric time series signal into a two-dimensional target two-dimensional time series diagram, which adds spatial information and makes the noise shape more prominent, which is conducive to the neural network capturing these noise information and improving the accuracy of model denoising.
[0083] Furthermore, the 2D convolutional neural network includes convolution blocks, normalization and ELU activation functions, such as Figure 3 By increasing the number of convolution blocks and normalization, the depth of the two-dimensional convolutional neural network can be continuously increased.
[0084] Among them, the convolution block refers to a module composed of multiple convolution layers, which can be implemented by a stacked structure containing 3×3 convolution kernels to extract correlation information from sub-band features. Normalization refers to the standardization of the convolution output, which can be implemented by a batch normalization layer to adjust the data distribution to accelerate model convergence. The ELU activation function refers to the exponential linear unit function, which can be implemented by α The activation function form of the parameters is used to introduce nonlinear transformations and alleviate the gradient disappearance problem.
[0085] Specifically, this method optimizes the stability of feature distributions by introducing a normalization layer. Combined with the nonlinear response of the ELU function in the negative range, this method enhances the model's ability to capture weak noise features in low signal-to-noise ratio signals. This effectively improves the accuracy of two-dimensional convolutional neural networks in resolving multi-scale noise features, while preserving the effective components of the original signal while achieving targeted suppression of strong human interference. The combined use of normalization and ELU further enhances the stability of model training, avoiding the vanishing gradient problem caused by activation function saturation, enabling the denoising model to maintain reliable denoising performance even in complex electromagnetic environments.
[0086] like Figures 2 to 17 Traditional denoising methods are difficult to achieve high-precision audio magnetotelluric time series denoising under strong human interference due to limitations in usage conditions and denoising performance. Neural networks are still unable to cope with strong human interference because they ignore the characteristics of human noise. As for human noise, it is not only located in the time series and reflects the temporal characteristics, but also often affects the electric and magnetic tracks of the audio magnetotelluric time series at the same time, reflecting the correlation characteristics. From the perspective of sparse representation, human noise is concentrated in sparse coefficients with large values, and also reflects the sparsity characteristics. Therefore, how to design a neural network denoising algorithm based on the correlation, temporal and sparsity characteristics of human noise is the key to achieving high-precision audio magnetotelluric time series denoising under strong human interference.
[0087] Therefore, one embodiment of the present application provides a method for removing human noise by combining a neural network and a wavelet transform. The method includes the following steps: Step S910: creating a sample set.
[0088] Before preparing the sample set, actual audio magnetotelluric time series with low noise and strong human interference were selected from the survey area, and noise-free simulated audio magnetotelluric time series were generated by forward simulation. Then, the audio magnetotelluric time series were cut using a sliding time window to prepare samples.
[0089] like Figures 6 to 9 The process of making samples by sliding window cutting audio magnetotelluric time series is shown. Figure 6 and Figure 8 represent the audio magnetotelluric time series with and without noise, respectively. Figure 7 and Figure 9 Respectively Figure 6 and Figure 8 The data segments extracted within the time window. Figures 6 to 9 They all include 5 parts, namely (a) to (e), Figure 6 For example, Figure 6 Part (a) is the electric path (E X ), Figure 6Part (b) is the electric path (E Y ), where the electric path (E X ) and the electric path (E Y ) in different directions, Figure 6 Part (c) is the track (H X ), Figure 6 The (d) part is the track (H Y ), Figure 6 The (e) part is the track (H Z ), where the unit for the electrical track is millivolts per kilometer (mV / km) and the unit for the magnetic track is nanotesla (nT). The sliding window is slid from left to right with a step size smaller than the sliding window length, simultaneously scanning the electrical and magnetic tracks of the audio magnetotelluric time series. The multi-component data segments within the sliding window are sequentially extracted. This approach ensures that the sample contains information from the electrical and magnetic tracks collected within the same time period, thus preserving the correlation characteristics of the human-induced noise. It also ensures that the previous and next data segments overlap, maintaining the temporal characteristics of the human-induced noise.
[0090] After completing the sliding window cutting, adaptive sparse transform is used to extract the human noise in the noisy data segment and superimpose it on the noise-free data segment to preliminarily form a sample-label pair, where the sample is the noise-free data segment after superimposing the noise, and the label is the noise-free data segment without superimposing the noise.
[0091] Step S920, Gram angle and field transformation; Gram angle and field transforms are then used as feature engineering techniques on the samples to enhance the morphology, amplitude, and spectral information of the human-induced noise in the samples. Gram angle and field transforms can increase the amount of information in the samples while preserving the temporal characteristics, making it easier for the neural network to capture the human-induced noise information during training. It first normalizes and scales the samples: (1); In formula (1), and Represent the samples before and after normalization scaling, Indicates the maximum value operation, and then uses the arc cosine function to Mapping to polar coordinates: (2); In formula (2), Represents samples on polar coordinates, for each sample point on polar coordinates The GASF transformation can be completed by cosine superposition to obtain a two-dimensional timing diagram : (3); Figure 10The representation of different forms of audio magnetotelluric time series signals in the sample (noise-free, harmonic interference, square wave interference, charge and discharge triangle wave interference) at each coordinate is shown. Figure 10 Part (a) to Figure 10 Part (d) in the equation is a Cartesian coordinate in different forms. Figure 10 Part (e) to Figure 10 The (h) part is the polar coordinate form of different forms. Figure 10 Part (i) to Figure 10 Part (l) in the figure shows different forms of Gram angle and field transform. It can be seen that Gram angle and field transform are very sensitive to different data forms. Based on the previously formed sample label pairs, after performing Gram angle and field transform on each sample, the final sample label pair can be obtained.
[0092] Step S930, two-dimensional wavelet transform; In the denoising of audio magnetotelluric time series, human-induced noise is concentrated in large sparse coefficients, reflecting sparsity characteristics. Therefore, a wavelet transform is combined with a Convolutional Neural Network (CNN). During network training, the denoising model captures the sparsity characteristics of human-induced noise in the form of wavelet sparse coefficients. Since the input to the network is a two-dimensional time series graph, a two-dimensional wavelet transform is used. Its forward transform expression is: (4); In formula (4), , , and Represent the wavelet sparse coefficients on the approximate subband, horizontal detail subband, vertical detail subband and diagonal detail subband respectively, and is a spatial variable, and denote the scale and translation basis functions, respectively. is the scale parameter, and is the translation parameter, Represents the inner product operation. The corresponding two-dimensional wavelet inverse transform is obtained by dual scaling basis function and dual translation basis functions accomplish: ; Step S940: Obtain a denoising model through CNN extended by two-dimensional wavelet transform.
[0093] The denoising model used is based on a two-dimensional convolutional neural network as the main structure and is expanded through wavelet transform. The model takes advantage of the sparsity characteristics of human noise, regards the wavelet forward transform as a downsampling replacement of the pooling layer, and the wavelet inverse transform as an upsampling replacement of linear interpolation. The wavelet forward and inverse transforms are used to connect each two-dimensional convolutional neural network, so that the model can complete the efficient screening of noise information in the form of wavelet sparse coefficients during training.
[0094] like Figure 3 ,The 2D convolutional neural network mainly consists of convolution blocks, ,normalization layers, and ELU activation functions.
[0095] After the first wavelet forward transform is performed on the input information, it is decomposed into approximate subbands, horizontal detail subbands, vertical detail subbands, and diagonal detail subbands of different frequency bands. A second wavelet forward transform is then performed to downsample the subbands of each frequency band. This continuous wavelet decomposition allows the neural network to analyze noise information in different frequency bands at multiple scales. Compared to the existing pooling layer of the neural network, the wavelet forward transform can better capture the spectral information of human-induced noise. Similarly, compared to conventional neural network upsampling methods such as linear interpolation, upsampling each frequency subband using an inverse wavelet transform can obtain network output without losing noise information, effectively transferring human-induced noise information within the model and achieving high-precision denoising. Finally, the previously prepared sample set is input into the model for training. Once the network's loss function converges, it can be used to denoise audio magnetotelluric time series.
[0096] Step S950: Use of the model.
[0097] The experimental data of this embodiment are provided below: (1) Simulated data experiments; In order to verify the effectiveness of this method, a simulated data experiment was conducted to compare the denoising performance of wavelet transform, DDTF, CNN and this embodiment. The signal-to-noise ratio (SNR) was used as the evaluation criterion in the simulated data experiment. The experimental results of the simulated data are shown in Figure 2. Figure 11 As shown in Table 1, Figure 11 Part (a) is the noise-free audio magnetotelluric time series. Figure 11 Part (b) shows the result of adding charging and discharging triangle wave interference to the noise-free audio magnetotelluric time series. Figure 11 Part (c) shows the denoising results of wavelet transform, DDTF, CNN, and this embodiment from top to bottom. Figure 11 Part (d) in is the corresponding denoising residual. Figure 11As can be seen from Table 1, among the comparison methods, the denoising performance of the wavelet transform is the worst. Not only does its denoising result have obvious residual human noise, but it also filters out the signal most severely, and its SNR is also the lowest among the comparison methods. The denoising performance of DDTF is better than that of the wavelet transform, but worse than the deep learning method, and it has difficulty dealing with the sudden changes in human noise. Both CNN and this embodiment have achieved good denoising results and can remove human noise very well. Compared with CNN, the SNR after denoising in this embodiment is higher, and the signal filtering is also smaller. Therefore, among the comparison methods, the denoising performance of this embodiment is the best.
[0098] Table 1
[0099] (2) Experiments with real data; In the actual data experiment, the denoising performance of wavelet transform, DDTF, CNN and this embodiment is further compared. The actual audio magnetotelluric time series was collected in Lusong, Anhui Province. One of the time series segments was selected for denoising comparison. The sampling rate of the data was 150. Figure 12 shown. Figure 12 Part (a) is the noisy audio magnetotelluric time series. Figure 12 Part (b) is the spectrum of the noisy audio magnetotelluric time series. Figure 12 Part (c) shows the denoising results of wavelet transform, DDTF, CNN, and this embodiment from top to bottom. Figure 12 Part (d) in the figure is the corresponding denoising result spectrum diagram. It can be seen that the noisy audio magnetotelluric time series fragment is severely interfered by the charging and discharging triangle wave, and the wavelet transform has difficulty in coping with the charging and discharging triangle wave interference. The noise spectrum can be clearly seen in the denoising result spectrum diagram. Although there is no obvious residual human noise in the DDTF denoising result, it can be seen from the spectrum diagram that it filters out the effective signal very seriously, especially the low-frequency part. The denoising results of CNN and this embodiment are relatively good, but through the comparison of the spectrum diagrams, it is found that there is residual noise interference in CNN, such as the 30-40 Hz part, and the denoising performance of this embodiment is better than the three comparison methods. While protecting the effective signal well, there is no obvious residual human noise.
[0100] In order to better evaluate the denoising results of the actual data, in addition to the apparent resistivity, phase curve and spectrum, the Nyquist diagram is also used for evaluation. The traditional data evaluation method evaluates the apparent resistivity and phase separately, ignoring the dispersion relationship between the real and imaginary parts of the impedance (i.e., the Hilbert transform relationship). Therefore, it is proposed to use the Nyquist diagram as a tool to show the inherent characteristics of the complex response of the AMT to measure the quality of the AMT measuring point. The Nyquist diagram uses the real part of the impedance as the horizontal axis and the imaginary part as the vertical axis. By plotting data points at different frequencies, it intuitively reflects the dispersion relationship of the impedance. That is to say, for a good quality measuring point, its apparent resistivity and phase curves should not only be smooth and continuous without drastic fluctuations, but also its impedance projection point should have a clockwise trend in the Nyquist diagram. Compared with the traditional evaluation method that uses the AMT polarization direction to evaluate the denoising quality of each frequency point, the use of the Nyquist diagram can judge the denoising quality of all frequency points at one time, which is more efficient. In the Nyquist diagram, the real part of the impedance in the Nyquist diagram is and the imaginary part The expression is: (5); In the above formula (5), represents the Cauchy principal value operation, 、 represents different frequency variables, Represents a constant matrix, and the expression is: (6); in, Compared with the formula (5) and The relationship between them is: (7); in Represents a constant matrix The real coefficients of the matrix, Represents a constant matrix The imaginary coefficients of the matrix, not in italics Represents an imaginary number.
[0101] in, represents the constant coefficient matrix. Introduced in formula (5) and Respectively The corresponding constant parameters and The corresponding constant parameters are expressed as follows: (8); in, represents different arbitrary real constants.
[0102] Based on the existing technology, this method further derives the formula (5) of the Nyquist diagram and applies it to AMT denoising for the first time.
[0103] Reference Figures 13 to 17 ,in Figure 13 Part (a) is the apparent resistivity curve of the magnetotelluric time series with high noise frequency. Figure 13 Part (b) is the phase curve of the magnetotelluric time series with high noise frequency. Figure 13 Parts (c) and (d) are Nyquist plots of the noisy high-frequency magnetotelluric time series in different directions (where Re(Z xy ) and Im(Z yx ) represent the real and imaginary parts of the impedance in different directions respectively). Figure 14 Parts (a) to (d) are Figure 13 Schematic diagram of parts (a) to (d) after wavelet transform denoising; Figure 15 Parts (a) to (d) are Figure 13 Schematic diagram of parts (a) to (d) after DDTF denoising; Figure 16 Parts (a) to (d) are Figure 13 Schematic diagram of parts (a) to (d) after CNN denoising; Figure 17 Parts (a) to (d) are Figure 13 Parts (a) to (d) in the figure are schematic diagrams of the denoising method of this embodiment; it can be seen from the figure that the measuring point is seriously interfered with by human noise, and there are obvious discontinuities and jumps in the apparent resistivity and phase curves, and the Nyquist plot does not show a clockwise trend. After denoising, the noise interference of the measuring point has been improved to varying degrees. Among them, the denoising result of the wavelet transform is the worst. The apparent resistivity and phase curves still have large distortion in the medium and low frequency bands, and its Nyquist plots in the medium and low frequency bands are also very disordered. The denoising result of DDTF is better than that of the wavelet transform. It only has some distorted frequencies in the apparent resistivity and phase curves in the yx direction, and its Nyquist plot can also roughly show a clockwise trend. The denoising result of CNN is better than the above two traditional denoising methods. Not only are there fewer frequency points of noise interference in the apparent resistivity and phase curves than in the wavelet transform and DDTF, but the disorder in the Nyquist plot is also less than that of these two methods. The denoising result of this embodiment performs best among the comparison methods. Not only is the distortion of the apparent resistivity and phase curves affected by human noise basically removed, but the Nyquist plot also shows a clear clockwise trend.
[0104] This embodiment has the following advantages: 1) In creating the sample set, we used a sliding time window and Gram's angle and field transform to simultaneously scan the electrical and magnetic tracks of the audio magnetotelluric time series to create multi-component samples. This allowed us to extract the correlation characteristics of human-induced noise while ensuring overlap between the previous and next data segments and maintaining the temporal characteristics. Finally, we used the Gram's angle and field transform to increase the amount of human-induced noise information in the samples while preserving the temporal characteristics, allowing the model to pay more attention to human-induced noise during training. 2) In the model design, CNN is used as the main framework, and the two-dimensional wavelet transform is used to expand the neural network. The sparse characteristics of human noise are extracted in the form of wavelet sparse coefficients. By decomposing the input information into high-frequency, medium-high-frequency, medium-low-frequency, and low-frequency wavelet sparse coefficients during model training, the potential noise spectrum information can be explored and the denoising accuracy can be improved.
[0105] like Figure 18 One embodiment of the present application provides a human noise removal system combining a neural network and a wavelet transform. The system includes a signal acquisition module, a two-dimensional graph generation module, and a noise removal module. The signal acquisition module 1100 is used to obtain the target audio magnetotelluric time series signal from which human noise is to be removed; The two-dimensional graph generation module 1200 is used to convert the target audio magnetotelluric time series signal into a target two-dimensional time series graph based on the Gram angle and field transformation; The noise removal module 1300 is used to input the target two-dimensional time series diagram into the denoising model, so as to decompose the sub-band features of multiple frequency bands from the target two-dimensional time series diagram based on the two-dimensional wavelet forward transform according to the denoising model, and extract the corresponding multiple sub-band intermediate features from the multiple sub-band features based on the two-dimensional convolutional neural network, and fuse the multiple sub-band intermediate features based on the two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human noise.
[0106] It should be noted that the embodiment of the human noise removal system combining a neural network and a wavelet transform and the embodiment of the human noise removal method combining a neural network and a wavelet transform are based on the same inventive concept. Therefore, the relevant contents of the embodiment of the human noise removal method combining a neural network and a wavelet transform are also applicable to the embodiment of the human noise removal system combining a neural network and a wavelet transform, and will not be repeated here.
[0107] The system provided in this embodiment converts a one-dimensional target audio magnetotelluric time series signal into a two-dimensional target two-dimensional time series graph, adds two-dimensional spatial information to the target audio magnetotelluric time series signal, makes the shape of the human-induced noise more prominent, and facilitates the denoising model to capture these human-induced noises; then, the denoising model decomposes multiple sub-band features of different frequency bands from the target two-dimensional time series graph based on a two-dimensional wavelet forward transform, extracts multiple corresponding sub-band intermediate features from the multiple sub-band features based on a two-dimensional convolutional neural network, and fuses the multiple sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human-induced noise. Here, a two-dimensional wavelet transform and a two-dimensional convolutional neural network are combined to decompose the input target two-dimensional time series graph into wavelet sparse sub-band features of different frequencies. In the denoising model, the two-dimensional convolutional neural network can analyze the spectral information of the human-induced noise in different frequency bands in the form of wavelet sparse coefficients, so that the denoising model can explore the sparse characteristics of the human-induced noise, thereby improving the high-precision denoising accuracy of the denoising model under strong human-induced noise interference.
[0108] Reference Figure 19 , an embodiment of the present application further provides an electronic device, the electronic device comprising: at least one memory; at least one processor; at least one program; The programs are stored in the memory, and the processor executes at least one program to implement the human noise removal method combining the neural network and wavelet transform described above.
[0109] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.
[0110] The electronic device according to the embodiment of the present application is described in detail below.
[0111] The processor 1600 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 can store function systems and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the program code is stored in the memory 1700 and is called by the processor 1600 to execute the human noise removal method combining a neural network and a wavelet transform in the embodiments of the present application.
[0112] Input / output interface 1800, used for information input and output; Communication interface 1900, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.); Bus 2000 , which transmits information between various components of the device (e.g., processor 1600 , memory 1700 , input / output interface 1800 , and communication interface 1900 ); The processor 1600 , the memory 1700 , the input / output interface 1800 , and the communication interface 1900 are connected to each other in communication within the device via the bus 2000 .
[0113] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned human noise removal method combining neural network and wavelet transform.
[0114] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0115] The embodiments described in this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0116] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0118] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0119] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0120] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0121] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0122] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0123] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0124] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.
[0125] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above-mentioned implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions without violating the spirit of the embodiments of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the embodiments of the present application.
Claims
1. A method for removing human noise by combining neural network and wavelet transform, characterized in that: The method comprises: Obtain the target audio magnetotelluric time series signal from which human noise is to be removed; Converting the target audio magnetotelluric time series signal into a target two-dimensional time series diagram based on Gram angle and field transformation; The target two-dimensional time series graph is input into a denoising model, so as to decompose a plurality of sub-band features of different frequency bands from the target two-dimensional time series graph based on a two-dimensional wavelet forward transform according to the denoising model, extract a plurality of corresponding sub-band intermediate features from the plurality of sub-band features based on a two-dimensional convolutional neural network, and fuse the plurality of sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human noise.
2. The method for removing human noise by combining neural network and wavelet transform according to claim 1, characterized in that: The denoising model is composed of the two-dimensional convolutional neural network, n-times the two-dimensional wavelet forward transform and n-times the two-dimensional wavelet inverse transform; n is a positive integer greater than 1; The method comprises: decomposing a plurality of sub-band features of different frequency bands from a target two-dimensional time series graph based on a two-dimensional wavelet forward transform, extracting a plurality of corresponding sub-band intermediate features from the plurality of sub-band features based on a two-dimensional convolutional neural network, and fusing the plurality of sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out human noise, including: Decomposing the target two-dimensional time series graph into m sub-band features of different frequency bands according to the first two-dimensional wavelet forward transform; m is 4; Extracting m corresponding first sub-band intermediate features from the m sub-band features according to the two-dimensional convolutional neural network; Decomposing the m first sub-band intermediate features into corresponding m second sub-band features according to the second two-dimensional wavelet forward transform; According to the two-dimensional convolutional neural network From the second sub-band features, the corresponding The second sub-band intermediate features; And so on, until we get the two-dimensional convolutional neural network from From the nth sub-band features, extract the corresponding nth sub-band intermediate features; According to the first two-dimensional wavelet inverse transform, Each m corresponding n-th sub-band intermediate features in the n-th sub-band intermediate features are fused to obtain n+1th sub-band features; According to the two-dimensional convolutional neural network From the n+1th sub-band features, extract the corresponding The n+1th sub-band intermediate features; And so on, until the m 2n-th sub-band intermediate features are extracted from the m 2n-th sub-band features according to the two-dimensional convolutional neural network, and the m 2n-th sub-band intermediate features are fused according to the n-th two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out the human noise.
3. The method for removing human noise by combining neural network and wavelet transform according to claim 2, characterized in that: The denoising model is trained in the following way: Acquire a time series sample; the time series sample is an audio magnetotelluric time series signal obtained by superimposing human noise on a noise-free audio magnetotelluric time series, and the label of the time series sample is the audio magnetotelluric time series signal without superimposed human noise; converting the time series samples into two-dimensional time series graph samples based on Gram angle and field transformation; The two-dimensional time series graph samples and their labels are input into the denoising model for training until the trained denoising model is obtained.
4. The method for removing human noise by combining neural network and wavelet transform according to claim 3, characterized in that: The obtaining of time series samples includes: Acquire real audio magnetotelluric time series containing human noise; Generating a noise-free simulated audio magnetotelluric time series based on forward simulation; wherein the simulated audio magnetotelluric time series and the actual audio magnetotelluric time series have the same number of electromagnetic channels; Sampling the actual audio magnetotelluric time series using a sliding time window to obtain an audio magnetotelluric time series containing human noise; wherein the step size of the sampling sliding time window is smaller than the length of the sliding time window; Sampling the simulated audio magnetotelluric time series using the sliding time window and the step size of the sliding time window to obtain a noise-free audio magnetotelluric time series; Human noise is extracted from the audio magnetotelluric time series containing human noise, and the extracted human noise is superimposed on the noise-free audio magnetotelluric time series to obtain a time series sample.
5. The method for removing human noise by combining neural network and wavelet transform according to claim 3, characterized in that: The converting of the time series samples into two-dimensional time series graph samples based on Gram angle and field transformation comprises: Normalizing and scaling the time series samples; Mapping the normalized and scaled time series samples onto polar coordinates; Converting the time series samples mapped onto polar coordinates into two-dimensional time series graph samples based on Gram angle and field transformation; The converting of the target audio magnetotelluric time sequence signal into a target two-dimensional time sequence diagram based on Gram angle and field transformation includes: Normalizing and scaling the target audio frequency magnetotelluric time series signal; Mapping the normalized and scaled target audio magnetotelluric time series signal to polar coordinates; Based on Gram angle and field transformation, the target audio frequency magnetotelluric time series signal mapped onto polar coordinates is converted into a target two-dimensional time series graph.
6. The method for removing human noise by combining neural network and wavelet transform according to claim 1, characterized in that: The two-dimensional convolutional neural network includes at least one convolution block, normalization and ELU activation function.
7. The method for removing human noise by combining neural network and wavelet transform according to claim 1, characterized in that: After obtaining the target audio magnetotelluric time series signal after filtering out human noise, the method further includes: Extract all frequency points from the target audio magnetotelluric time series signal after filtering out human noise; Calculating the degree of distortion of all the frequency points based on the Nyquist diagram; Evaluate the quality of the denoising model according to the distortion levels of all the frequency points; Wherein, the real part of the impedance in the Nyquist plot is and the imaginary part The expression is: ; ; ; ; in, 、 represents different frequency variables, represents a constant matrix, Represents a constant matrix The real coefficients of the matrix, Represents a constant matrix The imaginary coefficient of the matrix, represents an imaginary number, express A matrix with constant coefficients, represents the Cauchy principal value operation, express The corresponding constant parameter, express The corresponding constant parameter, Indicates the An arbitrary real constant.
8. A human noise removal system combining neural network and wavelet transform, characterized in that: The system comprises: A signal acquisition module is used to obtain the target audio magnetotelluric time series signal from which human noise is to be removed; a two-dimensional graph generating module, configured to convert the target audio frequency magnetotelluric time series signal into a target two-dimensional time series graph based on Gram angle and field transformation; The noise removal module is used to input the target two-dimensional time series graph into a denoising model, so as to decompose a plurality of sub-band features of different frequency bands from the target two-dimensional time series graph based on a two-dimensional wavelet forward transform according to the denoising model, extract a plurality of corresponding sub-band intermediate features from the plurality of sub-band features based on a two-dimensional convolutional neural network, and fuse the plurality of sub-band intermediate features based on a two-dimensional wavelet inverse transform to obtain the target audio magnetotelluric time series signal after filtering out human noise.
9. An electronic device, characterized in that: The method comprises at least one controller and a memory for communicating with the controller; the memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the at least one controller to enable the at least one controller to execute the method for removing human noise by combining a neural network and a wavelet transform as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method for removing human noise by combining a neural network and a wavelet transform as described in any one of claims 1 to 7.