Transient electromagnetic signal denoising method based on adaptive fuzzy optimization singular spectrum analysis
By combining CEEMD and PSO optimization with WFNC, the SSA reconstruction order is adaptively optimized, solving the denoising problem of transient electromagnetic signals in high-noise environments. This achieves effective separation of signal and noise, improving signal quality and analysis accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JIAOTONG UNIV
- Filing Date
- 2025-08-26
- Publication Date
- 2026-05-01
AI Technical Summary
In high-noise environments, denoising methods for transient electromagnetic signals struggle to effectively select the reconstruction order, leading to signal distortion or noise residue. This is especially true in urban environments where electromagnetic interference is severe, making it difficult for existing methods to effectively separate signals and noise.
After using complementary set empirical mode decomposition (CEEMD) and sample entropy to screen modes, the dominant modes of signal and noise are determined by particle swarm optimization (PSO). Combined with weighted fuzzy noise clustering (WFNC) algorithm, the order of reconstructed by adaptive optimization singular spectral analysis (SSA) is determined by fuzzy entropy and singular value difference spectrum, so as to achieve accurate separation of signal and noise.
It effectively removes transient electromagnetic signal distortion in high-noise environments, improves the signal-to-noise ratio, ensures the preservation of signal characteristics, reduces noise residue, and enhances the accuracy and reliability of signal analysis.
Smart Images

Figure CN121958747A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of signal processing technology in geophysical exploration, and in particular to a method for denoising transient electromagnetic (TEM) signals based on adaptive fuzzy optimization singular spectrum analysis (SSA). Background Technology
[0002] Transient electromagnetic imaging (TEM) has been widely used due to its advantages such as fast construction speed, high detection efficiency, and deep imaging depth. However, on the one hand, transient electromagnetic signals decay rapidly, with their dynamic range changing by 4 to 5 orders of magnitude within milliseconds. Therefore, weak signals are easily submerged in environmental electromagnetic noise, severely reducing the accuracy of electromagnetic inversion imaging. On the other hand, various electromagnetic interferences exist in urban environments, such as power frequency noise, broadband electromagnetic noise, and human-induced noise, further exacerbating the difficulty of recovering transient electromagnetic signals in the mid-to-late stages. Therefore, accurate denoising of transient electromagnetic signals is of great significance for achieving high-precision inversion imaging of urban underground spaces.
[0003] Many scholars have proposed various denoising methods for transient electromagnetic signals. Singular Spectrum Analysis (SSA) can more effectively separate signals and noise by decomposing and reconstructing the time-domain trajectory matrix, and it has unique advantages, especially in suppressing periodic interference. The key to noise removal using SSA lies in the selection of the reconstruction order. Golyandina et al. [N. Golyandina, V. Nekrutkin, and AAZhigljavsky, Analysis of time series structure: SSA and related techniques. CRC press, 2001] pointed out that finding the "inflection points" in the singular spectrum can distinguish the signal components corresponding to larger singular values from the noise components corresponding to smaller singular values. When the signal-to-noise ratio is high, signal denoising can be effectively achieved.
[0004] When the noise intensity is low or the type is relatively simple, the signal component dominates the noisy signal, and the above-mentioned denoising algorithm can effectively remove the noise. However, the time-frequency characteristics of non-steady-state nonlinear transient electromagnetic signals are more complex, and a single denoising method can easily cause signal distortion or noise residue. To solve the above problems, researchers have proposed a variety of joint denoising methods in recent years. Song et al. [D. Song, G. Feng, T. Qi, H. Wang, D. Pan, and L. Zhang, "A new combined transient electromagnetic noise reduction method of VMD-SVD-WTD based on improved dung beetle optimization algorithm with multi-strategy fusion," Measurement Science and Technology, vol. 36, no. 1, 2024, doi:10.1088 / 1361-6501 / ad7e42] proposed a dung beetle optimization algorithm (CRCDBO) based on cross-beetle optimization to determine the penalty factor α and the number of modes K for VMD decomposition. Then, for each mode obtained from VMD, SVD is used to remove components with energy less than the average singular value. Next, soft-threshold WTD is used for denoising, and finally, the modes are summed to obtain the denoised signal. While these wavelet-based denoising methods can analyze signals at different frequency scales, transient electromagnetic signals are dominated by high-frequency components in the early stages and low-frequency components in the later stages. Using WTD to remove components in a certain frequency band may lead to signal distortion. Wang et al. [P. Wang, J. Dong, L. Wang, and S. Qiao, "Signal Denoising Method Based on EEMD and SSA Processing for MEMS Vector Hydrophones," Micromachines (Basel), vol. 15, no. 10, Sep 24 2024, doi:10.3390 / mi15101183.] used EEMD for preliminary signal denoising, and then used SSA to further extract signal components from the denoised signal, effectively improving the accuracy and reliability of signal analysis. However, some IMFs obtained by EEMD decomposition may contain mixed sequences of signal and noise, and signal components are easily discarded during reconstruction. In addition, SSA may also cause distortion in the process of empirically selecting singular values. Furthermore, these denoising methods are still ineffective in removing broadband noise with the same frequency as the signal. In order to meet the requirements of transient electromagnetic detection in high-noise environments, there is an urgent need for a high-performance denoising method for low signal-to-noise ratio transient electromagnetic signals. Summary of the Invention
[0005] This invention addresses the challenge of selecting the reconstruction order when using singular spectrum analysis for denoising transient electromagnetic signals under high noise conditions. It utilizes an optimization algorithm to determine the denoising order when the signal components in the singular spectrum are not significant, and then uses the WFNC clustering method to achieve accurate reconstruction of singular values containing both noise and signal.
[0006] The objective of this invention is achieved through the following technical solution, as shown in the flowchart below. Figure 1 As shown, the detailed steps include:
[0007] Step 1: For noisy signals, use complementary set empirical mode decomposition (CEEMD). After adding several pairs of complementary white noise to the noisy signal, perform mirror extension and decomposition on the noisy signal at the end, and then truncate the obtained mode to the original length to obtain the final mode.
[0008] Step 2: Calculate the sample entropy (SE) for all modes, and classify the modes into signal dominant classes Φ based on the sample entropy. s With noise-dominant class Φ n Calculate the average entropy of all modal samples. Modes with sample entropy greater than the average are classified as noise-dominant modes, and the remaining modes are classified as signal-dominant modes.
[0009] Step 3: Solve the strict boundary of SSA reconstruction for each mode using the PSO optimization algorithm. Here, the fuzzy entropy FE of the signal is used as the fitness function for optimization. The dominant energy component in a noisy signal can be either signal or noise: when the signal component is dominant, the larger singular values in the singular spectrum correspond to the signal, while when noise is dominant, the smaller singular values in the singular spectrum are signal components. Therefore, the signal-dominant mode and the noise-dominant mode are solved in the first and second halves of the singular spectrum, respectively. The fitness function for the optimization problem is shown below:
[0010]
[0011] Where FuzzyEn(·) represents the fuzzy entropy of the signal, indicating the degree of disorder in the signal, and J represents the total number of modes. X=[x1,x2,…,x J ] represents the boundary of the strictly reconstructed order of each mode that needs to be solved, I′ j (x j ) is the j-th mode I j The result is reconstructed through singular value decomposition, and has
[0012]
[0013] Where U and V are the left and right eigenvectors obtained from singular value decomposition, and Φ s and Φn These are the signal-dominant mode set and the noise-dominant mode set obtained in step two, where L is the embedding window length of the SSA, i.e., the total number of singular values, and λ... i Let be the i-th singular value. A smaller objective function indicates a higher degree of signal disorder after denoising and a better denoising effect; conversely, a larger objective function indicates a poorer denoising effect. Here, minimizing the fuzzy entropy objective function can effectively remove random noise from the signal, improve signal order, and obtain the strict boundary X of all modal reconstruction orders. Strict =[x s1 x s2 , ..., x sJ ].
[0014] Step 4: Estimate the loose boundary for singular value reconstruction of each mode based on the singular value spectrum and singular value difference spectrum. The loose order of the signal-dominant mode starts at the first singular value and ends with a loose boundary; while the noise-dominant mode ends at the last singular value and starts with a loose boundary. The selection method is as follows:
[0015] ① Calculate the singular value spectrum Λ for the i-th mode. i and singular value difference spectrum Δ i ② Find Δ i Find the average value δ of all extreme points in the given set. avg If mode i is the dominant signal mode, then Δ i Medium greater than δ avg The points will form a set Φ1; otherwise, the set will be less than δ. avg ③ Construct a set Φ1 from the points; ③ Calculate Λ i Average value λ avg If mode i is the dominant signal mode, then the values of Λi greater than λ will be... avg The points will form a set Φ2; otherwise, it will be less than λ. avg Construct a set Φ2 from the points, and let the minimum value in Φ2 be λ. min The maximum value is λ max ④ If the mode is the signal-dominant mode, then find the closest λ in Φ1. min The singular values are used as the loose boundary for SSA reconstruction. If the mode is the noise-dominated mode, then the closest value to λ in Φ1 is found. max The singular values are used as the loose bounds for SSA reconstruction.
[0016] Finally, the relaxed boundary X of all modes is obtained. relaxed =[x r1 x r2 , ..., x rJ The region between the loose boundary and the strict boundary is called the fuzzy interval of the mode.
[0017] Step 5: Use the Weighted Fuzzy Noise Clustering (WFNC) algorithm to cluster the singular values within the fuzzy interval, classifying them into two categories: signal and noise, and assigning membership degrees. The objective function of WFNC is shown below:
[0018]
[0019] Where X = [x1, x2, ..., x n ] represents input data of length n, where the k-th data element is x. k =[x k1 ,x k2 ,…,x kp V = [v1, v2, ..., v] c ] represents c cluster centers, where the i-th cluster center v i =[v i1 ,v i2 ,…,v ip p is the data feature dimension, and δ is called the noise distance. U = [u ik ] c×n Let u represent the membership matrix, and u represent the membership degree matrix. ik This represents the membership degree of the k-th data point to the i-th cluster. W = [w ij ] c×p Let w represent the feature weight matrix, where w ij Let represent the weight of the j-th feature with respect to the i-th cluster center. The feature weights satisfy... Furthermore, the membership coefficient m and the weight coefficient β satisfy m>1 and β>1, respectively. The first term of the objective function is the weighted sum of the distances from each data point to each cluster center, representing the degree of clustering. However, during the clustering process, when some data points with significant measurement errors are too far from other points, it may cause the cluster centers to shift, affecting the clustering results. Therefore, a penalty term is introduced here to penalize remote data points.
[0020] The cluster centers and membership degrees of the clustering algorithm are updated using the Lagrange multiplier method:
[0021]
[0022] Where d ik This is the weighted distance between the k-th data point and the i-th cluster center. If c clusters are needed, the algorithm adds a cluster to capture outlier data points that do not belong to any explicitly defined cluster. In WFNC, this is the sum of the membership degrees of all classes. Unassigned membership This indicates the degree to which a point belongs to an outlier cluster. Outliers are usually abnormal points caused by systematic errors, etc. Their distance d >> δ from other data points, and their membership degree u → 0, thus reducing the degree to which the point belongs to c clusters. This leads to an increase in the penalty term in the fitness function, while the cluster centers are unaffected by the outlier. Therefore, this algorithm can effectively reduce the impact of outliers. The noise distance δ is used to change the degree to which data points are clustered into outlier clusters. When δ increases, fewer points are classified as outliers, while when δ decreases, more data points are classified as outliers. Generally, δ is 1 to 2 times the standard deviation of the data. The weight w is used to measure the influence of data features on clustering. A larger weight indicates a greater influence of the feature. The formula for updating the weight of the j-th feature for the i-th cluster center is as follows:
[0023]
[0024] in The distance between each data point x and the j-th feature of the i-th cluster center v is used. If the distance of a certain feature is too large, the weight of that feature is reduced to balance the effects between different features.
[0025] Step Six: Weight the singular values within the fuzzy interval using membership degrees as weights, and then combine this weighted sum with the singular values within the other strict orders to denoise and reconstruct the mode. For the i-th mode, the denoised mode is shown below:
[0026]
[0027] The denoised modes are summed to obtain the denoised signal.
[0028] The detailed implementation process of this method is as follows: Figure 2 As shown.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] This method effectively mitigates the distortion defects of SSA transient electromagnetic signal denoising in high electromagnetic noise environments and improves the signal-to-noise ratio (SNR) of the denoised signal. Compared to conventional SSA denoising methods, this method adaptively optimizes the selection of the SSA denoising reconstruction order based on the signal SNR, preventing signal distortion or noise residue caused by empirical selection. Furthermore, since some singular values in the singular spectrum of the noisy signal contain both signal and noise components, this method utilizes a WFNC-based singular value signal noise membership prediction method to extract the signal components from the singular values for signal reconstruction, further enhancing the SNR. Attached Figure Description
[0031] Figure 1 This is a flowchart of the transient electromagnetic signal denoising method based on adaptive fuzzy optimization singular spectrum analysis of the present invention;
[0032] Figure 2 This is a flowchart of an embodiment of the transient electromagnetic signal denoising method based on adaptive fuzzy optimization singular spectrum analysis of the present invention;
[0033] Figure 3 It is a time-domain waveform diagram of a noisy transient electromagnetic signal;
[0034] Figure 4 It is the IMF after CEEMD decomposition;
[0035] Figure 5 These are the entropy spectra of each signal mode;
[0036] Figure 6 It is the membership degree of the singular values in the fuzzy interval to the signal components;
[0037] Figure 7 It shows the time-domain waveforms of the transient electromagnetic signal before and after noise reduction. Detailed Implementation
[0038] The invention will be further described below with reference to the accompanying drawings. A transient electromagnetic signal denoising method based on adaptive fuzzy optimization singular spectrum analysis, such as... Figure 1-2 As shown, the main steps include the following:
[0039] Step 1: For Figure 3 The noisy signal shown was decomposed using Complementary Ensemble Empirical Mode Decomposition (CEEMD) to obtain 10 IMFs and residuals, as follows: Figure 4 As shown.
[0040] Step 2: Calculate the sample entropy (SE) of the 10 IMFs, such as... Figure 5 As shown, the average entropy of all samples is selected as the threshold, and modes with sample entropy less than the threshold are defined as the dominant signal mode Φ. s The rest are noise-dominant modes Φ n Therefore, modes 1 to 3 are noise-dominated modes, and modes 4 to 10 are signal-dominated modes.
[0041] Step 3: Perform SSA processing on the 10 modes, with the embedding window length L set to N / 3, where N is the sequence length. Here, the fuzzy entropy FE of the signal is used as the fitness function for optimization, and the PSO optimization algorithm is used to obtain the strict boundaries for singular value reconstruction of each mode. For different signal-to-noise ratios, the singular value components with high energy may be either signals or noise. When signal energy dominates, larger singular values in the singular spectrum correspond to signal components, while when noise energy dominates, smaller singular values in the singular spectrum correspond to signal components. Therefore, the signal-dominant mode is optimized using the first half of the singular spectrum, while the noise-dominant mode is optimized using the second half of the singular spectrum. The fitness function for the optimization problem is shown below:
[0042]
[0043] FuzzyEn(·) represents the fuzzy entropy of the signal, indicating the degree of confusion in the signal. X = [x1, x2, ..., x...] 10 ] represents the strict boundary for denoising and reconstruction of each mode, I′ j (x j ) is the j-th mode I j The result is reconstructed through singular value decomposition, and has
[0044]
[0045] Where U and V are the left and right eigenvectors obtained from singular value decomposition, and Φ s and Φ n Let L be the signal-dominant mode set and the noise-dominant mode set obtained in step two, respectively. Let L be the embedding window length of the SSA, i.e., the total number of singular values, and λ be the noise-dominant mode set. i Let be the i-th singular value. When the objective function is small, it indicates that the signal is less chaotic after denoising and the denoising effect is good; conversely, it indicates that the denoising effect is poor. Therefore, we optimize and solve the objective function to minimize the fuzzy entropy to obtain the strict boundary of all modal reconstruction orders.
[0046] Step 4: Calculate the singular values of each mode and reconstruct the loose boundary based on the singular value spectrum and singular value difference spectrum. If the mode is signal-dominant, the starting point of the loose order is the first singular value, and the ending point is the loose boundary; if the mode is noise-dominant, the ending point is the last singular value, and the starting point is the loose boundary. Here, the region between the strict boundary and the loose boundary of each mode is defined as the fuzzy interval.
[0047] Step 5: Use the Weighted Fuzzy Noise Clustering (WFNC) algorithm to cluster the singular values within the fuzzy interval, dividing them into two classes: signal and noise. The membership degree of each singular value belonging to one of these two classes is then given. The objective function of WFNC is shown below:
[0048]
[0049] Where X = [x1, x2, ..., x n ] represents input data of length n, where the k-th data element is x. k =[x k1 ,x k2 ,…,x kp V = [v1, v2, ..., v] c ] represents c cluster centers, where the i-th cluster center v i =[v i1 ,v i2 ,…,v ipp is the data feature dimension, and δ is called the noise distance. U = [u ik ] c×n Let u represent the membership matrix, and u represent the membership degree matrix. ik This represents the membership degree of the k-th data point to the i-th cluster. W = [w ij ] c×p Let w represent the feature weight matrix, where w ij Let represent the weight of the j-th feature with respect to the i-th cluster center. The feature weights satisfy... Furthermore, the membership coefficient m and the weight coefficient β satisfy m>1 and β>1, respectively. The first term of the objective function is the weighted sum of the distances from each data point to each cluster center, representing the degree of clustering. However, during the clustering process, when some data points with significant systematic errors are too far from other points, it may cause the cluster centers to shift, affecting the clustering results. Therefore, a penalty term is introduced here to penalize data points that deviate from the target cluster.
[0050] The cluster centers and membership degrees of the clustering algorithm are updated using the Lagrange multiplier method:
[0051]
[0052] Where d ik This is the weighted distance between the k-th data point and the i-th cluster center. Here, we cluster the data into two clusters: a noise cluster and a signal cluster. The algorithm then adds a new cluster to capture outliers. In WFNC, this is the sum of the membership degrees of all classes. Unassigned membership This indicates the degree to which a point belongs to an outlier cluster. Outliers are usually abnormal points caused by systematic errors, etc. Their distance d >> δ from other data points, and their membership degree u → 0, thus reducing the degree to which the point belongs to two clusters. This leads to an increase in the penalty term in the fitness function, while the cluster centers are unaffected by the point. Therefore, this algorithm can effectively reduce the impact of outliers. The role of the noise distance δ is to change the degree to which data points are clustered into outlier clusters. When δ increases, fewer points are classified as outliers; conversely, when δ decreases, more data points are classified as outliers.
[0053] The weight w is used to measure the influence of a data feature on clustering. A larger weight indicates that the feature has a greater influence. The formula for updating the weight of the j-th feature with respect to the i-th cluster center is as follows:
[0054]
[0055] in The distance between each data point x and the j-th feature of the i-th cluster center v is used. If the distance to a feature is too large, the weight of that feature will be reduced, thus balancing the effects between different features.
[0056] Taking the 7th mode as an example, after clustering, the membership degrees of the 3rd to 10th singular values to the signal class are obtained, such as... Figure 6 As shown, the membership degree of the fourth singular value belonging to the signal class is 0.102. We use this value as a weight to reconstruct the fourth singular value.
[0057] Step Six: Weight the singular values within the fuzzy interval using membership degrees as weights, and then combine this weighted sum with the singular values within the other strict orders to reconstruct the denoised signal. For the i-th mode, the denoised mode is shown below:
[0058]
[0059] The denoised modes are summed to obtain the denoised signal. Comparing the denoised sequence with the original sequence as follows: Figure 7 As shown, the method of this invention can effectively suppress early high-frequency burst noise and later power frequency interference, while preserving the signal's attenuation pattern and inflection point characteristics well, without obvious Gibbs oscillations or signal distortion.
Claims
1. A method for denoising transient electromagnetic signals based on adaptive fuzzy optimization singular spectrum analysis, characterized in that, Includes the following steps: Step 1: Perform complementary ensemble empirical mode decomposition (CEEMD) on the acquired noisy transient electromagnetic signal. After adding several pairs of complementary white noise to the noisy signal, first extend and decompose the end of the noisy signal, and then truncate the obtained modes to their original length to obtain several intrinsic mode functions (IMFs). Step 2: Calculate the sample entropy (SE) of all modes, and based on the average value of the sample entropy of all modes, divide all modes into a signal-dominant mode set and a noise-dominant mode set; Step 3: Perform Singular Spectrum Analysis (SSA) on all modes, construct a fitness function with the objective of minimizing the signal fuzzy entropy (FE), and use the Particle Swarm Optimization (PSO) algorithm to solve for the optimal SSA reconstruction order, i.e., the strict boundary, in parallel for each IMF; wherein, for IMFs in the signal-dominant mode set, the strict boundary is searched in the first half of the singular value spectrum; for IMFs in the noise-dominant mode set, the strict boundary is searched in the second half of the singular value spectrum. Step 4: For each IMF whose strict boundary has been obtained, calculate the relaxed order of the singular value reconstruction of each mode based on the distribution characteristics of its own singular value spectrum and singular value difference spectrum, and construct the fuzzy interval based on the relaxed order and the strict order. Step 5: Use the Weighted Fuzzy Noise Clustering (WFNC) algorithm to cluster the singular values within the fuzzy interval, classifying them into two categories: signal and noise, and giving the membership degree of the singular values to the signal and noise categories respectively; Step 6: Use membership degree as weight to weight the singular values in the fuzzy interval, and then combine it with the singular values in the other strict order to denoise and reconstruct the mode. Then add up all the reconstructed modes to obtain the denoised transient electromagnetic signal.
2. The transient electromagnetic signal denoising method based on adaptive fuzzy optimization singular spectrum analysis as described in claim 1, characterized in that: The fitness function formula in step three is shown below: Where FuzzyEn(·) represents the fuzzy entropy of the signal, represents the degree of disorder of the signal, J represents the total number of modes, and X = [x1, x2, ..., x...]. J ] represents the strict order boundary that needs to be optimized for each mode, I′ i (x i ) is the i-th mode I i The result is reconstructed through singular value decomposition, and has Where U and V are the left and right eigenvectors obtained from singular value decomposition, and Φ s and Φ n These are the signal-dominant mode set and the noise-dominant mode set obtained in step two, respectively. L is the embedding window length of the SSA, i.e., the total number of singular values, and λ is the noise-dominant mode set. i Let i be the i-th singular value.
3. The method according to claim 1 or 2, characterized in that, The method for determining the loose boundary in step four includes: For the dominant mode of the signal, the starting point of its loose boundary is the first singular value ordinal number, and the ending point is determined by its singular value difference spectrum; For the noise-dominant mode, the endpoint of its loose boundary is the last singular ordinal number, and the starting point is determined by its singular value difference spectrum.
4. The transient electromagnetic signal denoising method based on adaptive fuzzy optimization singular spectrum analysis as described in claim 1, characterized in that: The WFNC clustering algorithm in step five clusters singular values within the fuzzy interval into signal and noise classes and can provide the membership degree of each singular value to each class. The objective function of WFNC is as follows: Where X = [x1, x2, ..., x n ] represents input data of length n, where the k-th data element is x. k =[x k1 ,x k2 ,…,x kp V = [v1, v2, ..., v] c ] represents c cluster centers, where the i-th cluster center v i =[v i1 ,v i2 ,…,v ip ]; p is the data feature dimension, δ is called the noise distance. U = [u ik ] c×n Let u represent the membership matrix, and u represent the membership degree matrix. ik W represents the membership degree of the k-th data point to the i-th cluster, where W = [w ij ] c×p Let w represent the feature weight matrix, where w ij This represents the weight of the j-th feature with respect to the i-th cluster center, where the feature weights satisfy w ij ≤1, Furthermore, the membership coefficient m and the weight coefficient β satisfy m>1 and β>1, respectively. The first term of the objective function is the weighted sum of the distances from each data point to each cluster center, representing the degree of clustering. The second term is a penalty term, used to penalize outliers that are far from other data points.
5. The transient electromagnetic signal denoising method based on adaptive fuzzy optimization singular spectrum analysis as described in claim 4, characterized in that: The cluster centers and membership degrees of the clustering algorithm are updated using the Lagrange multiplier method: Where d ik Let w be the weighted distance from the k-th data point to the i-th cluster center. The weight w is used to measure the influence of the data feature on the clustering. A larger weight indicates that the feature has a greater influence. The formula for updating the weight of the j-th feature with respect to the i-th cluster center is as follows: in The distance between each data point x and the j-th feature of the i-th cluster center v is given.