Roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio test
By using multi-dimensional perceptual autoencoder and signal-to-noise ratio inspection methods in roadbed settlement data cleaning, the problem of difficulty in identifying noise and abnormal points is solved, and data cleaning with higher accuracy and reliability is achieved, ensuring data integrity and continuity.
Patent Information
- Application Number
- CN202510266443.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has difficulty in identifying noise and abnormal points in roadbed settlement data cleaning, resulting in uncertainty and misjudgment or misjudgment in data analysis.
The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio test is adopted, and the data is decomposed and reconstructed through a multi-dimensional perception autoencoder, and combined with the signal-to-noise ratio test and Gaussian mixed model, the abnormal data is accurately identified and deleted.
It improves the accuracy and reliability of data cleaning, avoids misjudgment and misjudgment, ensures data integrity and continuity, and provides a more reliable data processing solution for roadbed settlement analysis.
Smart Images

Figure CN120196860A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of civil engineering operation and maintenance, and particularly relates to a subgrade settlement data cleaning method based on reconstruction error and signal-to-noise ratio test. Background Art
[0002] In the construction of transportation infrastructure, China has shifted from the rapid construction stage to the stable growth stage. At present, the stock of transportation infrastructure is huge, facing a large number of maintenance problems. For highway and railway subgrades, due to their simple structure, good comprehensive economic benefits and excellent performance, they are widely used in the main lines of China. The settlement control of subgrades is crucial for the durability of pavement structures. Especially for high-speed railways, uneven settlement of subgrades will lead to a significant reduction in the running stability of trains and even the risk of derailment. Due to the granular characteristics of soil and the random characteristics of internal soil parameters, the service performance of structures during the operation period usually cannot be described by the design parameters during geological exploration. At present, the commonly used methods for analyzing the settlement state of subgrades mainly adopt data-driven methods. This method judges the settlement state of subgrades by analyzing the historical state and development trend on site, and can realize the identification of subgrade settlement stability under uncertain soil conditions, providing an effective method for subgrade settlement analysis.
[0003] A high-quality data set is the premise for implementing data-driven methods. In actual engineering, data is usually collected through IoT monitoring. Due to the complex service environment of subgrade structures, they are affected by the periodic load of trains and external natural environment disturbances, and the structural response has obvious dynamic characteristics. The structural state data sensed has errors caused by accidental events, which affects the effectiveness of data-driven methods. Identifying and clearing these disturbed data is one of the effective methods to ensure the reliability of analysis. At present, the main cleaning methods for engineering time series data include the 3-sigma rule method, box plot method, and machine learning methods, etc. However, in the actual application process, it is difficult for the 3-sigma rule to meet the requirement of normal distribution of data. The box plot method has certain limitations for data whose overall expression based on the median is unreliable. Machine learning methods in data cleaning applications mainly include supervised methods. Supervised methods require a certain number of known samples to train the model. However, for civil engineering, it is usually impossible to obtain enough data to train the model, which hinders the further application of this method.
[0004] Therefore, unsupervised learning methods with a small sample size requirement are proposed for subgrade settlement prediction, such as autoencoders and autoregressors, etc. This method provides an effective solution for subgrade settlement analysis. However, in the application process, due to the lack of a clear abnormal identification threshold, it is impossible to accurately locate the abnormal state, and there are often missed judgments or misjudgments in engineering. Summary of the Invention
[0005] In order to solve the above problems, the present invention proposes a method for cleaning subgrade settlement data based on reconstruction error and signal-to-noise ratio test.
[0006] The technical solution of the present invention is: A method for cleaning subgrade settlement data based on reconstruction error and signal-to-noise ratio test includes the following steps:
[0007] S1. Collect on-site monitoring data and decompose the on-site monitoring data to obtain a number of IMF data;
[0008] S2. Construct a multi-dimensional perception autoencoder and input each IMF data into the multi-dimensional perception autoencoder to obtain a number of reconstructed IMF sequences with complete lengths;
[0009] S3. According to a number of reconstructed IMF sequences with complete lengths, determine whether the reconstructed monitoring sequence is normal. If so, complete the data cleaning; otherwise, enter S4;
[0010] S4. Delete abnormal data and supplement missing values to complete the data cleaning.
[0011] Further, S2 includes the following sub-steps:
[0012] S21. Input each IMF data into the multi-dimensional perception autoencoder for reconstruction to obtain a reconstructed IMF sequence of the IMF data under the analysis window;
[0013] S22. Select the first time_win items from a number of IMF data and supplement them to each reconstructed IMF sequence to obtain a reconstructed IMF sequence with a complete length, where time_win represents the size of the analysis window.
[0014] Further, in S21, the multi-dimensional perception autoencoder includes a first one-dimensional convolutional layer, a first max-pooling layer, a second one-dimensional convolutional layer, a second max-pooling layer, an LSTM layer, a first fully-connected layer, and a second fully-connected layer connected in sequence.
[0015] Further, in S21, the calculation formula for the size time_win of the analysis window is:
[0016]
[0017] In the formula, n represents the total length of the data, and int(·) represents the rounding formula.
[0018] Further, in S21, the expression of the loss function Loss of the multi-dimensional perception autoencoder is:
[0019]
[0020] In the formula, y irepresents the value of the i-th reconstructed IMF sequence, represents the true value of the original sequence corresponding to the i-th reconstructed IMF sequence, smooth represents the minimum value given to ensure the mathematical meaning of the loss, and e represents the exponent.
[0021] Furthermore, in S22, the i-th reconstructed IMF sequence imf′ i has the following expression:
[0022] imf′ i = [y i,1 , y i,2 ,..., y i,time_win-1 , y′ i,time_win ,..., y′ i,n ;
[0023] In the formula, y i,1 represents the first data in the i-th imf component, y i,2 represents the second data in the i-th imf component, y i,time_win-1 represents the (time_win - 1)-th data in the i-th imf component, y′ i,time_win represents the time_win-th data in the i-th imf component, y′ i,n represents the n-th data in the i-th imf component, and time_win represents the analysis window size.
[0024] Furthermore, S3 includes the following sub-steps:
[0025] S31: Add several reconstructed IMF sequences of full length to generate a reconstructed monitoring sequence;
[0026] S32: Calculate the signal-to-noise ratio corresponding to the reconstructed monitoring sequence;
[0027] S33: According to the signal-to-noise ratio corresponding to the reconstructed monitoring sequence, determine whether the reconstructed monitoring sequence is normal. If it is, end; otherwise, enter S3.
[0028] Furthermore, in S33, the expression for determining whether the reconstructed monitoring sequence is normal is:
[0029]
[0030] In the formula, SNR m represents the signal-to-noise ratio calculated during the m-th cleaning process, and SNR m-1 represents the signal-to-noise ratio calculated during the (m - 1)-th cleaning process.
[0031] Furthermore, S4 includes the following sub-steps:
[0032] S41. Fit the frequency distribution of the reconstructed monitoring sequence using the Gaussian mixture model, select the bilateral 5% confidence interval as the error interval, mark the data falling within the bilateral 5% range as abnormal data and delete it;
[0033] S42. Supplement the missing values in the reconstructed monitoring sequence after deleting the abnormal data to complete data cleaning.
[0034] The beneficial effects of the present invention are as follows:
[0035] (1) Different from the traditional cleaning method, the present invention adopts a deep learning model based on a multi-dimensional perception autoencoder for the problems of noise and abnormal points in engineering time series data. The data is divided into different intrinsic mode functions through EMD decomposition, and combined with the reconstruction ability of the autoencoder, the value at each moment is reconstructed, thus effectively improving the accuracy and reliability of data cleaning;
[0036] (2) Compared with traditional methods such as the 3σ criterion method and the box plot method, the present invention does not require the data to conform to a specific distribution, has stronger adaptability and flexibility, and is particularly suitable for time series data with non-linearity and dynamic changes such as subgrade settlement. The abnormal state of the reconstructed sequence is determined through signal-to-noise ratio test, and the abnormal points are accurately identified and deleted by combining the Gaussian mixture model, making the data cleaning process more accurate and avoiding misjudgment and missed judgment;
[0037] (3) The present invention further supplements the missing values and reconstructs the original sequence, ensuring data integrity and continuity. Through this series of innovations, the present invention provides a more reliable data processing solution for subgrade settlement analysis, with significant engineering practical value. Description of the Drawings
[0038] Figure 1 It is a flow chart of the subgrade settlement data cleaning method based on reconstruction error and signal-to-noise ratio test;
[0039] Figure 2 It is a structural schematic diagram of the multi-dimensional perception autoencoder;
[0040] Figure 3 It is an EMD decomposition diagram;
[0041] Figure 4 It is a comparison diagram before and after noise reduction. Detailed Embodiments
[0042] The embodiments of the present invention will be further described below with reference to the drawings.
[0043] As Figure 1 shown, the present invention provides a subgrade settlement data cleaning method based on reconstruction error and signal-to-noise ratio test, including the following steps:
[0044] S1. Collect on-site monitoring data, decompose the on-site monitoring data to obtain several IMF data;
[0045] S2. Construct a multi-dimensional perception autoencoder, and input each IMF data into the multi-dimensional perception autoencoder to obtain several reconstructed IMF sequences with complete lengths;
[0046] S3. According to several reconstructed IMF sequences with complete lengths, determine whether the reconstructed monitoring sequence is normal. If so, complete data cleaning; otherwise, go to S4;
[0047] S4. Delete abnormal data and supplement missing values to complete data cleaning.
[0048] IMF refers to the intrinsic mode function.
[0049] In the embodiment of the present invention, S2 includes the following sub-steps:
[0050] S21. Input each IMF data into the multi-dimensional perception autoencoder for reconstruction to obtain a reconstructed IMF sequence of the IMF data under the analysis window;
[0051] S22. Select the first time_win items from several IMF data and supplement them to each reconstructed IMF sequence to obtain a reconstructed IMF sequence with a complete length, where time_win represents the size of the analysis window.
[0052] In the embodiment of the present invention, in S21, the multi-dimensional perception autoencoder includes a first one-dimensional convolutional layer, a first max-pooling layer, a second one-dimensional convolutional layer, a second max-pooling layer, an LSTM layer, a first fully-connected layer, and a second fully-connected layer connected in sequence.
[0053] In the embodiment of the present invention, an encoder-decoder architecture is adopted, where the encoder part includes convolutional and pooling layers to extract local features and gradually compress the time series information; the decoder part decodes the compressed features into the target output through the fully-connected layer.
[0054] Specifically: In the encoder part, first, local features of the time series are extracted through two layers of one-dimensional convolution. The number of input channels is increased to 64 in the first convolutional layer, and further increased to 128 in the second convolutional layer. After each layer of convolution, downsampling is achieved through max pooling to gradually reduce the time dimension. This part is mainly used to learn the local dependencies of the sequence. Subsequently, a one-way LSTM layer is used to further extract the long-term dependency features of the time series. The hidden state dimension of the LSTM layer is the latent dimension of the encoder, and an addition operation is used to fuse the last two hidden states to obtain a fixed-size feature representation. In the decoder part, the feature representation is gradually mapped to the final output through two fully connected networks. The first fully connected layer maps the latent feature dimension to 64 dimensions, and the second fully connected layer further maps it to a one-dimensional output. At the same time, the result is normalized between 0 and 1 through the Sigmoid activation function. The design of the decoder part ensures that the model can compress high-dimensional sequence features into a concise representation and decode it into an effective output, enabling it to capture both local features of the sequence and learn global dependencies, making it suitable for the analysis and prediction tasks of time series data. Between all neural network layers, a variable feature Sigmoid function is used for activation, and its calculation formula is:
[0055]
[0056] In the formula, σ(·) represents the activation function, w and b represent the learnable parameters in the activation function, x represents the data matrix, and e represents the exponent.
[0057] In the embodiment of the present invention, in S21, the calculation formula for the analysis window size time_win is:
[0058]
[0059] In the formula, n represents the total length of the data, and int(·) represents the rounding formula.
[0060] In the embodiment of the present invention, in S21, the expression of the loss function Loss of the multi-dimensional perception autoencoder is:
[0061]
[0062] In the formula, y i represents the value of the i-th reconstructed IMF sequence, represents the true value of the original sequence corresponding to the i-th reconstructed IMF sequence, smooth represents the minimum value given to ensure the mathematical meaning of the loss, and e represents the exponent.
[0063] In the embodiment of the present invention, in S22, the expression of the i-th reconstructed IMF sequence imf′ i with a complete length is:
[0064] imf′ i = [y i,1 , y i,2 , ..., y i,time_win-1 , y′ i,time_win , ..., y′ i,n ;
[0065] Wherein, y i,1 represents the first data in the i-th imf component, y i,2 represents the second data in the i-th imf component, y i,time_win-1 represents the (time_win - 1)-th data in the i-th imf component, y′ i,time_win represents the time_win-th data in the i-th imf component, y′ i,n represents the n-th data in the i-th imf component, and time_win represents the analysis window size.
[0066] In an embodiment of the present invention, S3 includes the following sub-steps:
[0067] S31. Add several reconstructed IMF sequences with complete lengths to generate a reconstructed monitoring sequence;
[0068] S32. Calculate the signal-to-noise ratio corresponding to the reconstructed monitoring sequence;
[0069] S33. According to the signal-to-noise ratio corresponding to the reconstructed monitoring sequence, determine whether the reconstructed monitoring sequence is normal. If so, end; otherwise, enter S3.
[0070] In an embodiment of the present invention, in S33, the expression for determining whether the reconstructed monitoring sequence is normal is:
[0071]
[0072] Wherein, SNR m represents the signal-to-noise ratio calculated during the m-th cleaning process, and SNR m-1 represents the signal-to-noise ratio calculated during the (m - 1)-th cleaning process. When m = 1, SNR0 = SNR1.
[0073] In an embodiment of the present invention, S4 includes the following sub-steps:
[0074] S41. Fit the frequency distribution of the reconstructed monitoring sequence using a Gaussian mixture model, select the two-sided 5% confidence interval as the error interval, and mark and delete the data falling within the two-sided 5% range as abnormal data;
[0075] S42. Perform missing value supplementation on the reconstructed monitoring sequence after deleting abnormal data to complete data cleaning.
[0076] The following is an illustration in combination with specific embodiments.
[0077] S1. Collect on-site monitoring data x = [x1, x2, …, x n , where x1 represents the first piece of data in the time series data sequence, x2 represents the second piece of data in the time series data sequence, and x n represents the nth piece of data in the time series data sequence, and the time intervals between the monitoring data are kept consistent, as shown in Table 1.
[0078] Table 1
[0079] Time 1s 2s 3s 4s … ns Data 1 2 3 4 … 7
[0080] In addition, it is required that the length of the monitoring data shall not be less than the data length required for the training of the autoencoder in S3;
[0081] S2. Decompose the data x = [x1, x2, …, x n using the EMD method. The time series signal X(t) after EMD decomposition can be expressed as:
[0082]
[0083] In the formula, IMF i (t) represents the ith intrinsic mode function, and r m (t) represents the final residual term. In the present invention, let IMF m+1 (t) = r m (t), m represents the total number of intrinsic mode functions obtained by decomposition, and IMF m+1 (t) represents the final residual sequence.
[0084] In the embodiment of the present invention, as Figure 3 shown, it is the IMF curve after EMD decomposition. The obtained IMF components are shown in Table 2.
[0085] Table 2
[0086] Time 1s 2s 3s 4s … ns IMF1 0.3 0.2 0.5 1.3 … 1.5 IMF2 0.1 0.14 0.16 0.2 … 3 … … … … … … … IMFm 0.5 0.3 2 3.3 … 5.1 IMFm + 1 0.01 0.003 0.02 0.33 … 0.001
[0087] S3. Establish a multi-dimensional perception autoencoder MulAE.
[0088] S4. Import each IMF data into the multi-dimensional perception autoencoder MulAE for curve analysis and reconstruction.
[0089] S5. Establish a curve abnormal state analysis method based on the signal-to-noise ratio to judge whether there is an abnormality in the sequence.
[0090] S6. If the data is abnormal, establish an abnormal data point recognition method based on the percentile of the reconstruction error, determine the position of the error point, and replace the error point based on the reconstruction curve; for the incomplete time series after deletion, use the reconstructed sequence to supplement the missing values after deletion; finally, regenerate the original sequence curve based on all the IMF curves.
[0091] In the embodiment of the present invention, as Figure 4 shown, it is a comparison chart between the regenerated data sequence and the original data sequence.
[0092] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification, characterized in that: The following steps are involved: S1. Collect on-site monitoring data and decompose the on-site monitoring data to obtain several IMF data; S2, construct a multidimensional perceptual autoencoder, and input each IMF data into the multidimensional perceptual autoencoder to obtain several reconstructed IMF sequences of full length; S3, judging whether the reconstructed monitoring sequence is normal according to several reconstructed IMF sequences of full length, if so, completing data cleaning, otherwise entering S4; S4. Delete abnormal data and fill in missing values to complete data cleaning.
2. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 1 is characterized in that: The S2 comprises the following sub-steps: S21, inputting each IMF data into a multidimensional perceptual autoencoder for reconstruction, and obtaining a reconstructed IMF sequence of the IMF data under the analysis window; S22. Select the first time_win item from a number of IMF data and add it to each reconstructed IMF sequence to obtain a reconstructed IMF sequence of full length, where time_win represents the analysis window size.
3. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 2 is characterized in that: In S21, the multidimensional perceptual autoencoder includes a first one-dimensional convolutional layer, a first maximum pooling layer, a second one-dimensional convolutional layer, a second maximum pooling layer, an LSTM layer, a first fully connected layer and a second fully connected layer which are connected in sequence.
4. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 2 is characterized in that: In S21, the calculation formula of the analysis window size time_win is: In the formula, n represents the total length of the data, and int(·) represents the rounding formula.
5. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 2 is characterized in that: In S21, the expression of the loss function Loss of the multidimensional perceptual autoencoder is: In the formula, y i represents the value of the i-th reconstructed IMF sequence, It represents the true value of the i-th reconstructed IMF sequence corresponding to the original sequence, smooth represents the minimum value given by mathematical meaning to ensure the loss, and e represents the exponent.
6. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 2 is characterized in that: In S22, the i-th full-length reconstructed IMF sequence imf′ i The expression is: imf′ i =[and i,1 ,and i,2 ,...,and i,time_win-1 ,and' i,time_win ,...,and' i,n ]; In the formula, y i,1 represents the first data in the i-th imf component, y i,2 represents the second data in the i-th imf component, y i,time_win-1 represents the time_win-1th data in the i-th imf component, y′ i,time_win Represents the time_win data in the i-th imf component, y′ i,n It represents the nth data in the ith IMF component, and time_win represents the analysis window size.
7. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 1 is characterized in that: The S3 comprises the following sub-steps: S31, adding a plurality of reconstructed IMF sequences of full length to generate a reconstructed monitoring sequence; S32, calculating the signal-to-noise ratio corresponding to the reconstructed monitoring sequence; S33: judging whether the reconstructed monitoring sequence is normal according to the signal-to-noise ratio corresponding to the reconstructed monitoring sequence, if so, the process ends, otherwise, the process proceeds to S3.
8. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 6 is characterized in that: In S33, the expression for judging whether the reconstructed monitoring sequence is normal is: In the formula, SNR m Represents the signal-to-noise ratio calculated during the mth cleaning process, SNR m-1 It represents the signal-to-noise ratio calculated during the m-1th cleaning process.
9. The roadbed settlement data cleaning method based on reconstruction error and signal-to-noise ratio verification according to claim 1 is characterized in that: The S4 comprises the following sub-steps: S41. Use the Gaussian mixture model to fit and reconstruct the frequency distribution of the monitoring sequence, and select a bilateral 5% confidence interval as the error interval, and mark the data falling within the bilateral 5% range as abnormal data and delete them; S42, supplement the missing values of the reconstructed monitoring sequence after deleting the abnormal data, and complete the data cleaning.