A method for predicting urban land subsidence based on an improved Transformer model
By constructing a multi-feature dataset and optimizing the input and output layers of the Transformer model, combined with a direct multi-step prediction method, the problem of the Transformer model failing to fully consider various environmental factors in urban land subsidence prediction is solved, achieving subsidence prediction with higher accuracy and better generalization ability.
Patent Information
- Application Number
- CN202511195919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing Transformer models fail to adequately consider the impact of various environmental factors in urban land subsidence prediction, resulting in insufficient prediction accuracy and generalization ability.
A multi-feature dataset was constructed, including land cover data, road network data, and precipitation data. The input and output layers of the Transformer model were optimized. A dual-scale dataset partitioning and direct multi-step prediction method were adopted, combined with the Adam W optimizer and GELU activation function to enhance the model's nonlinear processing capability and prediction accuracy.
It improves the accuracy and generalization ability of urban land subsidence prediction, can better capture the complex correlation characteristics of multiple environmental factors, realize direct multi-step prediction of urban land subsidence, and improve prediction accuracy and stability.
Smart Images

Figure CN120724159B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban land subsidence prediction technology, and in particular to an urban land subsidence prediction method based on an improved Transformer model. Background Technology
[0002] Urban land subsidence is a global environmental problem that has a serious impact on urban infrastructure, buildings, and residents' lives. Accurate subsidence prediction is crucial for accurate and timely subsidence management and for promoting rational urban planning and sustainable development.
[0003] In recent years, with the improvement of computing power and the rapid development of time-series InSAR technology, prediction models represented by deep learning have become the mainstream method for urban land subsidence prediction. The Transformer model solves the limitations of traditional network structures on training complex models, such as the inherent computational order of RNNs and the limited receptive field of convolutional kernels in CNNs. It can more effectively capture long-range dependencies and complex correlation features between elements in time-series data, and its application in time-series data prediction has become a research hotspot.
[0004] However, conventional Transformer models are trained using only time-series subsidence data as a single dataset, failing to consider the impacts of various human activities and natural environmental factors on land subsidence during the prediction process. This makes it difficult to adequately address the combined effects of multiple environmental factors in urban land subsidence prediction, resulting in limitations in both prediction accuracy and generalization ability. Summary of the Invention
[0005] The purpose of this invention is to provide an urban land subsidence prediction method based on an improved Transformer model, which solves the problem that existing methods do not fully consider the influence of multiple environmental factors during the prediction process. By constructing a multi-feature dataset containing multiple environmental factors and optimizing the input and output layers of the Transformer model, the model's ability to handle nonlinear complex features and express nonlinear trends is enhanced, thereby improving the accuracy and generalization ability of the model in predicting urban land subsidence.
[0006] To achieve the above objectives, this invention provides a method for predicting urban land subsidence based on an improved Transformer model, comprising the following steps:
[0007] Step S1: Construct a multi-feature dual-scale dataset including environmental factor data and land subsidence data;
[0008] Step S2: Construct the Transformer model, including the encoder and decoder;
[0009] Step S3: Embedding enhancement is performed on the encoder input layer in the Transformer model;
[0010] Step S4: Improve the decoder in the Transformer model;
[0011] Step S5: Optimize and adjust the improved Transformer model to achieve urban ground subsidence prediction.
[0012] Preferably, in step S1, a multi-feature dual-scale dataset including environmental factor data and land subsidence data is constructed, and the specific process is as follows:
[0013] Step S11: Based on the land subsidence data, standardize the data format and spatial resolution of each environmental factor; among which, the environmental factor data includes land cover data, road network data and precipitation data;
[0014] Step S12: Unify the time resolution among various environmental factors and ensure time alignment;
[0015] Step S13: Based on the feature data with unified spatiotemporal resolution, spatial location matching and temporal alignment obtained in the above process, create a dual-scale dataset;
[0016] Among them, the four types of characteristic data with unified spatiotemporal resolution, matching spatial location and time alignment include: ground subsidence data, land cover data, distance data between each point and the road network, and precipitation data.
[0017] Preferably, in step S11, the specific process of unifying the data format and spatial resolution of various environmental factors based on ground subsidence data includes:
[0018] Step S111: In ArcGIS software, the Euclidean distance tool is used to process the road network data and convert it into raster data representing the distance between each point in the study area and the road network.
[0019] Step S112: Spatial resampling is performed on three environmental factors: land cover data, distance data between land cover and road network, and precipitation data, to unify the spatial resolution to the resolution of the ground subsidence data.
[0020] Step S113: While maintaining a uniform resolution, set up the environment and use a snap-to-grid method to ensure that each grid of the environmental factor data and the ground settlement data accurately corresponds in spatial location.
[0021] Preferably, in step S12, the specific process of unifying the temporal resolution among various environmental factors and ensuring time alignment includes:
[0022] Based on the dates of each period of land subsidence data, directly look up the land cover data and the distance data between each point and the road network for that year to obtain land cover data and distance data between each point and the road network that are aligned with the time of each period of land subsidence data;
[0023] The monthly precipitation data obtained represents the cumulative precipitation in each month. Based on this data, the daily average precipitation in each month is calculated to obtain the cumulative precipitation data that is time-aligned with the ground subsidence data.
[0024] Among them, the data for each futures index is relative to the previous monitoring day.
[0025] Preferably, in step S13, a dual-scale dataset is created based on the obtained feature data with uniform spatiotemporal resolution, spatial location matching, and temporal alignment. The specific process includes:
[0026] First, before training the Transformer model, sample points are extracted from the raster data into an Excel spreadsheet file. During the dataset partitioning stage, a dual-scale strategy is adopted, dividing the training dataset and the test training set in a 7:3 ratio and then standardizing them.
[0027] Then, the training dataset, based on high-resolution feature data, acquires all four feature data by uniformly distributing effective random sampling points within the study area; the test dataset spatially resamples the acquired four feature data to a low resolution, extracting all raster data to achieve full coverage of the study area.
[0028] Finally, the Transformer model was trained using high-resolution training samples and tested using low-resolution global data.
[0029] Preferably, in step S3, the encoder input layer in the Transformer model is embedded and enhanced, and the specific process is as follows:
[0030] The encoder input layer in the improved Transformer model consists of three parts: a linear layer, a normalization layer, and ReLU activation.
[0031] (1) In the linear layer, the input four-dimensional feature data is projected to the high dimension, i.e. the hidden layer dimension space, to increase the expressive power of the features and learn the interaction relationship between multiple variables through the weight matrix;
[0032] (2) In the normalization layer, the hidden layer dimension of each sample is normalized to eliminate the difference in dimensions;
[0033] (3) In the ReLU activation part, the Transformer model is allowed to learn the nonlinear interaction relationship between features and establish dynamic coupling between multiple variables;
[0034] The basic principle of the ReLU activation function is as follows:
[0035] (1);
[0036] In the formula, Represents the ReLU activation function; Indicates the input value; when input... When the output is a linear value, Preserve the original information; when input When the output is 0, the neuron is completely suppressed.
[0037] Preferably, in step S4, the decoder input layer in the Transformer model is improved, specifically including:
[0038] A direct multi-step prediction method is adopted, which uses a dynamic queryer structure to dynamically generate the initial query vector of the target sequence input to the decoder based on the encoder output, so as to achieve future settlement prediction based on only four feature data. The specific process is as follows:
[0039] First, the memory matrix [batch size, input time step, hidden layer dimension] output by the encoder is averaged in the time dimension and compressed to obtain the global features [batch size, hidden layer dimension].
[0040] Then, the global features are mapped to [prediction time step, Transformer model hidden layer dimension] through a linear layer, and reshaped into [batch size, prediction time step, hidden layer dimension] dimensions to obtain the initial query vector of the target sequence and input into the decoder.
[0041] Preferably, in step S4, the decoder output layer in the Transformer model is enhanced, specifically including:
[0042] The improved decoder output layer consists of two linear layers, a normalization layer, and a GELU activation layer;
[0043] When the last decoder layer outputs high-dimensional features, i.e., the hidden layer dimension features, the first linear layer maps these high-dimensional features to the intermediate dimension and eliminates the dimensional differences through the normalization layer.
[0044] Then, the GELU activation function is used for nonlinear activation to capture the nonlinear correlation between mid-dimensional features, thereby maintaining the nonlinear trend between the sedimentation data of multiple time steps in the final output.
[0045] Finally, the mid-dimensional features after nonlinear activation are mapped to the low target dimension through the second linear layer, so as to achieve accurate output of direct multi-step settlement prediction results;
[0046] The basic principle of the GELU function is as follows:
[0047] (2);
[0048] (3);
[0049] In the formula, This represents the activation function of the Gaussian Error Linear Unit. It is the cumulative distribution function (CDF) of the standard Gaussian distribution.
[0050] Preferably, in step S5, the improved Transformer model is optimized and adjusted, specifically including:
[0051] The Adam W optimizer is used, and a dynamic learning rate adjustment strategy is employed to achieve convergence of the Transformer model.
[0052] The five hyperparameters of the Transformer model—hidden layer dimension, number of attention heads, number of encoding / decoding layers, batch size, and dropout rate—were determined using Bayesian optimization. The specific process includes:
[0053] First, a prior probability model of hyperparameters is established for the objective function. After each hyperparameter sampling, the output of the function is observed to update the prior knowledge.
[0054] Then, a posterior probability distribution is formed, and the Transformer model is iteratively updated and sampling points are selected to search for hyperparameter combinations within a finite number of function evaluations.
[0055] Therefore, this invention adopts the above-mentioned urban ground subsidence prediction method based on an improved Transformer model. It constructs a multi-feature dataset based on high-precision ground subsidence monitoring data and three environmental factors affecting its subsidence. It uses a dual-scale strategy to divide the training and testing datasets, optimizes and improves the Transformer deep learning method, and adopts a direct multi-step subsidence prediction method to realize urban ground subsidence prediction that takes into account the influence of multiple environmental factors.
[0056] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0057] Figure 1 This is a flowchart of the improved Transformer model of this invention;
[0058] Figure 2This is a flowchart illustrating the creation process of the feature dataset of this invention. Detailed Implementation
[0059] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0060] like Figure 1 As shown, a method for predicting urban land subsidence based on an improved Transformer model includes the following steps:
[0061] Step S1: Construct a multi-feature dual-scale dataset including environmental factor data and land subsidence data, such as... Figure 2 As shown.
[0062] Step S11: Using ground subsidence data as a benchmark, standardize the data format and spatial resolution of various environmental factors.
[0063] The environmental factor data includes land cover data, road network data, and precipitation data.
[0064] Step S111: In ArcGIS software, the Euclidean distance tool is used to process the road network data and convert it into raster data representing the distance between each point in the study area and the road network.
[0065] Step S112: Spatial resampling is performed on three environmental factors: land cover data, distance data between land cover and road network, and precipitation data, to unify the spatial resolution to the resolution of the ground subsidence data.
[0066] Step S113: While maintaining a uniform resolution, set up the environment and use a snap-to-grid method to ensure that each grid of the environmental factor data and the ground settlement data accurately corresponds in spatial location.
[0067] Step S12: Unify the time resolution among various environmental factors and ensure time alignment.
[0068] Under normal circumstances, land cover type and the distance between each point and the road network will not change drastically in a short period of time, and there is a lack of simple and effective scientific methods to accurately interpolate this type of data over time.
[0069] Therefore, based on the dates of the ground subsidence data for each period, the land cover data and the distance data between each point and the road network for that year can be directly retrieved to obtain land cover data and distance data between each point and the road network that are aligned with the time of the ground subsidence data for each period.
[0070] The acquired monthly precipitation data represents the cumulative precipitation for each month. Based on this data, the average daily precipitation for each month is calculated to obtain cumulative precipitation data aligned with the land subsidence data. The data for each period refers to the cumulative precipitation relative to the previous monitoring date.
[0071] Step S13: Based on the feature data with unified spatiotemporal resolution, spatial location matching and temporal alignment obtained in the above process, create a dual-scale dataset.
[0072] Among them, the four types of characteristic data with unified spatiotemporal resolution, matching spatial location and time alignment include: ground subsidence data, land cover data, distance data between each point and the road network, and precipitation data.
[0073] Unlike traditional tasks that focus on single-point prediction within areas of severe subsidence, this invention aims to achieve accurate overall prediction of large-scale ground subsidence in urban scenarios. Therefore, before training the Transformer model, sample points are extracted from raster data into an Excel spreadsheet file. During the dataset partitioning stage, a dual-scale strategy is adopted, dividing the training dataset and the test dataset in a 7:3 ratio, and then standardizing the dataset.
[0074] The training dataset, based on high-resolution feature data, acquires all four feature data types by uniformly distributing effective random sampling points within the study area (eliminating regions with missing data), thus enhancing the representativeness of the samples while ensuring the accuracy of the feature data. The test dataset, on the other hand, spatially resamples the acquired four feature data to a lower resolution, extracting all raster data and achieving full coverage of the study area under limited computing power.
[0075] Based on the above process, high-resolution training samples are used to train the Transformer model, and low-resolution global data are used to test the model. This allows the model to fully learn the accurate features of multivariate data in different regions during training, significantly improving the efficiency of large-scale settlement prediction while fully verifying the generalization ability of the Transformer model.
[0076] Step S2: Construct the Transformer model.
[0077] The Transformer model mainly consists of two parts: the encoder and the decoder. The specific structure is as follows:
[0078] Step S21: The encoder of the Transformer model includes an input layer, a positional encoder, and several encoding layers.
[0079] When training the Transformer model, four types of feature data (ground subsidence data, land cover data, distance data between each point and the road network, and precipitation data) are first converted into a high-dimensional embedding vector in the encoder input layer; then, a position encoding vector is added to the embedding vector through position encoding before it is input into the encoding layer.
[0080] Each coding layer consists of two parts: a multi-head attention network and a feedforward network.
[0081] (1) In multi-head attention, each input vector is first multiplied by a different weight matrix (obtained through training), mapping it to h subspaces to obtain h query, key, and value matrices, where h represents the number of attention heads. Then, the obtained h query, key, and value matrices are used to independently calculate the attention score for each head. Finally, the outputs of all attention heads are concatenated and fused through a linear layer.
[0082] (2) The feedforward network consists of two fully connected layers, which further enhances the Transformer model’s ability to learn nonlinear features by expanding the dimension of the hidden layers.
[0083] The above process enables the Transformer model to simultaneously capture complex features between different feature data and data at different time steps, thereby fully extracting the global features of multivariate time series data.
[0084] Step S22: The decoder of the Transformer model includes several decoding layers and an output layer.
[0085] Each decoding layer consists of three parts: Masked Multi-Head Attention, Encoder-Decoder Attention, and Feedforward Network.
[0086] (1) The masked multi-head attention part introduces a mask matrix (a matrix whose future time step position is infinitesimal) into the attention calculation, so that the attention weight of the future time step position is close to zero, thereby ensuring that when predicting the settlement value of the current time step, it is based only on the settlement data known in the past and cannot see the information of the future time step.
[0087] (2) The encoder-decoder attention part aligns the context memory (key and value) output by the encoder with the query of the decoder, and integrates the information above and below.
[0088] (3) The output of the last decoding layer in the decoder is used to obtain the prediction data for the future time step through linear mapping.
[0089] Furthermore, in the decoder and encoder of the Transformer model, the "fusion & normalization" unit after each part is for residual connections and layer normalization. The residual connection operation adds the input and output of each layer to alleviate the vanishing gradient problem in deep networks; layer normalization standardizes the activation values of each layer to ensure a consistent input distribution across layers, thereby improving the training efficiency and stability of the Transformer model.
[0090] Step S3: Embedding enhancement is performed on the encoder input layer in the Transformer model.
[0091] The encoder input layer in the improved Transformer model consists of three parts: a linear layer, a normalization layer, and ReLU (Rectified Linear Unit) activation.
[0092] (1) In the linear layer, the input four-dimensional feature data is projected to a high-dimensional (hidden layer dimension) space to increase the expressive power of the features and learn the interaction relationship between multiple variables through the weight matrix.
[0093] (2) In the normalization layer, the hidden layer dimension of each sample is normalized to eliminate the difference in dimensions and prevent certain features from dominating gradient updates due to excessively large numerical ranges.
[0094] (3) In the ReLU activation part, the Transformer model is allowed to learn the complex nonlinear interaction relationship between features, establish dynamic coupling between multiple variables, and suppress feature combinations that have no significant effect on settlement after linear mapping, thereby improving the robustness of the Transformer model and alleviating the gradient vanishing problem.
[0095] The basic principle of the ReLU activation function is as follows:
[0096] (1);
[0097] In the formula, This represents the ReLU (Rectified Linear Unit) activation function; This represents the input value. When input... When the output is a linear value, Preserve the original information; when input When the output is 0, the neuron is completely suppressed.
[0098] As shown in the above equation, there is no simple proportional relationship between the output and input of the ReLU activation function, breaking the limitation of linear superposition. Negative inputs are suppressed to 0, allowing only positive signals to pass, effectively reducing redundant computation; and the gradient in the positive interval is always 1, alleviating the gradient vanishing problem.
[0099] Step S4: Improve the decoder in the Transformer model.
[0100] Step S41: Improve the input layer of the decoder.
[0101] Traditional Transformer models employ a recursive multi-step prediction method in time series forecasting, outputting prediction results step by step, with each prediction dependent on the data from the previous step. However, in this invention's task of predicting future settlement at multiple steps based on historical settlement data and environmental factor data, the environmental factor data for future time steps is unknown. If the traditional recursive multi-step prediction method of the Transformer model is used, it would require predicting four feature data points for each future time step, significantly increasing the complexity of Transformer model training and making it impossible to verify the accuracy of the generated environmental factor data.
[0102] Therefore, this invention employs a direct multi-step prediction method. Based on a dynamic queryer structure, it dynamically generates the initial query vector of the target sequence that the decoder needs to input based on the encoder's output, thus achieving future settlement prediction based solely on four types of feature data. The specific process is as follows:
[0103] First, the memory matrix [batch size, input time step, hidden layer dimension] output by the encoder is averaged in the time dimension and compressed to obtain the global features [batch size, hidden layer dimension].
[0104] Then, the global features are mapped to [prediction time step, Transformer model hidden layer dimension] through a linear layer, and reshaped into [batch size, prediction time step, hidden layer dimension] dimensions to obtain the initial query vector of the target sequence and input into the decoder.
[0105] Step S42: Enhance the output layer of the decoder.
[0106] In direct multi-step prediction tasks, traditional Transformer models directly map the high-dimensional features output by the decoder to the target dimension through linear layers. On the one hand, this cannot remove irrelevant or redundant information that may be contained in the high-dimensional features, and on the other hand, it may ignore the interrelationships between the prediction outputs at different time steps.
[0107] To address this issue, the improved output layer comprises two linear layers, a normalization layer, and a GELU activation layer. After the last decoder layer outputs high-dimensional (hidden layer dimension) features, the first linear layer maps these features to an intermediate dimension, and the normalization layer eliminates dimensional differences. Then, a smoother GELU (Gaussian Error Linear Unit) function, compared to ReLU, is used for non-linear activation to capture the non-linear correlations between the intermediate-dimensional features, thus preserving the non-linear trend among the settlement data across multiple time steps in the final output. Finally, the second linear layer maps the non-linearly activated intermediate-dimensional features to a lower target dimension, achieving accurate output of direct multi-step settlement prediction results.
[0108] The basic principle of the GELU function is as follows:
[0109] (2);
[0110] (3);
[0111] In the formula, This represents the activation function of the Gaussian Error Linear Unit. It is the cumulative distribution function (CDF) of the standard Gaussian distribution.
[0112] From the above equation, we know that the GELU activation function is activated by the input value. Its probability of occurrence Multiplication allows for dynamic adjustment of activation intensity. When When it is large, GELU degenerates into a linear activation, similar to the positive interval states in the ReLU activation function; when When approaching 0, The smoothness of GELU makes the activation process gentler, preserving some negative information. Compared to the ReLU activation function, the GELU (Gaussian Error Linear Unit) activation function is a smooth, non-monotonic activation function with a non-zero derivative everywhere, alleviating the "neuron death" problem of the ReLU activation function. The GELU activation function adaptively adjusts the activation intensity through a gating mechanism (input multiplied by the cumulative distribution function of a Gaussian distribution), enabling more refined implementation of complex feature interactions.
[0113] Step S5: Optimize and adjust the improved Transformer model.
[0114] The Transformer model-based land subsidence prediction task, taking into account the continuous asymptotic physical characteristics of land subsidence, adopts the smoother and more continuous GELU as the activation function. Since the monitoring accuracy of InSAR-based land subsidence monitoring in severely subsidence areas may be reduced due to ground decoherence, the MSE loss function, which is more sensitive to large errors, is used.
[0115] Furthermore, the Adam W (Adaptive Moment Estimation with Weight Decay Decoupling) optimizer, which has stronger generalization performance than the traditional Adam (Adaptive Moment Estimation) optimizer, is employed. A dynamic learning rate adjustment strategy is used to accelerate the convergence of the Transformer model; specifically, the initial learning rate is set to 1×10⁻⁶. -3 If the validation loss does not improve during 20 consecutive iterations, then the learning rate is decayed by multiplying the learning rate by a decay factor of 0.5.
[0116] The five hyperparameters of the Transformer model—hidden layer dimension, number of attention heads, number of encoding / decoding layers, batch size, and dropout rate—were determined using Bayesian optimization. The specific process includes:
[0117] First, a prior probability model of hyperparameters is established for the objective function. After each hyperparameter sampling, the output of the function is observed to update the prior knowledge.
[0118] Then, a posterior probability distribution is formed, and the Transformer model is iteratively updated and sampling points are selected to efficiently search for the optimal combination of hyperparameters within a limited number of function evaluations.
[0119] Example 1
[0120] Predicting ground subsidence in a certain area of City J based on multi-feature datasets and an improved Transformer model.
[0121] I. Creation of multi-feature datasets.
[0122] First, in ArcGIS software, the Euclidean distance tool is used to process the road network vector data, converting it into raster data representing the distance between each point in the study area and the road network.
[0123] Then, spatial resampling was performed on the data of the three environmental factors to unify the spatial resolution to 30×30 m.
[0124] Finally, based on the dates of the 66 periods of ground subsidence data (excluding the first period where all monitoring data were 0), the land cover data and the distance data between each point and the road network for that year were directly retrieved; the cumulative precipitation data for each of the 66 monitoring periods were calculated based on the monthly cumulative precipitation data.
[0125] In the dataset partitioning stage, a dual-scale strategy was adopted. The training dataset was based on high-resolution (30×30 m) feature data. By uniformly distributing 10,000 effective random sampling points in the study area (excluding areas with missing data), the four feature parameters were fully obtained, which enhanced the representativeness of the samples while ensuring the accuracy of the feature variable data.
[0126] The test dataset resamples the four feature data spaces to a low resolution of 300×300 m, fully extracts all raster data, and achieves full coverage of the study area under limited computing power.
[0127] II. Hyperparameter optimization of Transformer model.
[0128] In a practical application task of predicting future land subsidence based on four types of feature data, the improved Transformer model had the smallest test error under the hyperparameter combination of 64 hidden layer dimensions, 4 attention heads, 2 encoding / decoding layers, 32 batch size, and 0.1 dropout rate. It was determined to be the optimal hyperparameter combination for the multivariate time series Transformer subsidence prediction model.
[0129] Based on the established multivariate time series Transformer subsidence prediction model and optimal hyperparameter combination settings, 55 periods of high-precision land subsidence and spatiotemporally aligned environmental factor data from January 3, 2018 to May 1, 2023 were used to predict 11 periods of land subsidence from June 6, 2023 to May 31, 2024. The prediction results of the multivariate time series Transformer subsidence prediction model (Transformer-Multivariate model) were compared and analyzed with the SBAS-InSAR monitoring results within the same monitoring period.
[0130] As shown in Table 1, the settlement results predicted by the Transformer-Multivariate model and the settlement results monitored by SBAS-InSAR remained consistent over 11 monitoring periods. This indicates that the prediction method proposed in this invention has a certain degree of reliability. A comparative analysis with Table 2 reveals that the RMSE and R-values for prediction using single settlement data and prediction using multi-feature datasets are significantly lower. 2The results showed that the latter not only had smaller errors but also higher fitting accuracy, indicating that using multi-feature datasets for model training can effectively improve the accuracy of subsidence prediction and reflect the true situation of urban ground subsidence.
[0131] Table 1 Comparison of Transformer model prediction results and SBAS-InSAR monitoring results
[0132] ;
[0133] Table 2 Comparison of Transformer model prediction results for single-factor and multi-feature datasets.
[0134] ;
[0135] The results show that the improved Transformer model constructed in this invention can fully extract the complex correlation features between multiple environmental factors. The prediction accuracy is not only better than the prediction results using single settlement data, but also close to the SBAS-InSAR monitoring results. This indicates that the present invention can not only achieve high-precision ground settlement prediction results, but also has higher prediction accuracy and better fitting effect than the prediction results using a single settlement dataset.
[0136] Example 2
[0137] A comparative analysis of urban land subsidence prediction using the improved Transformer model and three traditional recurrent neural network models commonly used for time series forecasting tasks—RNN, GRU, and LSTM—was conducted using the same training, validation, and test sets.
[0138] For the four deep learning models—Transformer, RNN, GRU, and LSTM—compared to those using single-feature datasets, the RMSE of predictions using multi-feature datasets decreased by 6.83%, 1.66%, 0.64%, and 0.21%, respectively, as shown in Table 3; R 2 The increases were 0.76%, 0.27%, 0.09%, and 0.01%, respectively, as shown in Table 4.
[0139] This demonstrates that using multi-feature datasets enables deep learning models to learn more key features influencing future subsidence trends during training, thereby improving subsidence prediction accuracy. Furthermore, the Transformer model's multi-head attention mechanism allows it to fully learn complex features among multiple variables across multiple subspaces; using multi-feature datasets results in a greater improvement in model prediction accuracy.
[0140] On the other hand, regardless of whether training is based on a single-feature dataset or a multi-feature dataset, the average RMSE rankings of the four deep learning models for each period are: Transformer < GRU < LSTM < RNN; the average R for each period 2 rankings are: Transformer > GRU > LSTM > RNN.
[0141] This shows that the accuracy of ground settlement prediction using deep learning methods is more affected by the model structure adopted. Compared with the three traditional recurrent neural network models, the self-attention mechanism of the Transformer model can effectively capture the long-range dependence relationships in the historical settlement data of time series, thus significantly improving the settlement prediction accuracy.
[0142] Generally speaking, the improved Transformer model constructed in the present invention is consistently superior to other models in terms of the two indicators of RMSE and R 2 The average RMSE of the 11-period ground settlement prediction results is 6.0387 mm, and the average R 2 reaches 0.9564, demonstrating excellent prediction performance.
[0143] Table 3 RMSE of the prediction results of each model
[0144] ;
[0145] Table 4 R of the prediction results of each model 2
[0146] ;
[0147] In addition, when compared with common recurrent neural network models, the settlement prediction accuracies of multiple models using multi-feature datasets have increased by 20.6%, 1.4%, and 7.9% respectively compared with the three recurrent neural network models of RNN, GRU, and LSTM. The improved Transformer model constructed in the present invention based on multi-feature datasets has higher prediction accuracy and better model fitting effect compared with other models. The present invention not only maintains relatively stable accuracy in multi-period ground settlement prediction, but also has certain advantages and development potential in the urban ground settlement prediction task.
[0148] Therefore, the present invention adopts the above-mentioned urban ground settlement prediction method based on an improved Transformer model, and through improving the Transformer model structure, direct multi-step prediction of urban ground settlement is realized based on multi-factor influencing factors, improving the accuracy and generalization ability of the Transformer model for urban ground settlement prediction.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for predicting urban land subsidence based on an improved Transformer model, characterized in that, Includes the following steps: Step S1: Construct a multi-feature dual-scale dataset including environmental factor data and land subsidence data. The specific process is as follows: Step S11: Based on the land subsidence data, standardize the data format and spatial resolution of each environmental factor; among which, the environmental factor data includes land cover data, road network data and precipitation data; Step S12: Unify the time resolution among various environmental factors and ensure time alignment; Step S13: Based on the feature data with unified spatiotemporal resolution, spatial location matching and temporal alignment obtained in the above process, create a dual-scale dataset; Among them, the four types of characteristic data with unified spatiotemporal resolution, spatial location matching and time alignment include: ground subsidence data, land cover data, distance data between each point and the road network, and precipitation data. Step S2: Construct the Transformer model, including the encoder and decoder; Step S3: Embedding enhancement is performed on the encoder input layer in the Transformer model. The specific process is as follows: The encoder input layer in the improved Transformer model consists of three parts: a linear layer, a normalization layer, and ReLU activation. (1) In the linear layer, the input four-dimensional feature data is projected to the high dimension, i.e. the hidden layer dimension space, to increase the expressive power of the features and learn the interaction relationship between multiple variables through the weight matrix; (2) In the normalization layer, the hidden layer dimension of each sample is normalized to eliminate the difference in dimensions; (3) In the ReLU activation part, the Transformer model is allowed to learn the nonlinear interaction relationship between features and establish dynamic coupling between multiple variables; The basic principle of the ReLU activation function is as follows: (1); In the formula, Represents the ReLU activation function; Indicates the input value; when input... When the output is a linear value, Preserve the original information; when input When the output is 0, the neuron is completely suppressed; Step S4: Improve the decoder in the Transformer model, specifically including: Step S41: Using a direct multi-step prediction method, based on a dynamic queryer structure, the initial query vector of the target sequence input to the decoder is dynamically generated according to the encoder output, realizing future settlement prediction based only on four feature data. The specific process is as follows: First, the memory matrix [batch size, input time step, hidden layer dimension] output by the encoder is averaged in the time dimension and compressed to obtain the global features [batch size, hidden layer dimension]. Then, the global features are mapped to [prediction time step, Transformer model hidden layer dimension] through a linear layer, and reshaped into [batch size, prediction time step, hidden layer dimension] dimensions to obtain the initial query vector of the target sequence and input into the decoder; Step S42: Enhance the decoder output layer in the Transformer model, specifically including: The improved decoder output layer consists of two linear layers, a normalization layer, and a GELU activation layer; When the last decoder layer outputs high-dimensional features, i.e., hidden layer dimensional features, the first linear layer maps these high-dimensional features to intermediate dimensions and eliminates dimensional differences through a normalization layer. Then, the GELU activation function is used for nonlinear activation to capture the nonlinear correlation between mid-dimensional features, thereby maintaining the nonlinear trend between the sedimentation data of multiple time steps in the final output. Finally, the mid-dimensional features after nonlinear activation are mapped to the low target dimension through the second linear layer, so as to achieve accurate output of direct multi-step settlement prediction results; The basic principle of the GELU function is as follows: (2); (3); In the formula, This represents the activation function of the Gaussian Error Linear Unit. It is the cumulative distribution function (CDF) of the standard Gaussian distribution; Step S5: Optimize and adjust the improved Transformer model to achieve urban land subsidence prediction, specifically including: The Adam W optimizer is used, and a dynamic learning rate adjustment strategy is employed to achieve convergence of the Transformer model. The five hyperparameters of the Transformer model—hidden layer dimension, number of attention heads, number of encoding / decoding layers, batch size, and dropout rate—were determined using Bayesian optimization. The specific process includes: First, a prior probability model of hyperparameters is established for the objective function. After each hyperparameter sampling, the output of the function is observed to update the prior knowledge. Then, a posterior probability distribution is formed, and the Transformer model is iteratively updated and sampling points are selected to search for hyperparameter combinations within a finite number of function evaluations.
2. The urban land subsidence prediction method based on an improved Transformer model according to claim 1, characterized in that, In step S11, the specific process of unifying the data format and spatial resolution of various environmental factors based on ground subsidence data includes: Step S111: In ArcGIS software, the Euclidean distance tool is used to process the road network data and convert it into raster data representing the distance between each point in the study area and the road network. Step S112: Spatial resampling is performed on three environmental factors: land cover data, distance data between land cover and road network, and precipitation data, to unify the spatial resolution to the resolution of the ground subsidence data. Step S113: While maintaining a uniform resolution, set up the environment and use a snap-to-grid method to ensure that each grid of the environmental factor data and the ground settlement data accurately corresponds in spatial location.
3. The urban land subsidence prediction method based on an improved Transformer model according to claim 1, characterized in that, In step S12, the specific process of unifying the temporal resolution among various environmental factors and ensuring time alignment includes: Based on the dates of each period of land subsidence data, directly look up the land cover data and the distance data between each point and the road network for that year to obtain land cover data and distance data between each point and the road network that are aligned with the time of each period of land subsidence data; The monthly precipitation data obtained represents the cumulative precipitation in each month. Based on this data, the daily average precipitation in each month is calculated to obtain the cumulative precipitation data that is time-aligned with the ground subsidence data. Among them, the data for each futures index is relative to the previous monitoring day.
4. The urban land subsidence prediction method based on an improved Transformer model according to claim 1, characterized in that, In step S13, based on the obtained feature data with uniform spatiotemporal resolution, spatial location matching, and temporal alignment, a dual-scale dataset is created. The specific process includes: First, before training the Transformer model, sample points are extracted from the raster data into an Excel spreadsheet file. During the dataset partitioning stage, a dual-scale strategy is adopted, dividing the training dataset and the test training set in a 7:3 ratio and then standardizing them. Then, the training dataset, based on high-resolution feature data, acquires all four feature data by uniformly distributing effective random sampling points within the study area; the test dataset spatially resamples the acquired four feature data to a low resolution, extracting all raster data to achieve full coverage of the study area. Finally, the Transformer model was trained using high-resolution training samples and tested using low-resolution global data.
Citation Information
Patent Citations
Transform-based InSAR technology permafrost region multivariable time sequence deformation prediction method and device
CN114966692A
Design method of cement mixing pile for controlling differential settlement of lake region expressway
CN115935725A