Shield tunneling machine duct piece maximum floating quantity prediction method based on improved TSMIXer model

By improving the TSMixer model, introducing a self-attention mechanism and GraphNet network structure, the problems of low prediction accuracy of the maximum upwelling of the pipe sheet in the shield machine construction are solved, and a higher accuracy and robust prediction effect is achieved.

CN120123969APending Publication Date: 2025-06-10CHINA RAILWAY SHISIJU GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510165806.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, during the construction of the shield machine, the prediction accuracy of the maximum uplift of the pipe sheet is low, the model adaptability is poor, and it is difficult to effectively model the timing dependence and spatial relationship of sensor data.

Method used

The TSMixer model is improved, and the self-attention mechanism and GraphNet network structure are introduced to capture the timing dependencies in time series data through the self-attention mechanism. The GraphNet network structure models the spatial dependencies between sensors through graph convolution operations.

Benefits of technology

It improves the accuracy and robustness of the prediction of the maximum uplift of the pipe sheet, enhances the model's ability to capture complex space-time dependencies, and improves the generalization ability of the model and adaptability in the construction environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123969A_ABST
    Figure CN120123969A_ABST
Patent Text Reader

Abstract

The invention provides a shield tunneling machine duct piece maximum floating quantity prediction method based on an improved TSMIX model. The method comprises the steps of data collection, wherein related data of a shield tunneling machine in the construction process are collected from a sensor; preprocessing the data; the method comprises the following steps: constructing an improved TStool model, and introducing a self-attention mechanism and a GraphNet network structure; data set division: dividing the preprocessed data into a training set, a verification set and a test set according to a proportion of 8: 1: 1, and adjusting hyper-parameters by using a cross validation method; model training: performing model training by adopting the divided training set; an improved TStool model is evaluated, and the constructed improved TStool model is evaluated by using the test set; and inputting to-be-predicted construction data into the trained improved TSMIXer model, and outputting the predicted maximum floating amount of the duct piece. According to the method, the maximum floating amount of the duct piece in the construction process of the shield tunneling machine can be accurately predicted through fine representation and fusion of multi-sensor data, and real-time and accurate construction decision support is provided for a construction party.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and tunnel construction, and in particular, is a method for predicting the maximum floating amount of a shield machine segment based on an improved TSMixer model. Background Art

[0002] During tunnel construction, the shield machine, as a widely used construction equipment, undertakes multiple tasks such as soil excavation and segment installation. However, during the advancement of the shield machine, the floating phenomenon of the segments is one of the important factors affecting the construction quality and safety. Segment floating refers to the vertical floating phenomenon of the tunnel segments during the construction process due to changes in factors such as stratum pressure, propulsion force, and grouting volume. If this phenomenon cannot be predicted and controlled in time, it will cause deformation of the tunnel structure and even affect the subsequent construction process. In severe cases, it may lead to construction safety accidents.

[0003] Traditionally, the prediction methods of pipe segment buoyancy are mostly based on empirical formulas or simple regression models. Although these methods can provide certain prediction results in some cases, due to the complexity of the construction environment, the variability of geological conditions and the nonlinear characteristics of the construction process, the existing methods have great limitations in accuracy and applicability. Especially in the actual construction process, involving complex data interactions and associations of multiple sensors, traditional methods often find it difficult to fully consider and capture the impact of these multiple factors on the buoyancy of the pipe segment. Therefore, how to establish a high-precision pipe segment buoyancy prediction model has become a technical problem that needs to be solved urgently in the field of tunnel construction.

[0004] With the rapid development of deep learning technology, time series prediction models have demonstrated excellent capabilities in processing time series data. Time series prediction models such as recurrent neural networks (RNN), long short-term memory networks (LSTM) and transformers have achieved remarkable results in many fields. These models can extract patterns from historical data and make accurate predictions by learning the temporal dependencies in time series data. However, traditional time series models often have problems such as low computational efficiency, model overfitting, and insufficient modeling capabilities for long time series dependencies when processing complex multivariate time series data. Especially in complex engineering environments such as tunnel construction, which involves the interaction and correlation of multiple sensor data, a single time series model often finds it difficult to capture all key factors.

[0005] To improve the accuracy and efficiency of time series prediction, the TSMixer model, as a new type of time series prediction method, has achieved good results in the field of time series data processing due to its efficient modeling ability and flexible feature expression. By mixing the time dimension information of multiple features, TSMixer can better adapt to the complex patterns in time series data. Although TSMixer has achieved good application effects in some fields, there are still certain deficiencies in dealing with complex multi-sensor data, especially in the modeling of spatial dependence relationships.

[0006] In the prediction task of the shield segment floating amount, in addition to the temporal dependence relationship of time series data, the spatial dependence relationship between sensors is also crucial. The data provided by each sensor is not only related to time but also affected by the spatial position and interaction between sensors. The traditional TSMixer model fails to fully consider the spatial relationship between sensors, which limits the performance of the model in multi-sensor data processing and restricts its accuracy and robustness in practical applications. Therefore, how to effectively introduce spatial dependence in time series data modeling and solve the potential mining problem in multi-sensor data processing has become the key to current technical research.

[0007] To address the above challenges and improve the TSMixer model, introducing the self-attention mechanism and the Graph Neural Network (GraphNet) to better handle the prediction task of the shield segment floating amount has become an urgent technical problem to be solved. Through the self-attention mechanism, the model can automatically adjust the relationship between each time step and adaptively weight the correlation between different sensor data, thereby improving the model's ability to capture complex spatio-temporal dependence relationships. In addition, combining the Graph Neural Network (GraphNet) to model the spatial dependence relationship between sensors can effectively improve the model's representation ability for multi-sensor data and further enhance the prediction accuracy. This improvement not only solves the deficiencies of the traditional model in dealing with long temporal dependencies and capturing complex spatial relationships but also effectively improves the generalization ability of the model.

[0008] Therefore, by introducing the self-attention mechanism and the Graph Neural Network (GraphNet), the present invention can more accurately capture spatio-temporal dependence relationships, improve the accuracy of predicting the shield segment floating amount, and solve the deficiencies of existing models in time series data processing and multi-sensor data modeling. These technical improvements will effectively enhance the safety and construction progress of tunnel construction and provide a new efficient prediction method for related fields. Summary of the Invention

[0009] Objective of the invention: To propose a method for predicting the maximum floating amount of shield tunnel segments based on an improved TSMixer model, so as to solve the problems in the prior art such as low prediction accuracy of the maximum floating amount of tunnel segments, poor model adaptability, and lack of effective modeling of the temporal dependence and spatial relationship of sensor data during the construction of shield machines.

[0010] Technical solution: A method for predicting the maximum floating amount of shield tunnel segments based on an improved TSMixer model, including:

[0011] S1. Data collection: Collect relevant data of the shield machine during construction from sensors, including at least parameters related to the propulsion of the shield machine: soil pressure, segment stress, and construction speed;

[0012] S2. Data preprocessing, including removing invalid or abnormal data to ensure the quality of the input data; filling in missing values in the data to avoid the impact of incomplete data on model training; normalizing the data of each sensor to eliminate the difference in dimensions; screening out variables with little influence on the prediction of the floating amount of tunnel segments through an adaptive feature selection algorithm, and retaining important influencing factors;

[0013] S3. Construct an improved TSMixer model, introduce a self-attention mechanism and a GraphNet network structure to effectively process the temporal dependence and spatial dependence in sensor data. The model includes: The self-attention mechanism captures the correlation between different time steps in time series data by calculating the relationship between query, key, and value vectors; The GraphNet network structure models the spatial relationship between sensor data through graph convolution operations to enhance the model's understanding of the spatial dependence between sensors;

[0014] S4. Dataset division: Divide the preprocessed data in step S2 into a training set, a validation set, and a test set according to a ratio of 8:1:1, and use the cross-validation method to adjust the hyperparameters to optimize the model training process, which is used to improve the generalization ability of the model on unknown data;

[0015] S5. Model training: Use the training set divided in step S4 to train the model;

[0016] S6. Evaluation of the improved TSMixer model: Use the test set in step S4 to evaluate the improved TSMixer model constructed in step S3. The evaluation indicators include the mean square error MSE, the mean absolute error MAE, and the R-square value R 2 , and conduct a comparative experiment with other comparative models. The comparative models include: the original TSMixer model, the iTransformer model, and the PatchMixer model;

[0017] S7. Input the construction data to be predicted into the trained improved TSMixer model, and output the predicted maximum segment floating amount, so as to provide accurate construction suggestions for the construction party.

[0018] According to a further improvement of the present invention, in S21, by detecting missing values, outliers and illogical records, invalid or abnormal data are removed, and the missing values are filled by the mean imputation method to avoid the negative impact of incomplete data on model training.

[0019] S22. Normalize the data of each sensor to eliminate the influence of dimensional differences on model training, so as to improve the convergence and stability of the model.

[0020] S23. Through the adaptive feature selection algorithm, variables with little influence on the prediction of the segment floating amount are screened out, and key features with significant effects on the prediction accuracy are retained.

[0021] S24. Generate a data set, which contains 584 rows of data, covering parameters collected by multiple sensors, including: the pressures of the 1st to 6th excavation chambers, rolling angle, cutter head speed, cutter head torque, the detection of the inlet pressure of the slurry feeding pump, and penetration; and the maximum segment floating amount recorded during construction is used as the label of the data set.

[0022] According to a further improvement of the present invention, the self-attention mechanism in step S3 is a mechanism for capturing the correlation between different time steps in the input sequence. When predicting the maximum segment floating amount, the self-attention mechanism assigns different attention weights to each time step, focuses on the time points with greater influence on the prediction result, and thus improves the ability to extract time series features. Specifically, it includes the following steps:

[0023] S3a1. Given the input sequence where T is the number of time steps and d is the feature dimension of each time step, map it to query, key and value vectors:

[0024] where, is a learnable parameter matrix, and d k is the embedding dimension;

[0025] S3a2. The dot product of the query Q and the key K represents the attention score, and then it is scaled and normalized:

[0026]

[0027] where, Softmax is used to normalize the weights of each time step so that the sum of the weights is 1, is used to prevent the numerical value from being too large;

[0028] S3a3. Use the multi-head attention mechanism to divide the query, key, and value into h subspaces and calculate them in parallel:

[0029] MultiHead(Q, K, V) = Concat(head 1 , …, head h )W O ;

[0030] where each head i = Attention(Q i , K i , V i ), and W O is the output projection matrix.

[0031] According to a further improvement of the present invention, the GraphNet network structure in step S3 models the spatial dependencies between sensors through graph convolution operations. The specific steps are as follows:

[0032] Step S3b1. Define the dependencies between different sensors: Consider the input data of the graph convolution network as a graph, where each node in the graph corresponds to a sensor, and the connection relationship between nodes is represented by the adjacency matrix ; Initially, the adjacency matrix is an identity matrix, indicating that the relationships between all sensors are initially independent. As the model is trained, the adjacency matrix will be adjusted as a learnable parameter to automatically learn the spatial relationships between sensors. The definition of the adjacency matrix is as follows:

[0033] where is an n pts × n pts identity matrix, and n pts represents the number of sensor parameters;

[0034] Step S3b2. Extract the feature information of the sensor nodes;

[0035] Step S3b3. Graph convolution operation: Aggregate the feature information of each sensor node through the adjacency matrix , that is, perform a weighted average of the features of each node and its adjacent nodes to obtain the updated node features. The mathematical formula for the graph convolution operation can be expressed as:

[0036] where H (l) is the node feature matrix of the first layer. For the input data, initially H (0) is the input sensor data; is the normalized adjacency matrix, which is used to define the connection relationship between nodes in the graph; W(l) is the learnable weight matrix of the first layer, which controls the transformation of node features; σ is the activation function, used to introduce non-linearity;

[0037] Step Sb4: Combine the features of each node with the features of its adjacent nodes to extract the spatial information of the nodes, and obtain the spatial dependence relationship between sensors.

[0038] According to a further improvement of the present invention, the GraphNet network includes two graph convolutional layers. The first graph convolution expands the input feature dimension from in_features to 2×in_features, and the second graph convolution reduces it to out_features, and finally outputs the prediction result; specifically as follows:

[0039] The first graph convolution: The feature matrix X with an input dimension of n pts ×in ― features is multiplied by the adjacency matrix to obtain the weighted node features; the feature dimension is expanded through a fully connected layer to obtain a new feature matrix X 0 , and the dimension becomes n pts ×2×in ― features; its calculation process is:

[0040] Among them, W 1 is the learnable weight matrix of the first graph convolution, and σ is the activation function, usually the ReLU function;

[0041] The second graph convolution: The feature matrix X with a dimension of n pts ×2×in ― features is 0 multiplied by the adjacency matrix again for weighted aggregation; the feature dimension is compressed to out_features through a fully connected layer to obtain the final output X 1 , and this output is the predicted value of the maximum uplift of the segment; the calculation process of this layer is:

[0042] Among them, W 2 is the learnable weight matrix of the second graph convolution.

[0043] According to a further improvement of the present invention, hyperparameter tuning is required during the training process of step S5. The specific training parameters include: the batch size is set to 4, the initial learning rate is set to 0.001, the optimizer is set to the Adam optimizer, the number of training epochs is set to 100, and the loss function is set to the mean squared error MSE; in each round of training, the input data is calculated through forward propagation to generate a prediction result; then, the error between the prediction result and the true value is calculated, and the model parameters are updated through backpropagation.

[0044] Beneficial effects: By introducing the self-attention mechanism, graph convolutional network, and adaptive feature selection algorithm, the present invention can effectively capture the temporal and spatial dependencies in the data and automatically screen out the most meaningful features, thereby improving the accuracy and robustness of the prediction; the improved TSMixer model in the present invention has strong adaptability and can provide accurate prediction results in a complex construction environment, effectively improving the safety and stability of construction; the present invention can accurately predict the maximum floating amount of the segment, which can help the construction party timely discover potential problems and take necessary countermeasures, thereby reducing the risks during construction and ensuring the safety and smooth progress of the construction process. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is the overall flowchart of the solution of the present invention.

[0046] Figure 2 is a schematic diagram of the architecture of the improved TSMixer model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] Referring to Figure 1 and Figure 2 the present invention provides a technical solution: a method for predicting the maximum floating amount of the segment of a shield machine based on an improved TSMixer model. As Figure 1 shown, it includes the following steps:

[0049] The first step is to collect relevant data during the construction of the shield machine from multiple sensors. Specifically, it includes a series of parameters related to the advancement of the shield machine, such as soil pressure, segment stress, construction speed, etc. These data reflect the working environment, soil layer state, and cutterhead operation situation faced by the shield machine during construction. The collected data needs to be preprocessed to ensure quality: First, by detecting missing values, outliers, and illogical records, invalid or abnormal data is removed, and the missing values are filled using the mean imputation method to avoid the negative impact of incomplete data on model training; Second, the data of each sensor is normalized to eliminate the influence of dimensional differences on model training, thereby improving the convergence and stability of the model; Finally, through an adaptive feature selection algorithm, variables with less influence on the prediction of segment floating amount are screened out, and key features that have a significant effect on the prediction accuracy are retained. The core idea of the adaptive feature selection algorithm is to dynamically evaluate the contribution of each feature to the prediction result through the model training process, and gradually screen out the most informative features, thereby improving the generalization ability and prediction accuracy of the model. Specifically, a method based on Gradient Boosting and importance scoring is used to rank and select features. This algorithm repeatedly trains the model and uses the error backpropagation of the model to evaluate the relative importance of each feature, and finally generates an adaptive feature selection mechanism. Finally, the processed dataset contains 584 rows of data, covering parameters collected by multiple sensors, including the pressure of the 6th excavation chamber, cutterhead speed, detection of the inlet pressure of the slurry pump for mud feeding, etc., as the input for subsequent model training and prediction.

[0050] The second step is to use an improved TSMixer model to predict the maximum floating amount of the segments. The improved TSMixer model combines the self-attention mechanism and the GraphNet network structure to effectively handle the temporal dependence and spatial dependence in the sensor data. The model structure is as Figure 2 shown.

[0051] The self-attention mechanism is implemented through the following steps: First, the input time series data (where T is the number of time steps and d is the feature dimension of each time step) is mapped to query Q, key K, and value V:

[0052] where, W Q 、W K 、W V are learnable parameter matrices;

[0053] Second, the attention scores are obtained by calculating the dot product between query Q and key K, and are normalized through the Softmax function:

[0054] where, dk is the embedding dimension;

[0055] Finally, through the multi-head attention mechanism, the query, key, and value are divided into multiple sub-spaces and calculated in parallel:

[0056] MultiHead(Q, K, V) = Concat(head 1 , …, head h )W O ;

[0057] where each head i = Attention(Q i , K i , V i ), and W O is the output projection matrix.

[0058] To handle the spatial relationships of sensor data, the improved TSMixer model introduces the GraphNet network structure. GraphNet effectively models the mutual relationships between sensors through graph convolution operations. In the present invention, the core idea of GraphNet is to regard sensors as nodes in a graph, and the relationships between sensors are represented by an adjacency matrix . Specifically, graph convolution operations are used to perform weighted averaging on sensor data, thereby extracting the spatial dependence information between nodes and enhancing the model's ability to process spatial features. In GraphNet, the adjacency matrix is initially an identity matrix, and as the training process progresses, the adjacency matrix will be adjusted as a learnable parameter. The calculation process of the graph convolution operation is as follows:

[0059] where H (l) is the node feature matrix of the l-th layer, W (l) is the learnable weight matrix of the l-th layer, controlling the transformation of node features, and σ is an activation function (such as ReLU) for introducing non-linearity. Through this process, the model can extract and fuse the spatial dependence information in sensor data layer by layer.

[0060] In the present invention, the GraphNet network includes two graph convolution layers. The role of each graph convolution layer is to process the input features and gradually extract spatial features. The first graph convolution layer expands the input feature dimension from in_features to 2×in_features, and the second graph convolution layer reduces it to out_features, and finally outputs the prediction result. Specifically, the working principle of the graph convolution layer of GraphNet is as follows:

[0061] The first graph convolution layer: The role of this layer is to expand the input feature matrix X (with dimension npts ×in ― features) is multiplied by the adjacency matrix to obtain weighted node features. Then, the feature dimension is expanded through a fully connected layer to obtain a new feature matrix X 0 , and the dimension becomes n pts ×2×in ― features. The calculation process of this layer is as follows:

[0062] where W 1 is the learnable weight matrix of the first graph convolution, and σ is the activation function, usually the ReLU function is selected.

[0063] Second graph convolution: In this layer, the feature matrix X 0 (with dimension n pts ×2×in ― features) is multiplied by the adjacency matrix again for weighted aggregation. Then, the feature dimension is compressed to out_features through a fully connected layer to obtain the final output X 1 , and this output is the predicted value of the maximum uplift of the segment. The calculation process of this layer is as follows:

[0064] where W 2 is the learnable weight matrix of the second graph convolution.

[0065] Third step, in order to ensure that the model can effectively learn from the data and have good generalization ability, hyperparameter tuning is required during the training process. The specific training parameters include: batch size: 4, initial learning rate: 0.001, optimizer: Adam optimizer, number of training epochs: 100, loss function: mean squared error (MSE). In each round of training, the input data is calculated through forward propagation to generate prediction results. Then, the error between the prediction result and the true value is calculated, and the model parameters are updated through backpropagation. During the training process, a validation set is also used to evaluate the model performance, and the learning rate and other hyperparameters are adjusted according to the validation set results to further improve the training effect and prediction accuracy of the model.

[0066] Fourthly, after the model training is completed, use the test set to evaluate the improved TSMixer model and conduct a comparative experiment with other time series prediction networks to further verify the superiority of the improved model. The compared models include the original TSMixer model, the iTransformer model, and the PatchMixer model, and these models are all representative models in the current time series prediction field. Through the comparative experiment, the performance of different models in the prediction task of the maximum uplift of the segment can be comprehensively evaluated, so as to select the best model. The results of the comparative experiment are shown in Table 1. It can be seen that the improvement strategy of TSMixer in the present invention helps to improve the prediction effect of the model on the maximum uplift of the shield segment, and all aspects of indicators are significantly better than the baseline network.

[0067] Table 1

[0068] Method MSE MAPE MSPE <![CDATA[R 2 > Improved TSMixer 0.1702 0.4685 0.4100 0.9066 TSMixer 0.2060 2.0094 49.7843 0.8870 iTransformer 0.2458 1.6342 13.8972 0.8651 PatchMixer 0.2063 1.4181 11.2738 0.8868

[0069] Fifthly, after the model training is completed, the construction party can input the construction data to be predicted into the trained improved TSMixer model. After being processed by the model, the predicted value of the maximum uplift of the segment is output. Through this prediction result, the construction party can predict the change trend of the segment uplift in advance and take necessary countermeasures to avoid unexpected uplift phenomena during the construction process and ensure the construction safety.

[0070] The method for predicting the maximum uplift of the shield segment based on the improved TSMixer model provided by the present invention can effectively improve the prediction accuracy of the maximum uplift of the segment and the robustness of the model by introducing the self-attention mechanism and the GraphNet network structure. By deeply mining the time series data and spatial data, this method can help the construction party identify construction risks in advance and ensure the safety and stability of the shield machine construction.

[0071] The above are all preferred examples of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting the maximum floating amount of shield machine segments based on an improved TSMixer model, characterized in that: The following steps are involved: S1. Data collection: Collect relevant data of the shield machine during the construction process from sensors, including at least parameters related to the advancement of the shield machine: soil layer pressure, segment stress and construction speed; S2. Data preprocessing, including removing invalid or abnormal data to ensure the quality of input data; filling missing values ​​in the data to avoid the impact of incomplete data on model training; normalizing the data of each sensor to eliminate dimensional differences; using an adaptive feature selection algorithm to filter out variables with little impact on the prediction of the floating volume of the pipe segment and retain important influencing factors; S3. Build an improved TSMixer model, introduce the self-attention mechanism and GraphNet network structure to effectively process the temporal dependency and spatial dependency in sensor data. The model includes: the self-attention mechanism captures the correlation between different time steps in the time series data by calculating the relationship between the query, key and value vectors; the GraphNet network structure models the spatial relationship between sensor data through graph convolution operations to enhance the model's understanding of the spatial dependency between sensors; S4, data set division, the pre-processed data in step S2 is divided into a training set, a validation set and a test set in a ratio of 8:1:1, and the hyperparameters are adjusted using the cross-validation method to optimize the model training process and improve the generalization ability of the model on unknown data; S5, model training, using the training set divided in step S4 to perform model training; S6, Improved TSMixer model evaluation, use the test set in step S4 to evaluate the improved TSMixer model constructed in step S3, the evaluation indicators include mean square error MSE, mean absolute error MAE, R square value R 2 , and conduct comparative experiments with other comparison models, including: the original TSMixer model, the iTransformer model, and the PatchMixer model; S7. Input the construction data to be predicted into the trained improved TSMixer model, and output the predicted maximum floating amount of the pipe segment to provide accurate construction suggestions to the construction party.

2. According to claim 1, a method for predicting the maximum floating amount of a shield machine segment based on an improved TSMixer model is characterized in that: The step S2 is specifically as follows: S21. Remove invalid or abnormal data by detecting missing values, outliers and illogical records, and use mean interpolation method to fill missing values ​​to avoid the negative impact of incomplete data on model training; S22, normalizing the data of each sensor to eliminate the influence of dimensional differences on model training, so as to improve the convergence and stability of the model; S23, through the adaptive feature selection algorithm, the variables with little influence on the prediction of the floating volume of the pipe segment are screened out, and the key features with significant effect on the prediction accuracy are retained; S24, generating a data set, which contains 584 rows of data, covering parameters collected by multiple sensors, including: pressure of excavation chambers No. 1 to No. 6, rolling angle, cutter head speed, cutter head torque, mud inlet pressure detection of slurry pump, and penetration; The maximum buoyancy of the segments recorded during construction is used as the label of the dataset.

3. According to claim 1, a method for predicting the maximum floating amount of a shield machine segment based on an improved TSMixer model is characterized in that: The self-attention mechanism in step S3 is a mechanism for capturing the correlation between different time steps in the input sequence. When predicting the maximum floating amount of the pipe segment, the self-attention mechanism allocates different attention weights to each time step, focusing on the time points that have a greater impact on the prediction results, thereby improving the ability to extract time series features. Specifically, the following steps are included: S3a1, given input sequence Where T is the number of time steps and d is the feature dimension of each time step, which is mapped into query, key and value vectors: Q = XW Q ,K=XW K ,V=XW V ; in, is the learnable parameter matrix, d k is the embedding dimension; S3a2, the dot product of the query Q and the key K represents the attention score, which is then scaled and normalized: Among them, Softmax is used to normalize the weight of each time step so that the sum of the weights is 1. Used to prevent the value from being too large; S3a3, using the multi-head attention mechanism, the query, key and value are divided into h subspaces and calculated in parallel: MultiHead(Q,K,V)=Concat(head1,…,head h )W O ; Among them, each head i =Attention(Q i ,K i ,V i ), W O is the output projection matrix.

4. According to claim 1, a method for predicting the maximum floating amount of a shield machine segment based on an improved TSMixer model is characterized in that: The GraphNet network structure in step S3 models the spatial dependency between sensors through graph convolution operations. The specific steps are as follows: Step S3b1, define the dependency relationship between different sensors: the input data of the graph convolutional network is regarded as a graph, each node in the graph corresponds to a sensor, and the connection relationship between the nodes is represented by the adjacency matrix To represent; Initially, the adjacency matrix It is a unit matrix, indicating that the relationship between all sensors is initially independent. As the model is trained, the adjacency matrix will be adjusted as a learnable parameter to automatically learn the spatial relationship between sensors. The definition of the adjacency matrix is ​​as follows: in, is a pts ×n pts The identity matrix, n pts Indicates the number of sensor parameters; Step S3b2, extracting characteristic information of sensor nodes; Step S3b3, graph convolution operation, through the adjacency matrix To aggregate the feature information of each sensor node, the feature of each node is weighted averaged with the feature of its adjacent nodes to obtain the updated node feature. The mathematical formula of the graph convolution operation can be expressed as: Among them, H (l) is the node feature matrix of the first layer. For the input data, initially H (0) is the input sensor data; is a normalized adjacency matrix, which is used to define the connection relationship between nodes in the graph; W (l) is the learnable weight matrix of the first layer, which controls the transformation of node features; σ is the activation function, which is used to introduce nonlinearity; Step Sb4: Combine the features of each node with the features of its adjacent nodes, extract the spatial information of the node, and obtain the spatial dependency relationship between sensors.

5. According to claim 4, a method for predicting the maximum floating amount of shield machine segments based on an improved TSMixer model is characterized in that: The GraphNet network contains two graph convolution layers. The first layer of graph convolution expands the input feature dimension from in_features to 2×in_features, and the second layer of graph convolution reduces it to out_features, and finally outputs the prediction result; the details are as follows: The first layer of graph convolution: the input dimension is n pts ×in ― The feature matrix X and adjacency matrix of features Multiply them together to get the weighted node features; expand the feature dimension through the fully connected layer to get a new feature matrix X0 with a dimension of n pts ×2×in ― features; its calculation process is: Among them, W1 is the learnable weight matrix of the first layer of graph convolution, σ is the activation function, and the ReLU function is usually selected; The second layer of graph convolution: dimension is n pts ×2×in ― The feature matrix X0 of features is again Multiply them together for weighted aggregation; compress the feature dimension to out_features through the fully connected layer to obtain the final output X1, which is the predicted value of the maximum floating amount of the pipe segment; the calculation process of this layer is: Among them, W2 is the learnable weight matrix of the second layer of graph convolution.

6. The method for predicting the maximum floating amount of a shield machine segment based on an improved TSMixer model according to claim 1 is characterized in that: In step S5, hyperparameter adjustment is required during the training process. The specific training parameters include: batch size is set to 4, initial learning rate is set to 0.001, optimizer is set to Adam optimizer, number of training rounds is set to 100, and loss function is set to mean square error (MSE). In each round of training, the input data is calculated through forward propagation to generate a prediction result. Then, the error between the prediction result and the true value is calculated, and the model parameters are updated through back propagation.