Soil damage parameter prediction method based on depth domain adaptive transfer learning
Through deep domain adaptation transfer learning, the attention mechanism and hybrid domain adaptation framework are used to align the soil loading condition characteristics, which solves the problem of insufficient generalization ability of the soil damage parameter prediction model under different working conditions, and achieves higher prediction accuracy and lower data acquisition cost.
Patent Information
- Application Number
- CN202510939191.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing data-driven soil damage parameter prediction model has insufficient generalization ability between different loading conditions and limited ability to extract time series features, resulting in low prediction accuracy.
A method based on deep domain adaptation transfer learning is adopted, and a long short-term memory recurrent neural network with an integrated attention mechanism is used to extract deep temporal features. The feature distributions of the source domain and the target domain are aligned through a hybrid domain adaptation framework of explicit distance measurement and implicit adversarial learning to generate domain-invariant feature representations.
It improves the generalization ability of the model under different loading conditions, improves the prediction accuracy of data-scarce target conditions, reduces the dependence on a large amount of labeled data in the target domain, and reduces experimental costs and time.
Smart Images

Figure CN120688369A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intersectional technology of geotechnical engineering and artificial intelligence, and specifically to a soil damage parameter prediction method based on deep domain adaptive transfer learning. Background Art
[0002] In the design and stability analysis of geotechnical engineering projects such as slopes, foundation pits, and tunnels, accurately describing and predicting the mechanical behavior of soil under complex stress paths, particularly the damage evolution of its internal structure, is crucial for ensuring project safety. Soil damage parameters are key internal variables for quantifying this damage state. Therefore, developing efficient and accurate soil damage parameter prediction techniques has significant theoretical value and engineering application prospects.
[0003] Existing technologies for predicting soil damage parameters primarily rely on methods driven by experimental data and numerical simulations. These methods utilize machine learning or deep learning models, such as multi-layer perceptrons (MLPs) and long short-term memory recurrent neural networks (LSTMs), to learn the nonlinear mapping relationship between the soil's stress-strain state and its damage parameters. These models are trained using extensive data acquired under specific loading conditions, enabling them to predict soil damage parameters under those conditions.
[0004] While existing technologies can predict soil failure parameters under specific conditions, they still have some shortcomings. The predictive performance of existing data-driven models is highly dependent on the loading conditions of the training data. When applied to new loading conditions with scarce data, their generalization ability is insufficient, and the prediction accuracy decreases significantly. This is because different soil loading paths, such as isotropic compression and conventional triaxial compression, produce stress and strain data with significant differences in statistical characteristics and distribution. Traditional supervised learning models are typically constructed based on the assumption that training and test data are independent and identically distributed. Therefore, feature mappings trained under one loading condition (source domain) are difficult to directly and effectively apply to another loading condition (target domain) due to domain shift. This limits the model's ability to generalize across conditions and forces researchers to conduct costly and time-consuming data collection for each new loading condition. Furthermore, existing models are limited in their ability to extract temporal features when dealing with the dynamic process of soil loading. Soil failure is a cumulative evolutionary process, and the importance of nodes at different time steps in the sequence is not equal. For example, the characteristic information of key physical nodes, such as the structural yield point, has a greater impact on the final prediction results. However, some existing models do not incorporate mechanisms to automatically focus on key time nodes when processing sequence data. As a result, the extracted feature representations fail to fully reflect the core dynamic information in the sequence, thus limiting further improvements in model prediction accuracy. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a soil damage parameter prediction method based on deep domain adaptive transfer learning, which solves the problem that the existing data-driven soil damage parameter prediction model has insufficient generalization ability and low prediction accuracy when the target working-condition data is scarce due to differences in data distribution between different loading conditions.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a soil damage parameter prediction method based on deep domain adaptive transfer learning, the method comprising the following steps:
[0007] Obtain soil loading process sequence data of the source domain and the target domain corresponding to different soil loading conditions;
[0008] A long short-term memory recurrent neural network with an integrated attention mechanism is used to process the soil loading process sequence data to extract deep temporal features that can automatically focus on key physical nodes in the loading process and generate a context vector;
[0009] A hybrid domain adaptation framework that combines explicit distance metrics and implicit adversarial learning is used to perform adversarial training on the long short-term memory recurrent neural network with an integrated attention mechanism. The adversarial training utilizes the context vector to simultaneously close the feature distributions of the source domain and the target domain at a macro level and align them at a micro level, while minimizing the supervised prediction loss.
[0010] The adversarially trained neural network is applied to the sequence data of the soil loading process to be tested, and the predicted values of the soil damage parameters are output.
[0011] In an optional technical solution, the soil loading process sequence data includes strain state parameters and stress state parameters. The strain state parameters include plastic body strain E v and plastic deviatoric strain ε s , which is calculated as follows:
[0012] E v =E x +E y +E z ;
[0013]
[0014] Where, E x 、E y 、E z With ε x , ε y , ε z They are strain components in different directions and represent strains in different directions respectively; E v is the plastic volume strain; εs is the plastic deviatoric strain.
[0015] The stress state parameters include the deviatoric stress q and the intermediate principal stress coefficient b, which are calculated as follows:
[0016]
[0017] Where q is the deviatoric stress; b is the intermediate principal stress coefficient; σ1 is the maximum principal stress; σ2 is the intermediate principal stress; and σ3 is the minimum principal stress.
[0018] In an optional technical solution, the long short-term memory recurrent neural network processes sequence data through a unit structure including an input gate, a forget gate, and an output gate. Within a time step t, the gates and states are updated as follows:
[0019] Input gate: i t =σ(W i ·[h t-1 ,x t ]+b i );
[0020] Forget gate: f t =σ(W f ·[h t-1 ,x t ]+b f );
[0021] Cell state: c t =f t c t-1 +i t tanh(W c ·[h t-1 ,x t ]+b c );
[0022] Output gate: o t =σ(W o ·[h t-1 ,x t ]+b o );
[0023] Hidden state: h t =o t ⊙tanh(c t );
[0024] Where i t is the input gate output at the current moment; f t is the forget gate output at the current moment; o t is the output gate output at the current moment; c t is the cell state at the current moment; h tis the hidden state at the current moment; x t is the input feature vector at the current moment; h t-1 is the hidden state at the previous moment; c t-1 is the cell state at the previous moment; W i 、W f 、W c 、W o are the weight matrices corresponding to the input gate, forget gate, cell state, and output gate respectively; b i 、b f 、b c 、b o are the bias vectors corresponding to the input gate, forget gate, cell state and output gate respectively; [h t-1 ,x t ] is the concatenation of the hidden state at the previous moment and the input feature vector at the current moment; σ is the sigmoid activation function; tanh is the hyperbolic tangent activation function.
[0025] The attention mechanism calculates an importance weight for the hidden state output by the long short-term memory recurrent neural network at each time step in the sequence, and performs a weighted summation on all hidden states according to the importance weight to generate a context vector with a fixed dimension.
[0026] In an optional technical solution, the hybrid domain adaptation framework achieves alignment of the source domain and target domain feature distributions by combining explicit distance metric with implicit adversarial learning. The explicit distance metric achieves macroscopically closer feature distributions by calculating the maximum mean difference of multiple kernels, and its distribution difference loss L mmd The calculation method is:
[0027]
[0028] Where, L mmd is the multi-core maximum mean difference loss; m is the number of samples in the source domain; s i is the i-th context vector sampled from the source domain; n is the number of samples in the target domain; t j is the j-th context vector sampled from the target domain; φ(·) is the kernel function mapped to the reproducing kernel Hilbert space; is the reproducing kernel Hilbert space; For the reproducing kernel Hilbert space The square of the norm in .
[0029] The implicit adversarial learning achieves microscopic alignment of feature distribution by introducing an adversarial domain classifier, which is used to distinguish whether the context vector comes from the source domain or the target domain.
[0030] In an optional technical solution, the adversarial training is implemented through a joint optimization objective, whose total loss L total The mathematical expression is:
[0031] L total =L pred +λL mmd +γL adv ;
[0032] Where, L pred is the supervised prediction loss calculated on the source domain; L total is the total loss of the joint optimization objective; L adv is the adversarial domain classification loss; L mmd is the multi-core maximum mean difference loss; λ is the hyperparameter used to balance the multi-core maximum mean difference loss; γ is the hyperparameter used to balance the adversarial domain classification loss.
[0033] During backpropagation to update the parameters of the neural network, the gradient generated by the adversarial domain classification loss is reversed through a gradient reversal layer, so that the network parameters of the feature extraction part are updated to minimize the supervised prediction loss and the multi-core maximum mean difference loss, while maximizing the adversarial domain classification loss, thereby generating domain-invariant feature representations.
[0034] In an optional technical solution, the step of outputting the predicted value of the soil damage parameter is specifically to input the context vector with fixed dimension generated by the attention mechanism into a multi-layer perceptron composed of multiple layers of fully connected layers. The multi-layer perceptron adopts a regularization strategy to prevent overfitting and finally outputs the predicted value of the soil damage parameter.
[0035] The present invention provides a soil damage parameter prediction method based on deep domain adaptive transfer learning.
[0036] It has the following beneficial effects:
[0037] 1. The present invention uses the multi-core maximum mean difference to perform global alignment at the macro level and the adversarial domain classifier to perform local alignment at the micro level, so that the model can learn the domain-invariant features between the source domain and the target domain, thereby improving the generalization ability of the model under different loading conditions and the prediction accuracy of target conditions with scarce data.
[0038] 2. The present invention reduces the dependence on a large amount of labeled data in the target domain and reduces the cost and time of experimental data acquisition by transferring knowledge from a data-rich source domain to a data-scarce target domain.
[0039] 3. The present invention adopts a long short-term memory recursive neural network with an integrated attention mechanism, which can effectively process the time series data during the soil loading process and automatically assign higher weights to the key node information in the sequence, thereby improving the effectiveness of feature extraction and reducing the final prediction error. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the method flow of the present invention;
[0041] Figure 2 Schematic diagram of the data acquisition and preprocessing process of the present invention;
[0042] Figure 3 Schematic diagram of correlation analysis between input features and output targets of the present invention;
[0043] Figure 4 Schematic diagram of the depth domain adaptation model structure of the present invention;
[0044] Figure 5 Schematic diagram of the training process of the hybrid domain adaptation framework of the present invention. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the specification of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0046] Please see the attached Figure 1 , Figure 1 The present invention provides a method for predicting soil damage parameters based on deep domain adaptive transfer learning according to an embodiment of the present invention. The method can be executed by a computer device and may include the following steps:
[0047] S101. Acquire soil loading process sequence data of a source domain and a target domain corresponding to different soil loading conditions.
[0048] S102. A long short-term memory recurrent neural network with an integrated attention mechanism is used to process the soil loading process sequence data to extract deep temporal features and generate context vectors.
[0049] S103. A hybrid domain adaptation framework that combines explicit distance metrics and implicit adversarial learning is used to perform adversarial training on long short-term memory recurrent neural networks with integrated attention mechanisms.
[0050] S104: Apply the adversarial training-completed neural network to the sequence data of the soil loading process to be tested, and output the predicted value of the soil damage parameter.
[0051] The technical solution provided in this embodiment can be implemented through a prediction system, which may include a data acquisition module 10 , a feature extraction module 20 , a domain adaptation training module 30 , and a parameter prediction module 40 .
[0052] The data acquisition module 10 is used to execute step S101. This module acquires soil loading process sequence data from the source domain and the target domain. The source domain data is a labeled dataset, and the target domain data is the dataset to be predicted. Its labels may or may not be used during the training phase.
[0053] The feature extraction module 20 is used to execute step S102. The main body of this module is a long short-term memory recurrent neural network, which is used to process the input sequence data. The network processes information in each time step t through a unit structure including an input gate, a forget gate, and an output gate. Its update method is:
[0054] Input gate: i t =σ(W i ·[h t-1 ,x t ]+b i );
[0055] Forget gate: f t =σ(W f ·[h t-1 ,x t ]+b f );
[0056] Cell state: c t =f t c t-1 +i t tanh(W c ·[h t-1 ,x t ]+b c );
[0057] Output gate: o t =σ(W o ·[h t-1 ,x t ]+b o );
[0058] Hidden state: h t =o t ⊙tanh(c t );
[0059] Where i tis the input gate output at the current moment; f t is the forget gate output at the current moment; o t is the output gate output at the current moment; c t is the cell state at the current moment; h t is the hidden state at the current moment; x t is the input feature vector at the current moment; h t-1 is the hidden state at the previous moment; c t-1 is the cell state at the previous moment; W i 、W f 、W c 、W o are the weight matrices corresponding to the input gate, forget gate, cell state, and output gate respectively; b i 、b f 、b c 、b o are the bias vectors corresponding to the input gate, forget gate, cell state and output gate respectively; [h t-1 ,x t ] is the concatenation of the hidden state at the previous moment and the input feature vector at the current moment; σ is the sigmoid activation function; tanh is the hyperbolic tangent activation function; ⊙ is the Hadamard product.
[0060] The feature extraction module 20 further integrates an attention mechanism. After the LSTM RNN processes the entire input sequence, it will obtain a hidden state {h1,h2,...,h T} is a sequence consisting of .
[0061] The attention mechanism calculates an importance weight for each hidden state in the sequence and takes a weighted sum of all hidden states to generate a fixed-dimensional context vector. This context vector aggregates information from the entire sequence and enhances information at key time steps.
[0062] The domain adaptation training module 30 is used to execute step S103. This module trains the feature extraction module 20 and the parameter prediction module 40 through a hybrid domain adaptation framework. The framework includes an explicit distance measurement unit and an implicit adversarial learning unit.
[0063] The explicit distance metric unit calculates the multi-kernel maximum mean difference (MK-MMD) to achieve a macroscopic closeness between the source domain and the target domain feature distribution. Its distribution difference loss L mmd The calculation method is:
[0064]
[0065] Where, L mmdis the multi-core maximum mean difference loss; m is the number of samples in the source domain; s i is the i-th context vector sampled from the source domain; n is the number of samples in the target domain; t j is the j-th context vector sampled from the target domain; φ(·) is the kernel function mapped to the reproducing kernel Hilbert space; is the reproducing kernel Hilbert space; For the reproducing kernel Hilbert space The square of the norm in .
[0066] The implicit adversarial learning unit achieves microscopic alignment of feature distribution by introducing an adversarial domain classifier and a gradient reversal layer (GRL). The domain classifier is used to distinguish whether the context vector generated by the feature extraction module 20 comes from the source domain or the target domain.
[0067] The domain adaptation training module 30 trains the entire model through a joint optimization objective consisting of supervised prediction loss, multi-core maximum mean difference loss and adversarial domain classification loss, with a total loss L total The mathematical expression is:
[0068] L total =L pred +λL mmd +γL adv ;
[0069] Where, L pred is the supervised prediction loss calculated on the source domain; L total is the total loss of the joint optimization objective; L adv is the adversarial domain classification loss; L mmd is the multi-core maximum mean difference loss; λ is the hyperparameter used to balance the multi-core maximum mean difference loss; γ is the hyperparameter used to balance the adversarial domain classification loss.
[0070] During backpropagation, the gradient reversal layer reverses the gradient generated by the adversarial domain classification loss and passes it to the feature extraction module 20. This operation causes the parameter update direction of the feature extraction module 20 to minimize the classification accuracy of the domain classifier, thereby prompting it to generate domain-invariant features.
[0071] Parameter prediction module 40, during the training phase, receives the context vector from feature extraction module 20 and outputs the predicted value to calculate L pred During the prediction phase of step S104, this module receives the context vector generated by the feature extraction module 20 after processing the sequence data to be tested, and outputs the final predicted value of the soil damage parameter. This module can be implemented as a multi-layer perceptron consisting of multiple fully connected layers.
[0072] Please see the attached Figure 2 , Figure 2 Figure 2 is a schematic diagram of the data acquisition and preprocessing process according to one embodiment of the present invention. In a preferred embodiment, the soil loading process sequence data is acquired through numerical simulation. Specifically, 3D discrete element simulation software, such as PFC7.0, is used to conduct multiple simulation experiments under different loading conditions.
[0073] The loading conditions included isotropic compression, constant stress ratio compression, conventional triaxial compression, and true triaxial compression. During each simulation, stress and strain state parameters, as well as the corresponding soil microscopic damage parameters, were continuously recorded over multiple time steps to form a time series dataset.
[0074] The time series dataset acquired under each loading condition is defined as an independent domain. When performing transfer learning, one or more domains are selected as source domains based on the prediction objective, and a domain to be predicted is selected as the target domain. The source domain contains samples labeled with known soil damage parameters, while the target domain is the set of samples to be predicted.
[0075] Before inputting into the model, the acquired sequence data is preprocessed. First, the dataset is checked for missing values or outliers, and these are addressed. One possible approach is to fill in missing values through interpolation or remove outliers. Next, the dataset within each domain is randomly divided into training and test sets, with a reasonable split ratio of 6:4.
[0076] To eliminate the impact of different physical dimensions on model training, the dataset is normalized. This process maps the numerical range of all input features to a unified interval, such as [0, 1] or [-1, 1]. This step helps accelerate model convergence and improve training stability.
[0077] Please see the attached Figure 3 , Figure 3 Figure 2 is a schematic diagram illustrating the correlation analysis between input features and output targets according to one embodiment of the present invention. In a preferred embodiment of the present invention, multidimensional input features are selected from acquired sequence data to characterize the physical state of soil during loading. The input features include strain state parameters and stress state parameters.
[0078] Strain state parameters include plastic volume strain and plastic deviatoric strain. Stress state parameters include mean effective stress, structural yield stress, deviatoric stress, and intermediate principal stress coefficient. The intermediate principal stress coefficient characterizes the relative position of the intermediate principal stress between the maximum and minimum principal stresses. The structural yield stress is the stress value at which the soil structure yields and is associated with a sharp change in the degree of damage within the soil.
[0079] In order to verify the validity of the selected input features, the grey relational analysis method was used to quantitatively evaluate the correlation between each input feature and the soil damage parameters as the output target.
[0080] like Figure 3 As shown, the results of this correlation analysis can be displayed as a heat map, where the value of each cell in the map represents the gray correlation between an input feature and the output target. The analysis results show that the gray correlation between all the input features selected in this example and the soil damage parameters is higher than 0.8, indicating that the selected features are highly correlated with the prediction target and can serve as effective input for subsequent prediction models.
[0081] Please see the attached Figure 4 , Figure 4 2 is a schematic diagram of the structure of a depth domain adaptation model according to an embodiment of the present invention. In a preferred embodiment of the present invention, the constructed depth domain adaptation model consists of a feature extraction module 20 and a parameter prediction module 40.
[0082] The feature extraction module 20 is primarily composed of a long short-term memory (LSTM) recurrent neural network (RNN) with an integrated attention mechanism. Given that the soil loading process data is typically sequential, with correlations between data points in previous and subsequent time steps, LSTM RNNs are effective for processing this type of data. Through its internal gated unit structure, LSTM RNNs capture and transmit long-term dependencies within the sequence.
[0083] An attention mechanism is integrated on top of the LSTM RNN. This attention mechanism receives the hidden state sequence output by the LSTM RNN at all time steps, calculates an importance weight for each hidden state in the sequence, and then performs a weighted sum of all hidden states based on this weight. This process ultimately generates a fixed-dimensional context vector, which serves as a deep abstract representation of the entire input sequence and is input to the parameter prediction module 40.
[0084] The parameter prediction module 40 is used to receive the context vector output from the feature extraction module 20 and output the final soil damage parameter prediction value. This module is implemented by a multi-layer perceptron (MLP).
[0085] In this embodiment, the multilayer perceptron consists of four fully connected layers. Rectified linear units (ReLUs) are used as activation functions between each fully connected layer. To prevent overfitting during model training, the Dropout regularization strategy is introduced in the training of the multilayer perceptron. This strategy temporarily sets the outputs of some neurons to zero with a certain probability during each training step, thereby reducing the coadaptability between neurons and enhancing the generalization ability of the model.
[0086] Please see the attached Figure 5 , Figure 5 Figure 2 is a schematic diagram of the training process for a hybrid domain adaptation framework according to one embodiment of the present invention. In a preferred embodiment of the present invention, a hybrid domain adaptation framework that combines explicit distance metrics with implicit adversarial learning is used to train the model. The training goal is to minimize the supervised prediction loss in the source domain while reducing the distribution difference between the source and target domains in the feature space.
[0087] The explicit distance metric is implemented through the Multi-kernel Maximum Mean Difference (MK-MMD). To avoid the dependence of a single kernel function on parameter selection, this embodiment uses multiple Gaussian kernel functions with different bandwidth parameters to more comprehensively quantify the inter-domain differences. Its optimization goal is:
[0088]
[0089] Where, is the square value of the maximum mean difference calculated using the kth kernel function; β k is the weight coefficient corresponding to the kth kernel function, the sum of which is 1; k is the total number of kernel functions used.
[0090] Implicit adversarial learning is implemented through an adversarial domain classifier and a gradient reversal layer (GRL). The domain classifier is a binary classifier that takes as input the context vector output by the feature extraction module and aims to determine whether the vector originates from the source domain or the target domain. The gradient reversal layer is connected between the feature extraction module and the adversarial domain classifier.
[0091] The entire model training process is driven by a joint optimization objective, which is composed of the supervised prediction loss L pred , multi-core maximum mean difference loss L mmd and adversarial domain classification loss L adv In the back propagation phase of training, when the gradient is transferred from the adversarial domain classification loss L adv During the return, the gradient reversal layer multiplies its gradient value by a negative constant before passing it to the feature extraction module.
[0092] This operation makes the parameter update of the feature extraction module aim to maximize the classification loss of the adversarial domain classifier, that is, to generate feature representations that make the domain classifier unable to distinguish the source domain. At the same time, the parameter update of the feature extraction module also aims to minimize the supervised prediction loss L pred and multi-core maximum mean difference loss L mmd Through this adversarial training method, the feature representation learned by the model is both task predictive and domain invariant.
[0093] In the specific training configuration, the Adam optimizer was used to perform gradient descent. Hyperparameters such as the learning rate and batch size were determined through cross-validation. Furthermore, an early stopping strategy was employed during training: the model's performance on the validation set was monitored, and training was terminated if performance did not improve over multiple epochs to prevent overfitting.
[0094] In a preferred embodiment of the present invention, after the model training is completed, the trained deep domain adaptation model is applied to the test set data of the target domain to predict soil damage parameters.
[0095] In order to quantitatively evaluate the prediction performance of the method of the present invention, the root mean square error (RMSE), relative absolute error (RAE) and coefficient of determination (R 2 ) as the evaluation index.
[0096] The root mean square error (RMSE) is used to evaluate the average error between the model's predicted value and the true value, and is calculated as follows:
[0097]
[0098] Where n is the number of samples in the test set; y i is the true soil damage parameter value of the i-th sample; is the predicted value output by the model for the i-th sample;
[0099] The relative absolute error (RAE) is used to compare the absolute error of the model to the absolute error produced by predicting using only the mean and is calculated as:
[0100]
[0101] Where n is the number of samples in the test set; y i is the true soil damage parameter value of the i-th sample; is the predicted value output by the model for the i-th sample; is the average value of all true samples;
[0102] Coefficient of determination (R 2 ) is used to measure the degree of fit between the model's predicted value and the true value, and is calculated as follows:
[0103]
[0104] Where n is the number of samples in the test set; y i is the true soil damage parameter value of the i-th sample; is the predicted value output by the model for the i-th sample; is the average of all true sample values.
[0105] In the above formula, the mean of the actual damage value, the method predicted value and the actual value are respectively represented. is the residual variance, which represents the variance of the dependent variable. This paper uses RMSE, RAE and R 2 Evaluating the prediction performance of a model from multiple dimensions can help us understand the advantages and disadvantages of the model more comprehensively and avoid misjudgment caused by a single indicator. RMSE provides the average size of the method error, and RAE gives the ratio of the method error to the average prediction method error. The smaller the value of the two, the smaller the prediction error of the method. 2 It shows the model's ability to explain the data. The larger the value, the better the method can fit the data.
[0106] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A soil damage parameter prediction method based on deep domain adaptive transfer learning, characterized in that: The following steps are involved: S1. Obtain soil loading process sequence data of the source domain and the target domain corresponding to different soil loading conditions; S2. Processing the soil loading process sequence data using a long short-term memory recurrent neural network integrated with an attention mechanism to extract deep temporal features that can automatically focus on key physical nodes in the loading process and generate a context vector; S3. Adopting a hybrid domain adaptation framework that combines explicit distance metrics and implicit adversarial learning to conduct adversarial training on the long short-term memory recurrent neural network with integrated attention mechanism. The adversarial training utilizes the context vector to simultaneously close the feature distributions of the source domain and the target domain at a macro level and align them at a micro level, while minimizing the supervised prediction loss. S4. Apply the adversarial trained neural network to the sequence data of the soil loading process to be tested, and output the predicted value of the soil damage parameter.
2. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 1 is characterized in that: The step of obtaining soil loading process sequence data of the source domain and the target domain corresponding to different soil loading conditions includes: performing numerical simulation by three-dimensional discrete element simulation software to obtain soil loading process sequence data of the source domain and the target domain corresponding to different soil loading conditions.
3. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 2 is characterized in that: The extraction step can automatically focus on the deep temporal features of the key physical nodes in the loading process, specifically by the long short-term memory recurrent neural network for the cell state c at each time step. t The update is implemented using the following expression: c t =f t ·c t-1 +i t ·tanh(W c ·[h t-1 ,x t ]+b c ); Where c t-1 is the cell state at the previous moment; x t h is the characteristic vector composed of the strain state parameter and the stress state parameter input at the current moment; t-1 is the hidden state at the previous moment; f t and i t are the outputs of the forget gate and input gate respectively; c t is the cell state at the current moment; tanh is the hyperbolic tangent activation function; W c is the weight matrix used to update the cell state; [h t-1 ,x t ] is the concatenation of the hidden state at the previous moment and the input feature vector at the current moment; b c is the bias vector used to update the cell state.
4. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 3 is characterized in that: The step of generating a context vector comprises: Calculating the importance weight of the hidden state output by the long short-term memory recurrent neural network at each time step in the sequence; Perform a weighted sum of all hidden states according to the importance weights.
5. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 4 is characterized in that: The explicit distance metric in the hybrid domain adaptation framework is to achieve macroscopically closer feature distribution by calculating the maximum mean difference of multiple kernels.
6. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 5 is characterized in that: The hybrid domain adaptation framework includes using multi-kernel maximum mean difference to measure the distribution difference between the source domain and the target domain, and its distribution difference loss L mmd The calculation method is as follows: Where, L mmd is the multi-core maximum mean difference loss; m is the number of samples in the source domain; s i is the i-th context vector sampled from the source domain; n is the number of samples in the target domain; t j is the j-th context vector sampled from the target domain; φ(·) is the kernel function mapped to the reproducing kernel Hilbert space; is the reproducing kernel Hilbert space; For the reproducing kernel Hilbert space The square of the norm in .
7. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 6 is characterized in that: The steps of performing adversarial training on the long short-term memory recurrent neural network integrated with the attention mechanism include: By introducing an adversarial domain classifier, the feature distribution is aligned at the micro level; The adversarial domain classifier is used to distinguish whether the context vector comes from a source domain or a target domain, and perform adversarial training with the neural network.
8. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 7 is characterized in that: The multi-core maximum mean difference works in synergy with the adversarial domain classifier, wherein the multi-core maximum mean difference is used to achieve global alignment of feature distributions of the source domain and the target domain; based on the global alignment, the adversarial domain classifier further eliminates subtle feature differences between the two domains that are difficult to measure through adversarial game.
9. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 8 is characterized in that: The step of performing adversarial training on the long short-term memory recurrent neural network integrated with the attention mechanism further includes: This is achieved through a joint optimization objective, which is composed of supervised prediction loss, multi-core maximum mean difference loss and adversarial domain classification loss, and its total loss L total The mathematical expression is: THE total =L pred +λL mmd +γL adv ; Where, L pred is the supervised prediction loss calculated on the source domain; L total is the total loss of the joint optimization objective; L adv is the adversarial domain classification loss; L mmd is the multi-core maximum mean difference loss; λ is the hyperparameter used to balance the multi-core maximum mean difference loss; γ is the hyperparameter used to balance the adversarial domain classification loss; In which, when backpropagating to update the parameters of the neural network, the gradient generated by the adversarial domain classification loss is reversed through a gradient reversal layer, so that the neural network maximizes the adversarial domain classification loss while minimizing the supervised prediction loss and the multi-core maximum mean difference loss.
10. The soil damage parameter prediction method based on deep domain adaptive transfer learning according to claim 9 is characterized in that: The step of outputting the predicted value of the soil damage parameter comprises: The context vector with fixed dimension generated by the attention mechanism is input into a multi-layer perceptron composed of multiple layers of fully connected layers. The multi-layer perceptron adopts a regularization strategy to prevent overfitting and finally outputs the predicted value of the soil damage parameter.