Bridge post-earthquake rapid evaluation method based on active learning XGBoost

The integration of active learning and XGBoost model with entropy-based sample selection and hyperparameter optimization addresses inefficiencies in traditional bridge safety assessment methods, providing a rapid and accurate evaluation of bridge safety post-earthquake.

CN120316633APending Publication Date: 2025-07-15LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510445439.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Traditional bridge post-seismic safety assessment methods are inefficient and subjective, making it difficult to quickly and accurately evaluate the safety status of bridges, and lack of data analysis, low model training efficiency, and limited evaluation accuracy.

Method used

Combining active learning strategies and XGBoost classification model, by calculating the uncertainty of the sample (based on entropy metric), the most representative samples for model training are automatically screened out, and hyperparameters are optimized through grid search to build an efficient bridge post-seismic evaluation method.

Benefits of technology

It significantly improves the efficiency and accuracy of post-seismic assessment of bridges, reduces data labeling costs, saves computing resources, provides scientific decision-making support, and provides reliable support for post-seismic bridge emergency repair and traffic recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316633A_ABST
    Figure CN120316633A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge post-earthquake rapid evaluation method based on active learning XGBoost. The method comprises the following steps: firstly, carrying out bridge modeling by using OpenSees, and establishing a machine learning data set; the method comprises the following steps: selecting a data set, dividing the data set into an initial training set, an unlabeled pool and a test set, selecting a most uncertain sample from the unlabeled pool through an active learning method, performing XGBoost model training in combination with the selected sample, optimizing hyper-parameters through grid search, and selecting an optimal XGBoost model as a final model; compared with a traditional method, the bridge post-earthquake safety assessment method has the advantages that the assessment efficiency is improved, high accuracy is ensured, and the bridge post-earthquake safety assessment method is wide in application prospect and practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for rapid post - earthquake assessment of bridges based on active learning XGBoost, belonging to the technical field of measuring bridge damage degree and post - earthquake safety assessment. Background Art

[0002] With the acceleration of the urbanization process, as key transportation infrastructure, the post - earthquake safety assessment of bridges has become particularly important. Traditional safety assessment methods mostly rely on manual inspection and simplified calculation models, which have problems such as low efficiency and strong subjectivity. Especially in the post - earthquake emergency situation, it is difficult to quickly and accurately assess the safety status of bridges. Therefore, traditional assessment methods can no longer meet the requirements of rapid and accurate post - earthquake assessment of modern bridges.

[0003] To address these challenges, in recent years, data - driven bridge safety assessment methods have gradually received attention, especially those combined with machine learning techniques. These methods can use the post - earthquake response data of bridges for efficient assessment, significantly improving the efficiency and accuracy of assessment. However, such methods still face problems such as insufficient data analysis, low model training efficiency, limited assessment accuracy, and failure to effectively screen out the most valuable samples for model training in traditional training, resulting in poor training effects.

[0004] To solve the above problems, the present invention proposes an innovative method for rapid post - earthquake assessment of bridges based on active learning XGBoost. The innovation of this method lies in combining the active learning strategy with the XGBoost classification model. By calculating the uncertainty of samples (based on entropy measurement), it automatically screens out the most representative samples for model training, gradually adding the samples with higher uncertainty in the unlabeled pool to the training set, significantly improving the quality of the training set, and thus effectively enhancing the model learning efficiency and assessment accuracy.

[0005] The introduction of active learning not only reduces the dependence on labeled samples, lowers the cost of data annotation, but also can accelerate the model training process and improve the assessment accuracy. Traditional methods usually require a large number of labeled samples to effectively train the model, while the present invention significantly improves the training effect and assessment accuracy with fewer labeled samples by intelligently screening the samples that contribute the most to the model. To further optimize the model performance, the present invention also combines grid search to finely tune the hyperparameters of XGBoost to ensure the efficiency and accuracy of post - earthquake bridge safety assessment. Summary of the Invention

[0006] The present invention proposes an XGBoost-based post-earthquake rapid assessment method for bridges using active learning, aiming to significantly improve the efficiency and accuracy of post-earthquake assessment. This method determines the post-earthquake state of the bridge through finite element modeling and constructs a comprehensive dataset containing bridge parameters and post-earthquake response data. Subsequently, random sampling is used to preliminarily divide the dataset to ensure that the sample distributions of the training set, unlabeled pool, and test set are representative, providing a high-quality data basis for model training.

[0007] In the model training stage, the innovation of the present invention lies in combining the active learning strategy with an improved version of the XGBoost algorithm. By calculating the uncertainty of samples (based on entropy measurement), samples that contribute the most to model training are preferentially selected. Through this active learning process, the present invention can intelligently select the samples with the most information from the unlabeled pool and gradually add them to the training set. This screening strategy effectively improves the quality of the training set, enabling the model to significantly improve the evaluation accuracy even when the number of labeled samples is small. Compared with traditional methods, active learning significantly reduces the dependence on labeled samples, reduces the data annotation cost, and avoids the interference of a large number of useless samples, thus greatly improving the model training efficiency.

[0008] In addition, grid search technology is used in the method to optimize the hyperparameters of XGBoost, further enhancing the classification ability and performance of the model, making the post-earthquake assessment results more accurate. Compared with traditional assessment methods, the present invention reduces the need for labeled samples through active learning, not only reducing the annotation cost but also saving computing resources, greatly improving the efficiency of post-earthquake assessment. Finally, using the trained XGBoost model, combined with the structural characteristics of the bridge and seismic motion data, the safety classification label of the post-earthquake bridge is quickly and accurately output. This method not only demonstrates excellent accuracy and efficiency in post-earthquake bridge safety assessment but also provides scientific and reliable decision-making support for post-earthquake bridge repair and traffic restoration, with broad application prospects and significant practical value.

[0009] The above object is achieved by the following technical solutions:

[0010] An XGBoost-based post-earthquake rapid assessment method for bridges using active learning, the method comprising the following steps:

[0011] S1. Dataset construction: Use OpenSees to model the bridge, extract relevant bridge parameters as input features for machine learning, and divide the bridge safety state categories according to the post-earthquake performance of the bridge as output labels, thereby constructing a machine learning dataset;

[0012] S2. Sample division: Divide the dataset constructed in step S1 into an initial training set, an unlabeled pool, and a test set at a ratio of 1:6:3 to ensure the consistency of class distribution in different datasets;

[0013] S3. XGBoost model training: Through the active learning method, select the most uncertain samples from the unlabeled pool, combine the selected samples for XGBoost model training, and optimize the hyperparameters through grid search;

[0014] S4. Optimal model determination: By comparing the performance of the model under different active learning iteration times and hyperparameter combinations, select the optimal XGBoost model as the final model;

[0015] S5. Post-earthquake safety assessment: Using the optimal XGBoost model selected in step S4, with bridge parameters and ground motion data as inputs, output the corresponding safety status classification labels as the basis for post-earthquake safety assessment of bridges.

[0016] Furthermore, in step S1, the relevant bridge parameters are extracted as input features for machine learning, and the bridge safety status categories are divided according to the post-earthquake performance of the bridge as output labels, thereby constructing a machine learning dataset. The specific steps are as follows:

[0017] S111, Input feature extraction: Extract the structural information of the bridge and the peak acceleration during the earthquake as input features of the dataset;

[0018] S112, Output label division: Adopt the ATC-13 standard to divide the post-earthquake safety status of the bridge into three levels: green, yellow, and red, representing different safety levels respectively. The specific classification method is as follows:

[0019] Use the second stiffness method to determine the yield state of the bridge, that is, when the bridge stiffness drops to 50% of the initial stiffness, it is defined as the yield state; when the stress of the bridge cover concrete drops to 0, it is defined as the spalling state;

[0020] According to the post-earthquake bridge state division rules:

[0021] If the bridge state is lower than the yield state, mark it as green, that is, safe and can be passed;

[0022] If the bridge state is between the yield state and the spalling state, mark it as yellow, that is, it can be passed under emergency conditions, but further inspection is required before putting it into daily traffic;

[0023] If the bridge state exceeds the spalling state, mark it as red, that is, dangerous and prohibited from passing;

[0024] S113, Dataset Construction: According to the above criteria, mark the post-earthquake status of each bridge sample as green, yellow, or red, and use these labels as the output labels of the dataset.

[0025] Furthermore, the dataset constructed in step S1 is divided into an initial training set, an unlabeled pool, and a test set in a ratio of 1:6:3. The specific information is as follows:

[0026] In the dataset constructed in step S113, using random sampling, divide the dataset into an initial training set, an unlabeled pool, and a test set in a ratio of 1:6:3, and ensure that the proportion of samples in each safety status level in each subset is roughly the same as that in the original dataset, so as to avoid the adverse impact of class imbalance on model training and evaluation results.

[0027] Furthermore, in step S3, through the active learning method, select the most uncertain samples from the unlabeled pool, combine the selected samples for XGBoost model training, and optimize the hyperparameters through grid search. The specific steps are as follows:

[0028] S311, Train the XGBoost model with the initial training set: Use the initial training set to train the XGBoost model. The XGBoost model fits by learning the features and labels of the initial training set to discover the patterns and relationships in the data;

[0029] S312, Calculate the sample uncertainty based on entropy value: Calculate the entropy of each unlabeled pool sample to evaluate the sample uncertainty. For each sample x in the unlabeled pool i , calculate its predicted probability p of belonging to each safety status level i,j , and then use the entropy formula to measure the uncertainty of this sample:

[0030]

[0031] where H(x i ) represents the entropy value of sample x i , M is the number of categories, taking 3, and p i,j is the predicted probability that sample x i belongs to category j;

[0032] S313, Select the most uncertain samples and update the training set: Add the top q most uncertain samples calculated in step 312 to the initial training set, and remove this sample from the unlabeled pool. The updated training set is:

[0033]

[0034] The updated unlabeled pool is:

[0035]

[0036] Among them, X train and y train are the features and labels of the initial training set respectively, and X pool and y pool are the sample features and labels in the unlabeled pool respectively; and are the features and labels of the i-th most uncertain sample respectively;

[0037] S314. Retrain the model and determine the loss function: Based on the updated training set and retrain the XGBoost model, and optimize the hyperparameters through grid search. The goal of grid search is to find the optimal hyperparameter combination Θ * , to minimize the logarithmic loss function L:

[0038]

[0039] where N is the total number of samples in the updated training set, M is the number of safety status levels, that is, M takes 3, and y i,m is the true class label of sample i, and p i,m is the predicted probability that sample i belongs to a certain safety status level m;

[0040] S315. Optimize hyperparameters by grid search: After determining the loss function, grid search selects the optimal hyperparameter Θ * , and determines it by minimizing the loss function on the validation set:

[0041]

[0042] where Θ represents the set of hyperparameters of XGBoost (learning rate, maximum depth, regularization parameter, etc.), Γ is the set of all possible hyperparameter combinations, )X val , y val ) is the validation set, and L(X val , y val , Θ) represents the loss of XGBoost on the validation set given the hyperparameter Θ. Through grid search, find the optimal parameter Θ * After that, use the optimal hyperparameter Θ * to train the XGBoost model;

[0043] S316. Repeat the active learning iteration: Usually when the number of samples reaches 100, active learning can achieve better results. Set the number of active learning iterations to n, where Repeat steps S312 to S315 (n - 1) times, and output the accuracy on the test set;

[0044] S317. Reset the number of iterations to k times. Repeat steps S311 to S316.

[0045] S318. Reset the number of iterations to l times. In the present invention, the total number of samples in the unlabeled pool is N1 ≥ 100; repeat steps S311 to S316.

[0046] Further, in step S4, by comparing the performance of the XGBoost model under different active learning iteration times and hyperparameter combinations, and selecting the optimal XGBoost model as the final model, the specific steps are as follows:

[0047] S411. Compare the active learning sample numbers of the three iteration times n, k, and l set in steps 316 - 318. By analyzing the number of samples selected under each iteration time, evaluate the influence of different iteration times on sample selection.

[0048] S412. Compare the accuracy rates of the model on the test set under the three iteration times n, k, and l. The accuracy rate calculation formula is:

[0049]

[0050] Where:

[0051] TP (True Positive): The number of samples with the true class being positive and the model prediction also being positive.

[0052] TN (True Negative): The number of samples with the true class being negative and the model prediction also being negative.

[0053] FP (False Positive): The number of samples with the true class being negative and the model wrongly predicting as positive.

[0054] FN (False Negative): The number of samples with the true class being positive and the model wrongly predicting as negative.

[0055] By comparing the accuracy rates under different iteration times, determine the iteration time that can most improve the model performance.

[0056] S413. Compare the confusion matrices of the XGBoost model under the three iteration times n, k, and l. Use the confusion matrix to conduct a detailed analysis of the classification effect of the XGBoost model, and calculate the precision rate and recall rate.

[0057]

[0058] Through these metrics, further evaluate the classification ability of the model at different iteration times;

[0059] S414. Compare the F1-score of XGBoost under three iteration times, and analyze the ROC curve and AUC value of the model. The calculation formula of F1-score is:

[0060]

[0061] Use the weighted F1-score:

[0062]

[0063] where ω m is the proportion of samples of class m, and F1-score m is the F1-score of class m;

[0064] S415. Based on the comprehensive performance evaluation results in steps S411 to S414, select the optimal iteration times and the corresponding XGBoost model, and export the final XGBoost model and its corresponding optimal hyperparameter configuration.

[0065] Furthermore, the confusion matrix described in step S413 has the following specific information:

[0066] For a three-class classification problem, the confusion matrix is usually a 3x3 matrix. Adding the accuracy of rows and columns forms a 4x4 matrix. In the 4x4 matrix: the rows represent the actual classes, and the columns represent the predicted classes; the elements on the diagonal represent the number of correctly classified samples of a certain class; the non-diagonal elements represent misclassifications.

[0067] Beneficial effects

[0068] The present invention proposes an XGBoost-based post-earthquake rapid assessment method for bridges based on active learning, which can efficiently and accurately evaluate the post-earthquake safety status of bridges. By introducing an active learning strategy, this method dynamically selects the samples with the highest uncertainty in the unlabeled pool and adds them to the training set, ensuring that each expansion of the training set can maximize the improvement of the model accuracy and avoid the addition of redundant samples, thereby optimizing the quality of training data and improving the learning efficiency of the model.

[0069] Compared with the traditional full-sample training method, the present invention only requires fewer labeled samples to achieve or even exceed the accuracy of full-sample training. At the same time, in the case of class imbalance, it can effectively enhance the learning ability of the model for minority-class samples and significantly improve the classification accuracy. In addition, active learning reduces the scale of the training set and the computational overhead by efficiently screening the most representative training samples, while maintaining high model performance, making it suitable for large-scale post-earthquake bridge safety assessment tasks.

[0070] In order to further improve the classification effect of the model, the present invention uses grid search to optimize the hyperparameters of the XGBoost model to ensure the best performance. Finally, based on the input of bridge parameters and ground motion data, the trained XGBoost model can quickly output the safety classification labels of bridges after an earthquake, ensuring the accuracy and feasibility of the evaluation results. Compared with traditional evaluation methods, the present invention significantly shortens the evaluation time and provides an efficient and intelligent post-earthquake safety evaluation solution for bridge clusters, providing scientific support for post-earthquake bridge repair and traffic restoration. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is a flowchart of a rapid post-earthquake evaluation method for bridges based on active learning XGBoost;

[0072] Figure 2 is a numerical simulation diagram for bridge modeling. Figure 2 In (a) is a detailed diagram of finite element beam and column modeling, (b) is a diagram of column cross-section and material model, (c) is a foundation model diagram, and (d) is an abutment model diagram;

[0073] Figure 3 is a post-earthquake label diagram of the bridge;

[0074] Figure 4 is a demonstration diagram of the active learning method. Figure 4 In (a) is an actual decision boundary diagram, (b) is a randomly selected sample diagram, and (c) is a classification diagram after target sampling;

[0075] Figure 5 is a comparison diagram of three iterative confusion matrix diagrams of the bridge after an earthquake. Figure 5 In (a) is the relative training set confusion matrix for 2 iterations, (b) is the test set confusion matrix for 2 iterations, (c) is the relative training set confusion matrix for 5 iterations, (d) is the test set confusion matrix for 5 iterations, (e) is the relative training set confusion matrix for 15 iterations, and (f) is the test set confusion matrix for 15 iterations;

[0076] Figure 6 is a ROC-AUC curve diagram for 2 iterations;

[0077] Figure 7 is a ROC-AUC curve diagram for 5 iterations;

[0078] Figure 8 is a ROC-AUC curve diagram for 15 iterations; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0079] The present invention will be further described in detail below in conjunction with embodiments and specific implementation manners. However, this should not be construed as limiting the scope of the present invention to the following embodiments. All technologies implemented based on the content of the present invention fall within the scope of the present invention.

[0080] Embodiment

[0081] Taking a classic two-span box girder bridge as an example, the method for rapid post-earthquake assessment of bridges based on active learning XGBoost in this embodiment includes the following steps:

[0082] S1. Dataset construction: Use OpenSees to build a bridge model, extract bridge-related parameters as input features for machine learning, and divide the bridge safety state categories according to the post-earthquake performance of the bridge as output labels, thereby constructing a machine learning dataset;

[0083] S2. Sample division: Divide the dataset constructed in step S1 into an initial training set, an unlabeled pool, and a test set according to a ratio of 1:6:3 to ensure the consistency of the class distribution in different datasets;

[0084] S3. XGBoost model training: Through the active learning method, select the most uncertain samples from the unlabeled pool, combine the selected samples for XGBoost model training, and optimize the hyperparameters through grid search;

[0085] S4. Optimal model determination: By comparing the performance of the model under different active learning iteration times and hyperparameter combinations, select the optimal XGBoost model as the final model;

[0086] S5. Post-earthquake safety assessment: Use the optimal XGBoost model selected in step S4, take bridge parameters and ground motion data as input, and output the corresponding safety state classification label as the basis for post-earthquake safety assessment of the bridge.

[0087] The use of OpenSees for bridge modeling described in step S1, the specific information is:

[0088] The present invention selects to combine an integral abutment and a single-column pier for modeling a two-span box girder bridge, and the modeling information is as Figure 2 .

[0089] The steps of extracting bridge-related parameters as input features for machine learning and dividing the bridge safety state categories according to the post-earthquake performance of the bridge as output labels, thereby constructing a machine learning dataset described in step S1 are as follows:

[0090] S111, Input Feature Extraction: Extract the structural information of the bridge and the peak acceleration during an earthquake as the input features of the dataset. The present invention selects the following bridge information as the input features of the XGBoost machine learning model: bridge span, deck width, concrete compressive strength, steel yield strength, net height of pier columns, longitudinal reinforcement ratio of columns, transverse reinforcement ratio of columns, translational stiffness, lateral rotational stiffness, lateral / longitudinal rotational stiffness ratio, height of abutment back, pile stiffness, type of backfill soil, mass coefficient, damping ratio, direction of seismic wave propagation, maximum peak acceleration.

[0091] S112, Output Label Division: Adopting the ATC-13 standard, divide the post-earthquake safety state of the bridge into three levels: green, yellow, and red, representing different safety levels respectively. The specific classification method is as follows:

[0092] Use the second stiffness method to determine the yield state of the bridge, that is, when the bridge stiffness drops to 50% of the initial stiffness, it is defined as the yield state; when the stress of the bridge cover concrete drops to 0, it is defined as the spalling state.

[0093] According to the post-earthquake bridge state division rules:

[0094] If the bridge state is lower than the yield state, mark it as green (safe, can pass);

[0095] If the bridge state is between the yield state and the spalling state, mark it as yellow (can pass under emergency conditions, but further inspection is required before putting it into daily traffic);

[0096] If the bridge state exceeds the spalling state, mark it as red (dangerous, prohibited from passing). (For example Figure 3 )

[0097] The output label is the safety state of the bridge, which are green, yellow, and red respectively.

[0098] S113, Dataset Construction: According to the above standards, mark the post-earthquake state of each bridge sample as green, yellow, or red, and use these labels as the output labels of the dataset.

[0099] The dataset constructed in step S1 is divided into an initial training set, an unlabeled pool, and a test set according to the ratio of 1:6:3 as described in step S2. The specific information is as follows:

[0100] In the dataset constructed in step S213, adopt the method of random sampling to divide the dataset into an initial training set, an unlabeled pool, and a test set according to the ratio of 1:6:3, and ensure that the sample ratio of each safety state level in each subset is roughly consistent with the original dataset, so as to avoid the adverse impact of class imbalance on the model training and evaluation results.

[0101] In the active learning method described in step S3, the most uncertain samples are selected from the unlabeled pool, and the XGBoost model is trained by combining the selected samples, and the hyperparameters are optimized by grid search. The specific steps are as follows:

[0102] S311, training the XGBoost model with the initial training set: Train the XGBoost model using the initial training set. The XGBoost model fits by learning the features and labels of the initial training set to discover patterns and relationships in the data;

[0103] S312, calculating the sample uncertainty based on entropy value: Calculate the entropy of each unlabeled pool sample to evaluate the sample uncertainty. For each sample x in the unlabeled pool i , calculate its predicted probability p of belonging to each security status level i,j , and then use the entropy formula to measure the uncertainty of this sample:

[0104]

[0105] where H(x i ) represents the entropy value of sample x i , M is the number of categories, taking 3, p i,j is the predicted probability that sample x i belongs to category j;

[0106] S313, selecting the most uncertain samples and updating the training set: Add the top q most uncertain samples calculated in step 312 to the initial training set, and remove this sample from the unlabeled pool. The updated training set is:

[0107]

[0108] The updated unlabeled pool is:

[0109]

[0110] where X train and y train are the features and labels of the initial training set respectively, X pool and y pool are the sample features and labels in the unlabeled pool respectively, and are the features and labels of the i-th most uncertain sample respectively;

[0111] S314, retraining the model and determining the loss function: Retrain the XGBoost model based on the updated training set and , and optimize the hyperparameters through grid search. The goal of grid search is to find the optimal hyperparameter combination Θ *, to minimize the logarithmic loss function L:

[0112]

[0113] where N is the total number of samples in the updated training set, M is the number of safety status levels, i.e., M takes 3, and y i,m is the true class label of sample i, and p i,m is the predicted probability that sample i belongs to a certain safety status level m;

[0114] S315, Grid search to optimize hyperparameters: After determining the loss function, grid search selects the optimal hyperparameters Θ * , which is determined by minimizing the loss function on the validation set:

[0115]

[0116] where Θ represents the set of hyperparameters of XGBoost (learning rate, maximum depth, regularization parameter, etc.), Γ is all possible combinations of hyperparameters, and (X val , y val ) is the validation set, and L(X val , y val , Θ) represents the loss of XGBoost on the validation set given the hyperparameters Θ. Through grid search, the optimal parameter Θ * is found, and then the optimal hyperparameters Θ * are used to train the XGBoost model;

[0117] S316, Set the number of active learning iterations n. Usually, when the number of samples is 100, active learning can achieve better results. Therefore, 20 samples are selected for training in each iteration. Select n = 5 iterations for experiments, and add 2 iterations and 15 iterations (i.e., all training samples) for comparison. The specific settings are as follows:

[0118] 2 iterations: A total of 40 active learning samples;

[0119] 5 iterations: A total of 100 active learning samples;

[0120] 15 iterations: A total of 288 active learning samples (i.e., all samples in the unlabeled pool are used for active learning).

[0121] Through these three different settings of the number of iterations, compare the experimental results.

[0122] The active learning method proposed by the present invention is for the three-classification problem, while Figure 4 schematically shows the application process in the basic two-label classification task. Figure 4(a) shows the actual decision boundary of the data when all labels are known. First, 4 samples are randomly selected to construct an initial machine learning model to demonstrate the basic principle of active learning Figure 4 (b). Due to the differences in models, the initial decision boundary (represented by a solid line) may be biased, and the vertical line is used to indicate its variation. Subsequently, the active learning algorithm selects 4 samples closest to the current decision boundary and obtains their true labels Figure 4 (c). The model is retrained using the initial samples and the newly sampled target samples to optimize the decision boundary, making it more in line with the data distribution than the initial model. Taking the Figure 4 data shown as an example, the decision boundary of the new model trained on the basis of 8 samples (4 random samples + 4 actively selected samples) can more accurately fit the actual classification situation, thereby improving the model performance.

[0123] Through grid search, the optimal parameter Θ * is found and finally used to train XGBoost.

[0124] The best model parameters for 5 iterations are as follows:

[0125] alpha: 1

[0126] colsample_bytree: 1

[0127] learning_rate: 0.2

[0128] max_depth: 5

[0129] min_child_weight: 4

[0130] n_estimators: 160

[0131] reg_lambda: 1.5

[0132] subsample: 0.8

[0133] In step S4, by comparing the performance of the model under different active learning iteration times and hyperparameter combinations, the optimal XGBoost model is selected as the final model. The specific steps are as follows:

[0134] S411. The total number of active learning samples for the 3 iteration times (2, 5, 15) in step S316 is 40, 100, and 288 respectively. In order to achieve a high evaluation accuracy rate, the goal of the present invention is to achieve the best performance using the least number of samples. According to the current number of samples, 2 and 5 iterations are superior in quantity to 15 iterations;

[0135] S412. For each number of iterations (2, 5, 15), the present invention calculated the accuracy rate of the model and other classification metrics on the test set. The formula for calculating the accuracy rate is as follows:

[0136]

[0137] Where:

[0138] TP (True Positive): The number of samples with the true class being positive and the model prediction also being positive;

[0139] TN (True Negative): The number of samples with the true class being negative and the model prediction also being negative;

[0140] FP (False Positive): The number of samples with the true class being negative but the model wrongly predicting as positive;

[0141] FN (False Negative): The number of samples with the true class being positive but the model wrongly predicting as negative;

[0142] At the 3 numbers of iterations, the accuracy rates of the model are respectively:

[0143] 2 iterations: 0.807

[0144] 5 iterations: 0.834

[0145] 15 iterations: 0.828

[0146] From these results, it can be seen that the 5 - iteration case achieved the highest accuracy rate with a relatively small number of samples. Thus, it is judged that the 5 - iteration case can best improve the model performance;

[0147] S413. To further evaluate the classification effect of the model, a confusion matrix was used to conduct a detailed analysis of the classification effects of the three iterations of the model, and the precision and recall were calculated:

[0148]

[0149] The confusion matrices of the present invention for 3 different numbers of iterations are as Figure 5 , in the confusion matrix, the present invention regarded the red label as an important selection criterion because the correct prediction of the red label is crucial for traffic closure and detailed inspection. From the perspectives of safety and traffic impact, accurately identifying the bridges with red labels is more critical than other labels. At these three numbers of iterations, the highest recall rate of the red label for the 2 - iteration and 5 - iteration cases is 0.710; considering the overall accuracy rate comprehensively, the 5 - iteration case has the highest accuracy rate and the number of samples used is much less than that of the 15 - iteration case. Therefore, the 5 - iteration case is considered the optimal choice;

[0150] S414. To balance precision and recall, calculate the F1-score:

[0151]

[0152] This invention is for three-class classification and uses the weighted F1-score:

[0153]

[0154] where ω m is the proportion of samples in class m; the F1-score results under three iterations are as follows:

[0155] F1-score for 2 iterations: 0.759

[0156] F1-score for 5 iterations: 0.833

[0157] F1-score for 15 iterations: 0.826

[0158] From the comparison of F1-scores, it can be seen that the performance of 5 iterations is the best.

[0159] Figures 6 - 8 The ROC curve (Receiver Operating Characteristic curve) at different iteration times is shown. The ROC curve reflects the performance of the classification model at different decision thresholds. The closer the curve is to the upper left corner, the better the performance of the model. The AUC value (Area Under the Curve) measures the overall discrimination ability of the model. The range of the AUC value is from 0 to 1, and the closer the value is to 1, the better the classification performance of the model. Specifically: the average AUC value for 2 iterations is 0.938, the average AUC value for 5 iterations is 0.937, and the average AUC value for 15 iterations is 0.948, indicating that the model can better distinguish positive and negative classes at different decision thresholds and has a high prediction ability. Considering the ROC curve and AUC value comprehensively, the effect of 15 iterations is better than that of 5 and 2 iterations, indicating that in terms of overall performance, 15 iterations perform better, but 5 iterations still show the best prediction effect because fewer samples are used.

[0160] S415. Considering accuracy, confusion matrix, F1-score, and ROC curve analysis comprehensively, the model with 5 iterations performs the best. Therefore, 5 iterations are finally selected as the optimal number of iterations; and the final XGBoost model and its best hyperparameter configuration are exported;

[0161] Furthermore, the specific information of the confusion matrix described in step S413 is as follows:

[0162] For a three-class classification problem, the confusion matrix is usually a 3x3 matrix, and adding the accuracies of the rows and columns forms a 4x4 matrix; in the 4x4 matrix: the rows represent the actual classes, and the columns represent the predicted classes; the elements on the diagonal represent the number of correct classifications of a certain class; the off-diagonal elements represent misclassifications.

Claims

1. An active learning-based XGBoost rapid post-earthquake assessment method for bridges, characterized in that, The method includes the following steps: S1. Dataset construction: Use OpenSees to model the bridge, extract bridge-related parameters as input features for machine learning, and divide the bridge safety status categories according to the post-earthquake performance of the bridge as output labels, thereby constructing a machine learning dataset; S2. Sample division: Divide the dataset constructed in step S1 into an initial training set, an unlabeled pool, and a test set at a ratio of 1:6:3 to ensure the consistency of the category distribution in different datasets; S3. XGBoost model training: Through the active learning method, select the most uncertain samples from the unlabeled pool, combine the selected samples for XGBoost model training, and optimize the hyperparameters through grid search; S4. Optimal model determination: By comparing the performance of the model under different active learning iteration times and hyperparameter combinations, select the optimal XGBoost model as the final model; S5. Post-earthquake safety assessment: Use the optimal XGBoost model selected in step S4, take bridge parameters and ground motion data as input, and output the corresponding safety status classification label as the basis for post-earthquake safety assessment of the bridge.

2. The method for rapid post-earthquake assessment of bridges based on active learning XGBoost according to claim 1, wherein The specific steps for extracting bridge-related parameters as input features for machine learning and dividing the bridge safety status categories according to the post-earthquake performance of the bridge as output labels in step S1 to construct a machine learning dataset are as follows: S111, Input feature extraction: Extract the structural information of the bridge and the peak acceleration during the earthquake as input features of the dataset; S112, Output label division: Adopt the ATC-13 standard to divide the post-earthquake safety status of the bridge into three levels: green, yellow, and red, representing different safety levels respectively. The specific classification method is as follows: Use the second stiffness method to determine the yield state of the bridge, that is, when the bridge stiffness drops to 50% of the initial stiffness, it is defined as the yield state; when the stress of the bridge cover concrete drops to 0, it is defined as the spalling state; According to the post-earthquake bridge state division rules: If the bridge state is lower than the yield state, it is marked as green, that is, safe and can be passed; If the bridge state is between the yield state and the spalling state, it is marked as yellow, that is, it can be passed under emergency conditions, but further inspection is required before putting it into daily traffic; If the bridge state exceeds the spalling state, it is marked as red, that is, dangerous and passage is prohibited; S113, Dataset construction: According to the above standards, mark the post-earthquake state of each bridge sample as green, yellow, or red, and use these labels as the output labels of the dataset.

3. The rapid post-earthquake assessment method of bridges based on active learning XGBoost according to claim 1, wherein The specific method for dividing the dataset constructed in step S1 into an initial training set, an unlabeled pool, and a test set at a ratio of 1:6:3 in step S2 is as follows: In the dataset constructed in step S113, use the random sampling method to divide the dataset into an initial training set, an unlabeled pool, and a test set at a ratio of 1:6:3, and ensure that the sample ratio of each safety status level in each subset is roughly the same as that of the original dataset, so as to avoid the adverse impact of class imbalance on the model training and evaluation results.

4. The method for rapid post-earthquake assessment of bridges based on active learning XGBoost according to claim 1, wherein In the active learning method described in step S3, the most uncertain samples are selected from the unlabeled pool, and the XGBoost model is trained by combining the selected samples, and the hyperparameters are optimized by grid search. The specific steps are as follows: S311, training the XGBoost model with the initial training set: Use the initial training set to train the XGBoost model. The XGBoost model fits by learning the features and labels of the initial training set to discover the patterns and relationships in the data; S312, Calculate sample uncertainty based on entropy value: Calculate the entropy of each unlabeled pool sample to evaluate the sample uncertainty. For each sample x in the unlabeled pool i , calculate the predicted probability p of it belonging to each security status level i,j , and then use the entropy formula to measure the uncertainty of this sample: Among them, H(x i ) represents the entropy value of the sample x i , M is the number of categories, taking 3, p i,j is the predicted probability that the sample x i belongs to the category j; S313, selecting the most uncertain samples and updating the training set: Add the top q most uncertain samples calculated in step 312 to the initial training set, and remove the samples from the unlabeled pool. The updated training set is: The updated unlabeled pool is: Among them, X train and y train are the features and labels of the initial training set respectively, and X pool and y pool are the sample features and labels in the unlabeled pool respectively, and are the features and labels of the i-th most uncertain sample respectively; S314, Retrain the model and determine the loss function: Based on the updated training set and Retrain the XGBoost model and optimize the hyperparameters through grid search. The goal of the grid search is to find the optimal hyperparameter combination Θ * to minimize the logarithmic loss function L: Among them, N is the total number of samples in the updated training set, M is the number of safety status levels, that is, M takes 3, and y i,m is the true class label of sample i, and p i,m is the predicted probability that sample i belongs to a certain safety status level m; S315, Grid search for optimizing hyperparameters: After determining the loss function, the grid search selects the optimal hyperparameters Θ * , which is determined by minimizing the loss function on the validation set: Among them, Θ represents the set of hyperparameters of XGBoost (learning rate, maximum depth, regularization parameter, etc.), Γ is all possible combinations of hyperparameters, (X val , y val ) is the validation set, and L(X val , y val , Θ) represents the loss of XGBoost on the validation set given the hyperparameters Θ. Through grid search, the optimal parameter Θ * is found, and then the optimal hyperparameters Θ * are used to train the XGBoost model; S316, Repeated active learning iteration: Generally, when the number of samples reaches 100, active learning can achieve better results. Set the number of active learning iterations as n, where Repeat steps S312 to S315 (n - 1) times and output the accuracy on the test set; S317. Reset the number of iterations to k times. Repeat steps S311 to S316. S318. Reset the number of iterations to l times. In the present invention, the total number of samples in the unlabeled pool is N1≥100; repeat steps S311 to S316.

5. The method for rapid post-earthquake assessment of bridges based on active learning XGBoost according to claim 4, wherein, In step S4, by comparing the performance of the XGBoost model under different active learning iteration times and hyperparameter combinations, the optimal XGBoost model is selected as the final model. The specific steps are as follows: S411, comparing the number of active learning samples for the three iteration times n, k, and l set in steps 316 - 318. By analyzing the number of samples selected under each iteration time, evaluate the impact of different iteration times on sample selection; S412, comparing the accuracy of the model on the test set for the three iteration times n, k, and l. The accuracy calculation formula is: Where: TP (True Positive): The number of samples with the true class being positive and the model prediction also being positive; TN (True Negative): The number of samples with the true class being negative and the model prediction also being negative; FP (False Positive): The number of samples with the true class being negative and the model wrongly predicting as positive; FN (False Negative): The number of samples with the true class being positive and the model wrongly predicting as negative; By comparing the accuracy under different iteration times, determine the iteration time that can best improve the model performance; S413, comparing the confusion matrices of the XGBoost model for the three iteration times n, k, and l. Use the confusion matrix to conduct a detailed analysis of the classification effect of the XGBoost model, and calculate the precision and recall: Through these metrics, further evaluate the classification ability of the model under different iteration times; S414, comparing the F1 - score of the XGBoost model for the three iteration times, and analyzing the ROC curve and AUC value of the model. The calculation formula of the F1 - score is: Using the weighted F1 - score: where ω m is the proportion of samples in class m, and F1-score m is the F1-score of class m; S415, based on the comprehensive performance evaluation results in steps S411 to S414, select the optimal iteration time and its corresponding XGBoost model, and export the final XGBoost model and its corresponding best hyperparameter configuration.

6. The post-earthquake rapid assessment method for bridges based on active learning XGBoost according to claim 5, characterized in that The confusion matrix described in step S413, the specific information is: For a three-class classification problem, the confusion matrix is usually a 3x3 matrix, and adding the accuracies of rows and columns forms a 4x4 matrix. In the 4x4 matrix: the rows represent the actual classes, and the columns represent the predicted classes; the elements on the diagonal represent the number of correctly classified instances of a certain class; the off-diagonal elements represent misclassifications.