Roadbed bearing capacity prediction method based on large language model and related equipment
By integrating the MCI model and the XGBoost residual compensation model into a subgrade resilient modulus prediction method, and combining the semantic understanding capabilities of a large language model, the problems of high cost and limited applicability of traditional methods are solved, achieving high-precision and rapid subgrade bearing capacity assessment and engineering practicality report generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH AT WEIHAI
- Filing Date
- 2026-01-12
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional methods for determining the elastic modulus of roadbeds are costly and time-consuming. Conventional resilient modulus prediction models have limited applicability and cannot meet the need for rapid and accurate assessment of roadbed performance. Furthermore, existing physical experience models lack sufficient prediction accuracy under complex soil conditions.
A hybrid model for predicting the resilient modulus of roadbed is constructed, which integrates the MCI model and the XGBoost residual compensation model. By combining the semantic understanding capabilities of the large language model and optimizing the hyperparameters through Bayesian optimization, intelligent analysis and residual compensation of roadbed parameters are achieved, generating high-precision predicted values of the resilient modulus of roadbed.
It improves the accuracy and adaptability of roadbed bearing capacity prediction, reduces reliance on complex tests, and generates prediction results with physical consistency and engineering applicability, enabling the generation of reliable engineering recommendation reports.
Smart Images

Figure CN122132948A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of roadbed bearing capacity prediction technology, and in particular to a roadbed bearing capacity prediction method and related equipment based on a large language model. Background Technology
[0002] Traditional laboratory methods for determining the elastic modulus of roadbeds are costly and time-consuming. Conventional resilient modulus prediction models often rely on a large number of experiments to determine their internal parameters, or the models have limited applicability and are only applicable to a single soil type. At the same time, related measurement processes, such as the determination of matrix water absorption rate, are also very cumbersome. These factors together restrict the widespread application of rapid and accurate assessment of roadbed performance in engineering practice.
[0003] To address the aforementioned challenges, some improved prediction models have emerged, such as the MCI model proposed by Chu et al. The parameters required in this model, such as the plasticity index and moisture content, are relatively easy to obtain from engineering surveys. Compared with complex direct measurements, this reduces the difficulty of data acquisition to some extent and provides a more convenient way to estimate the resilient modulus of roadbeds.
[0004] However, such models based on physical experience still have limitations. As a fixed-form empirical formula, the prediction accuracy of the MCI model is limited by the simplification assumptions and the failure to consider complex factors. In practical applications, it may produce systematic prediction bias or residuals, making it difficult to fully adapt to various complex and changing field soil conditions and stress states. Therefore, relying solely on such models cannot fully meet the needs of high-precision prediction.
[0005] In summary, it is necessary to develop a new prediction method that can utilize prior knowledge from physical models, adaptively learn and correct prediction biases, and intelligently understand engineering requirements, in order to further improve the accuracy, adaptability, and engineering applicability of predictions. Summary of the Invention
[0006] The technical problem to be solved by this invention is to address the shortcomings of existing technologies. Specifically, it provides a method and related equipment for predicting the bearing capacity of roadbeds based on a large language model, as detailed below: 1) In a first aspect, the present invention provides a method for predicting the bearing capacity of roadbeds based on a large language model, the specific technical solution of which is as follows: A hybrid model for predicting the resilient modulus of the subgrade was constructed, which integrates the MCI model and the XGBoost residual compensation model. The MCI model is used to calculate the initial predicted value of the resilient modulus of the subgrade based on the subgrade parameters; the XGBoost model is used to predict the residual of the initial predicted value of the resilient modulus of the subgrade. The XGBoost residual compensation model was trained using the training set, and the hyperparameters of the XGBoost residual compensation model were tuned using the Bayesian optimization method to obtain the trained XGBoost residual compensation model. Based on the semantic understanding capabilities of a large language model, user input is parsed and target roadbed parameters are extracted; Based on the target subgrade parameters, the trained XGBoost residual compensation model, and the MCI model, the final predicted value of the subgrade resilient modulus is generated.
[0007] The beneficial effects of the roadbed bearing capacity prediction method based on a large language model provided by this invention are as follows: By constructing a hybrid model for predicting the resilient modulus of subgrade using both the MCI model and the XGBoost residual compensation model, this approach effectively combines the interpretability of physical empirical models with the nonlinear fitting advantages of data-driven models. The MCI model provides initial predictions based on subgrade parameters, establishing a physically consistent prediction benchmark. The XGBoost residual compensation model specifically learns and predicts the residuals of the initial predictions, thereby specifically compensating for the systematic biases of the MCI model. This fusion architecture improves the overall accuracy of the final subgrade resilient modulus predictions and its adaptability to different soil conditions. Using Bayesian optimization to tune the hyperparameters of the XGBoost residual compensation model automatically searches for superior hyperparameter combinations, resulting in a more generalized and stable prediction performance, avoiding the blindness and inefficiency of manual parameter tuning. Introducing semantic understanding capabilities based on a large language model to parse user input and extract target subgrade parameters enables the method to handle natural language or semi-structured engineering descriptions, lowering the barrier to entry for specialized software and improving the efficiency and user-friendliness of human-computer interaction. Finally, based on the extracted target subgrade parameters, the trained XGBoost residual compensation model and MCI model are used in synergy to generate predicted values, forming a complete workflow from intelligent interaction to high-precision calculation. This method not only inherits the advantage of easy-to-obtain MCI model parameters, but also overcomes the shortcomings of large prediction deviations in traditional physical models and the inconvenience of applying conventional prediction methods through residual compensation and intelligent analysis. It achieves more reliable and convenient subgrade bearing capacity assessment results while reducing reliance on complex experiments.
[0008] Based on the above scheme, the roadbed bearing capacity prediction method based on a large language model of the present invention can be further improved as follows.
[0009] Based on the target subgrade parameters, the trained XGBoost residual compensation model, and the MCI model, the final predicted subgrade resilient modulus value is generated, including: The target subgrade parameters are input into the MCI model to calculate the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters. The target subgrade parameters and the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters are then input into the trained XGBoost model to predict the residual. The initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters is added to the predicted residual to obtain the final predicted value of the subgrade resilient modulus.
[0010] The beneficial effects of adopting the above-mentioned further scheme are as follows: inputting the target subgrade parameters into the MCI model directly yields an initial predicted value of the subgrade resilient modulus with physical meaning, establishing a reasonable prediction benchmark. Inputting the target subgrade parameters and this initial predicted value together into the trained XGBoost residual compensation model enables targeted prediction of specific deviations in the physical model under current conditions, i.e., residuals. Finally, adding the initial predicted value to the predicted residuals completes the accurate correction of the physical benchmark value. This step-by-step calculation and compensation process ensures that the final generated subgrade resilient modulus prediction value possesses both physical consistency and data-driven adaptability, thus obtaining a more accurate output result than a single model.
[0011] Furthermore, it also includes generating a report containing the final predicted resilient modulus of the subgrade and engineering recommendations based on the final predicted resilient modulus of the subgrade and the domain knowledge obtained by the retrieval enhancement generation technology.
[0012] The beneficial effects of adopting the above-mentioned further approach are as follows: By combining domain knowledge acquired through retrieval-enhanced generation technology, the final predicted value of the roadbed resilient modulus can be interpreted and evaluated within a standardized engineering context. The generated report not only presents the predicted values but also integrates authoritative knowledge from standards, literature, and historical data, thereby deriving evidence-based reliability analysis and specific engineering recommendations. This process transforms the prediction results from isolated data points into reports that can directly guide construction and design, enhancing the engineering practical value of the output and the efficiency of user decision-making.
[0013] Furthermore, the process of obtaining the training set includes: obtaining multiple samples, each sample including the original subgrade parameters and the actual value of the subgrade resilient modulus corresponding to the original subgrade parameters; for each sample, inputting the original subgrade parameters into the MCI model to obtain the initial predicted value of the subgrade resilient modulus, calculating the residual between the initial predicted value of the subgrade resilient modulus and the actual value of the subgrade resilient modulus, combining the original subgrade parameters and the residual to form the enhanced features of the sample, and using the enhanced features of all samples to form the training set. The original subgrade parameters include plasticity index, moisture content, confining pressure and deviatoric stress.
[0014] The beneficial effects of adopting the above-mentioned further scheme are as follows: By calculating the initial predicted value of the subgrade resilient modulus of the MCI model for each sample and obtaining the corresponding residual, the target variable that the machine learning model needs to learn and predict is clearly defined. Combining the original subgrade parameters with the residuals into enhanced features ensures that the training set not only contains basic soil state information but also quantitative information on the prediction bias of the physical model. This data construction method enables the XGBoost residual compensation model to directly learn the complex nonlinear relationship between the original parameters and the prediction error of the MCI model during training, thus laying a solid data foundation for subsequent accurate residual compensation.
[0015] 2) Secondly, the present invention also provides a roadbed bearing capacity prediction system based on a large language model, the specific technical solution of which is as follows: It includes a model building module, a model training module, a parameter extraction module, and a model application module; The model building module is used to: construct a hybrid model for predicting the resilient modulus of the subgrade. This hybrid model integrates the MCI model and the XGBoost residual compensation model. The MCI model is used to calculate the initial predicted value of the resilient modulus of the subgrade based on the subgrade parameters; the XGBoost model is used to predict the residual of the initial predicted value of the resilient modulus of the subgrade. The model training module is used to: train the XGBoost residual compensation model using the training set, and fine-tune the hyperparameters of the XGBoost residual compensation model using the Bayesian optimization method to obtain the trained XGBoost residual compensation model. The parameter extraction module is used to: parse user input based on the semantic understanding capabilities of a large language model and extract target roadbed parameters; The model application module is used to generate the final predicted value of the roadbed resilient modulus based on the target roadbed parameters, the trained XGBoost residual compensation model, and the MCI model.
[0016] Based on the above scheme, the roadbed bearing capacity prediction system based on a large language model of the present invention can be further improved as follows.
[0017] Furthermore, the model application module is specifically used to: input the target subgrade parameters into the MCI model, calculate the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters, input the target subgrade parameters and the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters into the trained XGBoost model to predict the residual, and add the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters to the predicted residual to obtain the final predicted value of the subgrade resilient modulus.
[0018] Furthermore, it also includes a report generation module, which is used to generate a report containing the final predicted subgrade resilient modulus and engineering recommendations based on the final predicted subgrade resilient modulus and the domain knowledge obtained by the retrieval enhancement generation technology.
[0019] Furthermore, it also includes a training set acquisition module, which is used for: Multiple samples were obtained, each sample including the original subgrade parameters and the actual value of the subgrade resilient modulus corresponding to the original subgrade parameters; For each sample, the original subgrade parameters are input into the MCI model to obtain the initial predicted value of the subgrade resilient modulus. The residual between the initial predicted value of the subgrade resilient modulus and the actual value of the subgrade resilient modulus is calculated. The original subgrade parameters and the residual are combined into the enhanced features of the sample. The enhanced features of all samples constitute the training set. The original subgrade parameters include plasticity index, moisture content, confining pressure and deviatoric stress.
[0020] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so that the electronic device implements any of the above-mentioned methods for predicting the bearing capacity of roadbed based on a large language model.
[0021] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned methods for predicting the bearing capacity of roadbeds based on a large language model.
[0022] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below: Figure 1 This is a flowchart illustrating a method for predicting the bearing capacity of a roadbed based on a large language model, according to an embodiment of the present invention. Figure 2 A schematic diagram of the workflow for a hybrid model for predicting the resilient modulus of roadbed by fusing the MCI model and the XGBoost residual compensation model.
[0024] Figure 3 This is a schematic diagram of the workflow of the intelligent agent framework for predicting the bearing capacity of roadbeds based on a large language model.
[0025] Figure 4This is a schematic diagram of a roadbed bearing capacity prediction system based on a large language model, according to an embodiment of the present invention. Detailed Implementation
[0026] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0027] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0028] like Figure 1 As shown in the figure, an embodiment of the present invention provides a method for predicting the bearing capacity of roadbed based on a large language model, which includes the following steps: S1. Construct a hybrid model for predicting the resilient modulus of the subgrade. The hybrid model for predicting the resilient modulus of the subgrade integrates the MCI model and the XGBoost residual compensation model. The MCI model is used to calculate the initial predicted value of the resilient modulus of the subgrade based on the subgrade parameters. The XGBoost model is used to predict the residual of the initial predicted value of the resilient modulus of the subgrade. S10. Collect sample data for model training and testing. Each sample must contain a set of original terminal parameters describing the subgrade condition, and the actual value of the corresponding subgrade resilient modulus measured experimentally. The original subgrade parameters specifically include plasticity index, moisture content, confining pressure, and deviatoric stress. The confining pressure is often denoted as […] in the model. The deviatoric stress is denoted as Furthermore, based on the plasticity index and moisture content, another key state parameter, namely the consistency index, can be calculated. These parameters together form the basis of the model's input features.
[0029] S11. For each sample, substitute its parameters such as plasticity index, water content, confining pressure, and deviatoric stress into the mathematical expression of the MCI model. The formula for the MCI model is: , This represents the roadbed resilient modulus predicted by the MCI model; Indicates the consistency index; Indicates confining pressure; This represents cyclic deviatoric stress, which, in the context of this model, corresponds to the deviatoric stress in the input features. Equivalent; Standard atmospheric pressure is a constant. , , , , These are all empirical coefficients determined by fitting a large amount of experimental data to the MCI model. The result calculated using this formula is denoted as... This is the initial predicted value of the resilient modulus of the current sample roadbed.
[0030] S12. Calculate the initial predicted value. The difference between the actual value of the roadbed resilient modulus and the known value is defined as the residual. Simultaneously, a new characteristic is derived from the original data: the ratio of deviatoric stress to confining pressure. Subsequently, the original input features, including plasticity index, moisture content, and consistency index, are processed. Confining pressure eccentric stress Together with derived features Together with the calculated residual values, they form a new, more informative enhanced feature set. This enhanced feature set not only includes the basic physical state of the soil but also the prediction bias information of the initial physical model, laying the foundation for subsequent data-driven model learning of complex error patterns.
[0031] S13. The enhanced feature set of all obtained samples is used as training data. At this point, the model input consists of various features from the enhanced feature set, while the model's prediction target is the residual value. Before training, the dataset needs to be divided into a part for model training and validation, and a part for final performance testing. During the model training and validation phase, Bayesian optimization is used to systematically tune the hyperparameters of the XGBost residual compensation model to find the parameter combination that best performs on the validation set. The model's learning ability and generalization performance are evaluated using five-fold cross-validation. This process ultimately produces a trained and optimized XGBost residual compensation model, which has the ability to accurately predict the residual value of the initial prediction of the MCI model based on the input roadbed parameters and their initial prediction state.
[0032] S14. After completing the aforementioned steps, the MCI model and the XGBoost residual compensation model are integrated into a unified hybrid model for predicting the resilient modulus of the subgrade. For a completely new input, the prediction process of the hybrid model is as follows: first, input the target subgrade parameters into the MCI model to calculate the initial predicted value of the subgrade resilient modulus. Meanwhile, these target roadbed parameters, along with the calculated initial predicted values, are constructed into enhanced features consistent with the format in S13, and input into the trained XGBost residual compensation model to predict the corresponding residual values. The final prediction result It is obtained by adding the initial predicted value to the predicted residual, i.e. This fusion approach allows the hybrid model to maintain the interpretability framework of the physical model while compensating for the inherent biases of the physical model under complex real-world conditions through data-driven methods, thereby improving the overall prediction accuracy.
[0033] The MCI model, short for Modified Consistency Index Model, is an empirical model for predicting the resilient modulus of subgrade based on soil physical state parameters. By establishing mathematical relationships between the subgrade resilient modulus and key parameters such as the consistency index, confining pressure, and deviatoric stress, this model provides a preliminary estimate of the physical mechanism of subgrade mechanical behavior. Its advantage lies in the fact that required parameters, such as the plasticity index and moisture content, are relatively easy to obtain from engineering surveys, providing a physically meaningful baseline for prediction.
[0034] The XGBoost residual compensation model is a machine learning model based on the gradient boosting decision tree algorithm. In this framework, it is specifically used to predict the deviation between the initial estimate and the true value of the MCI model, i.e., the residual. This model can automatically learn and capture complex nonlinear relationships and interactions from high-dimensional augmented datasets containing both original and derived features. These relationships may not be fully described by simplified physical models such as the MCI model. By training this model to accurately compensate for the residual, the accuracy and adaptability of the entire prediction method are effectively improved.
[0035] The resilient modulus of subgrade is a key mechanical parameter characterizing the elastic deformation capacity of subgrade soil under repeated loading. It directly reflects the bearing capacity and resistance to deformation of the subgrade. Accurate prediction of the resilient modulus of subgrade is of vital engineering significance for road structure design, construction quality control, and long-term performance evaluation. Although traditional laboratory measurement methods are accurate, they are costly and time-consuming, making it difficult to meet the rapid evaluation needs of large-scale engineering applications. Therefore, developing efficient and accurate prediction models has significant practical value.
[0036] S2. Train the XGBoost residual compensation model using the training set, and fine-tune the hyperparameters of the XGBoost residual compensation model using the Bayesian optimization method to obtain the trained XGBoost residual compensation model. The process of acquiring the training set includes: acquiring multiple samples, each sample including the original subgrade parameters and the actual value of the subgrade resilient modulus corresponding to the original subgrade parameters; for each sample, inputting the original subgrade parameters into the MCI model to obtain the initial predicted value of the subgrade resilient modulus, calculating the residual between the initial predicted value of the subgrade resilient modulus and the actual value of the subgrade resilient modulus, combining the original subgrade parameters and the residual to form the enhanced features of the sample, and using the enhanced features of all samples to form the training set. The original subgrade parameters include plasticity index, moisture content, confining pressure, and deviatoric stress. Specifically: S020. Collect a large number of representative samples from laboratory test reports, engineering survey databases, or publicly available research materials. Each sample is an independent data unit and must contain two parts of information: the first part is a set of original subgrade parameters describing the specific subgrade soil state and stress conditions; the second part is the actual value of the subgrade resilient modulus corresponding to the set of original subgrade parameters, directly measured through standard indoor tests, such as repeated loading triaxial tests. The specific composition of the original subgrade parameters must include the plasticity index, water content, confining pressure, and deviatoric stress. These parameters are the basis for all subsequent calculations.
[0037] S021. For each sample, its original subgrade parameters need to be input into the MCI model. Specifically, the plasticity index and moisture content need to be extracted from the sample to calculate the consistency index. Simultaneously, the confining pressure values in the sample are extracted. and deviatoric stress value Substitute these parameters into the standard formula of the MCI model for calculation. The formula for the MCI model is expressed as: ,in, This represents the initial predicted value of the roadbed resilient modulus output by the MCI model. For distinction in subsequent steps, this calculation result is denoted as... ; Indicates the consistency index; Indicates confining pressure; This represents the cyclic deviatoric stress, which, in the context of this model, is numerically different from the deviatoric stress provided by the sample. They are equal, therefore they can be used directly. Perform calculations; Standard atmospheric pressure is a known constant value. , , , , These are fixed empirical coefficients pre-determined by fitting a large amount of data to the MCI model. By executing this formula, a corresponding coefficient is obtained for each sample. .
[0038] S022. Calculate the residuals, defined as the difference between the predicted values of the physical model and the actual observed values. This is done after obtaining the initial predicted values of the roadbed resilient modulus for each sample. Next, the actual value of the known roadbed resilient modulus of the sample needs to be read from the sample data, denoted as . Residual The calculation formula is: , The value quantitatively represents the prediction bias of the MCI model for this sample. It will serve as the target variable that the XGBoost residual compensation model needs to learn and predict during training.
[0039] S023. To provide machine learning models with richer information to learn residual patterns, it is not sufficient to use only the original single feature. For each sample, multiple relevant features need to be combined into a feature vector. This feature vector includes: the original subgrade parameters of the sample itself, namely plasticity index, moisture content, and confining pressure. eccentric stress ; and important engineering characteristics derived from these parameters, such as the ratio of deviatoric stress to confining pressure. In addition, it also includes the initial predicted value of the roadbed resilient modulus. Finally, all the above features are compared with the residuals calculated in the previous step. These are linked together to form a complete, labeled data pair. This data pair is the enhanced feature representation of the sample, with the combined feature vector as its input and the residual as its output. .
[0040] S024. Repeat steps S020 to S024 to process all available samples consistently. Each sample undergoes the above process, transforming it into a uniformly formatted enhanced feature data pair. Collecting these enhanced feature data pairs for all samples constitutes a complete dataset. This dataset contains rich input features and a clear prediction target, namely the residual. This dataset can be directly used for subsequent training, validation, and hyperparameter tuning tasks of the XGBoost residual compensation model. To ensure the fairness of model evaluation, it is usually necessary to divide this complete dataset into non-overlapping training, validation, and test subsets according to a certain ratio.
[0041] The plasticity index is an important indicator describing the plasticity of fine-grained soils. Numerically, it is equal to the difference between the liquid limit water content and the plastic limit water content of the soil. The plasticity index reflects the clay mineral composition and activity in the soil and is a key parameter for soil classification and determining its engineering properties. The higher the plasticity index, the stronger the plasticity of the soil, and the more sensitive its mechanical properties are to changes in water content.
[0042] Moisture content refers to the ratio of the mass of water to the mass of solid particles in soil, usually expressed as a percentage. Moisture content is an extremely important state variable affecting the strength, deformation, and compaction characteristics of soil. The mechanical properties of subgrade soil, such as the resilient modulus, change significantly with changes in moisture content, making it an indispensable input parameter in the prediction of subgrade bearing capacity.
[0043] Confining pressure refers to the average stress acting on a soil element in all directions, which is usually provided by a hydraulic chamber in triaxial tests. Confining pressure simulates the surrounding constraint pressure on soil in a roadbed structure, and it has a significant impact on the soil's strength and modulus. Generally speaking, as the confining pressure increases, the soil's stiffness and bearing capacity will increase accordingly.
[0044] Deviatoric stress refers to the difference between the maximum and minimum principal stresses borne by a soil element. It represents the level of shear stress that causes the soil to change shape. When a roadbed is subjected to traffic loads, deviatoric stress is a dynamically changing quantity that directly relates to whether the soil will undergo plastic deformation or failure. It is a key dynamic parameter for calculating the resilient modulus of the roadbed.
[0045] The specific implementation process of S2 is as follows: S20. The dataset consisting of the enhanced features of all samples is explicitly divided into three parts: a training set, a validation set, and a test set. A fixed random seed, such as random seed 42, is used for the partitioning to ensure the repeatability of each experiment. The training set is used to directly adjust the model's internal weights; the validation set is used to evaluate the model's performance under different hyperparameter configurations during training to guide the optimization direction; the test set is used to finally evaluate the generalization ability of the selected model and is not used in the entire tuning process. For the XGBoost residual compensation model, the input of the XGBoost residual compensation model is the enhanced feature vector of each sample, including the initial predicted values of plasticity index, water content, confining pressure, deviatoric stress, and roadbed resilient modulus; its output target is a single continuous value, that is, the residual corresponding to the sample. The model's task is to learn a mapping function from complex features to residual values. .
[0046] S21 and XGBoost models have many hyperparameters that control their learning behavior and complexity. Bayesian optimization methods require searching a predefined parameter space. Key hyperparameters that need tuning typically include: the learning rate, which controls the magnitude of model weight updates in each iteration; the maximum tree depth, which limits the number of layers each decision tree can grow, affecting the model's ability to capture details; the subsampling rate, which controls the proportion of training data used to train each tree, helping to enhance randomness and prevent overfitting; the column sampling rate, which controls the proportion of features available when building each tree; and the minimum child node weight, which determines the minimum sum of sample weights required for the tree to continue splitting, affecting the tree's stopping condition. A reasonable range of values should be set for each hyperparameter; for example, the learning rate between 0.01 and 0.3, and the maximum depth between 3 and 10. This parameter space is the set of all possible combinations of hyperparameters and is the region explored by Bayesian optimization.
[0047] S22. Establish the objective function for Bayesian optimization. The objective function serves as a bridge connecting hyperparameter selection and model performance evaluation. The input to this function is a set of specific hyperparameter values, and the output is a scalar score used to evaluate the quality of these hyperparameters. Implementing this function requires the following steps: First, initialize an XGBoost regression model using the input hyperparameter configuration. Then, train this model using the training set data. Finally, use the trained model to predict on the validation set to obtain the predicted residual values. Compare the predicted values with the actual residual values on the validation set. To compare, calculate a pre-selected evaluation metric, such as negative root mean square error, using the following formula: .in, Indicates the number of samples in the validation set. This represents the sample index. A negative value is used here because the Bayesian optimization algorithm defaults to finding the maximum value of the function, and the root mean square error (RMSE) should be as small as possible; taking a negative value transforms the problem into finding the maximum value. This calculated score is the return value of the objective function.
[0048] S23. The Bayesian optimization algorithm uses a probabilistic surrogate model, typically a Gaussian process, to model the unknown relationship between the combination of hyperparameters and the objective function score. At the start of the first iteration, the probabilistic surrogate model has no prior knowledge, and the algorithm randomly selects several sets of hyperparameters for initial evaluation. Based on the results of these initial points, the probabilistic surrogate model establishes a preliminary predicted distribution about the objective function. In each subsequent iteration, based on a sampling function, such as the desired improvement, the algorithm recommends the next most promising point from the hyperparameter space to bring performance improvement. Then, the defined objective function is called with this recommended new set of hyperparameters to obtain a new performance score, and this data point is added to the observation set to update the probabilistic surrogate model. This "recommendation-evaluation-update" loop is repeated for a specified number of iterations, such as 100 times.
[0049] S24. After the Bayesian optimization loop reaches the preset maximum number of iterations, the combination that achieves the highest objective function score on the validation set is selected from all evaluated hyperparameter combinations; this is the optimal hyperparameter combination. Then, using this optimal hyperparameter combination, a completely new XGBoost residual compensation model is reinitialized. This time, the entire original training and validation sets are combined as the final training data to train the model completely. The essence of model training is to learn the weights and structure of the decision tree set by minimizing a specified loss function, such as mean squared error loss, through a gradient boosting algorithm. The objective function of XGBoost can be expressed as... ,in, It is a loss function used to calculate the predicted residuals. Compared with the true residual The error; It is a regularization term used to control the first... tree To reduce complexity and prevent overfitting; It is the total number of trees; This refers to the number of training samples. After this step, we obtain an XGBoost residual compensation model that has been fitted to all available training data with optimal hyperparameter configuration. This is the trained XGBoost residual compensation model.
[0050] S25. Apply the XGBoost residual compensation model trained in S24 to an independent test set that has never participated in the training and tuning process. Perform predictions on the test set and calculate metrics such as the coefficient of determination and root mean square error to objectively and fairly evaluate the model's final generalization performance. Simultaneously, the entire history of the Bayesian optimization process, including the hyperparameters of each iteration and the corresponding validation set scores, should be saved for analysis and retrospective purposes.
[0051] It should be noted that the specific reason why the MCI model does not require training is that the MCI model is a published and established empirical physical model, and its mathematical form and internal parameters, i.e., the formulas, are consistent. coefficients in The original authors determined the model through a one-time fitting using independent and extensive historical experimental data. In the framework proposed in this invention, the MCI model is introduced as a mature, fixed component providing physical benchmark predictions. Its role is to compute initial predictions, rather than relearning its parameters from the current, specific dataset. Therefore, the parameters of the MCI model remain unchanged throughout the construction of the hybrid model, requiring no training or tuning using the currently obtained training set.
[0052] S3. Based on the semantic understanding capability of a large language model, user input is parsed and target roadbed parameters are extracted. The specific implementation process is as follows: S30. Receive the user's initial request through an interactive interface, such as a chat window or form input box. User input can be free natural language description, such as a text description; it can also be semi-structured data, such as a table or list containing partial parameters. Capture the complete raw text or data structure of the user input and prepare to send it to the large language model for processing. For obvious formatting errors or garbled characters, the system will perform basic cleaning, but mainly relies on the robust understanding of irregular language by the large language model.
[0053] S31. To guide the large language model in accurate engineering parameter parsing, a structured prompt message is constructed. This prompt message integrates fixed system instructions, the current user query, and explicit output format requirements. The system instructions define the model's role as a geotechnical engineering analysis assistant and specify that the task is to extract key parameters from text for predicting the resilient modulus of the subgrade. The user query includes the obtained raw input. The output format requirements explicitly instruct the large language model to list the identified parameters in a specified structured format, such as JSON. A simplified prompt template is as follows: "You are a geotechnical engineering analysis assistant. Please identify and extract the parameters necessary for calculating the resilient modulus of the subgrade from the user's question. The parameters that must be extracted include: plasticity index, water content, confining pressure, and deviatoric stress. Please output only one JSON object, with the parameter name as the key and the identified numerical value and unit as the value. User input is: [Insert user input here]". This complete prompt is sent to the selected large language model service interface.
[0054] S32. Upon receiving the prompt, the large language model utilizes its semantic understanding capabilities to deeply analyze the meaning of the user's input text. The model identifies that the user's core intent is to request subgrade-related calculations and locates descriptive fragments related to the target parameters within the context. For example, given the input "Predict the bearing capacity of a type of clay with a liquid limit of 35, a plastic limit of 18, a water content of 22%, under a confining pressure of 100 kPa and a deviatoric stress of 50 kPa," the model needs to understand that "liquid limit 35, plastic limit 18" together implicitly contain the basis for calculating the "plasticity index" (35-18=17), and directly identify explicit values such as "water content 22%", "confining pressure 100 kPa", and "deviatoric stress 50 kPa". Even if the input expression is not entirely standardized, such as the user using "lateral pressure" to refer to "confining pressure", the model can still map synonyms or near-synonyms to standard technical terms based on the extensive corpus knowledge acquired during training.
[0055] S33. Large language models not only perform direct text matching but also execute simple logical reasoning and arithmetic operations to complete necessary information. As shown in the example in S32, the model needs to derive the "plasticity index" from the values of "liquid limit" and "plastic limit." Internally, it can perform a calculation: Plasticity index = Liquid limit - Plastic limit = 35 - 18 = 17. Simultaneously, the model checks the completeness of the extracted parameter list. If a necessary parameter (such as deviatoric stress) is missing from the user input, the model may mark the value of that field as "missing" or "not provided" in the output, or leave it blank in the JSON according to the prompt. In some advanced implementations, the model can also perform preliminary verification of the reasonableness of the extracted parameters based on retrieved domain knowledge, such as determining whether the moisture content is within a common range.
[0056] S34. Following the strict requirements prompted by the system, the large language model organizes the understanding and extraction results from the previous steps into a clean, structured data format without additional interpretation. The output strictly adheres to a predefined JSON structure, for example: {"Plasticity Index":{"Value":17,"Unit":"None"}, "Moisture Content":{"Value":22,"Unit":"%"}, "Confining Pressure":{"Value":100,"Unit":"kPa"}, "Deviatoric Stress":{"Value":50,"Unit":"kPa"}}. This JSON object contains only the extracted parameter names, values, and units, facilitating downstream programmatic parsing. Upon receiving this output, the system completes the conversion from unstructured user input to structured target subgrade parameters.
[0057] S35. Parse the JSON string returned by S34, extract the "value" field, and form a standard numerical vector [plasticity index value, moisture content value, confining pressure value, deviatoric stress value]. This vector is the target subgrade parameters required for calculation by the MCI model and XGBoost residual compensation model in subsequent steps. If the parsing finds that any required parameter values are missing, the system can trigger a feedback loop, such as calling the large language model again to generate a natural language follow-up question for the user, requesting the supplementation of specific parameters.
[0058] The semantic understanding capability of large language models refers to the ability of large deep neural network models, after pre-training on massive amounts of text, to understand the true meaning of human language words, phrases, sentences, and passages in specific contexts. This capability enables the model to go beyond simple keyword matching and perform complex tasks such as intent recognition, reference resolution, relation extraction, and common-sense reasoning. For example, in this task, this capability is manifested in recognizing "the water content of the soil sample is 25%" and "water content 25%" as expressing the same parameter, and understanding the correspondence between "confining pressure" and the symbol p in "under a confining pressure of 100 kPa". The large language model used is DeepSeek-V3.2. This is a specific example of a large language model with the aforementioned semantic understanding and logical reasoning capabilities. In the actual system implementation, by calling the application programming interface provided by DeepSeek-V3.2, the constructed prompt text is sent to the model, and the generated response text containing structured parameter information is received, thereby completing the parameter extraction task. Other large language models can also be selected according to the actual situation.
[0059] S4. Based on the target subgrade parameters, the trained XGBoost residual compensation model, and the MCI model, the final predicted subgrade resilient modulus value is generated, specifically: The target subgrade parameters are input into the MCI model to calculate the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters. The target subgrade parameters and the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters are then input into the trained XGBoost model to predict the residuals. The initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters is added to the predicted residuals to obtain the final predicted value of the subgrade resilient modulus. The specific implementation process is as follows: S40. Receive and confirm the extracted target subgrade parameter set. This set must contain all four parameters required for calculation: plasticity index, moisture content, confining pressure, and deviatoric stress. Before inputting it into the model, necessary standardization or unit unification is required to ensure that its numerical format is consistent with the data format used in the model training phase. For example, ensure that the units for confining pressure and deviatoric stress are both kPa, and the moisture content is a percentage value. Assign these parameters to their corresponding variables to prepare for subsequent calculations.
[0060] S41. Substitute the prepared target subgrade parameters into the fixed mathematical formula of the MCI model for calculation. First, calculate the consistency index based on the plasticity index and moisture content. Then, clarify the values of each variable in the formula: the confining pressure is denoted as... The deviatoric stress is denoted as In the MCI model formula, cyclic deviatoric stress With input deviatoric stress Equivalent. Standard atmosphere Use constant values (usually 101.3 kPa). Empirical coefficients for the model. , , , , The model uses fixed values determined and published by its original authors. The following formula is used for calculation: ,in, This represents the initial predicted value of the roadbed resilient modulus calculated by the MCI model for the current target roadbed parameters. This calculation process is deterministic, contains no randomness, and outputs a specific numerical value.
[0061] S42. The input to the trained XGBoost residual compensation model is a fixed-dimensional vector containing multiple features. This feature vector must be constructed strictly according to the format of the augmented feature set used during model training. Therefore, this feature vector needs to be generated based on the current target subgrade parameters and the obtained initial predicted values. The feature vector typically includes the following parts: the original target subgrade parameters (plasticity index, moisture content, confining pressure)... eccentric stress Key derived features, such as the ratio of deviatoric stress to confining pressure. ; and the initial predicted value of the roadbed resilient modulus generated by the MCI model. Arrange all features in a preset order to form a numerical array or vector, denoted as . .
[0062] S43. Convert the feature vectors constructed in S42. The input features are fed into the XGBoost residual compensation model, which has already been trained and hyperparameter tuned. Internally, this model processes and computes the input features through a series of gradient boosting decision trees. The model's forward propagation process traverses all base learners, ultimately outputting a continuous scalar value. This output value is the residual predicted by the model corresponding to the current input feature, denoted as . Its mathematical meaning is the prediction error of the XGBoost model compared to the MCI model. The estimated value, i.e. ,in This represents the trained XGBoost model function.
[0063] S44. Initial predicted value of roadbed resilient modulus With residual value Perform algebraic addition. Calculate the final predicted value. The formula is as follows: This addition operation fulfills the core idea of residual compensation. It provides benchmark predictions based on physical experience, while It provides error corrections for the benchmark under the current specific conditions. The sum of the two is... This is the final predicted value of the roadbed resilient modulus after optimization using the hybrid model. This value will serve as the core output of the entire framework and will be used for subsequent engineering analysis and report generation.
[0064] Optionally, the above technical solution also includes: S5. Based on the final predicted subgrade resilient modulus value and combined with domain knowledge acquired through retrieval-enhanced generation technology, generate a report containing the final predicted subgrade resilient modulus value and engineering recommendations. The specific implementation process is as follows: S50, after obtaining the final predicted value of the roadbed resilient modulus. Then, using this as a guide, one or more retrieval queries need to be automatically constructed. Query construction is not only based on... This value itself will also be combined with the original target subgrade parameters that generated the prediction, including plasticity index, moisture content, confining pressure, and deviatoric stress, as well as intermediate indicators that may be used in the calculation process, such as the consistency index. For example, the system might generate something like "subgrade resilient modulus". Construction control standards or confining pressure at MPa, plasticity index X, and moisture content Y%. The search intent is for "Design specifications for resilient modulus at Z kPa". These searches aim to find relevant design specifications, construction experience, similar cases, or precautions from the knowledge base.
[0065] S51. The system simultaneously sends the query constructed in S50 to multiple knowledge sources for retrieval. The first key knowledge source is a local vector database, which pre-stores and encodes integrated geotechnical engineering standards, such as the mechanics-empirical pavement design guideline, as well as a large number of historical experimental logs and engineering reports. The query text is converted into vectors, and a similarity search is performed in the vector database to find the text fragments most relevant to the current prediction scenario. The second knowledge source is an internet search engine interface, which the system synchronously calls to obtain the latest research results, technical announcements, or supplementary verification data for specific parameter ranges. All retrieved text fragments, regardless of their source, are accompanied by accurate citation information, including title, source, and paragraph.
[0066] S52. Combine the original input parameters with the calculated final predicted value of the roadbed resilient modulus. The retrieved text fragments, along with all cited search terms, are integrated into a structured contextual information package. This package is carefully organized and fed into a large language model, such as DeepSeek-V3.2. Simultaneously, detailed generation instructions are provided to the large language model, requiring it to write a report in the style of a professional engineer. The report must be based on the provided numerical results and deeply integrate the retrieved domain knowledge for analysis; any qualitative judgments or suggestions derived from the search text must be clearly marked with citation numbers in the report.
[0067] S53, the large language model, based on the rich context provided by S52, begins generating report text. The report follows a predefined structural template. The report first generates a project summary and a review of input parameters. Then, in the core analysis section, the model first lists the final predicted values for the roadbed resilient modulus. Then, the model cross-references retrieved domain knowledge, such as... The model compares the predicted values with the recommended ranges for similar soil types in the MEPDG guidelines to generate a reliability assessment. It analyzes whether the predicted values are higher, lower, or within the typical range compared to the general range, and cites relevant standard provisions to explain the potential engineering implications of such differences. Finally, the model integrates all information to generate specific, actionable engineering recommendations, such as "Based on the predicted resilient modulus being lower than the typical value, it is recommended to strengthen compaction control and ensure that the moisture content does not exceed Y%", ensuring that each key recommendation is linked to specific citations.
[0068] The draft report generated by the S54 large language model is in plain text format, containing chapter titles, lists, and citation marks. Upon receiving this text, the system backend invokes the document formatting engine. This engine converts the text into a more visually professional format, such as a PDF document or webpage. During formatting, the engine automatically converts citation marks into standard footnotes or endnotes, ensuring that the citation content accurately corresponds to the retrieved original text information. Finally, a complete engineering analysis report is generated, which users can directly download or view. The report not only presents the final predicted value of the subgrade resilient modulus but also provides in-depth, evidence-based analysis and construction guidance based on this value.
[0069] Among them, retrieval-enhanced generation technology is an artificial intelligence approach that combines information retrieval with text generation. Before generating text content, this technology first retrieves documents or information fragments relevant to the current task from an external knowledge base. Subsequently, this retrieved information is provided as context to a large language model to guide and enhance its generation process. This technology effectively improves the accuracy, timeliness, and factual reliability of the generated content because it builds the model's generation capabilities on validated external knowledge, rather than relying entirely on the model's own potentially incomplete or outdated internal parameterized knowledge.
[0070] Domain knowledge refers to the systematic information, rules, standards, best practices, and theoretical consensus formed through long-term practice and summarization within a specific professional field. In geotechnical engineering, domain knowledge specifically includes engineering classification standards for various soil types, design specifications for subgrade materials, construction quality control guidelines, typical mechanical parameter ranges, experience in treating common defects, and the latest academic research findings. This knowledge typically exists in the form of standard documents, academic papers, experimental reports, and engineering manuals, and serves as an important basis for professional judgment and decision-making.
[0071] The generated report, for example, includes the following content: 1) Project Overview: This analysis assesses the bearing capacity of a specific subgrade soil condition provided by the user. The core task is to predict the resilient modulus of the subgrade under these conditions and provide engineering insights based on domain knowledge.
[0072] 2) Input parameters: The target subgrade parameters obtained from user input include: ① Plasticity index: 17; ② Moisture content: 22%; ③ Confining pressure (…). ): 100kPa; ④ Deviatoric stress ( 50 kPa; 3) Model Calculation and Output: Based on the characteristics of the input parameters, the agent routes to the applicable MCI-XGBoost-RC hybrid prediction model, specifically including: ① MCI Model Baseline Prediction: Substituting the input parameters into the MCI model formula Perform calculations, where With input deviatoric stress Equivalently, the initial predicted value of the roadbed resilient modulus is obtained. ② Residual Compensation Prediction: The trained XGBoost residual compensation model is used to correct the initial prediction residuals, and the predicted residual values are obtained. ③ Final prediction result: Combining the above two steps, the final predicted value of the roadbed resilient modulus is obtained: .
[0073] 4) Performance Evaluation: The performance metrics of the mixed model used in this prediction on the independent test set include: ① Coefficient of determination R²: 0.96; ② Root mean square error RMSE: 3.1 MPa; ③ Mean absolute error MAE: 2.4 MPa; ④ The high R² value and low error value indicate that the model has high reliability in predicting this example.
[0074] 5) Comprehensive analysis, including: ① Result interpretation: The predicted resilient modulus of the subgrade is 89.8 MPa. Searching the local knowledge base (MEPDG guidelines and historical experimental data) using enhanced generative techniques reveals that for low-plasticity clay with similar plasticity indices and water content ranges, this predicted value is at a lower-middle level within the typical empirical range (usually 80-110 MPa). ② Reliability explanation: The current input parameters are all within the characteristic distribution range of the model training data, with no outliers observed. Therefore, this prediction is an interpolation prediction with high confidence. The model's high performance index (R²=0.96) further supports this judgment. ③ Engineering significance: This predicted value indicates that under the current confining pressure and deviatoric stress conditions, the subgrade soil possesses a moderately high elastic stiffness. However, since the water content (22%) is close to the upper limit of the optimum water content for this type of soil, attention should be paid to the potential decrease in modulus under actual dynamic loads.
[0075] 6) Recommended Measures: Based on the prediction results and domain knowledge, construction and design recommendations are proposed, including: ① Compaction Control: During construction, the degree of compaction should be strictly controlled to ensure that the moisture content after on-site compaction does not exceed 22% to guarantee the achievement of the predicted rebound modulus level. ② Drainage Enhancement: It is recommended to set up effective subgrade drainage facilities to prevent moisture intrusion during service, which could lead to an increase in moisture content and a decrease in modulus. ③ Verification Tests: For critical road sections, it is recommended to conduct portable falling weight deflectometer tests on-site to verify the consistency between the actual modulus and the predicted value. ④ Design Considerations: When designing the pavement structure layers, 89.8 MPa can be used as the input modulus value for the subgrade, and an appropriate modulus reduction factor should be considered to cope with long-term fatigue effects.
[0076] 7) Disclaimer: This report is automatically generated by an artificial intelligence system, and its content is based on input parameters and the currently integrated domain knowledge base. The report's conclusions and recommendations are for engineering professionals only and cannot replace the on-site judgment and detailed design of senior engineers. It is recommended to conduct necessary on-site verification and testing before making important engineering decisions.
[0077] Figure 2 The complete architecture and serialization workflow of the MCI-XGBoost-RC resilient modulus prediction hybrid model are described. The process begins with the extraction of MCI characteristics from the original MCI data. These input characteristics specifically include the consistency index. Confining pressure and equivalent to dynamic deviatoric stress deviatoric stress Then, the MCI fitting stage begins, where the MCI method is used to substitute the aforementioned features into the predetermined formula. The calculation yields a key MCI intermediate term, which is then output as the predicted MCI value of the roadbed resilient modulus, denoted as . Next, through calculation The difference between the measured value and the actual value in the laboratory is used to clearly define the residual target. Simultaneously, the ratio of deviatoric stress to confining pressure is also defined. This derived characteristic is related to the original multivariate characteristics—which encompass plasticity index (PI), moisture content, etc. Consistency Index Confining pressure eccentric stress Static confining pressure and dynamic deviatoric stress —and what I just received A systematic integration process is performed to construct an enhanced feature set with higher information dimensions. Following this, rigorous sample partitioning is conducted, dividing the entire dataset into a training / validation set (for model development) and a testing set (for final evaluation) using a fixed random seed of 42. During model training and optimization, the training / validation set is used to train various base models, including Support Vector Regression (SVR), XGBoost, AdaBoost, and Decision Tree Regression (DTR). Bayesian optimization is used to automatically search for and tune the hyperparameters of these models. During this process, five-fold cross-validation is used to repeatedly evaluate the generalization performance of the model under different parameter combinations, determining the optimal hyperparameters. Finally, the model with the optimal hyperparameters is applied to the testing set (which has not participated in any training or tuning) to predict the corresponding residual values. This predicted residual is then compared with the initial MCI prediction value. Algebraic addition is performed to complete the residual compensation step, thereby obtaining the final, more accurate predicted value of the roadbed resilient modulus from the hybrid model. The entire process culminates in a comprehensive model performance evaluation phase, using the coefficient of determination R², root mean square error (RMSE), mean absolute error (MAE), and A. 15 Various statistical indicators, such as indicators, are used to predict values. A quantitative comprehensive evaluation is conducted on the degree of agreement between the actual value and the true value.
[0078] Figure 3This paper demonstrates the collaborative workflow of an intelligent agent framework for predicting subgrade bearing capacity based on a large language model. The process begins with input reception, where the agent uses the DeepSeek-V3.2 large language model to parse user queries provided in natural language or structured data format, accurately extracting target subgrade parameters including plasticity index, water content, confining pressure, and deviatoric stress. Following this, the system concurrently initiates a knowledge enhancement module, accessing and enhancing the knowledge base (RAGKnowledge Base). The core of this knowledge base is a vector search database, embedding structured geotechnical engineering standards, numerous academic research papers, and detailed past experiment logs. Simultaneously, the agent invokes a web search tool, which performs three key tasks: validating the reasonableness and parameter range of the input parameters (RangeCheck), acquiring supplementary background knowledge for contextual supplementation of the current prediction scenario (ContextSupplement), and retrieving the latest state-of-the-art research results in the field (SOTA). To ensure the timeliness of recommendations, the process begins with knowledge retrieval. After retrieval, the process moves to the agent's decision-making and planning phase. Here, the agent acts as an intelligent model selector, performing logical reasoning based on extracted parameters and retrieved knowledge. For example, when input parameters such as plasticity index PI=6 and liquid limit LL=19.6 are present, the agent infers that the soil belongs to low-plasticity silt and matches it to a pre-trained subset, thus preferentially routing to the dedicated MCI-XGBoost-RC sub-model 2. Otherwise, it switches to the general hybrid model. Once a decision is made, the agent invokes the corresponding computational tool. This tool first runs the MCI method to perform MCI fitting on the input parameters, according to the formula... The predicted value of the roadbed resilient modulus (MCI) based on the physical benchmark is obtained. The tool then determines the residual target based on this. Subsequently, the XGBoost residual compensation module initiates the model training process. This process uses various base models such as Support Vector Regression (SVR), XGBoost, AdaBoost, and Decision Tree Regression (DTR) as candidates. It employs a Bayesian optimization strategy and five-fold cross-validation to evaluate and determine the optimal hyperparameters. Finally, it trains the optimal residual prediction model and performs residual compensation, adjusting the predicted residuals according to... According to and based on The final predicted value is obtained by adding the values together. In the final report generation phase, the agent, acting as a synthesizer, will, according to... Based on the numerical results, and by deeply integrating all cited domain knowledge (including normative clauses, case data, and the latest research conclusions) obtained from vector databases and web searches, a comprehensive engineering analysis report is generated that is structurally complete, includes quantitative predictions, reliability quantification assessments, and specific actionable construction recommendations with clear citations.
[0079] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation. The scheme after adjusting the order is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.
[0080] like Figure 4 As shown, an embodiment of the present invention provides a roadbed bearing capacity prediction system 200 based on a large language model, which includes a model building module 201, a model training module 202, a parameter extraction module 203, and a model application module 204. The model building module 201 is used to: build a hybrid model for predicting the resilient modulus of the subgrade. The hybrid model for predicting the resilient modulus of the subgrade integrates the MCI model and the XGBoost residual compensation model. The MCI model is used to calculate the initial predicted value of the resilient modulus of the subgrade based on the subgrade parameters. The XGBoost model is used to predict the residual of the initial predicted value of the resilient modulus of the subgrade. The model training module 202 is used to: train the XGBoost residual compensation model using the training set, and fine-tune the hyperparameters of the XGBoost residual compensation model using the Bayesian optimization method to obtain the trained XGBoost residual compensation model. The parameter extraction module 203 is used to: parse user input based on the semantic understanding capability of a large language model and extract target roadbed parameters; The model application module 204 is used to generate the final predicted value of the roadbed resilient modulus based on the target roadbed parameters, the trained XGBoost residual compensation model and MCI model.
[0081] Optionally, in the above technical solution, the model application module 204 is specifically used to: input the target subgrade parameters into the MCI model, calculate the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters, input the target subgrade parameters and the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters into the trained XGBoost model to predict the residual, and add the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters to the predicted residual to obtain the final predicted value of the subgrade resilient modulus.
[0082] Optionally, the above technical solution also includes a report generation module, which is used to generate a report containing the final predicted value of the roadbed resilient modulus and engineering recommendations based on the final predicted value of the roadbed resilient modulus and the domain knowledge obtained by the retrieval enhancement generation technology.
[0083] Optionally, the above technical solution further includes a training set acquisition module, which is used for: Multiple samples were obtained, each sample including the original subgrade parameters and the actual value of the subgrade resilient modulus corresponding to the original subgrade parameters; For each sample, the original subgrade parameters are input into the MCI model to obtain the initial predicted value of the subgrade resilient modulus. The residual between the initial predicted value of the subgrade resilient modulus and the actual value of the subgrade resilient modulus is calculated. The original subgrade parameters and the residual are combined into the enhanced features of the sample. The enhanced features of all samples constitute the training set. The original subgrade parameters include plasticity index, moisture content, confining pressure and deviatoric stress.
[0084] It should be noted that the beneficial effects of the roadbed bearing capacity prediction system 200 based on a large language model provided in the above embodiments are the same as those of the roadbed bearing capacity prediction method based on a large language model, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.
[0085] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned methods for predicting the bearing capacity of roadbeds based on a large language model.
[0086] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned methods for predicting the bearing capacity of roadbeds based on a large language model.
[0087] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
[0088] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for predicting the bearing capacity of roadbeds based on a large language model, characterized in that, include: A hybrid model for predicting the resilient modulus of the roadbed is constructed, which integrates the MCI model and the XGBoost residual compensation model. The MCI model is used to calculate the initial predicted value of the resilient modulus of the roadbed based on the roadbed parameters; the XGBoost model is used to predict the residual of the initial predicted value of the resilient modulus of the roadbed. The XGBoost residual compensation model is trained using the training set, and the hyperparameters of the XGBoost residual compensation model are tuned using the Bayesian optimization method to obtain the trained XGBoost residual compensation model. Based on the semantic understanding capabilities of a large language model, user input is parsed and target roadbed parameters are extracted; Based on the target subgrade parameters, the trained XGBoost residual compensation model, and the MCI model, the final predicted value of the subgrade resilient modulus is generated.
2. The method for predicting the bearing capacity of roadbed based on a large language model according to claim 1, characterized in that, Based on the target subgrade parameters, the trained XGBoost residual compensation model, and the MCI model, the final predicted subgrade resilient modulus value is generated, including: The target subgrade parameters are input into the MCI model to calculate the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters. The target subgrade parameters and the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters are then input into the trained XGBoost model to predict the residual. The initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters is added to the predicted residual to obtain the final predicted value of the subgrade resilient modulus.
3. The method for predicting the bearing capacity of roadbed based on a large language model according to claim 1 or 2, characterized in that, Also includes: Based on the final predicted value of the roadbed resilient modulus, and combined with the domain knowledge obtained by the retrieval enhancement generation technology, a report is generated that includes the final predicted value of the roadbed resilient modulus and engineering recommendations.
4. A method for predicting the bearing capacity of roadbed based on a large language model according to claim 1 or 2, characterized in that, The process of obtaining the training set includes: Multiple samples were obtained, each sample including the original subgrade parameters and the actual value of the subgrade resilient modulus corresponding to the original subgrade parameters; For each sample, the original subgrade parameters are input into the MCI model to obtain the initial predicted value of the subgrade resilient modulus. The residual between the initial predicted value of the subgrade resilient modulus and the actual value of the subgrade resilient modulus is calculated. The original subgrade parameters and the residual are combined to form the enhanced features of the sample. The enhanced features of all samples constitute the training set. The original subgrade parameters include plasticity index, moisture content, confining pressure, and deviatoric stress.
5. A roadbed bearing capacity prediction system based on a large language model, characterized in that, It includes a model building module, a model training module, a parameter extraction module, and a model application module; The model building module is used to: construct a hybrid model for predicting the resilient modulus of the subgrade, which integrates the MCI model and the XGBoost residual compensation model; the MCI model is used to calculate the initial predicted value of the resilient modulus of the subgrade based on the subgrade parameters; and the XGBoost model is used to predict the residual of the initial predicted value of the resilient modulus of the subgrade. The model training module is used to: train the XGBoost residual compensation model using the training set, and fine-tune the hyperparameters of the XGBoost residual compensation model using the Bayesian optimization method to obtain the trained XGBoost residual compensation model. The parameter extraction module is used to: parse user input based on the semantic understanding capability of a large language model and extract target roadbed parameters; The model application module is used to generate the final predicted value of the roadbed resilient modulus based on the target roadbed parameters, the trained XGBoost residual compensation model, and the MCI model.
6. The roadbed bearing capacity prediction system based on a large language model according to claim 5, characterized in that, The model application module is specifically used to: input the target subgrade parameters into the MCI model, calculate the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters, input the target subgrade parameters and the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters into the trained XGBoost model to predict the residual, and add the initial predicted value of the subgrade resilient modulus corresponding to the target subgrade parameters and the predicted residual to obtain the final predicted value of the subgrade resilient modulus.
7. A roadbed bearing capacity prediction system based on a large language model according to claim 5 or 6, characterized in that, It also includes a report generation module, which is used to generate a report containing the final predicted subgrade resilient modulus value and engineering recommendations based on the final predicted subgrade resilient modulus value and the domain knowledge obtained by the retrieval enhancement generation technology.
8. A roadbed bearing capacity prediction system based on a large language model according to claim 5 or 6, characterized in that, It also includes a training set acquisition module, which is used for: Multiple samples were obtained, each sample including the original subgrade parameters and the actual value of the subgrade resilient modulus corresponding to the original subgrade parameters; For each sample, the original subgrade parameters are input into the MCI model to obtain the initial predicted value of the subgrade resilient modulus. The residual between the initial predicted value of the subgrade resilient modulus and the actual value of the subgrade resilient modulus is calculated. The original subgrade parameters and the residual are combined to form the enhanced features of the sample. The enhanced features of all samples constitute the training set. The original subgrade parameters include plasticity index, moisture content, confining pressure, and deviatoric stress.
9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for predicting the bearing capacity of a roadbed based on a large language model as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for predicting the bearing capacity of a roadbed based on a large language model as described in any one of claims 1 to 4.