A machine learning-based method and system for predicting high-temperature viscosity of environmental sediments

By constructing a high-temperature viscosity prediction model for environmental sediments using machine learning methods, the problem of inaccurate prediction in existing technologies has been solved, achieving rapid and accurate high-temperature viscosity prediction. This provides data support for aviation safety assessment and explores the underlying mechanism.

CN120745467BActive Publication Date: 2025-12-12TIANMUSHAN LABORATORY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511263991.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-12
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing prediction models for high-temperature viscosity of environmental sediments lack systematic experimental data support, and the intrinsic relationship between characteristic parameters such as composition, particle size, crystal content, and moisture content and high-temperature viscosity is unclear, resulting in inaccurate predictions.

Method used

Machine learning methods were employed to collect ambient temperature characteristic parameters and high temperature viscosity data from environmental sediment samples. Data preprocessing, feature selection, and standardization were performed to construct a feature dataset suitable for machine learning models. Multiple prediction models were then trained, and the best-performing prediction model was finally selected for high temperature viscosity prediction.

Benefits of technology

This technology enables rapid and accurate prediction of the high-temperature viscosity of environmental sediments, providing crucial data support for aviation safety assessments and opening up new avenues for exploring the intrinsic mechanisms of high-temperature viscosity changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745467B_ABST
    Figure CN120745467B_ABST
Patent Text Reader

Abstract

The application provides a kind of environment deposit high temperature viscosity prediction method and system based on machine learning, the method includes determining the normal temperature characteristic parameter of each environment deposit sample and the high temperature viscosity value under different temperatures, constitutes model training original data set;The original data set is preprocessed and feature screening is carried out, and the physical quantity feature with high correlation with high temperature viscosity is obtained to construct final feature data set;Final feature data set is divided into training set, validation set and test set, and a plurality of preset machine learning models are trained in parallel based on the training set;Based on the validation set and according to the preset performance evaluation index, the optimal prediction model is selected from the performance evaluation index;The normal temperature characteristic data of the environment deposit to be measured is input into the optimal prediction model, the high temperature viscosity value of the deposit at the target temperature is predicted, and the generalization performance of different optimal prediction models is evaluated based on the test set;Realize the rapid, accurate and efficient prediction of the high temperature viscosity of environment deposit.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent prediction of material performance, and particularly relates to an environment deposit high-temperature viscosity prediction method and system based on machine learning. BACKGROUND

[0002] When the environment deposits such as sand, dust and volcanic ash enter the engine, they will quickly reach their melting point and begin to melt. The melted environment deposits wet the surface of the ceramic layer and diffuse along the grain boundaries, pores and micro-cracks inside the ceramic layer through capillary action, gradually penetrating into the metal bonding layer. After the engine stops, the surface temperature of the ceramic layer drops sharply, and the environment deposits cool down and shrink inside. Due to the significant difference in thermal expansion coefficient between the environment deposits and the ceramic layer, the cooled environment deposits generate stress inside the ceramic layer. With the continuous generation and release of stress, the thermal barrier coating will crack in various forms, eventually leading to peeling and failure.

[0003] The high-temperature viscosity of environment deposits is a key factor in determining their high-temperature erosion behavior, significantly affecting their interaction with the ceramic layer such as wetting, spreading and penetration. The high-temperature viscosity of environment deposits is significantly affected by their composition, particle size, crystal content and moisture content, etc. Thus, research on the construction of high-temperature viscosity prediction models for environment deposits has emerged. However, current high-temperature viscosity prediction models for environment deposits mainly focus on composition regulation, microstructure optimization and surface modification, and the understanding of the erosion mechanism of environment deposits is not comprehensive enough, especially lacking systematic experimental data support. In addition, the internal relationship between the characteristic parameters of environment deposits such as composition, particle size, crystal content and moisture content and their high-temperature viscosity is not clear. SUMMARY

[0004] Based on the above background, the purpose of the present application is to provide an environment deposit high-temperature viscosity prediction method and system based on machine learning to quickly and accurately predict the viscosity of environment deposits at different temperatures.

[0005] To achieve the above purpose, the present application adopts the following technical solutions:

[0006] In a first aspect, the present application provides an environment deposit high-temperature viscosity prediction method based on machine learning, comprising,

[0007] Collecting a plurality of environment deposit samples and measuring the characteristic parameters of each environment deposit sample at room temperature and the corresponding high-temperature viscosity values at different temperatures to form a model training original data set;

[0008] The model training original data set is preprocessed, including missing value processing and abnormal value processing, and feature screening is performed based on actual physical meaning and correlation between features, physical quantity features having high correlation with high temperature viscosity are obtained and standardized processing is performed, and a final feature data set suitable for machine learning model training is constructed;

[0009] The final feature data set is divided into a training set, a validation set and a test set, and a plurality of preset machine learning models are trained based on the training set; the comprehensive performance of each candidate model after training is evaluated based on the validation set and according to a preset performance evaluation index, and the optimal prediction model is selected from the candidate models;

[0010] The normal temperature feature data of the environment sediment to be measured is input into the selected optimal prediction model, the high temperature viscosity prediction value of the environment sediment at the target temperature is output, and the generalization performance of different optimal prediction models is evaluated based on the test set.

[0011] In a second aspect, the present application provides an environment sediment high temperature viscosity prediction system based on machine learning, comprising,

[0012] The original data set construction module is used for collecting a plurality of environment sediment samples, and measuring the characteristic parameters of each environment sediment sample at normal temperature and the corresponding high temperature viscosity values at different temperatures, to form a model training original data set;

[0013] The feature screening module is used for preprocessing the model training original data set, performing feature screening based on actual physical meaning and correlation between features, obtaining physical quantity features having high correlation with high temperature viscosity and performing standardized processing, and constructing a final feature data set suitable for machine learning model training;

[0014] The prediction model construction module is used for dividing the final feature data set into a training set, a validation set and a test set, training a plurality of preset machine learning models based on the training set; the comprehensive performance of each candidate model after training is evaluated based on the validation set and according to a preset performance evaluation index, and the optimal prediction model is selected from the candidate models;

[0015] The prediction output module is used for inputting the normal temperature feature data of the environment sediment to be measured into the selected optimal prediction model, outputting the high temperature viscosity prediction value of the environment sediment at the target temperature, and evaluating the generalization performance of different optimal prediction models based on the test set.

[0016] The present application has the following beneficial effects:

[0017] The method and system for predicting high-temperature viscosity of environmental deposits based on machine learning provided by the application utilize machine learning technology, are based on normal-temperature characteristics and high-temperature viscosity data of environmental deposits, perform data analysis and feature engineering on the normal-temperature characteristic data, utilize various types of machine learning models to construct a prediction model and perform optimization to obtain the prediction model, and innovatively realize rapid, accurate and efficient prediction of the high-temperature viscosity of environmental deposits. The method has simplicity and interpretability, not only provides key data and method support for practical applications such as aviation safety evaluation, but also opens up a new way for exploring the internal mechanism of the change of the high-temperature viscosity of environmental deposits.

[0018] In addition, the method not only takes the content of major elements and temperature as features to participate in the training of the prediction model, but also takes the content of trace elements, particle size characteristics, moisture content and crystal content into account to affect the viscosity, comprehensively considers the influencing factors of the high-temperature viscosity of environmental deposits, and realizes rapid and accurate prediction of the viscosity of environmental deposits at different temperatures. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0020] Figure 1 A flowchart of a method for predicting high-temperature viscosity of environmental deposits based on machine learning provided by an embodiment of the application is shown in the figure.

[0021] Figure 2 A schematic diagram of data collection content in step S1 is shown in the figure.

[0022] Figure 3 A schematic diagram of a method and process for data preprocessing and feature selection in step S2 is shown in the figure.

[0023] Figure 4 A schematic diagram of using a box plot to process outliers in step S2 is shown in the figure.

[0024] Figure 5 A schematic diagram of the result of normal distribution test using the Shapiro-Wilk method in step S2 is shown in the figure.

[0025] Figure 6 A correlation coefficient heat map between particle size characteristics in the feature engineering process in step S2 is shown in the figure.

[0026] Figure 7 A schematic diagram of performance data of the eight optimal prediction models in step S4 on the test set is shown in the figure.

[0027] Figure 8 A schematic diagram of a machine learning-based high-temperature viscosity prediction system for environmental deposits is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to further understand the present application, the preferred embodiments of the present application are described below in conjunction with the embodiments, but it should be understood that these descriptions are only for further illustrating the features and advantages of the present application, and are not limitations on the claims of the present application.

[0029] Embodiment 1

[0030] Reference Figure 1 The present embodiment provides a machine learning-based high-temperature viscosity prediction method for environmental deposits, comprising the following steps:

[0031] Step S1: Collecting a plurality of environmental deposit samples, and measuring the characteristic parameters of each sample at room temperature and the high-temperature viscosity values at different temperatures to form a model training original data set;

[0032] Step S2: Preprocessing the model training original data set, and performing feature screening based on the correlation between actual physical significance and characteristics to obtain physical quantity characteristics with high correlation with high-temperature viscosity and perform standardization processing to build a final feature data set suitable for machine learning model training;

[0033] Step S3: Dividing the final feature data set into a training set, a validation set and a test set, training a plurality of preset machine learning models based on the training set; based on the validation set and according to a preset performance evaluation index, performing comprehensive performance evaluation on each trained model to select the performance-optimal prediction model therefrom;

[0034] Step S4: Inputting the room temperature characteristic data of the environmental deposit to be measured into the selected optimal prediction model, outputting the high-temperature viscosity prediction value of the deposit at the target temperature; and evaluating the generalization performance of different models based on the test set.

[0035] Specifically,

[0036] The collection of room temperature characteristic data and high-temperature viscosity of environmental deposits in step S1 to form the data set for model training includes,

[0037] S101: Collecting environmental deposits in different regions, including different types such as sand, dust and volcanic ash.

[0038] S102: Measuring different room temperature characteristics and high-temperature viscosity values of environmental deposits, referring to Figure 2The normal temperature characteristics of the environmental sediment include trace element content, moisture content, crystal content, particle size characteristics and major element content, and measuring different normal temperature characteristics specifically includes:

[0039] Trace element content: The trace element content of the environmental sediment sample is tested by using a sample dissolution method. The specific process includes: before testing on the machine, the sample needs to be pretreated, first ground into powder and dried, the powder sample is weighed and placed in a sample dissolution bomb, HNO3 and HF are slowly added in turn, and the sample dissolution bomb is evaporated for multiple times. Finally, the solution is transferred into a polyethylene plastic bottle and diluted with HNO3. After a stable solution is formed, it is sent into an inductively coupled plasma mass spectrometer for analysis to obtain the content of each trace element of the sample.

[0040] Moisture content: The moisture content of the environmental sediment sample is measured by using a thermal gravimetric analysis method, and the specific process includes: using a thermal analyzer, when the sample is heated to (105±2) ℃, the water is completely evaporated, and the mass loss at this temperature is taken as the water mass in the sample.

[0041] Crystal content: The crystal content of the environmental sediment sample is measured by using an X-ray diffractometer, and the specific process includes: taking the sample powder, grinding it in a grinder, and sieving the ground powder. Then the powder sample is placed in the X-ray diffractometer for testing to obtain the crystal content of the environmental sediment sample.

[0042] Particle size characteristics: The particle size characteristics of the environmental sediment sample are measured by using a laser particle size analyzer and a particle size test special software. The sampling process includes taking a "laboratory sample" from a "large amount of powder", and taking a "test sample" from the "laboratory sample". After sampling is completed, the test sample is placed in the laser particle size analyzer for testing, and the particle size characteristics of the environmental sediment sample are obtained.

[0043] Major element content: The major element content of the environmental sediment sample is measured by using an X-ray fluorescence spectrometer, and the experimental process mainly includes loss on ignition determination, glass sheet preparation and on-machine testing.

[0044] The measurement results are:

[0045] The normal temperature characteristics of the environmental sediment include major element content, trace element content, crystal content, moisture content and particle size characteristics, specifically including:

[0046] The major element content includes Ca, Mg, Al, Si, Na, K, P, Ti, Fe and Mn ten element content data.

[0047] The trace element content includes Li, Be, Sc, V, Cr, Co and forty-three element content data.

[0048] The crystal content includes two data of a crystal percentage content and an amorphous percentage content.

[0049] The moisture content includes one data of a moisture percentage content.

[0050] The particle size characteristics include seven characteristic data of D10, D50, D90, a number average particle size, a length average particle size, an area average particle size, and a volume average particle size.

[0051] The high-temperature viscosity of the environmental sediment sample at different temperatures is measured by using a rotary viscometer. Specifically, a standard volume of the CMAS sample is prepared. Before testing, the sample is placed in a crucible, and the parameters of the viscometer are checked and set in sequence, including the heating rate and holding time at different temperature stages and the cooling rate. After the measurement is completed, the data is exported, including the time, temperature, and corresponding high-temperature viscosity at different temperatures.

[0052] S103: The five normal-temperature characteristics, temperature, and corresponding high-temperature viscosity of the environmental sediment obtained above are collectively constructed into a model training original data set, wherein the normal-temperature characteristics and temperature are used as characteristics for training, and the high-temperature viscosity is used as the prediction target of the model.

[0053] Referring to Figure 3 , the model training original data set is preprocessed in step S2, the features are selected based on the actual physical meaning and the correlation between the features, the physical quantity features with high correlation with the high-temperature viscosity are obtained and standardized, and a final feature data set suitable for training of a machine learning model is constructed, including,

[0054] S201: To construct a high-quality training data set suitable for subsequent machine learning models, the present application proposes a targeted data preprocessing strategy. The strategy includes missing value processing and outlier processing; specifically as follows:

[0055] (1) The outliers existing in the data set are identified and removed to improve the accuracy and reliability of the data. Specifically, it includes:

[0056] Referring to Figure 4 , the present application adopts a box plot (Box-Plot) to perform distribution visualization analysis on the feature data. In the box plot, the data distribution of each feature is clearly plotted, and the data boundaries are determined based on statistical principles. Specifically, the data points outside the upper and lower limit ranges in the graph are defined as outliers, the upper limit calculation formula is: Q1-1.5×IQR, and the lower limit calculation formula is: Q3+1.5×IQR. Wherein, the calculation formula of the interquartile range (IQR) is: IQR=Q3-Q1, Q1 represents the 25% quantile, and Q3 represents the 75% quantile. Taking Figure 4 as an example, the red points clearly identify the abnormal data points in the crystal content data.

[0057] The box plot is used for visual analysis of feature data, and abnormal data points exceeding the preset boundary range are accurately identified and located; and sample data containing the abnormal values are systematically removed, so that the data cleaning process is efficiently completed.

[0058] (2) The missing values in the data set are filled to ensure the integrity of the data set. Specifically, it includes:

[0059] For the missing values in the collected data, the application adopts an intelligent filling method based on a random forest model for processing. The core of the method is to predict the missing values by using the internal correlation of the data, and specifically includes the following steps:

[0060] Step A: Data set division. Traverse the data set, and divide the samples into two subsets according to whether each feature has missing values: the first subset (training set): the subset contains complete samples without missing values on the target feature. This subset is used to construct and train the missing value prediction model. The second subset (to be predicted set): the subset contains samples with missing values on the target feature. This subset will be the prediction object of the model.

[0061] Step B: Model training. The first subset (training set) is used to train the random forest model. During the training process, the model can effectively capture the complex nonlinear relationship between features by learning the voting results of multiple decision trees, thereby constructing a high-precision missing value prediction model.

[0062] Step C: Missing value prediction and filling. The trained random forest model is applied to the second subset (to be predicted set). For each sample in the subset, the model predicts the missing value of the target feature according to its non-missing feature values. Finally, the model-predicted values are filled into the corresponding positions in the original data set, thereby completing the repair of the missing values.

[0063] Further, the random forest is an ensemble learning algorithm based on decision trees, that is, multiple decision trees are combined together, and each time the data set is randomly selected with replacement, and part of the features are randomly selected as input, so as to combine the random forest, and the specific steps are as follows:

[0064] The training set is sampled with replacement from the training set, and a new training subset is formed by sampling a fixed number of times each time; a few features are randomly selected from all features; a complete decision tree is learned using the new training set and the randomly selected features; the above steps are repeated to construct a large number of decision trees to form a random forest.

[0065] The decision tree shows the model of decision rules and classification results in a tree data structure, converts the known data that seems to be disordered and chaotic into a tree model that can predict unknown data, each path from the root node to the leaf node represents a rule of decision, and the information gain and gain ratio are used to find the optimal partition attribute, which can be expressed as:

[0066] Information gain: ;

[0067] ;

[0068] ;

[0069] wherein, is the proportion of different feature attributes, is the total sample set, is the sample set after partition, and v is the total number of subsets after partition, and y corresponds to the data label.

[0070] Gain ratio: ;

[0071] wherein v is the total number of subsets after partition.

[0072] Further, in order to improve the generalization performance of the model and prevent overfitting, the pruning optimization technique is used when the decision tree model is constructed. By evaluating and removing redundant branches in the decision tree that contribute weakly to the prediction accuracy, the structure of the tree is optimized. Specifically, based on the preset cost complexity standard, the non-leaf nodes are pruned from bottom to top. This optimization process ensures that the model captures the key laws of data while avoiding overfitting to the training data.

[0073] S202: Based on the actual physical meaning and the correlation between the characteristics, the physical quantity characteristics of the environmental sediments are obtained, including,

[0074] The particle size characteristics of the environmental sediments include seven characteristic data. The correlation between the characteristics and the correlation between the characteristics and the viscosity are evaluated. The characteristics with high correlation coefficient between the characteristics are redundant characteristics. The corresponding characteristics are selected by the correlation coefficient value of the redundant characteristics and the high-temperature viscosity of the environmental sediments, the characteristics with low correlation are removed, and the characteristics with high correlation with the high-temperature viscosity are retained. Specifically, the characteristic screening method includes the following steps:

[0075] S2021: Common correlation evaluation methods include Pearson correlation coefficient and Spearman correlation coefficient, but Pearson correlation coefficient usually requires data distribution to conform to normal distribution, while Spearman coefficient does not assume data priori. In order to determine the correlation evaluation method of the present application, the Shapiro-Wilk method is used to perform normal distribution test on the input features, and the Shapiro-Wilk test is a method for testing whether the data comes from a normal distribution. It can be expressed as:

[0076]

[0077] In the formula: represents the value of each sample; is the weight corresponding to each sample obtained from the normal distribution expected value table; is the average value of the data set.

[0078] The Shapiro-Wilk test gives a test statistic and a corresponding p-value as a judgment of whether the feature data conforms to a normal distribution. The p-value represents the probability of observing an extreme situation under the assumption that the null hypothesis is true. If the p-value is less than the pre-set significance level, which is 0.05 in this embodiment, it is considered that the data does not conform to the normal distribution; otherwise, if the p-value is greater than the significance level, it is considered that the data may conform to the normal distribution. Referring to Figure 5 , more than 50% of the features do not conform to the normal distribution, and the Pearson coefficient is not suitable for quantifying the correlation of the data in the present application, so the Spearman coefficient is used to compare the correlation between the data in the subsequent steps of the present application.

[0079] S2022: Use spearman to analyze the correlation between the granularity features, which can be expressed as:

[0080]

[0081] In the formula: and represent two variables to be analyzed for correlation; and represent the average values of the two variables in the data set. The value reflects the correlation between the features, and the higher the correlation between the features, the more redundant it is, and needs to be screened. By comparing the correlation between the features and the high-temperature viscosity of the environmental sediment, the features with low correlation are removed, and the features with high correlation with the high-temperature viscosity are retained.

[0082] Specifically, the spearman correlation coefficient between the particle size characteristics is calculated, and when the absolute value of the correlation coefficient between the characteristics is greater than 0.6, it is determined that there is strong correlation between the two characteristics, and there is feature redundancy. See Figure 6 There is a strong correlation between D50, D90 and volume average particle size, and there is a strong correlation between D10, area average particle size, number average particle size and length average particle size. The highest correlation between D50, D90 and volume average particle size and high temperature viscosity is D50, and the highest correlation between D10, area average particle size, number average particle size and length average particle size and high temperature viscosity is D10, so D10 and D50 are selected as the physical characteristic quantity participating in training among the particle size characteristics.

[0083] S2023: Based on the actual physical meaning, in an embodiment of the present application, feature screening is performed between different descriptions of the same normal temperature characteristics. The sample crystal content is described by “relative crystal content” and “relative amorphous content”. Since there is a clear linear relationship between the two descriptors, this embodiment selects “relative crystal content” as the model training feature. The water content of the sample is only represented by “moisture content”, so feature screening is not required. At the same time, considering the important physical meaning of the high temperature viscosity of the major elements and trace elements in the environmental sediments, the major elements and trace elements are not subjected to feature screening. Finally, D10, D50, relative crystal content, moisture content, major elements and trace elements are selected as the model input features.

[0084] S203: Standardize the selected physical quantity characteristics and the high temperature viscosity data of the environmental sediments;

[0085] Standardization is a key preprocessing step for machine learning modeling, which can eliminate the dimensional differences between different characteristics, avoid the dominance of features with large numerical ranges, enhance the regularization constraint effect on weights, and meet the implicit assumption of linear model on feature comparability, thereby comprehensively improving the performance, stability and interpretability of the model. The physical quantity characteristics and the high temperature viscosity of the environmental sediments are standardized, including,

[0086] The selected physical quantity characteristics and the high temperature viscosity data of the environmental sediments are standardized using the Z-Score method, which can be represented as:

[0087]

[0088] wherein is the mean of the selected feature data, is the variance of the selected feature data, is the selected physical quantity characteristic data;

[0089] The feature dataset after Z-Score standardization is in accordance with a normal distribution with a mean of 0 and a variance of 1.

[0090] In step S3, the normalized final feature dataset is randomly divided into a training set, a validation set and a test set according to a preset proportion. Based on the training set, a plurality of preset machine learning models are trained. Based on the validation set and according to a preset performance evaluation index, the comprehensive performance of each trained model is evaluated, and the performance-optimal prediction model is selected from the trained models, including,

[0091] S301: Selecting a plurality of machine learning models for training, including Support Vector Machine (SVM), Multilayer Perceptron (MLP), K-Nearest Neighbors (KNN), AdaBoost, Bagging, Random Forest (RF), Gradient Boosting Decision Tree (GBDT) and Elastic Net.

[0092] S302: Hyperparameter optimization of the constructed plurality of machine learning models. Including,

[0093] Step A: Using grid search method to search for optimal parameters, predefining search range and step size of hyperparameters for each candidate model, including but not limited to maximum depth of random forest, number of hidden layer neurons and kernel function of support vector machine, etc.

[0094] Step B: Constructing a combined grid of hyperparameters, and performing model training and evaluation for each combination;

[0095] Step C: For each hyperparameter combination, using k-fold cross-validation method to evaluate model performance and calculate mean square error and determination coefficient on the validation set. By comparing the performance indicators under different hyperparameter combinations, the best-performing hyperparameter combination is finally selected. Specifically,

[0096] The hyperparameters are selected using 5-fold cross-validation, and the training set and validation set are selected according to the preset ratio of 4:1, that is, 4 groups of training set data and 1 group of validation set data are selected, and the model is trained by the training set data. The mean square error and the coefficient of determination are calculated on the samples reserved in the fold. This process is repeated 5 times, each time using a different fold as the validation set, which will eventually give 5 estimates of the validation set error. By calculating the average of the 5 validation set errors, the validation set error and the coefficient of determination of the model are obtained as the model performance criterion, so as to select the hyperparameters. Update the parameters of each model to obtain the best hyperparameters.

[0097] S303: Compare the best performance indicators obtained by all candidate models after hyperparameter optimization, and select the candidate model with the best comprehensive performance, that is, the strongest generalization ability, as the final prediction model according to the preset model selection standard, for performing the prediction task of the high-temperature viscosity of the environmental sediment. In this embodiment, the preset model selection standard is the lowest average validation set mean square error and the highest average validation set coefficient of determination in the 5-fold cross-validation.

[0098] Step S4: Input the normal temperature characteristic data of the environmental sediment to be tested into the selected optimal prediction model, output the high-temperature viscosity prediction value of the environmental sediment at the target temperature, and evaluate the generalization performance of the model based on the test set. Referring to Figure 7 , the mean square error and the coefficient of determination of each model are calculated on the test set to evaluate the generalization performance of the model.

[0099] Table 1 Mean square error and coefficient of determination of different models on the test set

[0100] Mean Squared Error Coefficient of Determination SVM 0.0198 0.9821 RF 0.0580 0.9474 MLP 0.0217 0.9804 KNN 31.0730 0.8903 AdaBoost 0.1215 0.8899 GBDT 0.0202 0.9817 Bagging 0.0607 0.9450 Elastic Net 0.0653 0.9408

[0101] From Table 1, it can be seen that the R² of the other six models is above 0.90 except for the AdaBoost and KNN models, and the coefficient of determination values of the SVM, MLP, and GBDT models are above 0.98, among which the R² value of the SVM is the largest, reaching 0.9821, and the coefficient of determination values of the other three models are also above 0.94. As for the mean square error, it generally follows the rule that the larger the coefficient of determination value, the smaller the corresponding MSE value, among which the smallest MSE value is 0.0198 for the SVM model, and the largest is 31.0730 for the KNN model, which is much larger than other models. It is worth noting that although the R² values of AdaBoost and KNN are similar, the mean square error values differ by hundreds of times. The SVM, MLP, and GBDT models have excellent generalization performance in the viscosity prediction of environmental sediment, and can accurately predict the viscosity of environmental sediment with different components at different temperatures.

[0102] Example 2

[0103] Referring to Figure 8 The embodiment provides a machine learning-based environment sediment high-temperature viscosity prediction system, comprising,

[0104] An original data set construction module is configured to collect a plurality of environment sediment samples, and measure characteristic parameters of each environment sediment sample at normal temperature and corresponding high-temperature viscosity values at different temperatures, to form a model training original data set;

[0105] A feature screening module is configured to pre-process the model training original data set, screen features based on actual physical meanings and correlations between features, obtain physical quantity features having high correlation with high-temperature viscosity, and perform standardization processing, to construct a final feature data set suitable for machine learning model training;

[0106] A prediction model construction module is configured to divide the final feature data set into a training set and a validation set, train a plurality of preset machine learning models by using the training set, perform comprehensive performance evaluation on each candidate model after training based on the validation set according to a preset performance evaluation index, and screen out a performance-optimal prediction model from the candidate models;

[0107] A prediction output module is configured to input normal-temperature characteristic data of an environment sediment to be measured into the screened optimal prediction model, and output a high-temperature viscosity prediction value of the sediment at a target temperature.

[0108] The above description of the embodiments is only used to help understand the method and its core idea of the present application. It should be noted that, for those skilled in the art, without departing from the principle of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0109] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing specific embodiments only and is not intended to be limiting. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0110] It should be noted that some of the example embodiments are described as processes that are depicted as flow diagrams. Although the flow diagrams describe the processes as a sequential process, many of the steps can be performed in parallel, concurrently or simultaneously. In addition, the order of the steps can be re-arranged. A process can be terminated when its operations are completed, but could also occur over longer periods of time. Processes might or might not have additional steps not shown in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc.

Claims

1. A method for predicting the high-temperature viscosity of environmental sediments based on machine learning, characterized in that, include, Multiple environmental sediment samples were collected, and the characteristic parameters of each environmental sediment sample at room temperature and the corresponding high-temperature viscosity values ​​at different temperatures were measured to form the original dataset for model training. The room temperature characteristics of the environmental sediments include trace element content, moisture content, crystal content, grain size characteristics, and major element content. The original dataset for model training is preprocessed, including handling missing values ​​and outliers. Feature selection is performed based on actual physical meaning and correlation between features to obtain physical quantity features that are highly correlated with high-temperature viscosity. These features are then standardized to construct the final feature dataset suitable for machine learning model training. The final feature dataset is divided into a training set, a validation set, and a test set. Multiple preset machine learning models are trained based on the training set. Based on the validation set and according to the preset performance evaluation indicators, a comprehensive performance evaluation is performed on each candidate model after training, and the prediction model with the best performance is selected. The ambient temperature characteristic data of the sediment in the test environment are input into the selected optimal prediction model, and the high temperature viscosity prediction value of the sediment at the target temperature is output. The generalization performance of different optimal prediction models is evaluated based on the test set.

2. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 1, characterized in that, The ambient temperature characteristics of the environmental sediments include trace element content, moisture content, crystal content, grain size characteristics, and major element content, and the corresponding data measurements include, The trace element content of environmental sediment samples was tested using the dissolution method. Thermogravimetric analysis was used to measure the moisture content of environmental sediment samples using a thermal analyzer. The crystal content of environmental sediment samples was determined using X-ray diffraction. The particle size characteristics of environmental sediment samples were measured using a laser particle size analyzer and dedicated particle size testing software. The major elemental composition of environmental sediment samples was determined using X-ray fluorescence spectrometry.

3. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 2, characterized in that, The content of major elements includes the content data of Ca, Mg, Al, Si, Na, K, P, Ti, Fe, and Mn. The trace element content includes data on the content of Li, Be, Sc, V, Cr, and Co elements; The crystal content includes data on the percentage content of crystals and the percentage content of amorphous materials; The moisture content includes percentage moisture content data; The particle size characteristics include D10, D50, D90, number-average particle size, length-average particle size, area-average particle size, and volume-average particle size characteristic data.

4. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 1, characterized in that, The missing value processing is based on the random forest model for missing value prediction and imputation; the outlier processing is based on the use of box plots to visualize and analyze the feature data, identify and locate outlier data points that exceed the preset boundary range.

5. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 4, characterized in that, The missing value prediction and imputation based on the random forest model includes, The dataset is traversed, and the samples are divided into a training set and a prediction set according to whether there are missing values ​​in the features. The training set contains complete samples without missing values ​​in the target features and is used to build and train the missing value prediction model. The prediction set contains samples with missing values ​​in the target features and is used as the prediction objects of the model. The random forest model is trained using the training set. During the training process, the model learns the voting results of multiple decision trees, captures the non-linear relationship between features, and constructs a missing value prediction model.

6. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 1, characterized in that, The feature selection based on the correlation between actual physical meaning and characteristics yields physical quantity features highly correlated with high-temperature viscosity, including: The input features are subjected to a normal distribution test, and then the correlation coefficient is calculated using Spearman. Redundant features are identified based on the correlation between features, and the redundant features are screened by combining the correlation coefficient between the features and the high temperature viscosity of environmental sediments and the physical meaning represented by the features, so as to obtain the physical feature quantities of environmental sediments.

7. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 1, characterized in that, The various preset machine learning models include Support Vector Machine, Multilayer Perceptron, K-Nearest Neighbors, AdaBoost, Bagging, Random Forest, Gradient Boosting Decision Tree, and / or Elastic Net.

8. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 1, characterized in that, The training of multiple preset machine learning models based on the training set also includes, The optimal parameters are searched using a grid search method. The search range and step size of the hyperparameters are predefined for each candidate model. The hyperparameters include, but are not limited to, the maximum depth of the random forest, the number of neurons in the hidden layer, and the kernel function of the support vector machine. Construct a combined grid of hyperparameters, and train and evaluate the model for each combination; For each hyperparameter combination, the k-fold cross-validation method is used to evaluate the model performance, and its mean squared error and coefficient of determination on the validation set are calculated. By comparing the performance indicators under different hyperparameter combinations, the best-performing hyperparameter combination is finally selected.

9. The method for predicting high-temperature viscosity of environmental sediments based on machine learning according to claim 1, characterized in that, The preset performance evaluation metrics are the lowest mean squared error of the average validation set and the highest mean coefficient of determination of the average validation set.

10. A machine learning-based system for predicting the high-temperature viscosity of environmental sediments, characterized in that, include, The original dataset construction module is used to collect various environmental sediment samples and measure the characteristic parameters of each environmental sediment sample at room temperature and the corresponding high temperature viscosity values ​​at different temperatures to form the original dataset for model training. The room temperature characteristics of the environmental sediments include trace element content, moisture content, crystal content, particle size characteristics and major element content. The feature selection module is used to preprocess the original dataset for model training, including handling missing values ​​and outliers. It also selects features based on their actual physical meaning and the correlation between features to obtain physical quantity features that are highly correlated with high-temperature viscosity and performs standardization processing to construct the final feature dataset suitable for machine learning model training. The prediction model building module is used to divide the final feature dataset into a training set, a validation set, and a test set, and to train various preset machine learning models using the training set. Based on the validation set, and according to the preset performance evaluation indicators, a comprehensive performance evaluation is performed on each candidate model after training, and the prediction model with the best performance is selected; and the generalization performance of different models is evaluated using the test set. The prediction output module is used to input the ambient temperature characteristic data of the sediment in the test environment into the selected optimal prediction model, and output the predicted high temperature viscosity value of the sediment at the target temperature.

Citation Information

Patent Citations

  • Sea algae cause analyzing and concentration predicting method based on machine learning, and system

    CN110379463A

  • Method and device for correcting content of Sr element in sediment

    CN114220490A