Quartz crucible high-temperature deformation resistance rate prediction method based on machine learning
Through machine learning-based methods, important features are screened and constructed, and combined with deep learning models, the accuracy problem of high-temperature deformation resistance prediction of quartz crucibles is solved, efficient and accurate prediction results are achieved, production processes are optimized and costs are reduced.
Patent Information
- Application Number
- CN202411965551.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The prior art is difficult to accurately predict the deformation resistance of quartz crucibles under high temperature environments, which affects the efficiency and product quality of aluminum alloy smelting process.
Using a machine learning-based method, we collect and preprocess the historical data of the quartz crucible, screen important features and construct interactive features, and combine deep learning models to train the prediction model to achieve accurate prediction of the high-temperature deformation resistance of the quartz crucible.
It significantly improves the accuracy and efficiency of high-temperature deformation resistance prediction of quartz crucibles, helps optimize production processes, reduce costs, and improve product quality.
Smart Images

Figure CN119989871A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of crucible deformation resistance prediction, and more specifically, to a method for predicting high-temperature deformation resistance rate of a quartz crucible based on machine learning. Background Art
[0002] As a high-temperature container, quartz crucible plays a vital role in high-temperature processing such as metal smelting, glass manufacturing, and ceramic firing. Its performance directly affects the quality and production efficiency of the product. Quartz crucibles need to have excellent high-temperature deformation resistance to ensure that their shape and size stability are maintained in high-temperature environments to prevent container breakage or product contamination caused by deformation.
[0003] In the smelting process of aluminum, the application of quartz crucible is particularly critical. The melting point of aluminum is about 660 degrees Celsius, and in the actual smelting process, the temperature in the furnace often needs to reach a higher level to ensure the full melting of aluminum and the removal of impurities. In such a high temperature environment, the crucible material must be able to withstand the combined effects of thermal and mechanical stresses, and quartz crucibles are an ideal choice due to their high melting point, good thermal stability and chemical stability.
[0004] The deformation resistance of quartz crucible is a key indicator to measure its performance at high temperature. Low deformation resistance means that the crucible is prone to deformation at high temperature, which may shorten the service life of the crucible, increase production costs, and even affect product quality. Therefore, accurately predicting the deformation resistance is an effective way to achieve more accurate material ratios, improve efficiency and product quality; therefore, with the development of aluminum alloy smelting technology, how to accurately predict the deformation resistance of quartz crucibles, especially the deformation resistance performance under different aluminum contents, has become a core issue in improving production efficiency, reducing costs, and optimizing the ratio of crucible materials. Summary of the invention
[0005] The present invention provides a method for predicting the high-temperature deformation resistance of a quartz crucible based on machine learning, which aims to achieve accurate prediction of the deformation resistance of a quartz crucible.
[0006] A method for predicting high temperature deformation resistance of a quartz crucible based on machine learning, characterized in that it comprises the following steps:
[0007] Step 1: Collect historical data related to the deformation resistance of quartz crucible, including aluminum alloy composition data, crucible physical property data, crucible use conditions and deformation resistance data, and the deformation resistance data is used as the true value label of the input data when training the corresponding prediction model;
[0008] Step 2: Preprocess the collected historical data, including missing value processing, outlier detection and processing, data normalization, and category data encoding to obtain preprocessed data;
[0009] Step 3: Based on the preprocessed data, Pearson correlation coefficient, LASSO regression and principal component analysis are performed step by step to select important features;
[0010] Step 4: construct the interaction features between smelting conditions and physical properties based on the preprocessed data to obtain the constructed new features;
[0011] Step 5: Train the prediction model constructed based on the deep learning model based on the constructed new features and the selected important features to obtain a trained prediction model;
[0012] Step 6: Based on the trained prediction model and real-time data, the high-temperature deformation resistance of the quartz crucible is predicted to obtain the prediction result.
[0013] The present invention realizes an efficient and accurate method for predicting the high-temperature deformation resistance of a quartz crucible by comprehensively applying machine learning and deep learning technologies. Firstly, key historical data are collected from actual production, covering information such as aluminum alloy composition, physical properties of the crucible, usage conditions and deformation resistance, thereby providing a real and comprehensive training basis for the model. After data preprocessing, the quality and consistency of the data set are guaranteed. By means of statistical methods such as the Pearson correlation coefficient, LASSO regression and principal component analysis, features that have a significant impact on the deformation resistance are screened out, and feature engineering is performed to construct new interactive features, which can reflect the complex relationship between smelting conditions and the physical properties of the crucible. The introduction of the deep learning model enables the prediction model to learn the deep-level patterns in the data, thereby significantly improving the prediction accuracy. Finally, the trained model can accurately predict the high-temperature deformation resistance of the quartz crucible based on real-time data, thereby providing a scientific basis for production decision-making, and helping to optimize the production process and reduce costs.
[0014] Preferably, the steps of processing the pre-processed data of the Pearson correlation coefficient are as follows:
[0015] Calculate the Pearson correlation coefficient between each feature and the corresponding deformation resistance:
[0016]
[0017] Where: Represents the correlation coefficient between feature X and anti-deformation rate Y; X i and Y i Represents the characteristic value and target variable value of the i-th sample; and represents the mean of feature X and anti-deformation rate Y; n represents the number of samples;
[0018] If the correlation coefficient between a feature and the target variable is r X,YIf the absolute value of is lower than the set first threshold, the feature is removed to obtain the first data set;
[0019] Based on the first data set, the Pearson correlation coefficient is calculated for each feature:
[0020]
[0021] Where: represents the correlation coefficient between feature X and feature Y′, X i and Y′ i Respectively represent the values of feature X and feature Y′ of the i-th sample; and Represent the mean of feature X and feature Y′ respectively;
[0022] If the Pearson correlation coefficient r between any feature pair X,Y′ If the absolute value of is higher than a preset second threshold, any one feature in the feature pair is removed to obtain a second data set.
[0023] Preferably, the steps of performing feature screening based on the LASSO regression are as follows:
[0024] Cross-validation: Divide the data set obtained by screening based on the Pearson correlation coefficient into several subsets, select one subset as the validation set each time, and the other subsets as the training set, calculate the mean square error under each regularization parameter λ, and use the regularization parameter λ with the smallest mean square error;
[0025] Model training: Based on the data set screened by the Pearson correlation coefficient and the selected regularization parameter λ, the LASSO regression model is trained. The feature X and the anti-deformation rate Y in the feature set are used to train the model by minimizing the loss function of the LASSO regression model. The gradient descent method is used to solve the optimization problem:
[0026]
[0027] Where: Y i represents the anti-deformation rate of the i-th sample; X ij represents the jth eigenvalue of the i-th sample; β j represents the regression coefficient of the jth feature; λ represents the regularization parameter; represents the regression coefficient vector obtained through training, which contains the weights of all features and represents the contribution of each feature to the deformation rate, where p is the total number of features; β0 represents the intercept term of the regression model; β j represents the regression coefficient of the jth feature, where j∈(1,2,…,p);
[0028] Feature screening: If β j =0, then remove the corresponding feature; βj ≠0, the corresponding features are retained; and the third data set is obtained based on feature screening.
[0029] Preferably, the specific steps of the principal component analysis are as follows:
[0030] Calculate the covariance matrix: Let the data matrix in the data set obtained after feature screening based on LASSO regression be X norm , of size n×p, where n is the number of samples and p is the number of features, the covariance is calculated based on the data matrix:
[0031]
[0032] Where: Σ is the p×p covariance matrix, each element σ ij Represents the covariance between the i-th feature and the j-th feature; Represents the data matrix X norm The transpose of
[0033] Compute eigenvalues and eigenvectors: Compute the eigenvalues and eigenvectors of the covariance matrix Σ:
[0034] Σv i =λ i v i ;
[0035] Where: i represents the i-th eigenvalue of the covariance matrix, which represents the variance in the feature direction; v i represents the i-th eigenvector of the covariance matrix, indicating the feature direction;
[0036] Construct dimensionality reduction space: for eigenvalues λ1,λ2,…λ p Sort by size, select the eigenvectors corresponding to the first k eigenvalues, and form a matrix V based on the eigenvectors corresponding to the first k eigenvalues k , the size is p×k, where p represents the number of features in the data set obtained after feature screening based on LASSO regression, and k represents the number of principal components selected;
[0037] Map the data to the reduced dimensionality space: Use matrix multiplication to calculate the coordinates of each data sample in the reduced dimensionality space:
[0038] Z=X norm V k ;
[0039] Where: Z represents the score matrix after dimensionality reduction, with a size of n×k, and each row represents the coordinate of a sample in the dimensionality reduction space;
[0040] Principal component selection: Calculate the ratio of the eigenvalue corresponding to each principal component to the sum of the total eigenvalues to obtain the contribution ratio of each principal component to the data variance:
[0041]
[0042] Where: G i Represents the contribution of the i-th principal component;
[0043] Calculate the cumulative contribution: Calculate the cumulative contribution of the first k principal components;
[0044]
[0045] Where: L k Represents the cumulative contribution of the first k principal components;
[0046] The calculated cumulative contribution is compared with the contribution threshold. If the cumulative contribution is higher than the threshold, the fourth data set is constructed with the currently selected first k principal components. If the cumulative contribution is lower than the threshold, the k value is increased until the cumulative contribution is higher than the threshold, and the selected principal component is obtained, and the fourth data set is constructed based on the selected principal components.
[0047] Preferably, the interaction characteristics between the smelting conditions and the physical properties include the interaction characteristics between the aluminum content and the crucible expansion coefficient and the load characteristics between the smelting time and the temperature; wherein the interaction characteristics between the aluminum content and the crucible expansion coefficient are constructed as follows:
[0048] interaction feature =log(1+A1 content )×(expansion coefficient ) β ;
[0049] Where: β is an adjustable parameter, reflecting the influence of the expansion coefficient on the interaction; log(1+A1 content ) indicates logarithmic transformation of aluminum content; expansion coefficient Indicates the expansion coefficient of the crucible; A1 content Indicates the aluminum content in aluminum alloy; interaction feature It represents the interaction characteristics between aluminum content and crucible expansion coefficient, which is a continuous value;
[0050] The load characteristics of the smelting time and temperature are constructed as follows:
[0051]
[0052] In the formula: melting timerepresents the melting time; temperature represents the melting temperature; λ represents an adjustable hyperparameter that controls the decay rate of the melting time on the load effect; e represents the base of the natural logarithm; load feature It represents the load characteristics of melting time and temperature and is a continuous value.
[0053] Preferably, the prediction model includes an input layer, a feature selection layer, a deep neural network layer, a regularization and optimization layer, and an output layer;
[0054] The input layer is used to input the constructed new features and the selected important features, wherein the input is a complete feature vector concatenated from the constructed new features and the selected important features;
[0055] The feature selection layer: the input feature vector passes through a self-attention mechanism, calculates the correlation between each feature to assign weights to each feature, weights the features based on the assigned weights to obtain weighted features, and adds a Dropout layer after the self-attention mechanism, wherein the Dropout layer discards some features with a predetermined probability to obtain retained weighted features, and passes the retained weighted features to the deep neural network layer;
[0056] The deep neural network layer includes multiple hidden layers, the output of each hidden layer is nonlinearly transformed through a ReLU activation function, and a Batch Normalization layer is added after each hidden layer for normalization, and a Dropout layer is added after the BatchNormalization layer to discard some features with a predetermined probability to obtain the retained deep features;
[0057] The regularization and optimization layer is used to impose L2 regularization constraints on the weights of the deep neural network layer to prevent the model from overfitting;
[0058] Output layer: Based on the output of the deep neural network layer, linear activation is used to output the predicted value of the anti-deformation rate.
[0059] Preferably, the loss function of the prediction model is as follows:
[0060]
[0061] Where: represents the mean square error loss; λ l2 represents the regularization hyperparameter of the deep neural network layer; W attn Represents the weight matrix in the self-attention mechanism; λ attn represents the regularization hyperparameter in the self-attention mechanism; p dropout,dnn represents the discard rate of the Dropout layer in the deep neural network layer; p dropout,attnrepresents the discard rate of the Dropout layer in the feature selection layer; θ j represents the jth parameter in the network; M represents the total number of parameters; Represents the total loss function.
[0062] The beneficial effects of the present invention include:
[0063] The present invention realizes an efficient and accurate method for predicting the high-temperature deformation resistance of a quartz crucible by comprehensively applying machine learning and deep learning technologies. Firstly, key historical data are collected from actual production, covering information such as aluminum alloy composition, physical properties of the crucible, usage conditions and deformation resistance, thereby providing a real and comprehensive training basis for the model. After data preprocessing, the quality and consistency of the data set are guaranteed. By means of statistical methods such as the Pearson correlation coefficient, LASSO regression and principal component analysis, features that have a significant impact on the deformation resistance are screened out, and feature engineering is performed to construct new interactive features, which can reflect the complex relationship between smelting conditions and the physical properties of the crucible. The introduction of the deep learning model enables the prediction model to learn the deep-level patterns in the data, thereby significantly improving the prediction accuracy. Finally, the trained model can accurately predict the high-temperature deformation resistance of the quartz crucible based on real-time data, thereby providing a scientific basis for production decision-making, and helping to optimize the production process and reduce costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0065] Figure 1 An overall step block diagram provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0066] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0067] See also Figure 1 As shown, a method for predicting the high temperature deformation resistance of a quartz crucible based on machine learning is characterized in that it comprises the following steps:
[0068] Step 1: Collect historical data related to the deformation resistance of quartz crucible, including aluminum alloy composition data, crucible physical property data, crucible use conditions and deformation resistance data, and the deformation resistance data is used as the true value label of the input data when training the corresponding prediction model;
[0069] The aluminum alloy composition data includes aluminum content, silicon content, iron content and content of other elements (such as copper content, magnesium content, etc.); wherein the aluminum alloy composition data is obtained based on chemical analysis technology (such as X-ray fluorescence spectroscopy, laser induced breakdown spectroscopy analysis, etc.) or obtained through a formula;
[0070] The crucible physical property data include:
[0071] Density: The density of the crucible material affects its heat capacity and thermal conductivity;
[0072] Thermal expansion coefficient: describes the expansion characteristics of the crucible material at high temperatures;
[0073] Compressive strength: the ability of the crucible material to resist compression under external forces;
[0074] Thermal conductivity: The thermal conductivity of the crucible material affects its temperature distribution and thermal stability;
[0075] Specific heat capacity: The amount of heat absorbed per unit mass of the crucible material, which affects the heat capacity of the crucible;
[0076] Oxidation resistance: The oxidation resistance of the crucible material in a high temperature environment affects its service life;
[0077] Obtained through standard physical testing methods, such as using thermal expansion instruments, thermal conductivity testers, compression testing equipment, etc.
[0078] The crucible use condition data includes:
[0079] Temperature: The temperature range that the crucible withstands during use, especially the operating temperature under high temperature conditions.
[0080] Use cycle: The number of times or duration of use of the crucible affects the fatigue and deformation capacity of the material.
[0081] Type of smelting material: The type of smelting metal (such as aluminum alloy, steel, copper, etc.) will affect the stress distribution of the crucible.
[0082] Melting speed: The speed at which the metal melts is closely related to the crucible's ability to resist deformation.
[0083] Crucible use environment: Environmental factors such as oxygen concentration in the furnace, furnace atmosphere (oxidizing or reducing atmosphere) and so on will also affect the performance of the crucible.
[0084] The anti-deformation rate data is obtained through experiments and currently available recorded data.
[0085] Step 2: Preprocess the collected historical data, including missing value processing, outlier detection and processing, data normalization, and category data encoding to obtain preprocessed data;
[0086] We use the median to fill missing values. For outlier detection, we calculate the mean and standard deviation of the data, define the normal range (such as data exceeding the mean ± 3 times the standard deviation is considered an outlier), and remove or replace the outliers.
[0087] The normalization adopts Z-score standardization: the data is converted into a distribution with a mean of 0 and a standard deviation of 1 by subtracting the mean and dividing by the standard deviation;
[0088] Categorical data encoding: For data containing categories (such as categorical information in aluminum alloy composition, smelting material type in crucible use conditions, etc.), categorical encoding is performed, such as using one-hot encoding: converting categorical variables into binary matrices, and each category becomes a binary feature. For example, if "aluminum alloy type" has 3 categories, it can be converted into 3 binary columns.
[0089] The above-mentioned preprocessing technical means are all conventional technical means in the field, and according to the above description, those skilled in the art can completely restore the specific steps of the above-mentioned preprocessing according to the above-mentioned technical means. Therefore, how to implement it specifically will not be repeated in this embodiment.
[0090] Step 3: Based on the preprocessed data, Pearson correlation coefficient, LASSO regression and principal component analysis are performed step by step to select important features;
[0091] As a possible implementation of this embodiment, the steps of processing the pre-processed data using the Pearson correlation coefficient are as follows:
[0092] Calculate the Pearson correlation coefficient between each feature and the corresponding deformation resistance:
[0093]
[0094] Where: Represents the correlation coefficient between feature X and anti-deformation rate Y; X i and Y i Represents the characteristic value and target variable value of the i-th sample; and represents the mean of feature X and anti-deformation rate Y; n represents the number of samples;
[0095] If the correlation coefficient between a feature and the target variable is rX,Y If the absolute value of is lower than the set first threshold, the feature is removed to obtain the first data set;
[0096] For example, the first threshold is set to 0.2, and we have five samples in total, each sample contains five aluminum contents, silicon contents and corresponding deformation resistances;
[0097] First, the mean value of aluminum content, the mean value of silicon content and the mean value of deformation resistance are calculated; the mean value of aluminum content and the corresponding mean value of deformation resistance are substituted into the calculation formula of the Pearson correlation coefficient between the above characteristics and the corresponding deformation resistance, and then the aluminum content of each sample and the corresponding deformation resistance of each sample are substituted into the above formula to calculate the Pearson correlation coefficient between the aluminum content and the deformation resistance; the Pearson correlation coefficient between the silicon content and the corresponding deformation resistance can also be calculated with reference to the aluminum content, that is, the calculation logic is the same;
[0098] Finally, if the absolute value of the Pearson correlation coefficient between the silicon content and the corresponding deformation resistance is less than 0.2, it will be eliminated; if it is greater, it will be retained, until all features are calculated in the above way, and finally all retained features are obtained to form the first data set.
[0099] Based on the first data set, the Pearson correlation coefficient is calculated for each feature:
[0100]
[0101] Where: represents the correlation coefficient between feature X and feature Y′, X i and Y′ i Respectively represent the values of feature X and feature Y′ of the i-th sample; and Represent the mean of feature X and feature Y′ respectively;
[0102] If the Pearson correlation coefficient r between any feature pair X,Y′ If the absolute value of is higher than a preset second threshold, any one feature in the feature pair is removed to obtain a second data set;
[0103] The above calculation logic is actually the same as the Pearson correlation coefficient between the feature and the corresponding anti-deformation rate, but the data is different. The Pearson correlation coefficient is calculated between each feature, that is, between the silicon content and the aluminum content. Assume that the second threshold we set is 0.8; if the absolute value of the calculated Pearson correlation coefficient of the silicon content and the aluminum content is greater than 0.8, then any one of the features can be eliminated; by eliminating one of the two highly correlated features, the problem of multicollinearity is significantly reduced, and the stability and generalization ability of the model are improved; and the information provided by highly correlated features is redundant, retaining the features will increase the complexity of the model, resulting in a long model training time and increased computing resource consumption, and by eliminating one of these highly correlated features, we reduce the number of parameters and the training cost of the model, and after elimination, the model will focus more on important features.
[0104] As a possible implementation of this embodiment, the steps of performing feature screening based on the LASSO regression are as follows:
[0105] Cross-validation: Divide the second data set obtained by screening based on the Pearson correlation coefficient into several subsets, select one subset as the validation set each time, and the other subsets as the training set, calculate the mean square error under each regularization parameter λ, and use the regularization parameter λ with the smallest mean square error;
[0106] Model training: Based on the second data set obtained by screening the Pearson correlation coefficient and the selected regularization parameter λ, the LASSO regression model is trained. The feature X and the anti-deformation rate Y in the feature set are used to train the model by minimizing the loss function of the LASSO regression model. The gradient descent method is used to solve the optimization problem:
[0107]
[0108] Where: Y i represents the anti-deformation rate of the i-th sample; X ij represents the jth eigenvalue of the i-th sample; β j represents the regression coefficient of the jth feature; λ represents the regularization parameter; represents the regression coefficient vector obtained through training, which contains the weights of all features and represents the contribution of each feature to the deformation rate, where p is the total number of features; β0 represents the intercept term of the regression model; β j represents the regression coefficient of the jth feature, where j∈(1,2,…,p);
[0109] Feature screening: If β j =0, then remove the corresponding feature; β j ≠0, the corresponding features are retained; and the third data set is obtained based on feature screening;
[0110] Exemplary:
[0111] The form of the LASSO regression model is as follows:
[0112]
[0113] Where: represents the predicted deformation resistance; X1, X2, …, X p Represent the values of different features respectively; X p represents the intercept term;
[0114] Assume that we have the following simplified data set, including 3 features (X1, X2, X3) and the target variable Y (anti-deformation rate), the specific values are shown in Table 1;
[0115] Table 1
[0116] Sample No. X1 X2 X3 Y 1 1 2 3 4 2 2 3 4 5 3 3 4 5 6
[0117] The steps for feature screening based on the LASSO regression are as follows:
[0118] Initialization: Select an initial regularization parameter λ;
[0119] Loss function calculation: For each regularization parameter λ, calculate the value of the loss function, for example, for λ = 0.1:
[0120] Loss(λ)=(4-(b0+b1·1+b2·2+b3·3)) 2 +(5-(b0+b1·2+b2·3+b3·4)) 2 +(5-(b0+b1·3+b2·4+b3·5)) 2 +0.1(|b1|+|b2|+|b3|);
[0121] Gradient descent method: Use gradient descent method to update coefficients b0, b1, b2, b3 to minimize the loss function Loss(λ);
[0122] Cross-validation: Repeat the above process of initialization to gradient descent, each choosing a different lambda value, and calculate the mean square error on the validation set;
[0123] Select the optimal λ: find the λ value that minimizes the mean square error on the validation set;
[0124] Final model: Use the optimal λ value and the entire dataset to train the final model and calculate the coefficient β j To filter features.
[0125] As a possible implementation of this embodiment, the specific steps of the principal component analysis are as follows:
[0126] Calculate the covariance matrix: Let the data matrix of the third data set obtained after feature screening based on LASSO regression be X norm , of size n×p, where n is the number of samples and p is the number of features, the covariance is calculated based on the data matrix:
[0127]
[0128] Where: Σ is the p×p covariance matrix, each element σ ij Represents the covariance between the i-th feature and the j-th feature; Represents the data matrix X norm The transpose of
[0129] Compute eigenvalues and eigenvectors: Compute the eigenvalues and eigenvectors of the covariance matrix Σ:
[0130] Σv i =λ i v i ;
[0131] Where: i represents the i-th eigenvalue of the covariance matrix, which represents the variance in the feature direction; v i represents the i-th eigenvector of the covariance matrix, represents the characteristic direction; the solution of the eigenvalue and the eigenvector is solved by numerical calculation software, such as NumPy, MATLAB, Mathematica, etc., and how to use these software for solving belongs to the conventional technical means in this field, so it will not be repeated.
[0132] Construct dimensionality reduction space: for eigenvalues λ1,λ2,…λ p Sort by size, select the eigenvectors corresponding to the first k eigenvalues, and form a matrix V based on the eigenvectors corresponding to the first k eigenvalues k , the size is p×k, where p represents the number of features in the data set obtained after feature screening based on LASSO regression, and k represents the number of principal components selected;
[0133] Map the data to the reduced dimensionality space: Use matrix multiplication to calculate the coordinates of each data sample in the reduced dimensionality space:
[0134] Z=X norm V k ;
[0135] Where: Z represents the score matrix after dimensionality reduction, with a size of n×k, and each row represents the coordinate of a sample in the dimensionality reduction space;
[0136] Principal component selection: Calculate the ratio of the eigenvalue corresponding to each principal component to the sum of the total eigenvalues to obtain the contribution ratio of each principal component to the data variance:
[0137]
[0138] Where: G i Represents the contribution of the i-th principal component;
[0139] Calculate the cumulative contribution: Calculate the cumulative contribution of the first k principal components;
[0140]
[0141] Where: L k Represents the cumulative contribution of the first k principal components;
[0142] The calculated cumulative contribution is compared with the contribution threshold. If the cumulative contribution is higher than the threshold, the fourth data set is constructed with the currently selected first k principal components; if the cumulative contribution is lower than the threshold, the k value is increased until the cumulative contribution is higher than the threshold, and the selected principal component is obtained, and the fourth data set is constructed based on the selected components; for example, the contribution threshold is set to 0.95; k is 8, if the cumulative contribution of the principal components of the 8 corresponding features is greater than 0.95, there is no need to add features; if the cumulative contribution is lower than 0.95, k is calculated by adding 1 in sequence until the contribution is greater than 0.95.
[0143] In actual production, quickly and accurately predicting the deformation resistance of quartz crucibles is crucial to optimizing production processes, reducing downtime, and improving production efficiency. In this embodiment, by calculating the correlation coefficient between the feature and the target variable (deformation resistance), the features that are highly correlated with the target variable can be preliminarily screened out; it is helpful to remove irrelevant or weakly correlated features and reduce data dimensions, thereby reducing the risk of overfitting of the model and improving the accuracy of prediction; LASSO regression imposes constraints on feature weights by introducing regularization terms, which can further screen out features that have a significant impact on the deformation resistance, while compressing or excluding unimportant features; it not only helps to simplify the model, but also improves the ability to identify key features, thereby improving the model. The ability to capture and predict complex relationships; and by calculating the Pearson correlation coefficient between feature pairs, highly correlated feature pairs can be identified and eliminated to avoid the model from learning redundant information, thereby optimizing the feature combination and improving the generalization ability of the model; finally, by selecting the principal component with the largest explained variance, the number of features can be reduced while retaining important information, thereby reducing the complexity and computational cost of the model. Therefore, in this embodiment, the feature selection technical solutions of Pearson correlation coefficient, LASSO regression and principal component analysis are gradually carried out. In the prediction of high-temperature deformation resistance of quartz crucible, the accuracy, efficiency and interpretability of the prediction model are significantly improved by optimizing the feature quality and model structure, providing strong technical support for the production process.
[0144] Step 4: construct the interaction features between smelting conditions and physical properties based on the preprocessed data to obtain the constructed new features;
[0145] As a possible implementation of this embodiment, the interaction characteristics between the smelting conditions and the physical properties include the interaction characteristics between the aluminum content and the crucible expansion coefficient and the load characteristics of the smelting time and temperature; wherein the interaction characteristics between the aluminum content and the crucible expansion coefficient are constructed as follows:
[0146] interaction feature =log(1+A1 content )×(expansion coefficient ) β ;
[0147] Where: β is an adjustable parameter, reflecting the influence of the expansion coefficient on the interaction; log(1+A1 content ) indicates logarithmic transformation of aluminum content; expansion coefficient Indicates the expansion coefficient of the crucible; A1 content Indicates the aluminum content in aluminum alloy; interaction feature It represents the interaction characteristics between aluminum content and crucible expansion coefficient, which is a continuous value;
[0148] The load characteristics of the smelting time and temperature are constructed as follows:
[0149]
[0150] In the formula: melting time represents the melting time; temperature represents the melting temperature; λ represents an adjustable hyperparameter that controls the decay rate of the melting time on the load effect; e represents the base of the natural logarithm; load feature It represents the load characteristics of melting time and temperature and is a continuous value.
[0151] Step 5: Train the prediction model constructed based on the deep learning model based on the constructed new features and the selected important features to obtain a trained prediction model;
[0152] As a possible implementation of this embodiment, the prediction model includes an input layer, a feature selection layer, a deep neural network layer, a regularization and optimization layer, and an output layer;
[0153] The input layer is used to input the constructed new features and the selected important features, wherein the input is a complete feature vector concatenated from the constructed new features and the selected important features;
[0154] The feature selection layer: the input feature vector passes through a self-attention mechanism, calculates the correlation between each feature to assign weights to each feature, weights the features based on the assigned weights to obtain weighted features, and adds a Dropout layer after the self-attention mechanism, wherein the Dropout layer discards some features with a predetermined probability (0.2 or 0.3) to obtain the retained weighted features, and passes the retained weighted features to the deep neural network layer;
[0155] The deep neural network layer includes multiple hidden layers, the output of each hidden layer is nonlinearly transformed through a ReLU activation function, and a Batch Normalization layer is added after each hidden layer for normalization, and a Dropout layer is added after the BatchNormalization layer to discard some features with a predetermined probability to obtain the retained deep features;
[0156] For example, three hidden layers are used, the first hidden layer includes 128 neurons and uses the ReLU activation function; the second hidden layer uses 64 neurons and uses the ReLU activation function; the third hidden layer uses 32 neurons and uses the ReLU activation function;
[0157] After the output of each hidden layer, a Batch Normalization layer is added to standardize the output of each hidden layer to ensure that the mean of the output is 0 and the variance is 1. Batch Normalization helps speed up training and reduce the network's sensitivity to initialization:
[0158]
[0159] Where: h BN represents the output vector of the Batch Normalization layer; h represents the output feature vector of the hidden layer; μ represents the mean of the current batch feature; σ represents the standard deviation of the current batch feature; γ and β represent learnable parameters;
[0160] There is a Dropout layer after each hidden layer to prevent overfitting. For example, the Dropout probability is set to 0.3, which means that 30% of the neurons will be randomly dropped during training.
[0161] The regularization and optimization layer is used to impose L2 regularization constraints on the weights of the deep neural network layer to prevent the model from overfitting;
[0162] The regularization is achieved by adding the sum of squared weights to the loss function:
[0163]
[0164] Where: L2loss represents the loss caused by regularization; λ represents the regularization coefficient; W i Represents the i-th weight of the model; see loss function for details;
[0165] In the process of training the model, the Adam optimizer is used to optimize the model by combining gradient descent and adaptive learning rate adjustment;
[0166] Output layer: Based on the output of the deep neural network layer, linear activation is used to output the predicted value of the anti-deformation rate.
[0167] The loss function of the prediction model is as follows:
[0168]
[0169] Where: represents the mean square error loss; λ l2 represents the regularization hyperparameter of the deep neural network layer; W attn Represents the weight matrix in the self-attention mechanism; λ attn represents the regularization hyperparameter in the self-attention mechanism; p dropout,dnn represents the discard rate of the Dropout layer in the deep neural network layer; pdropout,attn represents the discard rate of the Dropout layer in the feature selection layer; θ j represents the jth parameter in the network; M represents the total number of parameters; Represents the total loss function. Represents the regularization loss in the attention mechanism; It represents the regularization term in deep neural networks, which penalizes the square of all parameters in the network, that is, each hidden layer includes a weight and a bias, and the square of the weight and bias of each hidden layer is penalized.
[0170] In this embodiment, through the self-attention mechanism, the model can automatically learn and identify the correlation between input features, assign different weights to each feature, thereby highlighting important features, suppressing unimportant features, and improving the efficiency and quality of feature representation; a Dropout layer is introduced after each hidden layer of the feature selection layer and the deep neural network layer, and a certain proportion of neurons are randomly discarded to reduce the model's excessive dependence on training data and reduce the risk of overfitting.
[0171] The specific steps of optimizing the model using the Adam optimizer combined with gradient descent and adaptive learning rate adjustment are as follows:
[0172] Initialization: For each parameter θ of the prediction model (such as the weight matrix in the self-attention mechanism, the weight matrix and bias term in the deep neural network layer, and the learnable parameters γ and β in the batch normalization layer), initialize its value to a small random number from a normal distribution or a uniform distribution; set the learning rate η, which represents the step size of the parameter update during the model training process; set two decay rates β1 and β2, which are used to calculate the exponential decay rate of the first-order moment estimate and the second-order moment estimate of the gradient, respectively; set a small constant ∈, which is 10 -8 ;
[0173] Training: For each iteration t:
[0174] Forward propagation: forward propagation of input data through the model to obtain the predicted value;
[0175] Calculate loss: Calculate the loss function
[0176] Back propagation: Perform back propagation to calculate the loss function with respect to each parameter θ and gradient g t ;
[0177] Update the first-order moment estimate m t : Update the first-order moment estimate (i.e., the exponentially weighted moving average of the gradient) based on the gradient and decay rate β1:
[0178] mt =β1·m t-1 +(1-β1)·g t ;
[0179] Where: m t represents the order moment estimate of the t-th iteration; m t-1 represents the order moment estimate of the t-1th iteration; g t represents the gradient of the tth iteration;
[0180] Update the second-order moment estimate v t : Update the second-order moment estimate (i.e., the exponentially weighted moving average of the square of the gradient) based on the gradient and decay rate β2:
[0181]
[0182] Where: v t represents the second-order moment estimate of the t-th iteration; v t-1 represents the second-order moment estimate of the t-1th iteration;
[0183] Bias Correction: Perform bias correction on the first- and second-order moment estimates:
[0184]
[0185] Where: represents the bias-corrected first-order moment estimate; represents the bias-corrected second-order moment estimate;
[0186] Update parameters: Use the bias-corrected estimates to update the parameters θ:
[0187]
[0188] Where: θ t represents the parameters before updating; θ t+1 represents the updated parameters;
[0189] Repeated iteration: Repeat the steps from forward propagation to parameter update until the stopping condition is met, such as reaching the predetermined number of iterations, the value of the loss function no longer decreases significantly, etc.
[0190] Step 6: Based on the trained prediction model and real-time data, the high-temperature deformation resistance of the quartz crucible is predicted to obtain the prediction result.
[0191] The present invention realizes an efficient and accurate method for predicting the high-temperature deformation resistance of a quartz crucible by comprehensively applying machine learning and deep learning technologies. Firstly, key historical data are collected from actual production, covering information such as aluminum alloy composition, physical properties of the crucible, usage conditions and deformation resistance, thereby providing a real and comprehensive training basis for the model. After data preprocessing, the quality and consistency of the data set are guaranteed. By means of statistical methods such as the Pearson correlation coefficient, LASSO regression and principal component analysis, features that have a significant impact on the deformation resistance are screened out, and feature engineering is performed to construct new interactive features, which can reflect the complex relationship between smelting conditions and the physical properties of the crucible. The introduction of the deep learning model enables the prediction model to learn the deep-level patterns in the data, thereby significantly improving the prediction accuracy. Finally, the trained model can accurately predict the high-temperature deformation resistance of the quartz crucible based on real-time data, thereby providing a scientific basis for production decision-making, and helping to optimize the production process and reduce costs.
[0192] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for predicting the high temperature deformation resistance of a quartz crucible based on machine learning, characterized in that: The following steps are involved: Step 1: Collect historical data related to the deformation resistance of quartz crucible, including aluminum alloy composition data, crucible physical property data, crucible use conditions and deformation resistance data, and the deformation resistance data is used as the true value label of the input data when training the corresponding prediction model; Step 2: Preprocess the collected historical data, including missing value processing, outlier detection and processing, data normalization, and category data encoding to obtain preprocessed data; Step 3: Based on the preprocessed data, Pearson correlation coefficient, LASSO regression and principal component analysis are performed step by step to select important features; Step 4: construct the interaction features between smelting conditions and physical properties based on the preprocessed data to obtain the constructed new features; Step 5: Train the prediction model constructed based on the deep learning model based on the constructed new features and the selected important features to obtain a trained prediction model; Step 6: Based on the trained prediction model and real-time data, the high-temperature deformation resistance of the quartz crucible is predicted to obtain the prediction result.
2. The method for predicting high temperature deformation resistance of quartz crucible based on machine learning according to claim 1, characterized in that: The steps of processing the pre-processed data of the Pearson correlation coefficient are as follows: Calculate the Pearson correlation coefficient between each feature and the corresponding deformation resistance: Where: Represents the correlation coefficient between feature X and anti-deformation rate Y; X i and Y i Represents the characteristic value and target variable value of the i-th sample; and represents the mean of feature X and anti-deformation rate Y; n represents the number of samples; If the correlation coefficient between a feature and the target variable is r X,Y If the absolute value of is lower than the set first threshold, the feature is removed to obtain the first data set; Based on the first data set, the Pearson correlation coefficient is calculated for each feature: Where: represents the correlation coefficient between feature X and feature Y′, X i and Y′ i Respectively represent the values of feature X and feature Y′ of the i-th sample; and Represent the mean of feature X and feature Y′ respectively; If the Pearson correlation coefficient r between any feature pair X,Y′ If the absolute value of is higher than a preset second threshold, any one feature in the feature pair is removed to obtain a second data set.
3. The method for predicting high temperature deformation resistance of quartz crucible based on machine learning according to claim 1, characterized in that: The steps for feature screening based on the LASSO regression are as follows: Cross-validation: Divide the data set obtained by screening based on the Pearson correlation coefficient into several subsets, select one subset as the validation set each time, and the other subsets as the training set, calculate the mean square error under each regularization parameter λ, and use the regularization parameter λ with the smallest mean square error; Model training: Based on the data set screened by the Pearson correlation coefficient and the selected regularization parameter λ, the LASSO regression model is trained. The feature X and the anti-deformation rate Y in the feature set are used to train the model by minimizing the loss function of the LASSO regression model. The gradient descent method is used to solve the optimization problem: Where: Y i represents the anti-deformation rate of the i-th sample; X ij represents the jth eigenvalue of the i-th sample; β j represents the regression coefficient of the jth feature; λ represents the regularization parameter; represents the regression coefficient vector obtained through training, which contains the weights of all features and represents the contribution of each feature to the anti-deformation rate, where p is the total number of features; β0 represents the intercept term of the regression model; β j represents the regression coefficient of the jth feature, where j∈(1, 2, ..., p); Feature screening: If β j =0, then remove the corresponding feature; β j ≠0, the corresponding features are retained; A third data set is obtained based on feature screening.
4. The method for predicting high temperature deformation resistance of quartz crucible based on machine learning according to claim 1, characterized in that: The specific steps of the principal component analysis are as follows: Calculate the covariance matrix: Let the data matrix in the data set obtained after feature screening based on LASSO regression be X norm , of size n×p, where n is the number of samples and p is the number of features, the covariance is calculated based on the data matrix: Where: ∑ is the p×p covariance matrix, each element σ ij Represents the covariance between the i-th feature and the j-th feature; Represents the data matrix X norm The transpose of Calculate eigenvalues and eigenvectors: Calculate the eigenvalues and eigenvectors of the covariance matrix ∑: ∑v i =λ i v i ; Where: i represents the i-th eigenvalue of the covariance matrix, which represents the variance in the feature direction; v i represents the i-th eigenvector of the covariance matrix, indicating the feature direction; Construct a reduced dimensionality space: for eigenvalues λ1, λ2, ...λ p Sort by size, select the eigenvectors corresponding to the first k eigenvalues, and form a matrix V based on the eigenvectors corresponding to the first k eigenvalues k , the size is p×k, where p represents the number of features in the data set obtained after feature screening based on LASSO regression, and k represents the number of principal components selected; Map the data to the reduced dimensionality space: Use matrix multiplication to calculate the coordinates of each data sample in the reduced dimensionality space: Z=X norm V k ; Where: Z represents the score matrix after dimensionality reduction, with a size of n×k, and each row represents the coordinate of a sample in the dimensionality reduction space; Principal component selection: Calculate the ratio of the eigenvalue corresponding to each principal component to the sum of the total eigenvalues to obtain the contribution ratio of each principal component to the data variance: Where: G i Represents the contribution of the i-th principal component; Calculate the cumulative contribution: Calculate the cumulative contribution of the first k principal components; Where: L k Represents the cumulative contribution of the first k principal components; The calculated cumulative contribution is compared with the contribution threshold. If the cumulative contribution is higher than the threshold, the fourth data set is constructed with the currently selected first k principal components. If the cumulative contribution is lower than the threshold, the k value is increased until the cumulative contribution is higher than the threshold, and the selected principal component is obtained, and the fourth data set is constructed based on the selected principal components.
5. The method for predicting high temperature deformation resistance of quartz crucible based on machine learning according to claim 1, characterized in that: The interaction characteristics between the smelting conditions and the physical properties include the interaction characteristics between the aluminum content and the crucible expansion coefficient and the load characteristics of the smelting time and temperature; wherein the interaction characteristics between the aluminum content and the crucible expansion coefficient are constructed as follows: interaction feature =log(1+A1 content )×(expansion coefficient ) β ; Where: β is an adjustable parameter, reflecting the influence of the expansion coefficient on the interaction; log(1+A1 content ) indicates logarithmic transformation of aluminum content; expansion coefficient Indicates the expansion coefficient of the crucible; A1 content Indicates the aluminum content in aluminum alloy; interaction feature It represents the interaction characteristics between aluminum content and crucible expansion coefficient, which is a continuous value; The load characteristics of the smelting time and temperature are constructed as follows: In the formula: melting time represents the melting time; temperature represents the melting temperature; λ represents an adjustable hyperparameter that controls the decay rate of the melting time on the load effect; e represents the base of the natural logarithm; load feature It represents the load characteristics of melting time and temperature and is a continuous value.
6. The method for predicting high temperature deformation resistance of quartz crucible based on machine learning according to claim 1, characterized in that: The prediction model includes an input layer, a feature selection layer, a deep neural network layer, a regularization and optimization layer, and an output layer; The input layer is used to input the constructed new features and the selected important features, wherein the input is a complete feature vector concatenated from the constructed new features and the selected important features; The feature selection layer: the input feature vector passes through a self-attention mechanism, calculates the correlation between each feature to assign weights to each feature, weights the features based on the assigned weights to obtain weighted features, and adds a Dropout layer after the self-attention mechanism, wherein the Dropout layer discards some features with a predetermined probability to obtain retained weighted features, and passes the retained weighted features to the deep neural network layer; The deep neural network layer includes multiple hidden layers, the output of each hidden layer is nonlinearly transformed through a ReLU activation function, and a Batch Normalization layer is added after each hidden layer for normalization, and a Dropout layer is added after the BatchNormalization layer to discard some features with a predetermined probability to obtain the retained deep features; The regularization and optimization layer is used to impose L2 regularization constraints on the weights of the deep neural network layer to prevent the model from overfitting; Output layer: Based on the output of the deep neural network layer, linear activation is used to output the predicted value of the anti-deformation rate.
7. The method for predicting high temperature deformation resistance of quartz crucible based on machine learning according to claim 6, characterized in that: The loss function of the prediction model is as follows: Where: represents the mean square error loss; λ l2 represents the regularization hyperparameter of the deep neural network layer; W attn Represents the weight matrix in the self-attention mechanism; λ attn represents the regularization hyperparameter in the self-attention mechanism; p dropout,dm represents the discard rate of the Dropout layer in the deep neural network layer; p dropout,attn represents the discard rate of the Dropout layer in the feature selection layer; θ j represents the jth parameter in the network; M represents the total number of parameters; Represents the total loss function.
Citation Information
Patent Citations
Quartz glass crucible and production method therefor
CN109563639A
Multi-task learning piano polyphonic music automatic notation method based on local attention
CN116665704A
Method for predicting thickness of glass cold shell of cold crucible and early warning method of cold crucible
CN117776488A
Data prediction device and quartz glass crucible manufacturing system
JP2022093298A
Data prediction device and quartz glass crucible manufacturing system
JP2023088369A
Cited By
Rapid evaluation method for use performance of quartz crucible
CN120314133A