Prediction method of burst pressure of composite hydrogen storage cylinders based on small sample machine learning
Through the small sample machine learning method, a blast pressure prediction model of composite hydrogen storage cylinders is established using generative adversarial network and XGBoost technology, solving the accuracy of blast pressure prediction of composite hydrogen storage cylinders, and achieving rapid and efficient design iteration.
Patent Information
- Application Number
- CN202510387260.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art is difficult to accurately predict the blasting pressure of composite hydrogen storage cylinders, which makes the design and manufacturing process time-consuming and labor-intensive, and the predicted value and test value are very different, which requires repeated iteration.
A small sample machine learning method is used to generate virtual data through the generation of adversarial networks, combined with XGBoost extreme gradient enhancement and regularization technology, a blasting pressure prediction model for composite hydrogen storage cylinders is established, and the radial basis function is used for correction.
It realizes fast and accurate prediction based on a small amount of experimental data, reduces the number of experiments, improves the robustness and generalization ability of the prediction model, and avoids the error of complex numerical simulations.
Smart Images

Figure CN119903762B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydrogen storage cylinder bursting pressure prediction and machine learning, and in particular to a method for predicting the bursting pressure of a composite hydrogen storage cylinder based on small sample machine learning. Background Art
[0002] Composite hydrogen storage cylinders, as key equipment for gaseous hydrogen storage, have undergone extensive research and development in recent years. They consist of an inner liner, valve seat, and a carbon fiber-resin-based composite wrapping layer. The composite layer primarily bears the internal pressure load. The burst pressure of composite hydrogen storage cylinders is a key performance indicator in structural design and manufacturing, crucial for ensuring cylinder safety and reliability. However, the design and manufacturing process involves multiple steps, such as carbon fiber wrapping and resin curing, and numerous influencing factors, making accurate prediction of burst pressure challenging.
[0003] At present, the prediction of the burst pressure of composite hydrogen storage cylinders mainly relies on finite element simulation technology. The pressure value corresponding to structural failure is calculated through simplified numerical models and idealized assumptions, which often ignores the many random factors in the actual cylinder manufacturing process. As a result, the predicted values obtained through time-consuming modeling calculations are significantly different from the experimental values. Repeated experimental iterations are required to find the optimal design and process. Summary of the Invention
[0004] In order to overcome the above-mentioned defects in the prior art, the present invention provides a method for predicting the bursting pressure of composite hydrogen storage cylinders based on small sample machine learning. The integrated learning model is trained based on a small amount of burst test data to achieve rapid and accurate prediction of the cylinder bursting pressure.
[0005] To achieve the above object, the present invention adopts the following technical solutions, including:
[0006] The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning includes the following steps:
[0007] S1, obtaining the original data of the explosion test of the composite hydrogen storage cylinder, the original data including characteristic data and the corresponding actual explosion pressure;
[0008] S2, preprocessing the original data to form a real database;
[0009] S3: Train a generative adversarial network based on a real database, and use the trained generative adversarial network to generate virtual data several times the amount of original data to establish a virtual database;
[0010] S4, based on the virtual database, XGBoost extreme gradient boosting training is performed to preliminarily obtain the bursting pressure prediction model;
[0011] S5, correcting the preliminary bursting pressure prediction model based on the real database to obtain the final bursting pressure prediction model.
[0012] Preferably, in step S4, Lasso and Ridge regularization are introduced, and combined with K-fold cross validation technology, XGBoost extreme gradient boosting training of the model is performed.
[0013] Preferably, the specific method of step S4 is as follows:
[0014] Build an XGBoost extreme gradient boosting model. The model is composed of multiple decision trees. The number of decision trees is used as a hyperparameter. The optimal number of decision trees is determined through grid search. The optimal number of decision trees is obtained by minimizing the mean square error under K-fold cross validation.
[0015] After determining the optimal number of decision trees, set the search space for other hyperparameters, use the Bayesian optimization method to optimize various hyperparameter combinations, and obtain the optimal hyperparameter combination by minimizing the mean square error under K-fold cross validation;
[0016] Based on the determined optimal hyperparameter combination, the training set in the virtual database is learned, and the objective function Obj of the model training is:
[0017] ;
[0018] ;
[0019] ;
[0020] in, loss is the loss function; y i For the virtual database i The burst pressure value of each sample, For the i The predicted burst pressure of samples; i 、 n are the sample number and the total number of samples in the training set respectively; It is a regularization term, including Lasso regularization and Ridge regularization. Lasso regularization adds weights to the objective function. w j Absolute value penalty coefficient α Implement feature selection and sparsity; Ridge regularization through penalty coefficient l Limit the weight of the entire model w j Size; Hyperparameters c Used to control the strength of the regularization term.
[0021] Preferably, the data in the virtual database is randomly divided into K subsets. For each iterative training and testing, one of the subsets is selected as the test set, and the other K-1 subsets are combined as the training set for K-fold cross validation.
[0022] Preferably, an early stopping technique is introduced during the XGBoost extreme gradient boosting training process of the model, specifically: setting the initial value of the parameter T. If the performance of the trained model on the test set does not improve, the value of T is increased by 1, and the model training is continued; if the performance improves, T is reset and the model parameters are updated; when T reaches the preset value, the model training is stopped.
[0023] Preferably, in step S5, a radial basis function (RBF) method is used to perform adaptive proportional factor adjustment, and the initially obtained burst pressure prediction model is corrected in combination with a real database.
[0024] Preferably, the specific method of step S5 is as follows:
[0025] Calculate each sample in the real database x RBF kernel function value:
[0026] ;
[0027] in, F is the RBF kernel function; c represents the center point in the real database, which is the mean value of the real database; s is the width parameter;
[0028] ;
[0029] in, R is the data range;
[0030] ;
[0031] in, R k For the k The data range of the feature, that is, k The difference between the maximum and minimum values of a feature; m k For the k The weight of each feature is normalized by the importance value of each feature, i.e., the SHAP value, as the corresponding weight; k 、 M are the feature number and the total number of features respectively;
[0032] The objective function is the mean square error MSE :
[0033] ;
[0034] in, y i’ The real database i’ The true value of the burst pressure of the samples; For the i’ The burst pressure prediction value of samples, N is the number of samples in the real database; β is the scale factor;
[0035] The objective function is minimized through the optimization algorithm to obtain the optimized proportional factor β , used to calibrate the preliminary burst pressure prediction model:
[0036] ;
[0037] in, is the burst pressure prediction value of the burst pressure prediction model obtained preliminarily; is the corrected burst pressure prediction value, that is, the final burst pressure prediction value of the burst pressure prediction model.
[0038] Preferably, in step S5, the SHAP value of each feature is calculated using the SHAP library in the Python platform.
[0039] Preferably, in step S1, the characteristic data of the composite hydrogen storage cylinder include one or more of the head ellipsoid ratio, winding tension, winding pressure, inner liner thickness, inner liner aspect ratio, maximum curing temperature, thickness ratio of spiral winding layer to hoop winding layer, maximum winding angle of spiral winding, total winding time, and fiber volume fraction of the composite material layer.
[0040] Preferably, in step S2, robust normalization processing is performed on the feature data in the original data:
[0041] ;
[0042] Among them, X norm is the feature data after robust normalization; median( X ) is the feature data X median; IQR( X ) is the feature data X interquartile range.
[0043] The advantages of the present invention are:
[0044] (1) Considering that burst tests of hydrogen storage cylinders are time-consuming, labor-intensive, and expensive, an integrated machine learning approach can be used to analyze a small amount of test data to efficiently and accurately predict the burst pressure of the designed cylinders, accelerating the burst pressure prediction and design iteration process for composite hydrogen storage cylinders. The method of the present invention trains an integrated learning model based on a small amount of burst test data to achieve rapid and accurate prediction of the cylinder burst pressure.
[0045] (2) Based on a small amount of composite material hydrogen storage cylinder design and manufacturing related data and explosion test data, the association hidden behind the data is established to fully utilize the existing test data and quickly predict the explosion pressure, avoiding complex and distorted numerical simulation work.
[0046] (3) The key factors and data in the entire design and manufacturing process of hydrogen storage cylinders are recorded and incorporated into the database of the machine learning model to achieve accurate prediction of the bursting pressure.
[0047] (4) By integrating machine learning and cross-validation methods, the robustness and generalization ability of the model based on small sample data can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of the method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning.
[0049] Figure 2 Comparison chart of SHAP values for each feature. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0051] Depend on Figure 1 As shown, the method for predicting the bursting pressure of a composite hydrogen storage cylinder based on small sample machine learning of the present invention includes the following steps:
[0052] S1. Collect relevant data (i.e., raw data) related to the design and manufacturing of the composite hydrogen storage cylinder liner, the filament winding and curing process, and the cylinder burst test. This includes characteristic data and the corresponding actual burst pressure. Characteristic factors are factors that affect the cylinder burst strength, such as geometric characteristics, material parameters, layup scheme, and curing system.
[0053] S2. Normalize and preprocess the original data and establish a real database .
[0054] S3, based on real database Training Generative Adversarial Networks (GANs) to learn from real databases The spatial distribution characteristics of the original data are then used to generate new virtual data that is several times the amount of the original data using the trained GANs to establish a virtual database. , and have similar data distribution characteristics and statistical properties, in particular, The virtual data in the dataset can have better diversity, making up for the problem of too little single feature information in the experimental data, thereby improving the effect of model training and prediction.
[0055] S4, based on virtual database XGBoost extreme gradient boosting training was performed, and Lasso and Ridge regularization and early stopping techniques were introduced. Combined with K-fold cross-validation technology, this effectively prevented model overfitting and improved model generalization capabilities. A preliminary burst pressure prediction model for composite hydrogen storage cylinders was obtained. .
[0056] S5. Use radial basis function RBF method to adjust the adaptive scale factor, combined with the real database The initial burst pressure prediction model Correction was performed to obtain the final burst pressure prediction model , It has better prediction accuracy and generalization ability, and can quickly give accurate prediction values of burst pressure for parameter combinations within the design space of new composite hydrogen storage cylinders.
[0057] Example 1
[0058] The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning is as follows:
[0059] S11, obtaining the raw data of the composite hydrogen storage cylinder explosion test, the raw data including characteristic data and the corresponding actual explosion pressure. The characteristic data includes the head ellipsoid ratio k, winding tension F, winding pressure P, liner thickness h, liner aspect ratio L / D, maximum curing temperature T, spiral winding layer to hoop winding layer thickness ratio n, spiral winding maximum winding angle θ m , total winding time t, fiber volume fraction V of the composite material layer f Characteristic data was obtained by monitoring the winding and curing processes of hydrogen storage cylinders. The corresponding label is the burst pressure value (actual burst pressure) measured by the hydraulic burst test for each composite hydrogen storage cylinder. A total of 50 sample data were obtained.
[0060] S12, preprocess the raw data and build a real database of small samples.
[0061] Normalize the feature data to eliminate the impact of dimension and improve the convergence speed of the algorithm. In order to ensure the robustness to outliers in the data, normalize the feature data. X Perform robust normalization:
[0062] ;
[0063] Among them, X norm is the normalized feature data; median( X ) is the feature data X median; IQR( X ) is the feature data X The interquartile range refers to the difference between the third quartile and the first quartile after the data set is divided into four equal parts. It reflects the degree of dispersion of the middle 50% of the data set and is not sensitive to extreme values.
[0064] S13, based on the real database, training generative adversarial networks (GANs) to learn the spatial distribution characteristics of the original data in the real database, and then using the trained GANs to generate new virtual data that is several times the amount of original data to establish a virtual database.
[0065] Specifically, the number of virtual data generated by GANs is 10 times the number of original data, totaling 500 sample data, constituting a virtual database.
[0066] The data in the virtual database is randomly divided into K subsets. For each iteration of training and validation, one of the subsets is selected as the test set, and the other K-1 subsets are combined as the training set for subsequent K-fold cross-validation. Here, K is 10. During cross-validation, the dataset is divided into 10 equal parts, each containing 50 samples. 9 of these parts are selected as the training set, and the remaining part is used as the test set to evaluate the performance of the model.
[0067] S14, based on the virtual database, XGBoost extreme gradient boosting training was performed, Lasso and Ridge regularization and early stopping techniques were introduced, combined with K-fold cross-validation technology to effectively prevent model overfitting and improve model generalization ability. A preliminary burst pressure prediction model for composite hydrogen storage cylinders was obtained, which was recorded as the preliminary model.
[0068] Build an XGBoost extreme gradient boosting model, specifically composed of multiple decision trees. The number of decision trees is a key hyperparameter. The optimal number of decision trees can be determined through methods such as grid search or combined with the dataset size to achieve optimal model complexity and generalization. Here, a grid search is used to search for the number of decision trees within the range of 50 to 400 with a step size of 50. The optimal number of decision trees obtained by minimizing the mean squared error (MSE) under 10-fold cross-validation is 100.
[0069] After determining the number of decision trees, we set the search space for other hyperparameters, including the maximum depth of the tree (3-10), learning rate (0.01-0.3), minimum loss reduction (0-5), minimum weight of child nodes (1-10), subsampling ratio (0.5-1), column sampling ratio of each tree (0.5-1), etc. We used the Bayesian optimization tool on the Python platform to train the model for various hyperparameter combinations, and determined the optimal hyperparameter combination based on minimizing the mean square error under 10-fold cross validation.
[0070] Using the hyperparameters determined above, we conduct data learning on the training set in the virtual database. The objective function Obj is:
[0071] ;
[0072] ;
[0073] ;
[0074] in, loss is the loss function. For regression problems, the mean square error is used here to measure the gap between the model prediction value and the actual value; y i For the virtual database i The burst pressure value of each sample, For the i The predicted burst pressure of samples; i 、 n are the sample number and the total number of samples respectively. It is a regularization term, including Lasso regularization and Ridge regularization. Lasso regularization adds weights to the objective function. w j Absolute value penalty coefficient α Realize feature selection and sparsity to reduce the weight of some unimportant features; Ridge regularization uses penalty coefficients l Limit the weight of the entire model w j Size, to prevent the model from being overly dependent on individual features; hyperparameters c Used to control the strength of the regularization term.
[0075] Because the amount of training data for small-sample machine learning is small, it cannot fully represent the data distribution within the entire design range of composite hydrogen storage cylinders. The actual cylinder data set may contain certain features that have little effect on the burst pressure. In addition, the model is repeatedly iteratively trained on the small sample data set and thus learns the noise in the data. Many factors may cause the model to overfit and have poor generalization ability. That is, the prediction on the small sample data set is very accurate, but it cannot accurately predict the burst pressure of a newly designed cylinder. Regularization term The introduction of controls the complexity of the model, prevents the model from overfitting, and helps the model to accurately predict the bursting pressure of newly designed gas cylinders.
[0076] Using a virtual database, the XGBoost extreme gradient boosting model was trained for regression, yielding a preliminary burst pressure prediction model (preliminary model). Two types of regularization and K-fold cross-validation were used to effectively mitigate the risk of model overfitting. In K-fold cross-validation, the average mean squared error (MSE) of K validations was used to evaluate overall model performance. Early stopping was used to further prevent overfitting. The parameter T (number of iterations) was initially set to 0. If model performance on the test set did not improve, training was continued with T+1. If performance improved, T was reset to 0 and the model parameters were updated. Training was stopped when T reached a preset value (e.g., 10). The trained preliminary model was saved.
[0077] By using mean square error MSE, root mean square error RMSE, mean absolute error MAE and R 2 Four indicators are used to evaluate the prediction accuracy of the preliminary model from different perspectives. After training the XGBoost extreme gradient boosting on the virtual database, the MSE and RMSE of the preliminary model are less than 0.01, MAE is less than 0.03, and R 2 If it is greater than 0.85, the model prediction accuracy is high.
[0078] The SHAP library is used in the Python platform to calculate the SHAP value (importance value) of each feature. The SHAP value is used to quantify the contribution of each feature to the model prediction and is visualized through SHAP summary diagrams, such as Figure 2 As shown in the figure, it is found that the influence degree on the blasting pressure of composite materials is V from high to low. f ,n,F,θ m , k, T, P, t, h and L / D, where V f , n, and F have a greater impact on the bursting pressure, while h and L / D have almost no effect on the bursting pressure.
[0079] S15, using the radial basis function RBF method to perform adaptive scaling factor adjustment, input sample x From the real database, calculate the RBF kernel function value of each sample F ( x ):
[0080] ;
[0081] in, F is the RBF kernel function; c represents the center point in the real database, which is the mean value of the real database; s is the width parameter;
[0082] ;
[0083] in, R The data range is calculated by normalizing the SHAP value of each feature as the weight:
[0084] ;
[0085] in, R k For the k The data range of the feature, that is, k The difference between the maximum and minimum values of a feature; m k For the k The weight of each feature, which is normalized by the feature SHAP value; k 、 M are the feature number and the total number of features respectively.
[0086] The objective function is the mean square error MSE :
[0087] ;
[0088] in, N is the number of samples in the real database; β is the scale factor; y i’ The real database i’ The true value of the burst pressure of the samples; For the i’ The burst pressure prediction value of each sample.
[0089] Minimize the objective function through optimization algorithms such as gradient descent to obtain the optimized proportional factor β Burst pressure predictions used to calibrate the preliminary model:
[0090] ;
[0091] in, is the burst pressure prediction value of the preliminary model; is the corrected burst pressure prediction value, i.e., the final burst pressure prediction value of the burst pressure prediction model;
[0092] The initial burst pressure prediction model was calibrated using a real-world database to obtain a final model. This model exhibits improved prediction accuracy and generalization, enabling rapid and accurate predictions of burst pressure for new parameter combinations within the design space of composite hydrogen storage cylinders. As the composite hydrogen storage cylinder tests iterate, new test data is dynamically added to the model for further improved prediction performance.
[0093] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning, characterized in that: The following steps are involved: S1, obtaining the original data of the explosion test of the composite hydrogen storage cylinder, the original data including characteristic data and the corresponding actual explosion pressure; S2, preprocessing the original data to form a real database; S3: Train a generative adversarial network based on a real database, and use the trained generative adversarial network to generate virtual data several times the amount of original data to establish a virtual database; S4, based on the virtual database, XGBoost extreme gradient boosting training is performed to preliminarily obtain the bursting pressure prediction model; S5, correcting the preliminary burst pressure prediction model based on the real database to obtain the final burst pressure prediction model; In step S5, the radial basis function (RBF) method is used to perform adaptive scaling factor adjustment, and the initially obtained burst pressure prediction model is corrected in combination with the real database; The specific method of step S5 is as follows: Calculate each sample in the real database x RBF kernel function value: ; in, is the RBF kernel function; c represents the center point in the real database, which is the mean value of the real database; σ is the width parameter; ; in, R is the data range; ; in, R k For the k The data range of the feature, that is, k The difference between the maximum and minimum values of a feature; m k For the k The weight of each feature is normalized by the importance value of each feature, i.e., the SHAP value, as the corresponding weight; k 、 M are the feature number and the total number of features respectively; The objective function is the mean square error MSE : ; in, y i’ The real database i’ The true value of the burst pressure of the samples; For the i’ The burst pressure prediction value of samples, N is the number of samples in the real database; β is the scale factor; The objective function is minimized through the optimization algorithm to obtain the optimized proportional factor β , used to calibrate the preliminary burst pressure prediction model: ; in, is the burst pressure prediction value of the burst pressure prediction model obtained preliminarily; is the corrected burst pressure prediction value, i.e., the final burst pressure prediction value of the burst pressure prediction model; In step S1, the characteristic data of the composite hydrogen storage cylinder includes one or more of the following: head ellipsoid ratio, winding tension, winding pressure, liner thickness, liner aspect ratio, maximum curing temperature, thickness ratio of spiral winding layer to hoop winding layer, maximum winding angle of spiral winding, total winding time, and fiber volume fraction of the composite material layer; Collect relevant data, namely original data, on the design and manufacturing of the inner liner of the composite hydrogen storage cylinder, the fiber winding and curing process, and the cylinder explosion test.
2. The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning according to claim 1 is characterized in that: In step S4, Lasso and Ridge regularization are introduced, and combined with K-fold cross-validation technology, XGBoost extreme gradient boosting training of the model is performed.
3. The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning according to claim 2 is characterized in that: The specific method of step S4 is as follows: Build an XGBoost extreme gradient boosting model. The model is composed of multiple decision trees. The number of decision trees is used as a hyperparameter. The optimal number of decision trees is determined through grid search. The optimal number of decision trees is obtained by minimizing the mean square error under K-fold cross validation. After determining the optimal number of decision trees, set the search space for other hyperparameters, use the Bayesian optimization method to optimize various hyperparameter combinations, and obtain the optimal hyperparameter combination by minimizing the mean square error under K-fold cross validation; Based on the determined optimal hyperparameter combination, the training set in the virtual database is learned, and the objective function Obj of the model training is: ; ; ; in, loss is the loss function; y i For the virtual database i The burst pressure value of each sample, For the i The predicted burst pressure of samples; i 、 n are the sample number and the total number of samples in the training set respectively; It is a regularization term, including Lasso regularization and Ridge regularization. Lasso regularization adds weights to the objective function. w j Absolute value penalty coefficient α Implement feature selection and sparsity; Ridge regularization through penalty coefficient λ Limit the weight of the entire model w j Size; Hyperparameters γ Used to control the strength of the regularization term.
4. The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning according to claim 2 or 3, characterized in that: The data in the virtual database is randomly divided into K subsets. For each iterative training and testing, one of the subsets is selected as the test set, and the other K-1 subsets are combined as the training set for K-fold cross-validation.
5. The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning according to claim 2 or 3, characterized in that: During the XGBoost extreme gradient boosting training process of the model, an early stopping technique is also introduced. Specifically, the initial value of the parameter T is set. If the performance of the trained model on the test set does not improve, the value of T is increased by 1 and the model training continues; if the performance improves, T is reset and the model parameters are updated; when T reaches the preset value, the model training is stopped.
6. The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning according to claim 1 is characterized in that: In step S5, the SHAP value of each feature is calculated using the SHAP library in the Python platform.
7. The method for predicting the burst pressure of composite hydrogen storage cylinders based on small sample machine learning according to claim 1 is characterized in that: In step S2, robust normalization is performed on the feature data in the original data: ; Among them, X norm is the feature data after robust normalization; median( X ) is the feature data X median; IQR( X ) is the feature data X interquartile range.
Citation Information
Patent Citations
Carbon fiber fully-wound gas cylinder bursting pressure prediction method based on machine vision and deep learning
CN118332851A
Automatic data analysis system for water pressure blasting
CN118863811A