A method for machine learning prediction of triglyceride heat capacity combining sparse feature selection
By employing a two-step sparse feature selection and XGBoost model, the problems of data sparsity and high dimensionality in the prediction of triglyceride molecular heat capacity are solved, achieving high-precision and high-efficiency prediction, simplifying feature engineering, and providing real-time prediction and high-throughput screening capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUODIAN NANNING POWER GENERATION CO LTD
- Filing Date
- 2026-04-03
- Publication Date
- 2026-07-24
AI Technical Summary
In existing technologies, the prediction of the heat capacity of triglyceride molecules suffers from sparse experimental data and low efficiency. Traditional theoretical models rely on 'prior' parameters and lack accuracy. QSPR modeling faces the 'curse of dimensionality'. Traditional feature selection strategies cannot effectively handle high-dimensional data, resulting in poor model generalization ability and high computational cost.
A two-step sparse feature selection strategy is adopted, which combines the Pearson correlation coefficient method and the recursive feature elimination method to screen out sparse key features from the high-dimensional feature matrix and construct an XGBoost model for high-precision prediction, including data preparation, feature selection and model optimization.
It achieves a combination of high accuracy and high efficiency, with a model determination coefficient as high as 0.9880, a root mean square error of 44.99 J·mol⁻¹·K⁻¹, and a mean absolute error of 30.07 J·mol⁻¹·K⁻¹. It solves the problems of 'curse of dimensionality' and overfitting, and provides real-time prediction and high-throughput screening capabilities.
Smart Images

Figure CN122455165A_ABST