A method for machine learning prediction of triglyceride heat capacity combining sparse feature selection

By employing a two-step sparse feature selection and XGBoost model, the problems of data sparsity and high dimensionality in the prediction of triglyceride molecular heat capacity are solved, achieving high-precision and high-efficiency prediction, simplifying feature engineering, and providing real-time prediction and high-throughput screening capabilities.

CN122455165APending Publication Date: 2026-07-24GUODIAN NANNING POWER GENERATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUODIAN NANNING POWER GENERATION CO LTD
Filing Date
2026-04-03
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, the prediction of the heat capacity of triglyceride molecules suffers from sparse experimental data and low efficiency. Traditional theoretical models rely on 'prior' parameters and lack accuracy. QSPR modeling faces the 'curse of dimensionality'. Traditional feature selection strategies cannot effectively handle high-dimensional data, resulting in poor model generalization ability and high computational cost.

Method used

A two-step sparse feature selection strategy is adopted, which combines the Pearson correlation coefficient method and the recursive feature elimination method to screen out sparse key features from the high-dimensional feature matrix and construct an XGBoost model for high-precision prediction, including data preparation, feature selection and model optimization.

Benefits of technology

It achieves a combination of high accuracy and high efficiency, with a model determination coefficient as high as 0.9880, a root mean square error of 44.99 J·mol⁻¹·K⁻¹, and a mean absolute error of 30.07 J·mol⁻¹·K⁻¹. It solves the problems of 'curse of dimensionality' and overfitting, and provides real-time prediction and high-throughput screening capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122455165A_ABST
    Figure CN122455165A_ABST
Patent Text Reader

Abstract

The application discloses a kind of triglyceride heat capacity machine learning prediction methods combined with sparse feature selection, it is related to material genetic engineering and artificial intelligence cross technical field, including the following steps: step one: data preparation and initial high-dimensional feature matrix construction: temperature is as a physical variable, with the high-dimensional molecular descriptor vector calculated and merged, jointly constitute an initial feature matrix, as the original input of subsequent all processing steps;Step two: two-step sparse feature selection strategy is used to filter out the sparse key feature combination most relevant to heat capacity property and the lowest information redundancy from the initial high-dimensional feature matrix constructed in the previous step;Step three: construction and optimization of high-precision prediction model;Step four: systematic evaluation of model performance;Step five: prediction application.The present application solves the technical problems of poor model generalization ability, insufficient precision and sharp decline in prediction ability outside the training set in the prior art.
Need to check novelty before this filing date? Find Prior Art