Artificial intelligence grading classification-regression combination prediction method for multi-class materials
By employing an AI-based hierarchical classification-regression combined prediction method for multiple material categories, and using interval partitioning and model combination strategies, this approach addresses the shortcomings of traditional models in terms of fitting accuracy and generalization ability across multiple material categories, achieving higher accuracy and robustness in bandgap prediction.
Patent Information
- Application Number
- CN202511269720.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Traditional single regression models struggle to balance overall fitting accuracy and generalization ability across multiple data categories, especially in complex datasets with multimodal distributions and significant heteroscedasticity, where they are prone to systematic biases at interval boundaries and in sparse regions.
An AI-based classification-regression combined prediction method for multiple material categories is adopted. This method involves interval partitioning, feature extraction, and a combination of classification and regression models, including SVM classification models and multi-layer neural network regression models. Incremental learning optimization is used to implement a strategy of classification before regression, thereby improving the model's fitting ability and generalization.
It significantly improves the accuracy and robustness of material bandgap prediction, especially when the distribution of high and low bandgap samples is uneven. It can more accurately predict samples in different intervals, providing good generalization and interpretability.
Smart Images

Figure CN121171399A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of materials informatics and artificial intelligence, and in particular to an artificial intelligence-based hierarchical classification-regression combination prediction method for multiple categories of materials. Background Technology
[0002] The band gap is a core indicator for measuring the electronic structure and functional properties of materials. However, due to the highly nonlinear influence of element combination, crystal structure and local environment, traditional single regression models often struggle to balance global fitting accuracy and generalization ability.
[0003] In complex datasets with diverse distributions, multimodal distributions, and significant heteroscedasticity, uniformly trained global regression networks are prone to systematic biases at interval boundaries and in sparse regions. Therefore, a hybrid learning framework that simultaneously addresses both global separability and local fitability is needed to improve the accuracy and stability of bandgap prediction. Summary of the Invention
[0004] The purpose of this invention is to provide an artificial intelligence-based classification-regression combination prediction method for multiple types of materials, which can improve the accuracy and robustness of material bandgap prediction.
[0005] To achieve the above objectives, this invention provides an artificial intelligence-based classification-regression combination prediction method for multiple categories of materials, comprising the following steps:
[0006] Step S1: Obtain a historical sample dataset containing information on the chemical composition and crystal structure of the material. Divide the band gap value into multiple non-overlapping intervals based on the principle of equal width, and generate a corresponding interval label for each sample.
[0007] Step S2: Extract elemental attribute statistical features and crystal geometric features from chemical composition and crystal structure information, and obtain the key feature set after removing redundant features;
[0008] Step S3: Train an SVM classification model based on the key feature set and the corresponding interval labels to predict the band gap interval of the test sample and output the class label and the confidence score of each interval.
[0009] Step S4: Divide the historical sample dataset into multiple subsets according to interval labels, and train a multi-layer neural network regression model independently for each interval subset;
[0010] Step S5: For the sample to be predicted, call the classification model to obtain its band gap interval and confidence level; if the highest confidence level is higher than the preset threshold, call the multi-layer neural network regression model of the band gap interval to output the predicted value; if the highest confidence level is lower than the preset threshold, perform parallel regression of multiple intervals and perform weighted fusion according to the confidence level of each band gap interval.
[0011] Step S6: Add the misclassified samples or samples with regression errors exceeding the set threshold generated during the prediction process to the training set, and retrain the SVM classification model - multilayer neural network regression model.
[0012] Preferably, the interval division in step S1 is as follows: the samples with a band gap value of 0 eV are divided into the first interval, the samples with a band gap value of (0, 1] eV are divided into the second interval, the samples with a band gap value of (1, 2] eV are divided into the third interval, the samples with a band gap value of (2, 3] eV are divided into the fourth interval, and the samples with a band gap value greater than 3 eV are divided into the fifth interval.
[0013] Preferably, in step S1, each sample in the historical sample dataset includes material chemical composition and crystal structure information. The material chemical composition includes chemical formula, element types and atomic ratios, and the crystal structure information includes lattice constant, volume, space group number, symmetry parameter and coordination number.
[0014] Preferably, the elemental property statistical characteristics in step S2 include: average atomic number, number of valence electrons, electronegativity, atomic radius, first ionization energy, and atomic volume; the crystal geometric characteristics include: lattice parameters, unit cell volume, space group number, symmetry parameters, and coordination environment.
[0015] Preferably, the method for removing redundant features in step S2 is to calculate the Pearson correlation coefficient between features and remove one of the features with a Pearson correlation coefficient greater than 0.95.
[0016] Preferably, the kernel function of the SVM classification model in step S3 is a radial basis function, with hyperparameters C = 10 and γ = 0.01.
[0017] Preferably, the structure of the multilayer neural network model in step S4 is as follows: input layer, hidden layer 1, hidden layer 2, hidden layer 3, and output layer; wherein, hidden layer 1 contains 128 neurons, hidden layer 2 contains 64 neurons, hidden layer 3 contains 32 neurons, all hidden layers use the ReLU activation function, and the output layer contains 1 neuron and uses the linear activation function; the loss function during training of the multilayer neural network model is mean squared error, the optimizer is Adam, the learning rate is set to 0.001, the batch size is set to 128, and the maximum number of iterations is set to 5000.
[0018] Preferably, the preset threshold in step S5 is 0.6; the specific formula for weighted fusion is:
[0019]
[0020] in, p represents the final fused predicted value, K represents the number of intervals, and p represents the number of intervals. i This represents the predicted probability for the i-th interval. This represents the bandgap prediction value output by the neural network model for the i-th interval.
[0021] Preferably, the incremental learning optimization in step S6 specifically includes:
[0022] Error sample collection: During the prediction process, misclassified samples and samples whose regression prediction errors exceed the set threshold are collected and stored in the feedback database;
[0023] Optimization trigger conditions: When the number of new samples in the feedback database exceeds the preset number, or when a preset time interval has elapsed since the last optimization, the hyperparameter optimization process will be automatically triggered.
[0024] Joint hyperparameter optimization: Using a Bayesian optimization framework, the hyperparameters of the SVM classification model and the multilayer neural network regression model in each interval are jointly adjusted with the goal of minimizing the cross-validation mean absolute error or maximizing the classification accuracy and F1-score.
[0025] Bayesian optimization uses a Gaussian process as a surrogate model and selects the parameter set to be evaluated based on the desired improvement as the sampling strategy.
[0026] The optimized hyperparameters include: the penalty coefficient C and kernel width γ of the SVM classification model, and the learning rate, number of hidden layer neurons, and L2 regularization coefficient of the neural network regression model.
[0027] Model Update: After the optimization process is completed, the SVM classification model and the multi-layer neural network regression model for each interval are retrained using the optimal hyperparameter combination to update the model performance.
[0028] Therefore, this invention employs the aforementioned AI-based classification-regression combined prediction method for multi-category materials. Based on a "classification first, regression later" strategy, it significantly improves the model's fitting ability to samples in different bandgap ranges, especially when faced with uneven distribution of high and low bandgap samples. Furthermore, this method possesses good generalization and interpretability, providing a more accurate and reliable solution for material bandgap prediction. Attached Figure Description
[0029] Figure 1 This is a flowchart of an artificial intelligence-based classification-regression combined prediction method for multiple types of materials according to the present invention;
[0030] Figure 2 This is a schematic diagram illustrating the bandgap region division of the material of the present invention;
[0031] Figure 3 A comparison chart of the prediction performance of neural networks without interval partitioning;
[0032] Figure 4This is a comparison chart of the predictive performance of the regression models for each interval after the band gap is divided. Detailed Implementation
[0033] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0034] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0035] Example 1
[0036] The material structure information used in this embodiment comes from the CIF file of the crystalline material. The dataset contains 30,000 historical sample data with known band gap values, sourced from the MaterialProjects database. Each historical sample data includes the material's chemical composition (chemical formula, element type, atomic ratio) and crystal structure information (lattice constant, volume, space group number, symmetry parameter, coordination number, etc.). The CIF file is parsed and feature-engineered using Python's Matminer package to extract elemental attribute statistical features (including average atomic number, valence electron number, electronegativity, atomic radius, first ionization energy, atomic volume) and crystal geometric features (including lattice parameters, unit cell volume, space group number, symmetry parameter, coordination environment). Statistical methods (such as mean, standard deviation, range, maximum, minimum, median, etc.) are used to summarize the statistical features of each element in the dataset, generating a feature set representing the entire chemical formula, totaling hundreds of original features. Data cleaning is then performed to remove missing values and physically abnormal samples.
[0037] like Figure 1 As shown, an artificial intelligence-based classification-regression combined prediction method for multi-category materials includes the following steps:
[0038] Step S1: Interval partitioning.
[0039] like Figure 2 As shown, based on the bandgap value range, all samples are divided into 5 non-overlapping intervals:
[0040] The samples with a band gap of 0 eV were divided into the first interval, the samples with a band gap of (0, 1] eV were divided into the second interval, the samples with a band gap of (1, 2] eV were divided into the third interval, the samples with a band gap of (2, 3] eV were divided into the fourth interval, and the samples with a band gap greater than 3 eV were divided into the fifth interval.
[0041] The interval division adopts the principle of equal width to make the number of samples in each interval roughly balanced, thereby reducing the performance degradation of the model caused by class imbalance.
[0042] Step S2: Feature extraction.
[0043] All features were standardized, and redundant features with a Pearson correlation coefficient greater than 0.95 were removed to reduce multicollinearity among features. Principal component analysis (PCA) was then used to reduce the dimensionality of the high-dimensional feature space, retaining 95% of the cumulative variance, ultimately reducing the feature dimension to approximately 50 principal components. The retained principal components after dimensionality reduction are linear combinations of the original multidimensional physical features, no longer possessing a single physical meaning, but retaining the main information reflecting the material composition and structural variability in the original data. This effectively reduces model complexity, suppresses overfitting, and improves learning efficiency.
[0044] Step S3: Interval classification.
[0045] A Support Vector Machine (SVM) classification model was trained using the feature vectors after dimensionality reduction via PCA. The kernel function was a radial basis function (RBF), and the hyperparameters were selected as C=10 and γ=0.01 through five-fold cross-validation. As shown in Table 1, on the test set, the SVM classification model achieved a bandgap interval classification accuracy of approximately 65%, which is better than Random Forest (approximately 60%) and Gradient Boosting Tree (approximately 58%).
[0046] Table 1 Test Results
[0047] Model Name Classification accuracy Cross-validation average F1 SVM 65.0% 0.631 Random Forest 60.1% 0.598 Gradient boosting tree 56.3% 0.581
[0048] Step S4: Interval regression.
[0049] The training set was split into 5 subsets according to interval labels, and independent multilayer neural network regression models were trained for each subset.
[0050] The network structure is as follows: Input layer (corresponding to the number of features after PCA dimensionality reduction) — Hidden layer 1 (128 neurons, ReLU activation) — Hidden layer 2 (64 neurons, ReLU activation) — Hidden layer 3 (32 neurons, ReLU activation) — Output layer (1 neuron, linear activation). The loss function is mean squared error (MSE), the optimizer is Adam (learning rate 0.001, batch size 128), the maximum number of iterations is 5000, and the model converges in approximately 600 iterations.
[0051] Step S5: Combined reasoning.
[0052] In the prediction phase, the SVM classification model is first used to predict the band gap interval labels. When the classification confidence is ≥0.6, the neural network regression model of the corresponding interval is directly called to output the accurate band gap value. When the confidence is <0.6, multi-interval parallel regression is triggered, and weighted fusion is performed according to the confidence of each interval to improve the prediction stability.
[0053] The specific formula for weighted fusion is:
[0054]
[0055] in, p represents the final fused predicted value, K represents the number of intervals, which is 5 in this invention; i This represents the predicted probability of the i-th interval, expressed as the confidence level output by the SVM. This represents the bandgap prediction value output by the neural network model for the i-th interval.
[0056] In the actual prediction phase, if the classifier's interval judgment of the sample does not have sufficient confidence (e.g., less than 0.6), the system will call all (or several high-confidence) neural network regression models in parallel and output the bandgap prediction values respectively. And use the interval probability p corresponding to the sample i As a weighting factor, calculate the final fused output.
[0057] Step S6: Performance evaluation and feedback optimization.
[0058] like Figure 3 As shown, the R-value of the global multilayer neural network model without interval partitioning on the test set is... 2 It is only about 0.45, and the MAE is about 0.60 eV.
[0059] like Figure 4 As shown, after adopting the "classification first, regression later" method in this embodiment, the average R-squared value of the five interval regression models is... 2 The MAE increased to approximately 0.60, while the average MAE decreased to approximately 0.40 eV. The MAE distribution across all intervals ranged from 0.38 to 0.42 eV. R 2 It is distributed between 0.58 and 0.62.
[0060] After prediction is completed, the system will automatically record misclassified samples (incorrectly labeled predicted intervals) and high-error samples (bandgap prediction error MAE > 0.6 eV) to the feedback database. When the set conditions are met—such as the number of newly added error samples exceeding 500, or the time interval exceeding 7 days—the system will automatically trigger the Bayesian optimization process to jointly adjust the hyperparameters of the classification model and each interval regression model.
[0061] The optimization objective is to minimize the mean absolute error (MAE) under cross-validation, or to maximize classification accuracy and the F1 score. The optimization method is based on a Bayesian optimization framework, using a Gaussian process to construct a surrogate model and fitting the objective function with the current parameters and historical evaluation values. By using expected improvement as a sampling strategy, the system selects the next set of parameters to be evaluated in each round, gradually approaching the optimal solution.
[0062] Typical optimization parameters include the C-value and kernel width γ in support vector machines, and the learning rate, hidden layer size, and L2 regularization coefficient in neural networks. The entire optimization process is usually set to no more than 30 rounds, with the objective function value evaluated in each round based on 5-fold cross-validation. After optimization, the system saves the optimal parameter combination and retrains the model, completing a performance self-update to ensure that the model's prediction accuracy on newly added material data continues to improve.
[0063] An AI-based classification-regression combined prediction system for multi-category materials includes:
[0064] Data preprocessing module: used to acquire historical sample datasets containing information on the chemical composition and crystal structure of materials, and divide the band gap values into multiple non-overlapping intervals according to the principle of equal width, generating corresponding interval labels for each sample; at the same time, extracting elemental attribute statistical features and crystal geometric features from the chemical composition and crystal structure information, and obtaining the key feature set after removing redundant features;
[0065] Classification model training module: Based on the key feature set and the corresponding interval labels, train the SVM classification model to predict the band gap interval of the test sample, and output the class label and the confidence of each band gap interval;
[0066] Regression model training module: Divide the historical sample dataset into multiple subsets according to interval labels, and train a multi-layer neural network regression model independently for each interval subset;
[0067] Prediction module: For a sample to be predicted, the classification model is called to obtain its band gap interval and confidence level; if the highest confidence level is higher than the preset threshold, the multi-layer neural network regression model of the band gap interval is called to output the predicted value; if the highest confidence level is lower than the preset threshold, multi-interval parallel regression is performed, and weighted fusion is performed according to the confidence level of each band gap interval.
[0068] Incremental learning optimization module: Adds misclassified samples or samples with regression errors exceeding a set threshold generated during the prediction process to the training set, and retrains the SVM classification model and the multilayer neural network regression model.
[0069] It is worth noting that all contents not described in detail in this invention are existing technologies and are well known to those skilled in the art.
[0070] Therefore, the present invention employs the above-mentioned artificial intelligence classification-regression combination prediction method and system for multiple types of materials, which can effectively improve the accuracy and robustness of bandgap prediction, especially when the distribution of high and low bandgap samples is uneven, providing an accurate and reliable solution for material bandgap prediction.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-class material oriented artificial intelligence hierarchical classification-regression combined prediction method, characterized in that, Includes the following steps: Step S1: Obtain a historical sample dataset containing information on the chemical composition and crystal structure of the material. Divide the band gap value into multiple non-overlapping intervals based on the principle of equal width, and generate a corresponding interval label for each sample. Step S2: Extract elemental attribute statistical features and crystal geometric features from chemical composition and crystal structure information, and obtain the key feature set after removing redundant features; Step S3: Train an SVM classification model based on the key feature set and the corresponding interval labels to predict the band gap interval of the test sample and output the class label and the confidence score of each interval. Step S4: Divide the historical sample dataset into multiple subsets according to interval labels, and train a multi-layer neural network regression model independently for each interval subset; Step S5: For the sample to be predicted, call the classification model to obtain its band gap interval and confidence level; if the highest confidence level is higher than the preset threshold, call the multilayer neural network regression model of the band gap interval to output the predicted value. If the highest confidence level is lower than the preset threshold, perform parallel regression across multiple intervals and weighted fusion based on the confidence level of each band gap interval; Step S6: Add the misclassified samples or samples with regression errors exceeding the set threshold generated during the prediction process to the training set, and retrain the SVM classification model - multilayer neural network regression model.
2. The artificial intelligence hierarchical classification-regression combined prediction method for multi-class materials according to claim 1, characterized in that, The specific interval division in step S1 is as follows: the samples with a band gap value of 0 eV are divided into the first interval, the samples with a band gap value of (0, 1] eV are divided into the second interval, the samples with a band gap value of (1, 2] eV are divided into the third interval, the samples with a band gap value of (2, 3] eV are divided into the fourth interval, and the samples with a band gap value greater than 3 eV are divided into the fifth interval.
3. The artificial intelligence hierarchical classification-regression combined prediction method for multi-class materials according to claim 1, characterized in that, In step S1, each sample in the historical sample dataset includes material chemical composition and crystal structure information. The material chemical composition includes chemical formula, element types and atomic ratios, and the crystal structure information includes lattice constant, volume, space group number, symmetry parameter and coordination number.
4. The artificial intelligence hierarchical classification-regression combined prediction method for multi-class materials according to claim 1, characterized in that, The elemental property statistical characteristics in step S2 include: average atomic number, number of valence electrons, electronegativity, atomic radius, first ionization energy, and atomic volume; the crystal geometric characteristics include: lattice parameters, unit cell volume, space group number, symmetry parameters, and coordination environment.
5. The artificial intelligence hierarchical classification-regression combined prediction method for multi-class materials according to claim 1, characterized in that, The method for removing redundant features in step S2 is to calculate the Pearson correlation coefficient between features and remove one of the features with a Pearson correlation coefficient greater than 0.
95.
6. The artificial intelligence hierarchical classification-regression combined prediction method for multi-class materials according to claim 1, characterized in that, The kernel function of the SVM classification model in step S3 is the radial basis function, with hyperparameters C = 10 and γ = 0.
01.
7. The artificial intelligence-based classification-regression combined prediction method for multi-category materials according to claim 1, characterized in that, The structure of the multilayer neural network model in step S4 is as follows: input layer, hidden layer 1, hidden layer 2, hidden layer 3, and output layer; wherein, hidden layer 1 contains 128 neurons, hidden layer 2 contains 64 neurons, hidden layer 3 contains 32 neurons, and all hidden layers use the ReLU activation function; the output layer contains 1 neuron and uses the linear activation function; the loss function during training of the multilayer neural network model is mean squared error, the optimizer is Adam, the learning rate is set to 0.001, the batch size is set to 128, and the maximum number of iterations is set to 5000.
8. The artificial intelligence-based classification-regression combined prediction method for multi-category materials according to claim 1, characterized in that, The preset threshold in step S5 is 0.6; the specific formula for weighted fusion is: in, p represents the final fused predicted value, K represents the number of intervals, and p represents the number of intervals. i This represents the predicted probability for the i-th interval. This represents the bandgap prediction value output by the neural network model for the i-th interval.
9. The artificial intelligence-based classification-regression combined prediction method for multi-category materials according to claim 1, characterized in that, The incremental learning optimization in step S6 is specifically as follows: During the prediction process, misclassified samples and samples with regression prediction errors exceeding a set threshold are collected and stored in the feedback database. When the number of new samples in the feedback database exceeds the preset number, or when a preset time interval has elapsed since the last optimization, the hyperparameter optimization process will be automatically triggered. A Bayesian optimization framework is adopted to jointly adjust the hyperparameters of the SVM classification model and the multilayer neural network regression model in each interval, with the goal of minimizing the cross-validation mean absolute error or maximizing the classification accuracy and F1-score. Bayesian optimization uses a Gaussian process as a surrogate model and selects the parameter set to be evaluated based on the desired improvement as the sampling strategy. The optimized hyperparameters include: the penalty coefficient C and kernel width γ of the SVM classification model, and the learning rate, number of hidden layer neurons, and L2 regularization coefficient of the neural network regression model. After the optimization process is completed, the SVM classification model and the multilayer neural network regression model for each interval are retrained using the optimal hyperparameter combination to update the model performance.
Citation Information
Patent Citations
Construction method of TiH binary system machine learning potential
CN119578586A
One-dimensional single-port phonon crystal structure for enhancing acoustic signal
CN120071886A
Crystal structure condition generation method combining large model and diffusion model
CN120412828A
Solid electrolyte intelligent inverse design method fusing graph neural network and confidence analysis
CN120452642A
Machine learning based methods of analysing drug-like molecules
US20220383992A1
Cited By
Fluorescence in-situ hybridization image cross-interval accurate counting method and application thereof
CN122157257A
Precise Counting Method Across Intervals in Fluorescence In Situ Hybridization Images and Its Applications
CN122157257B