A method for improving material property prediction of a multilayer perceptron model

CN122822150APending Publication Date: 2026-09-25PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610786090.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

作为第三代太阳能电池的吸光层主体材料,卤化物钙钛矿具有优异的光伏性能,但其庞大的化学元素替换空间使得通过传统实验或计算模拟的手段筛选新材料成本高昂且耗时长

Benefits of technology

1.更强的通用学习能力:通过在残差连接中引入可学习的标量缩放因子,使网络能够自适应地调整不同层次特征传递的贡献度,增强了模型在神经网络中保留有效信息和学习复杂非线性关系的能力,缓解了传统MLP中的梯度消失问题,特别适合处理材料性质预测中复杂的构效关系。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822150A_ABST
    Figure CN122822150A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of material property prediction method for improving multilayer perceptron model.The present application constructs a residual connection multilayer perceptron model added residual connection structure, and the characteristic vector of material is learned and mapped by the model, and finally the prediction value of material target property is output.The method introduces learnable scalar scaling factor in residual connection, so that the network can adaptively adjust the contribution degree of feature transmission at different levels, enhance the ability of model to retain effective information and learn complex nonlinear relationship in neural network, alleviate the gradient vanishing problem in traditional MLP, especially suitable for processing complex structure-activity relationship in material property prediction.The method provides an efficient solution for high-throughput prediction and screening of new materials, especially halide perovskite photovoltaic material performance, and accelerates the discovery process of new materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of materials science and machine learning, specifically to an improved method for predicting material properties using a multilayer perceptron model. Background Technology

[0002] In materials science, developing efficient and accurate machine learning methods is crucial for predicting the properties of new materials and accelerating their discovery. However, in practical applications, materials datasets generally suffer from core challenges such as limited size, high feature dimensionality, and strong nonlinear structure-property relationships. Limited data volume makes it difficult to fully train complex models, high-dimensional features are prone to overfitting with limited samples, and complex nonlinear relationships place higher demands on the model's representation learning capabilities. This situation leads to traditional lightweight neural network models such as Multi-layer Perceptrons (MLPs) having inherent limitations in their model structure (e.g., low gradient propagation efficiency and weak feature reuse capabilities), making it difficult for them to outperform carefully optimized ensemble learning algorithms (such as XGBoost) in predictive performance, thus limiting their learning capabilities.

[0003] The aforementioned problems are particularly prominent in the development of perovskite photovoltaic materials. As the main material for the light-absorbing layer of third-generation solar cells, halide perovskites possess excellent photovoltaic performance. However, their vast space for chemical element substitution makes screening new materials through traditional experiments or computational simulations costly and time-consuming. Although existing research has attempted to apply various machine learning models, including traditional MLPs, to predict key material properties, typical preliminary work shows that traditional MLPs, due to insufficient feature learning capabilities, have failed to effectively address the complex structure-property relationships and the easy loss of key features, resulting in significant room for improvement in prediction accuracy. Sometimes, their accuracy is even lower than that of algorithms such as XGBoost, which limits their practical application value in the discovery of high-performance materials. Therefore, it is urgent to develop a targeted improved MLP model to enhance its robustness and accuracy in typical scenarios for predicting material properties. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention aims to provide an improved method for predicting material properties using a multilayer perceptron (MLP) model. This method introduces a learnable scalar scaling factor into the residual connections, enabling the network to adaptively adjust the contribution of feature transmission at different levels. This enhances the model's ability to retain effective information and learn complex nonlinear relationships within the neural network, mitigating the gradient vanishing problem in traditional MLPs. It is particularly suitable for handling complex structure-property relationships in material property prediction. This method provides an efficient solution for high-throughput prediction and screening of the performance of novel materials, especially halide perovskite photovoltaic materials, accelerating the discovery process of new materials.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for predicting material properties using an improved multilayer perceptron model, characterized by the following steps: Step 1: Obtain the material property feature dataset; classify and process the numerical features with missing values, including performing interpolation imputation operation on missing values ​​that have interpolation prerequisites, and performing preset backup missing value processing operation on missing values ​​that cannot be interpolated and imputed.

[0006] Step 2: Randomly divide the feature dataset processed in Step 1 into a training set, a validation set, and a test set; For example, the dataset can be randomly divided into training, validation, and test sets in an appropriate ratio (e.g., 80%:10%:10%) for model training, tuning, and final performance evaluation. Step 3: Introduce the residual connection structure into the multilayer perceptron to construct a residual connection multilayer perceptron model; Step 4: Input the training set described in Step 2 into the residual connected multilayer perceptron model obtained in Step 3 for multiple rounds of iterative training; during each round of iterative training, the residual connected multilayer perceptron model calculates the predicted value through forward propagation; L1 loss is used as the loss function to measure the prediction error, and the model parameters (including the weights and biases of each fully connected layer, etc.) are updated through backpropagation algorithm combined with Adam optimizer. Step 5: Use the residual connection multilayer perceptron model trained in Step 4 to predict material properties.

[0007] Based on the above plan, The material property feature dataset mentioned in step 1 includes: the chemical formula, constituent elements, constituent groups, and target property parameters corresponding to the constituent elements or groups for each material sample; The missing value handling method described in step 1 is zero-padding.

[0008] Based on the above plan, The training set mentioned in step 2 needs to be processed as follows: evaluate the variance of all numerical features in the training set, delete features with variance below a preset threshold, and retain the selected feature subset; standardize the feature subset, and calculate and save the mean and standard deviation of each feature.

[0009] The validation and test sets retain only the same features selected from the training set, and variance evaluation is not performed again. The validation and test sets are standardized using the standardized parameters saved from the training set to eliminate the impact of differences in the dimensions of different features on model training.

[0010] Based on the above plan, The residual connection multilayer perceptron (REMLP) model described in step 3 consists of an input layer, three hidden layers, a residual connection structure, and an output layer; the residual connection structure has a learnable scaling factor and is located between the input layer and the second hidden layer. The number of neurons in the input layer is equal to the number of features in the training set after step 2, and the number of neurons in the output layer is 1. This is suitable for single-target regression tasks. The number of neurons in each hidden layer can be adjusted according to the complexity and size of the dataset. The residual connection projects the input layer features through a linear layer to the output dimension of the second hidden layer, multiplies it by the learnable scaling factor α, and then adds it element-wise to the output of the second hidden layer to form an enhanced feature representation, which serves as the input to the third hidden layer. The residual path bypasses the first hidden layer to avoid losing shallow features.

[0011] The element-wise addition is as follows: the activated features output by the second hidden layer and the projected features globally scaled by the scaling factor α are respectively arithmetically added on the corresponding components of the same dimension index. The linear layer performs an identity mapping when the input dimension and the output dimension of the second hidden layer are the same.

[0012] The scaling factor α allows the model to adaptively adjust the contribution strength of the residual terms during training. ReLU is used as the activation function in all hidden layers, while no activation function is used in the output layer due to the regression task.

[0013] Based on the above plan, In step 4, during the training of the residual connection multilayer perceptron model, the number of neurons, training epochs, initial learning rate, and batch size of each of the three hidden layers are determined through hyperparameter search.

[0014] A storage medium, characterized in that the storage medium stores computer-executable instructions, which, when invoked by a computer, enable the computer to implement the material property prediction method of the improved multilayer perceptron model described above.

[0015] An electronic device is characterized by comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the material property prediction method of the improved multilayer perceptron model described above.

[0016] The material property prediction method of the improved multilayer perceptron model described in this invention has the following advantages: 1. Enhanced general learning capability: By introducing a learnable scalar scaling factor into the residual connections, the network can adaptively adjust the contribution of feature transmission at different levels, enhancing the model's ability to retain effective information and learn complex nonlinear relationships in the neural network. This alleviates the gradient vanishing problem in traditional MLPs and is particularly suitable for handling complex structure-property relationships in material property prediction.

[0017] 2. Higher prediction accuracy and better generalization: The improved model demonstrates better prediction accuracy than traditional MLP and XGBoost algorithms in various material property prediction tasks, and has good generalization performance.

[0018] 3. Significant Application Value: The REMLP-based material property prediction method provides an efficient solution for high-throughput prediction and screening of material properties. For example, in the research and development of halide perovskite photovoltaic materials, it can quickly and accurately predict key property parameters of perovskite materials, accelerating the discovery process of new materials. Attached Figure Description

[0019] The present invention includes the following figures: Figure 1 The REMLP model structure described in this invention Figure 2 This is a comparison chart of the R², MAE, and RMSE evaluation results of the REMLP model and the MLP model described in this invention on the test set in Example 1; Figure 3 This is a comparison chart of the R², MAE, and RMSE evaluation results of the REMLP model and the MLPi model described in this invention on the test set in Example 1; Figure 4 This is a comparison chart of the R², MAE, and RMSE evaluation results of the REMLP model and the MLP model described in this invention on the test set in Example 2; Figure 5 This is a comparison chart of the R², MAE, and RMSE evaluation results of the REMLP model and the MLPi model described in this invention on the test set in Example 2; Figure 6 This is a comparison chart of the R², MAE, and RMSE evaluation results of the REMLP model and the MLP model described in this invention on the test set in Example 3; Figure 7 This is a comparison chart of the R², MAE, and RMSE evaluation results of the REMLP model and the MLPi model described in this invention on the test set in Example 3. Detailed Implementation

[0020] The implementation process of the method of the present invention will be further described in detail below with reference to the embodiments.

[0021] Example 1: Prediction of the top energy level of the valence band in halide perovskites Step 1: Data Preparation and Feature Construction 1923 groups of ABX3 type and... were collected from publicly available perovskite material property datasets (Data source: Nakajima T, Sawada K. Discovery of Pb-Free Perovskite Solar Cells via High-Throughput Simulation on the K Computer. The Journal of Physical Chemistry Letters, 2017, 8(19): 4826-4831. DOI: 10.1021 / acs.jpclett.7b02203.). Valence band maximum (VBM) data for halide perovskite materials. Each data point includes the chemical formula of the perovskite material, the types of elements (or groups) at the A, B, and X positions, and the material's valence band maximum (VBM). Based on materials science knowledge, an initial set of features was constructed, including the relative atomic mass and polarizability of the A-site group, the ionic radius of the B-site element, and the electronegativity of the X-site element, resulting in 53 features (A-site cation mass, number of A-site cation atoms, A-site cation polarizability, A-site cation radius, A-site cation type, B'-site element atomic number, B'-site element group number, B'-site element period number, B'-site element first ionization energy, B'-site element van der Waals radius, B'-site element electronegativity, B'-site element atomic mass, B'-site element electron affinity, B'-site element valence electron number, B''-site element atomic number, B''-site element group number, B''-site element period number, B''-site element first ionization energy, B''-site element van der Waals radius, B''-site element electronegativity, B''-site element atomic mass, B''-site element electron affinity, B''-site element valence electron number, X-site element atomic number, X-site element group number, X-site element period ... period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number, X-site element period number The dataset includes the following parameters: number of atoms, first ionization energy of element X, van der Waals radius of element X, electronegativity of element X, atomic mass of element X, electron affinity of element X, number of valence electrons of element X, covalent radius of element X, boiling point of element B', boiling point of element B'', boiling point of element X, ionic radius of element B', ionic radius of element B'', ionic radius of element X, ionic radius ratio of element B' to element B'', ionic radius ratio of element B to cation A, proportion of transition metals, topological polar surface area of ​​cation A, number of hydrogen bond acceptors of cation A, number of hydrogen bond donors of cation A, number of aliphatic rings of cation A, number of aromatic rings of cation A, proportion of sp3 hybrid carbon atoms of cation A, number of heavy atoms of cation A, average mass of single atom of cation A, molar refractive index of cation A, radius ratio of element B to element X, and total mass. Zero-padding is performed on missing numerical features that account for less than 5% of the dataset. The dataset was randomly divided into training, validation, and test sets in a 80%:10%:10% ratio. Then, the feature variance was evaluated on the training set, and features with variances below 0.0001 were removed (in this example, these are the number of X-position element groups, the number of A-position cationic aliphatic rings, the number of A-position cationic aromatic rings, and the number of valence electrons at the X-position, totaling four items), resulting in 49 features. The validation and test sets retained only the same features selected from the training set. Finally, the training set was standardized, and the mean and standard deviation of each feature were saved. The validation and test sets were processed using the same standardization parameters to eliminate the influence of dimensions.

[0022] Step 2: Construct the REMLP model This embodiment constructs a REMLP model containing one input layer, three hidden layers, and one output layer, and introduces a residual connection structure with a learnable scaling factor.

[0023] Model architecture: Input layer: The number of neurons is 49, which is consistent with the feature dimension obtained in step 1.

[0024] Hidden layers: There are three layers in total. The specific number of neurons in each hidden layer (hidden1, hidden2, hidden3) is set as a hyperparameter and will be optimized and determined in subsequent steps (step 3).

[0025] Learnable scalable residual connections: Residual connections with learnable scaling factors are established between the input layer and the second hidden layer. Specifically, the input layer features are projected onto the output dimension of the second hidden layer via a linear layer, multiplied by a trainable scalar scaling factor α (initially 0.5), and then element-wise added to the output of the second hidden layer as the input to the third hidden layer. This design allows the model to adaptively adjust the fusion strength of the original features, enhancing feature reuse and mitigating gradient problems.

[0026] Output layer: has 1 neuron and is used to output the regression prediction value of the top energy level of the valence band.

[0027] Activation function: ReLU is used in all hidden layers, but no activation function is used in the output layer (regression task).

[0028] Step 3: Model Training and Optimization On the training set processed in step 1, the REMLP model constructed in step 2 is trained. The model calculates predictions using forward propagation, uses L1 loss as the loss function to measure prediction error, and updates model parameters (including weights and biases of each fully connected layer) using backpropagation combined with the Adam optimizer. The number of neurons, training epochs, initial learning rate, and batch size of each of the three hidden layers are determined through hyperparameter search. The optimal range for the number of hidden layer neurons is [16, 32, 64, 128, 256], the optimal range for the number of training epochs is 300 to 2400 (step size 300), the optimal range for the initial learning rate is 0.001 to 0.01, and the optimal range for the batch size is [64, 128, 256, 512, 1024]. The TPE optimization method in Hyperopt is used, with the goal of minimizing the MAE on the validation set, and 80 iterations are performed in the hyperparameter combination space to search for the optimal hyperparameter combination.

[0029] For a fair comparison, both the MLP and REMLP models underwent independent 80-step iterative optimization using the TPE method in Hyperopt within the same hyperparameter search space, with the goal of minimizing the MAE on their respective validation sets. The optimized hyperparameter combinations are as follows: MLP: {'hidden1': 256, 'hidden2': 32, 'hidden3': 64, 'learning_rate':0.004, 'epochs': 1500, 'batch_size': 64} REMLP: {'hidden1': 256, 'hidden2': 64, 'hidden3': 128, 'learning_rate': 0.009, 'epochs': 900,'batch_size': 128} Step 4: Model Testing and Performance Evaluation Using R², MAE, and RMSE as the main evaluation metrics, the training and validation sets from step 3 were combined into a new training set. The performance of the trained optimal model was evaluated using a test set that was not used in training and validation. The average results of 10 independent tests showed that REMLP achieved an R² of 0.8528, an MAE of 0.1443 eV, and an RMSE of 0.2078 eV on the test set. In contrast, the traditional MLP (without improved residuals) achieved an R² of 0.8424, an MAE of 0.1541 eV, and an RMSE of 0.2152 eV on the test set. The box plots of the test results are shown below. Figure 2 As shown.

[0030] To demonstrate that the performance improvement does not stem from changes in the hyperparameter combination, this embodiment also examines an MLP model (denoted as MLPi) with hyperparameters consistent with REMLP. This model's test set R² is 0.8333, MAE is 0.1570 eV, and RMSE is 0.2208 eV, which is inferior to REMLP. The test results are as follows: Figure 3 As shown.

[0031] Comparing the performance of REMLP with the traditional XGBoost method, the results in the previous paper showed that the XGBoost model, which also underwent 80-step TPE method iterative optimization to determine the hyperparameter combination, had a test set R² of 0.8481, MAE of 0.1490 eV, and RMSE of 0.2230 eV (Data source: Ye Y, Li R, et al. Machine learning for energyband prediction of halide perovskites. Materials Futures, 2025, 4(3):035601.DOI: 10.1088 / 2752-5724 / adeead.).

[0032] The comprehensive comparison results show that the REMLP model of this invention has better prediction accuracy than traditional MLP and XGBoost.

[0033] Example 2: Prediction of the bottom energy level of the conduction band in halide perovskites Step 1: Data Preparation and Feature Construction 1923 groups of ABX3 type and... were collected from publicly available perovskite material property datasets (Data source: Nakajima T, Sawada K. Discovery of Pb-Free Perovskite Solar Cells via High-Throughput Simulation on the K Computer. The Journal of Physical Chemistry Letters, 2017, 8(19): 4826-4831. DOI: 10.1021 / acs.jpclett.7b02203.). This dataset contains data on the conduction band minimum (CBM) of perovskite halide materials. Each data point includes the chemical formula of the perovskite material, the types of elements (or groups) at the A, B, and X positions, and the material's CBM. An initial feature set was constructed based on materials science knowledge, including the relative atomic mass and polarizability of the A-site group, the ionic radius of the B-site element, and the electronegativity of the X-site element, resulting in 53 features. Missing numerical features (representing less than 5% of the dataset) were padded with zeros. The dataset was then randomly divided into training, validation, and test sets at a ratio of 80%:10%:10%. The feature variance was then evaluated on the training set, and features with variances below 0.0001 were removed, resulting in 49 features. In this embodiment, both REMLP and MLP use a 29-dimensional feature subset optimized by feature optimization. This feature subset was determined by performing correlation analysis and filtering out redundant features on the 49-dimensional features of the training set in conjunction with the conduction band bottom energy level prediction task. Previous experiments have confirmed that this is the optimal feature dimension for traditional MLP models predicting conduction band bottom energy levels; therefore, a comparative experiment was conducted under this unified feature space. The validation and test sets retain only the same features selected in the training set. Finally, the training set is standardized, and the mean and standard deviation of each feature are saved. The validation and test sets are processed using the same standardization parameters to eliminate the influence of dimensions.

[0034] Step 2: Construct the REMLP model This embodiment constructs a REMLP model containing one input layer, three hidden layers, and one output layer, and introduces a residual connection structure with a learnable scaling factor.

[0035] Model architecture: Input layer: The number of neurons is 29, consistent with the feature dimensions obtained in step 1.

[0036] Hidden layers: There are three layers in total. The specific number of neurons in each hidden layer (hidden1, hidden2, hidden3) is set as a hyperparameter and will be optimized and determined in subsequent steps (step 3).

[0037] Learnable scalable residual connections: Residual connections with learnable scaling factors are established between the input layer and the second hidden layer. Specifically, the input layer features are projected onto the output dimension of the second hidden layer via a linear layer, multiplied by a trainable scalar scaling factor α (initially 0.5), and then element-wise added to the output of the second hidden layer as the input to the third hidden layer. This design allows the model to adaptively adjust the fusion strength of the original features, enhancing feature reuse and mitigating gradient problems.

[0038] Output layer: has 1 neuron and is used to output the regression prediction value of the conduction band bottom level.

[0039] Activation function: ReLU is used in all hidden layers, but no activation function is used in the output layer (regression task).

[0040] Step 3: Model Training and Optimization On the training set processed in step 1, the REMLP model constructed in step 2 is trained. The model calculates predictions using forward propagation, uses L1 loss as the loss function to measure prediction error, and updates model parameters (including weights and biases of each fully connected layer) using backpropagation combined with the Adam optimizer. The number of neurons, training epochs, initial learning rate, and batch size of each of the three hidden layers are determined through hyperparameter search. The optimal range for the number of hidden layer neurons is [16, 32, 64, 128, 256], the optimal range for the number of training epochs is 300 to 2400 (step size 300), the optimal range for the initial learning rate is 0.001 to 0.01, and the optimal range for the batch size is [64, 128, 256, 512, 1024]. The TPE optimization method in Hyperopt is used, with the goal of minimizing the MAE on the validation set, and 80 iterations are performed in the hyperparameter combination space to search for the optimal hyperparameter combination.

[0041] For a fair comparison, both the MLP and REMLP models underwent independent 80-step iterative optimization using the TPE method in Hyperopt within the same hyperparameter search space, with the goal of minimizing the MAE on their respective validation sets. The optimized hyperparameter combinations are as follows: MLP: {'hidden1': 128, 'hidden2': 256, 'hidden3': 32, 'learning_rate':0.008, 'epochs': 1500, 'batch_size': 64} REMLP: {'hidden1': 256, 'hidden2': 128, 'hidden3': 64, 'learning_rate': 0.009, 'epochs': 900,'batch_size': 128} Step 4: Model Testing and Performance Evaluation Using R², MAE, and RMSE as the main evaluation metrics, the training and validation sets from step 3 were combined into a new training set. The performance of the trained optimal model was evaluated using a test set that was not used in training and validation. The average results of 10 independent tests showed that REMLP achieved an R² of 0.8503, an MAE of 0.1410 eV, and an RMSE of 0.2083 eV on the test set. In contrast, the traditional MLP (without improved residuals) achieved an R² of 0.8238, an MAE of 0.1573 eV, and an RMSE of 0.2289 eV on the test set. The box plots of the test results are shown below. Figure 4 As shown.

[0042] To demonstrate that the performance improvement does not stem from changes in the hyperparameter combination, this embodiment also examines an MLP model (denoted as MLPi) with hyperparameters consistent with REMLP. This model's test set R² is 0.8243, MAE is 0.1576 eV, and RMSE is 0.2290 eV, which is inferior to REMLP. The test results are as follows: Figure 5 As shown.

[0043] Comparing the performance of REMLP with the traditional XGBoost method, the results in the previous paper showed that the XGBoost model, which also underwent 80-step TPE method iterative optimization to determine the hyperparameter combination, had a test set R² of 0.8298, a MAE of 0.1510 eV, and an RMSE of 0.2275 eV (Data source: Ye Y, Li R, et al. Machine learning for energyband prediction of halide perovskites. Materials Futures, 2025, 4(3):035601.DOI: 10.1088 / 2752-5724 / adeead.).

[0044] The comprehensive comparison results show that the REMLP model of this invention has better prediction accuracy than traditional MLP and XGBoost.

[0045] Example 3: Prediction of band gap in halide perovskites Step 1: Data Preparation and Feature Construction 1923 groups of ABX3 type and... were collected from publicly available perovskite material property datasets (Data source: Nakajima T, Sawada K. Discovery of Pb-Free Perovskite Solar Cells via High-Throughput Simulation on the K Computer. The Journal of Physical Chemistry Letters, 2017, 8(19): 4826-4831. DOI: 10.1021 / acs.jpclett.7b02203.). Bandgap data for halide perovskite materials were collected. Each data point included the chemical formula of the perovskite material, the types of elements (or groups) at the A, B, and X positions, and the bandgap. An initial feature set was constructed based on materials science knowledge, including the relative atomic mass and polarizability of the A-site group, the ionic radius of the B-site element, and the electronegativity of the X-site element, resulting in 53 features. Missing numerical features (less than 5% of the total data) were padded with zeros. The dataset was randomly divided into training, validation, and test sets at a ratio of 80%:10%:10%. The variance of the features was evaluated on the training set, and features with variances below 0.0001 were removed, resulting in 49 features. The validation and test sets retained only the same features selected from the training set. Finally, the training set was standardized, and the mean and standard deviation of each feature were saved. The validation and test sets were processed using the same standardization parameters to eliminate the influence of dimensions.

[0046] Step 2: Construct the REMLP model This embodiment constructs a REMLP model containing one input layer, three hidden layers, and one output layer, and introduces a residual connection structure with a learnable scaling factor.

[0047] Model architecture: Input layer: The number of neurons is 49, which is consistent with the feature dimension obtained in step 1.

[0048] Hidden layers: There are three layers in total. The specific number of neurons in each hidden layer (hidden1, hidden2, hidden3) is set as a hyperparameter and will be optimized and determined in subsequent steps (step 3).

[0049] Learnable scalable residual connections: Residual connections with learnable scaling factors are established between the input layer and the second hidden layer. Specifically, the input layer features are projected onto the output dimension of the second hidden layer via a linear layer, multiplied by a trainable scalar scaling factor α (initially 0.5), and then element-wise added to the output of the second hidden layer as the input to the third hidden layer. This design allows the model to adaptively adjust the fusion strength of the original features, enhancing feature reuse and mitigating gradient problems.

[0050] Output layer: has 1 neuron and is used to output the regression prediction value of the band gap size.

[0051] Activation function: ReLU is used in all hidden layers, but no activation function is used in the output layer (regression task).

[0052] Step 3: Model Training and Optimization On the training set processed in step 1, the REMLP model constructed in step 2 is trained. The model calculates predictions using forward propagation, uses L1 loss as the loss function to measure prediction error, and updates model parameters (including weights and biases of each fully connected layer) using backpropagation combined with the Adam optimizer. The number of neurons, training epochs, initial learning rate, and batch size of each of the three hidden layers are determined through hyperparameter search. The optimal range for the number of hidden layer neurons is [16, 32, 64, 128, 256], the optimal range for the number of training epochs is 300 to 2400 (step size 300), the optimal range for the initial learning rate is 0.001 to 0.01, and the optimal range for the batch size is [64, 128, 256, 512, 1024]. The TPE optimization method in Hyperopt is used, with the goal of minimizing the MAE on the validation set, and 80 iterations are performed in the hyperparameter combination space to search for the optimal hyperparameter combination.

[0053] For a fair comparison, both the MLP and REMLP models underwent independent 80-step iterative optimization using the TPE method in Hyperopt within the same hyperparameter search space, with the goal of minimizing the MAE on their respective validation sets. The optimized hyperparameter combinations are as follows: MLP: {'hidden1': 256, 'hidden2': 256, 'hidden3': 256, 'learning_rate': 0.009, 'epochs': 1500, 'batch_size':1024} REMLP: {'hidden1': 256, 'hidden2': 256, 'hidden3': 256, 'learning_rate': 0.01, 'epochs': 1200,'batch_size': 256} Step 4: Model Testing and Performance Evaluation Using R², MAE, and RMSE as the main evaluation metrics, the training and validation sets from step 3 were combined into a new training set. The performance of the trained optimal model was evaluated using a test set that was not used in training and validation. The average results of 10 independent tests showed that REMLP achieved an R² of 0.8142, an MAE of 0.2735 eV, and an RMSE of 0.3983 eV on the test set. In contrast, the traditional MLP (without improved residuals) achieved an R² of 0.7994, an MAE of 0.2855 eV, and an RMSE of 0.4120 eV on the test set. The box plots of the test results are shown below. Figure 6 As shown.

[0054] To demonstrate that the performance improvement does not stem from changes in the hyperparameter combination, this embodiment also examines an MLP model (denoted as MLPi) with hyperparameters consistent with REMLP. This model's test set R² is 0.7966, MAE is 0.2880 eV, and RMSE is 0.4192 eV, which is inferior to REMLP. The test results are as follows: Figure 7 As shown.

[0055] Comparing the performance of REMLP with the traditional XGBoost method, the results in the previous paper showed that the XGBoost model, which also determined the hyperparameter combination through 80-step TPE method iteration optimization, had a test set R² of 0.8008, MAE of 0.2848 eV, and RMSE of 0.4047 eV (Data source: Ye Y, Li R, et al. Machine learning for energyband prediction of halide perovskites. Materials Futures, 2025, 4(3):035601.DOI: 10.1088 / 2752-5724 / adeead.).

[0056] The comprehensive comparison results show that the REMLP model of this invention has better prediction accuracy than traditional MLP and XGBoost.

[0057] It should be noted that the above embodiments are only used to illustrate the technical concept and specific implementation process of the present invention, and are not intended to limit the present invention. The material property prediction method based on the improved multilayer perceptron model provided by the present invention has universality, and its network architecture details and training strategies can be flexibly adjusted according to specific task requirements. It should be understood that the specific implementation of the model, including but not limited to the number of network layers and neurons, the position of residual connections (such as connecting to different hidden layers), the form of adaptive scaling factors in residual connections, the type of activation function, the loss function and optimizer selection, etc., can all be adapted and optimized according to the dataset size of the target material system and the characteristics of the prediction task. Any technical solution that adopts the core concept of the present invention, namely, introducing a residual connection structure with a learnable scaling factor into the multilayer perceptron to enhance the learning ability of material property features, should be included within the protection scope of the present invention.

[0058] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

Claims

1. A method for predicting material properties using an improved multilayer perceptron model, characterized in that, Includes the following steps: Step 1: Obtain the material property feature dataset and handle missing values ​​for numerical features that cannot be filled by interpolation; Step 2: Randomly divide the feature dataset processed in Step 1 into a training set, a validation set, and a test set; Step 3: Introduce the residual connection structure into the multilayer perceptron to construct a residual connection multilayer perceptron model; Step 4: Input the training set described in Step 2 into the residual connected multilayer perceptron model obtained in Step 3 for multiple rounds of iterative training; during each round of iterative training, the residual connected multilayer perceptron model calculates the predicted value through forward propagation; Step 5: Use the residual connection multilayer perceptron model trained in Step 4 to predict material properties.

2. The material property prediction method for an improved multilayer perceptron model as described in claim 1, characterized in that: The material property feature dataset mentioned in step 1 includes: the chemical formula, constituent elements, constituent groups, and target property parameters corresponding to the constituent elements or groups for each material sample; The missing value handling method described in step 1 is zero-padding.

3. The material property prediction method for an improved multilayer perceptron model as described in claim 1, characterized in that: The training set mentioned in step 2 needs to be processed as follows: evaluate the variance of all numerical features in the training set, delete features with variance below a preset threshold, and retain the selected feature subset; calculate and save the mean and standard deviation of each feature in the feature subset, and use the above mean and standard deviation to standardize the feature subset.

4. The material property prediction method for an improved multilayer perceptron model as described in claim 1, characterized in that: The residual connection multilayer perceptron model described in step 3 consists of an input layer, three hidden layers, a residual connection structure, and an output layer; the residual connection structure has a learnable scaling factor and is located between the input layer and the second hidden layer. The residual connection is used to project the input layer features through the linear layer to the output dimension of the second hidden layer, multiply it by the learnable scaling factor α, and then add it element-wise to the output of the second hidden layer as the input of the third hidden layer.

5. The material property prediction method for an improved multilayer perceptron model as described in claim 4, characterized in that: In step 4, during the training of the residual connection multilayer perceptron model, the number of neurons, training epochs, initial learning rate, and batch size of each of the three hidden layers are determined through hyperparameter search.

6. A storage medium, characterized in that, The storage medium stores computer-executable instructions, which, when invoked by a computer, cause the computer to implement the material property prediction method of the improved multilayer perceptron model as described in any one of claims 1 to 5.

7. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the material property prediction method of the improved multilayer perceptron model according to any one of claims 1 to 5.