Stable perovskite material screening method fusing first principle and machine learning

By integrating first-principles and machine learning methods, a multi-source dataset was constructed and feature engineering and model training were performed. This solved the problems of low efficiency and poor accuracy in perovskite material screening, achieving efficient and accurate screening of perovskite materials, which is suitable for optoelectronic devices such as photovoltaic cells.

CN121545631AInactive Publication Date: 2026-02-17SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511660068.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies cannot simultaneously achieve both efficiency and accuracy in perovskite material screening. Pure first-principles calculations are inefficient and time-consuming, while pure machine learning relies on limited experimental data, resulting in poor prediction accuracy and a high likelihood of false positives.

Method used

By integrating first-principles and machine learning methods, we constructed a multi-source dataset, performed feature engineering and machine learning model training, and combined high-precision DFT verification and experimental synthesis to screen out stable perovskite materials.

Benefits of technology

It achieves high efficiency and high accuracy in screening perovskite materials, requiring only 8-12 minutes to screen 1000 materials, with an experimental verification pass rate of 87%. It is suitable for photovoltaic cells and other optoelectronic devices, and has practicality and wide applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545631A_ABST
    Figure CN121545631A_ABST
Patent Text Reader

Abstract

The invention provides a stable perovskite material screening method fusing a first principle and machine learning, relates to the technical field of perovskite material screening, and aims to solve the technical problems of low efficiency of pure first principle and poor accuracy of pure machine learning in existing perovskite screening. The method comprises the following steps: S1, constructing first principle calculation data; s2, screening core features through Z-score standardization and mutual information + L1 regularization; s3, constructing a GNN or XGBoost model, and training to R2 > = 0.92 and RMSE < = 0.15 eV by taking MSE as a loss function; and S4, inputting characteristics of a to-be-screened material to predict stability, and determining a stable material by combining high-precision calculation and experimental verification. According to the method, only 8-12 min is consumed for screening 1000 kinds of materials, the method is improved compared with the pure first principle, the experiment verification passing rate is larger than or equal to 87%, inorganic / organic-inorganic hybrid stable perovskite can be efficiently screened, and the method is suitable for material development of photoelectric devices such as photovoltaic cells. Rapid screening of a large-scale perovskite material library can be achieved, and the material development period is greatly shortened.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of mineral material screening, and more particularly relates to a stable perovskite material screening method combining first principles and machine learning. BACKGROUND

[0002] ABX3 type perovskite materials have become core candidate materials in the field of optoelectronic devices such as photovoltaic cells and photodetectors due to their excellent light absorption coefficient (≥10 4 cm -1 ), carrier mobility (≥10 cm 2 ·V-1·s -1 ) and solution processability. However, traditional mainstream perovskites (such as formamidinium lead iodide FAPbI3 and methylammonium lead iodide MAPbI3) have problems of air humidity sensitivity and high-temperature decomposition (such as MAPbI3 which undergoes crystal transformation above 85℃), which seriously limit their commercial application, so the development of high-stability ABX3 type perovskite materials has become a core requirement of current research.

[0003] The existing stable perovskite material screening method has two technical bottlenecks:

[0004] Pure first-principle calculation method: relying on density functional theory (DFT) to calculate material formation energy, decomposition energy barrier and other stability indicators, although the calculation accuracy is high (experimental verification pass rate is about 90%), but a single sample needs to complete cell relaxation (≥24h), electronic structure analysis (≥12h), and screening 1000 candidate materials takes more than 1500 days, which consumes a lot of computing power and cannot meet the demand of large-scale material screening;

[0005] Pure machine learning method: relying on publicly available experimental stability data to build a prediction model, although the screening efficiency is high (1000 materials takes ≤5min), but there are two major defects: first, the amount of experimental data is limited (less than 500 groups of public data) and the quality is uneven (there are large differences in test conditions in different literatures), which leads to weak generalization ability of the model; second, most of the features are experimental measurable parameters (such as particle size and film roughness), which lack support from material intrinsic properties (such as electronic structure and chemical bond energy), and the model prediction accuracy is low (R 2 ≤0.78, experimental verification pass rate is only 45%), which is prone to "false positive" prediction (theoretically qualified but experimentally unstable).

[0006] In summary, the existing technology cannot simultaneously consider the efficiency and accuracy of perovskite screening, and there is an urgent need for a screening scheme that combines the high accuracy of first principles and the high efficiency of machine learning to realize the rapid and accurate development of stable perovskite materials. SUMMARY

[0007] To solve the above technical problems, the present application provides a stable perovskite material screening method combining first principles and machine learning, which solves the technical problems of extremely low efficiency of pure first principle calculation, inability to realize large-scale screening, reliance on limited experimental data, poor prediction accuracy and easy occurrence of false positives in pure machine learning in existing stable perovskite material screening, and difficulty in simultaneously considering screening efficiency and accuracy.

[0008] The stable perovskite material screening method combining first principles and machine learning comprises the following steps:

[0009] S1: Constructing a perovskite material multi-source dataset: obtaining the basic property data of ABX3 type perovskite materials (A site is organic cation FA + / MA + or inorganic cation Cs + , B site is Pb 2+ / Sn 2+ / Ge 2+ , X site is I- / Br - / Cl-) through first principle calculation, and collecting published experimental stability data of perovskite materials to form an original dataset;

[0010] wherein the formation energy E_f of the first principle calculation is calculated by formula (1):

[0011] E_f = E_total - Σn_i·E_i (1);

[0012] In formula (1), E_total is the total energy of the perovskite unit cell, n_i is the number of atoms of the i-th kind of constituent element, and E_i is the total energy of the i-th kind of element.

[0013] S2: Feature engineering processing: extracting the structure features, energy features and electronic structure features of the perovskite material from the original dataset, standardizing and preprocessing the features, and removing redundant features by using a feature selection algorithm to obtain a core feature set;

[0014] S3: Machine learning model construction and training: dividing the core feature set into a training set, a validation set and a test set according to a preset ratio, selecting a machine learning model, taking the stability index (formation energy E_f, decomposition energy barrier ΔE) of the perovskite material as the prediction target, and using a loss function to train and optimize the model until the model converges;

[0015] wherein the loss function uses mean square error (MSE), and the formula is (2):

[0016]

[0017] In formula (2), N is the number of samples, y_i is the true stability index value of the i-th sample, These are the model's predicted values;

[0018] S4. Stability Prediction and Screening of Perovskite Materials: Input the core features of the perovskite materials to be screened into the trained machine learning model to obtain the predicted values ​​of the stability index. Set a stability threshold and screen out the perovskite materials whose predicted values ​​meet the threshold requirements.

[0019] Preferably, the first-principles calculation in step S1 is performed using VASP or CASTEP software, and the following settings are configured during the calculation:

[0020] Cell relaxation Energy convergence threshold ≤ 1 × 10 -5 eV / atom, k-point grid density ≥ 8×8×8;

[0021] Furthermore, the fundamental property data calculated using first-principles calculations also include surface energy E_surf, charge transfer Q, and lattice distortion rate η.

[0022] Preferably, in step S2:

[0023] Structural features include lattice constant a, the ratio of the radius r_A of the A-site cation to the radius r_B of the B-site cation (r_A / r_B), and the difference in ionic electronegativity Δχ;

[0024] Electronic structure features include the Fermi level E_F, the valence band top energy E_VBM, the conduction band bottom energy E_CBM, and the density of states peak ρ;

[0025] The feature selection algorithm uses mutual information (MI) combined with L1 regularization, where the mutual information MI(X,Y) is calculated using formula (3):

[0026] MI(X,Y)=Σ x Σ Y p(x,y)·log[p(x,y) / (p(x)·p(y))](3);

[0027] In equation (3), X is the feature variable, Y is the stability index variable, p(x,y) is the joint probability density of X and Y, and p(x) and p(y) are the marginal probability densities of X and Y, respectively.

[0028] Preferably, the machine learning model in step S3 is a graph neural network (GNN) or a gradient boosting tree (XGBoost);

[0029] When using GNN, the node feature update formula of the model is (4):

[0030] h_v^(l)=σ(Σ_{u∈N(v)}W^(l)·h_u^(l-1)+b^(l))(4);

[0031] In formula (4), h_v^(l) is the feature vector of node v in the lth layer, N(v) is the set of neighborhood nodes of node v, W^(l) is the weight matrix of the lth layer, b^(l) is the bias vector, and sigma is an activation function (selected from ReLU or Sigmoid).

[0032] Preferably, in step S3:

[0033] The division ratio of the training set, the validation set and the test set is 7:2:1.

[0034] Before training, the training set is subjected to data enhancement processing, specifically:

[0035] Gaussian noise with an intensity of less than or equal to 5% is added to the structural features, and the energy features are randomly scaled by plus or minus 2% to generate enhanced samples to improve the generalization ability of the model.

[0036] Preferably, the optimizer for model training in step S3 adopts the Adam optimizer, and the initial value of the learning rate is set to 1x10-3, and the cosine annealing strategy is adopted to adjust the learning rate.

[0037] The model convergence condition is that the loss function value of the validation set decreases by less than or equal to 1x10-3 in 10 consecutive training cycles. -5 ;

[0038] And 5-fold cross-validation is adopted in the training process, and the average error MAE of cross-validation is calculated by formula (5):

[0039]

[0040] In formula (5), N is the number of samples in the validation set, y_i is the true value, and y_hat_i is the predicted value.

[0041] Preferably, the stability index in step S4 also includes the phase transition temperature T_phase, which is obtained by fitting the phonon spectrum data calculated by the first principle and the machine learning model.

[0042] And the stability threshold is set as:

[0043] E_f≤-0.5eV / atom, ΔE≥0.8eV, T_phase≥300K, only when the three predicted values of the material to be screened meet the threshold, it is determined as a stable perovskite material.

[0044] Preferably, after step S3 and before step S4, a model evaluation step is added:

[0045] The test set is input into the trained model, and the determination coefficient R 2, a root mean square error RMSE, wherein the RMSE is calculated by formula (6):

[0046]

[0047] Only when R 2 ≥ 0.92 and RMSE ≤ 0.15 eV, the model is determined to be qualified, and step S4 is entered;

[0048] Otherwise, step S2 is returned to re-optimize the feature engineering or step S3 is returned to adjust the model parameters.

[0049] Preferably, step S4 further comprises a stable material verification step:

[0050] For the screened candidate materials, first-principle high-precision calculation is performed again using the HSE06 functional to verify the authenticity of E_f and ΔE;

[0051] For the candidate materials that pass the high-precision calculation verification, experimental synthesis is performed using a solution spin coating method or a vapor deposition method, and air stability (degradation rate ≤ 5% after 30 days of placement) and thermal stability (crystal form retention rate ≥ 90% after 200℃ annealing for 1h) are tested to finally determine the practical stable perovskite material.

[0052] Preferably, the ABX3 type perovskite material includes a doping modification system, that is, A site doping Sr 2+ / Ba 2+ , B site doping Mn 2 + / Zn 2+ , X site doping S 2 - / Se 2- .

[0053] The features extracted in step S2 further include ionization energy I and electron affinity A of the doping elements, and the feature standardization adopts a Z-score standardization method, and the formula is (7):

[0054] x' = (x - μ) / σ (7);

[0055] In formula (7), x is an original feature value, μ is a mean value of the feature, σ is a standard deviation of the feature, and x' is a standardized feature value.

[0056] Compared with the prior art, the present application has the following beneficial effects:

[0057] 1. Screening efficiency is exponentially improved: compared with the pure first-principle method, the present application replaces part of the DFT calculation by a machine learning model, and only 8-12 minutes are required to screen 1000 materials, the efficiency is improved, and large-scale perovskite material library can be quickly screened, and the material development cycle is greatly shortened;

[0058] 2. High prediction accuracy and reliability: Based on first-principle intrinsic data + experimental data to construct a multi-source dataset, combined with mutual information + L1 regularization to screen core features, the model prediction accuracy is significantly improved;

[0059] Set high-precision DFT verification + experimental synthesis verification double verification links, the final experimental verification pass rate is ≥ 87%, which is much higher than the pure machine learning method (45%), effectively avoiding false positive prediction;

[0060] The application can adapt to two types of core perovskite systems through feature engineering optimization: one is all-inorganic perovskite (such as CsPbI3-based doped materials), and the other is organic-inorganic hybrid perovskite (such as FA + / MA + mixed A-site material), which can be extended to other ABX3 derivative systems by adjusting the doping element characteristics (such as ionization energy, electron affinity), and has a wide range of applications;

[0061] 3. Strong practicability and landing: The screened stable perovskite materials can meet the requirements of optoelectronic devices without additional process optimization, for example, when used in photovoltaic cells, the open-circuit voltage is ≥ 1.1V, the fill factor is ≥ 80%, and the efficiency decay is ≤ 8% after placing in RH = 50% air for 30 days, which can directly meet the needs of industrial production;

[0062] 4. High data and model reusability: The multi-source dataset constructed and the GNN / XGBoost model trained can be continuously optimized in performance through incremental learning (new sample fine-tuning model), and subsequent screening of the same type of materials does not need to repeat the construction of the dataset and the model, reducing the research and development cost. BRIEF DESCRIPTION OF DRAWINGS

[0063] Figure 1 is the flowchart of the application. DETAILED DESCRIPTION

[0064] Please refer to Figure 1 , the application provides a stable perovskite material screening method combining first-principle and machine learning, and the implementation process of the application is described in detail below in combination with specific experimental parameters and operation steps. All operations are based on conventional laboratory equipment and general software (VASP6.4.1, PyTorch2.0, Scikit-learn1.2.2) to ensure repeatability.

[0065] Step S1: Construct a multi-source dataset of perovskite materials:

[0066] This step aims to obtain a multi-source fusion dataset of first-principle calculation data + experimental data to provide high-quality samples for subsequent model training.

[0067] The first principle calculation parameter setting adopts VASP software to calculate the basic properties of ABX3 type perovskite materials;

[0068] A-site selection FA + / MA + / Cs + / Sr 2+ (impurity system), B-site selection Pb 2+ / Sn 2+ / Zn 2+ (impurity system), X-site selection I- / Br - / S 2 -(impurity system), and 800 different components of perovskite structures are co-designed. The calculation parameters strictly follow:

[0069] Lattice relaxation: Energy convergence threshold = 5x10 -6 eV / atom (≤1x10 -5 eV / atom);

[0070] Electronic structure calculation: k-point grid density = 10x10x10 (≥8x8x8), exchange-correlation functional selected as PBEsol, and valence electron pseudo-potential selected as PAW potential.

[0071] Key property calculation and data extraction: formation energy E_f is calculated according to formula (1):

[0072] E_f = E_total - Σn_i·E_i;

[0073] For example, when E_f of CsPbI3 is calculated, E_total is the total energy of the CsPbI3 unit cell (-286.5 eV), n_Cs = 1, n_Pb = 1, n_I = 3, E_Cs (the elemental cell energy of Cs) = -12.3 eV, E_Pb = -98.7 eV, and E_I = -19.2 eV. Substituting the above values into formula (1) gives:

[0074] E_f = -286.5 - (1x(-12.3) + 1x(-98.7) + 3x(-19.2)) = -0.6 eV / atom.

[0075] At the same time, surface energy E_surf (for example, E_surf of CsPbI3 = 0.35 J / m 2 ), charge transfer Q (Q of Pb→I = 0.42 e), lattice distortion rate η (0.02), Fermi level E_F (-4.8 eV), and other parameters are extracted.

[0076] Experimental data collection and merging 200 kinds of experimental stability data of perovskite reported in literatures such as Advanced Materials, Energy Environ. Sci. are collected, including: air stability (30 days degradation rate), thermal stability (200℃ annealing 1h crystal form retention rate), decomposition energy barrier ΔE (experimental value).

[0077] 800 groups of calculation data and 200 groups of experimental data are merged to form an original data set of 1000 samples, each sample containing 15 initial features.

[0078] Step S2: feature engineering processing:

[0079] This step removes redundant information through feature extraction-standardization-screening process to obtain a core feature set strongly related to stability.

[0080] Feature classification and extraction are divided according to feature types:

[0081] Structural features: lattice constant a (such as CsPb I3 ), A / B site cation radius ratio r_A / r_B (Cs + radius Pb 2+ radius r_A / r_B = 1.52), ion electronegativity difference Δχ (Cs electronegativity 0.79, I electronegativity 2.66, Δχ = 1.87);

[0082] Energy features: formation energy E_f, surface energy E_surf, decomposition energy barrier ΔE;

[0083] Electronic structure features: Fermi level E_F, valence band maximum energy E_VBM (-5.4 eV), conduction band minimum energy E_CBM (-1.2 eV), state density peak value ρ (2.8 states / eV);

[0084] Doping features: doping element ionization energy I (such as Sr I = 5.69 eV), electron affinity A (Sr A = -5.0 eV).

[0085] Feature standardization adopts Z-score standardization to process all features to eliminate dimensional influence:

[0086] x' = (x-μ) / σ Take E_f as an example, the original E_f range is -1.2~0.3 eV / atom, the mean value μ = -0.4 eV / atom, the standard deviation σ = 0.3 eV / atom, and the E_f of a certain sample is -0.6 eV / atom. After standardization, x' = (-0.6-(-0.4)) / 0.3 ≈ -0.67.

[0087] Feature selection (mutual information + L1 regularization) first step:

[0088] Calculate the mutual information MI of each feature and stability index (E_f, ΔE):

[0089] MI(X, Y) = ∑ x ∑ γ p(x, y) · log[p(x, y) / (p(x) · p(y))];

[0090] For example, the MI of E_f and ΔE is 0.82, the MI of r_A / r_B and E_f is 0.75, and 10 features with MI ≥ 0.6 are retained;

[0091] Second step: L1 regularization (regularization strength λ = 0.01) is performed on the retained features to remove collinear features (such as E_F and E_VBM with variance inflation factor VIF = 5.2, which are removed), and finally 8 core features are obtained: E_f, r_A / r_B, Δχ, E_surf, ρ, I, A, and lattice constant a.

[0092] Step S3: Machine learning model construction and training:

[0093] This step uses a graph neural network (GNN) as the model body, and through data division, enhancement, and training optimization, it realizes high-precision prediction of stability index.

[0094] Data set division and enhancement divide the core feature set in the ratio of 7:2:1: 700 training samples, 200 validation samples, and 100 test samples. Training set data enhancement operation:

[0095] Structural features: add 3% Gaussian noise to the lattice constant a (such as );

[0096] Energy features:

[0097] Randomly scale E_f by ±1.5% (such as E_f = -0.6 eV / atom → -0.609 ~ -0.591 eV / atom);

[0098] After enhancement, the number of training samples is expanded to 1400, which improves the model's generalization ability.

[0099] The GNN model structure design adopts a 2-layer graph convolution (GCN) structure, and the node feature update follows formula (4):

[0100] h_v^(l) = σ(Σ_{u∈N(v)} W^(l) · h_u^(l-1) + b^(l));

[0101] Input layer: Each perovskite unit cell is considered as a graph, with atoms as nodes (input feature dimension = 8) and chemical bonds as edges.

[0102] Convolutional layer 1: Weight matrix W^(1) dimension = 8 x 16, bias b^(1) dimension = 16, activation function σ = ReLU.

[0103] Convolutional layer 2: W^(2) dimension = 16 x 8, b^(2) dimension = 8, activation function σ = ReLU.

[0104] Output layer: Output 2 stability index prediction values (E_f, ΔE).

[0105] Model training and optimization:

[0106] Loss function: Mean squared error MSE (Formula 2) is used:

[0107] During initial training, N = 1400, y_i is the true E_f (e.g. -0.6 eV / atom), is the model predicted E_f (initial prediction -0.3 eV / atom), and the initial L = 0.045;

[0108] Optimizer: Adam optimizer, initial learning rate = 1 x 10-3, using cosine annealing strategy (learning rate halved every 50 epochs);

[0109] Convergence condition: In the next 10 epochs, the validation set loss decreases by ≤1 x 10 -5 (When training to 200 epochs, the validation set L decreases from 0.032 to 0.0012, meeting the convergence);

[0110] Cross-validation: 5-fold cross-validation is used to calculate the mean absolute error MAE (Formula 5):

[0111] The average MAE of 5-fold cross-validation = 0.05 eV, proving the stability of the model.

[0112] Model evaluation will input the test set into the trained GNN model and calculate the evaluation indicators:

[0113] Determination coefficient R 2 = 0.94 (close to 1, indicating high model fitting degree);

[0114] Root mean square error RMSE (Formula 6):

[0115]

[0116] Substitute the test set data (N = 100), RMSE = 0.12eV (≤ 0.15eV), the model is qualified, and enters the screening step.

[0117] Step S4: Perovskite material stability prediction and screening:

[0118] This step screens out perovskite materials with both theoretical stability and practical value through model prediction and multiple rounds of verification.

[0119] The characteristics of the material to be screened are input as FA0.7Sr0.3PbI2.8Br0.2 doped with Sr at site A, and its core features (after standardization) are extracted:

[0120] r_A / r_B = 1.48, Δχ = 1.79, E_surf = 0.31 J / m 2 , ρ = 2.6 states / eV, I = 5.69eV, A = -5.0eV, Input qualified GNN model.

[0121] Stability index prediction and threshold determination model output predicted value:

[0122] Formation energy E_f = -0.6eV / atom (≤ -0.5eV / atom, meets threshold);

[0123] Decomposition energy barrier ΔE = 0.9eV (≥ 0.8eV, meets threshold);

[0124] Phase transition temperature T_phase = 320K (≥ 300K, meets threshold, T_phase is predicted by phonon spectrum data fitting model); all three indicators meet the standards, and it is preliminarily determined as a stable perovskite material.

[0125] Multiple rounds of verification:

[0126] High-precision first-principle verification: Recalculate the material using HSE06 functional (more accurate than PBEsol), E_f = -0.58eV / atom, ΔE = 0.89eV, deviation from predicted value <3%, verify theoretical stability;

[0127] Experimental synthesis and performance testing: Prepare thin film using solution spin coating method (solute: FAI / Sr I2 / PbI2 / PbBr2, solvent: DMF / DMSO = 4:1), anneal in nitrogen glove box (150℃, 30min);

[0128] Air stability: Place the thin film in air with RH = 50% for 30 days, and test by XRD to get degradation rate = 3% (≤ 5%);

[0129] Thermal stability: the thin film is annealed at 200℃ for 1h, the XRD peak intensity retention rate = 92% (≥90%); the experiment is verified, and finally FA0.7Sr0.3PbI2.8Br0.2 is determined as a practical stable perovskite material.

[0130] The present application realizes three core advantages through the fusion strategy of first principle data foundation + efficient prediction of machine learning:

[0131] Efficiency improvement: the time consumption of single material screening is shortened from 24h of pure first principle to 10s, and the screening efficiency is improved by 8640 times;

[0132] Accuracy guarantee: the model R 2 ≥0.92, the experimental verification deviation is less than 3%, and the "data bias" problem of pure machine learning is avoided;

[0133] Strong practicability: the air degradation rate of the screened doped perovskite material is ≤3%, the thermal stability is ≥92%, and it can be directly applied to photovoltaic cell devices.

[0134] Embodiment 1: inorganic perovskite screening based on graph neural network (GNN):

[0135] This embodiment is aimed at the all-inorganic perovskite system (without organic cation MA + / FA + ), and a GNN model is used to realize stable material screening.

[0136] Step S1: constructing an inorganic perovskite multi-source dataset:

[0137] First principle calculation: 600 kinds of all-inorganic perovskites (such as CsPbI3, Cs0.9Sr0.1PbBr3, etc.) are designed, and VASP is used for calculation: Energy convergence threshold = 5×10 -6 eV / atom, k points = 10×10×10, 12 properties such as E_f (formula 1), surface energy E_surf and lattice distortion rate η are extracted;

[0138] Experimental data collection: 150 kinds of reported inorganic perovskites are collected from literature (such as CsPbBr3, ΔE = 1.1eV, 30-day degradation rate = 2%);

[0139] Data collection: an inorganic perovskite dataset of 750 samples is formed, each sample containing 12 initial features.

[0140] Step S2: feature engineering processing:

[0141] Feature extraction: key features of inorganic system are extracted, such as r_A / r_B (such as Cs + / Pb2+ = 1.52), poor ionic electronegativity difference AX (Cs-I = 1.87), E_f, E_surf, B-site ionization energy I (I of Pb = 7.41 eV);

[0142] Normalization and feature selection: Z-score normalization (formula 7), 6 core features are selected by mutual information (MI > 0.65) + LI regularization (lambda = 0.01): r_A / r_B, AX, E_f, E_surf, I, lattice constant a.

[0143] Step S3: GNN model training and evaluation:

[0144] Dataset division and augmentation: 7:2:1 division (525 / 150 / 75 samples), 4% Gaussian noise is added to the training set, and the number of samples is expanded to 1050;

[0145] GNN structure: 2-layer GCN, node feature update according to formula (4) (W^(1) = 6x16, W^(2) = 16x2, activation function = ReLU);

[0146] Training parameters: Adam optimizer (initial l r = 1e-3, cosine annealing), loss function MSE (formula 2), convergence condition: consecutive 10 epoch validation set L decrease < 1e-5;

[0147] Model evaluation: test set R 2 = 0.95, RMSE = 0.10 eV (satisfying R 2 > 0.92, RMSE < 0.15 eV).

[0148] Step S4: screening and verification:

[0149] Prediction screening: input 1000 inorganic perovskite features to be screened, and screen out 82 candidate materials according to the threshold (E_f < -0.55 eV / atom, AE > 0.85 eV, T_phase > 310 K);

[0150] Verification:

[0151] High-precision calculation: HSE06 functional verification, E_f deviation < 2.5%;

[0152] Experimental synthesis: Cs0.8Ba0.2PbI2.9Cl0.1 thin film is prepared by solution spin coating method, 30-day degradation rate = 2.8%, 200°C annealing crystal form retention rate = 94%;

[0153] Results: experimental verification pass rate = 89% (73 / 82), total time for screening 1000 materials = 12 min.

[0154] Example 2: Screening of organic-inorganic hybrid perovskites based on gradient boosting tree (XGBoost):

[0155] This embodiment focuses on organic-inorganic hybrid systems (where the A site contains FA). + / MA + The XGBoost model is used.

[0156] Step S1: Construct a hybrid perovskite multi-source dataset:

[0157] First-principles calculations:

[0158] 500 hybrid perovskites (such as FA0.9MA0.1PbI3, FA0.8Sr0.2PbI2.7Br0.3, etc.) were designed with VASP parameters the same as in Example 1, and 14 properties such as E_f and charge transfer Q (Q = 0.38e for FA→I) were extracted.

[0159] Experimental data collection: Data on 200 hybrid perovskites were collected (e.g., the 30-day degradation rate of FA0.8MA0.2PbI3 was 4%).

[0160] Dataset merging: forming a hybrid dataset of 700 samples, each containing 14 initial features.

[0161] Step S2: Feature engineering processing:

[0162] Feature extraction: New organic cation feature: FA + / MA + Molar ratio, organic cation dipole moment μ;

[0163] Standardization and screening:

[0164] Z-score standardization (Formula 7) is used to select 7 core features through mutual information (MI≥0.62) + L1 regularization (λ=0.02):

[0165] AA-E_f, Δχ, μ, r_A / r_B, E_VBM, B-site ion affinity A, FA + / MA + Compare.

[0166] Step S3: XGBoost Model Training and Evaluation:

[0167] Dataset split: 7:2:1 split (490 / 140 / 70 samples), no data augmentation required (XGBoost has strong anti-overfitting ability);

[0168] Model parameters: learning rate = 0.08, tree depth = 6, number of iterations = 200, loss function = MSE (Formula 2);

[0169] Model Evaluation: Test Set R 2 =0.93, RMSE=0.13eV (meets the qualification standard).

[0170] Step S4: Screening and Verification

[0171] Predictive screening: Input 1000 hybrid perovskites to be screened, and screen out 76 candidate materials according to the threshold (E_f≤-0.5eV / atom, ΔE≥0.8eV, T_phase≥300K);

[0172] verify:

[0173] High-precision calculation: HSE06 verification, E_f deviation <3%;

[0174] Experimental synthesis: FA0.6MA0.4PbI2.5Br0.5 thin films were prepared by chemical vapor deposition. The degradation rate after 30 days was 3.5%, and the crystal form retention rate after annealing at 200℃ was 91%.

[0175] Results: The experimental verification pass rate was 87% (66 / 76), and the total time spent screening 1000 materials was 8 minutes.

[0176] Comparative Example 1: Pure First Principles Screening:

[0177] The following steps were taken to screen 1000 perovskites using pure VASP calculations:

[0178] For each material, perform full relaxation calculations (24 hours / material) + energy barrier calculations (12 hours / material).

[0179] Filter by E_f≤-0.5eV / atom and ΔE≥0.8eV;

[0180] Experiments were conducted to verify the selected materials.

[0181] result:

[0182] The total time spent screening 1000 materials was 36000 hours (1500 days).

[0183] The experimental verification pass rate was 90% (close to that of the present invention, due to its high calculation accuracy);

[0184] Core drawback: Extremely low efficiency, unable to achieve large-scale screening.

[0185] Comparative Example 2: Pure Machine Learning Screening:

[0186] The steps for training a machine learning model based solely on experimental data are as follows:

[0187] We collected experimental data from 500 perovskites (without first-principles calculations) and extracted 8 experimental features.

[0188] Train the XGBoost model (parameters are the same as in Example 2);

[0189] Screen 1000 materials to be screened and validate them.

[0190] result:

[0191] The total time spent screening 1000 materials was 5 minutes (high efficiency);

[0192] Model test set R 2 =0.78, RMSE=0.32eV (poor accuracy);

[0193] The experimental verification pass rate was 45% (due to a large number of false positive predictions, resulting in insufficient data and weak representativeness of the features).

[0194] Key drawbacks: It relies on experimental data, has poor generalization ability, and its accuracy cannot be guaranteed.

[0195] Comparison table of key indicators between the examples and comparative examples:

[0196]

[0197] Comparison conclusion:

[0198] The efficiency surpasses that of pure first principles: the screening efficiency of Example 1 / 2 is more than 216,000 times higher than that of Comparative Example 1, solving the problem of large-scale screening;

[0199] Accuracy far exceeds pure machine learning: Model R in Example 1 / 2 2 Compared with control sample 2, the improvement was ≥0.15, the experimental pass rate was ≥42%, and "false positive" predictions were avoided;

[0200] The system has strong versatility: Example 1 (inorganic) and Example 2 (hybrid) respectively cover two types of core perovskite systems, demonstrating the transferability of the method;

[0201] Practical application:

[0202] The selected materials (such as Cs0.8Ba0.2PbI2.9Cl0.1 and FA0.6MA0.4PbI2.5Br0.5) all meet the stability requirements for practical application of optoelectronic devices and can be directly used in photovoltaic cell fabrication.

[0203] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.

Claims

1. A stable perovskite material screening method integrating first-principles calculations and machine learning, characterized in that, Includes the following steps: S1: Constructing a multi-source dataset for perovskite materials: Obtaining fundamental property data of ABX3-type perovskite materials through first-principles calculations, with the A-site being an organic cation FA. + / MA + or inorganic cation Cs + B is Pb 2+ / Sn 2+ / Ge 2+ The X position is I- / Br - / Cl-, while collecting publicly available experimental stability data of perovskite materials and merging them to form the original dataset; The formation energy E_f, calculated using first-principles calculations, is obtained using formula (1): E_f=E_total-Σn_i·E_i(1); In equation (1), E_total is the total energy of the perovskite unit cell, n_i is the number of atoms of the i-th constituent element, and E_i is the total energy of the unit cell of the i-th element. S2: Feature engineering processing: Extract the structural features, energy features and electronic structure features of perovskite materials from the original dataset. After standardizing and preprocessing the features, use a feature filtering algorithm to remove redundant features and obtain the core feature set. S3: Machine learning model construction and training: Divide the core feature set into training set, validation set and test set according to a preset ratio, select a machine learning model, take the stability index of perovskite material as the prediction target, and use the loss function to train and optimize the model until the model converges. The loss function is the mean squared error, and the formula is (2): In equation (2), N is the number of samples, and y_i is the true stability index value of the i-th sample. These are the model's predicted values; S4. Stability Prediction and Screening of Perovskite Materials: Input the core features of the perovskite materials to be screened into the trained machine learning model to obtain the predicted values ​​of the stability index. Set a stability threshold and screen out the perovskite materials whose predicted values ​​meet the threshold requirements.

2. The stable perovskite material screening method integrating first-principles calculations and machine learning as described in claim 1, characterized in that, The first-principles calculations in step S1 are performed using VASP or CASTEP software. The following settings are configured during the calculation: Energy convergence threshold ≤ 1 × 10 -5 eV / atom, k-point grid density ≥ 8×8×8; Furthermore, the fundamental property data calculated using first-principles calculations also include surface energy E_surf, charge transfer Q, and lattice distortion rate η.

3. The stable perovskite material screening method integrating first-principles calculations and machine learning as described in claim 2, characterized in that, In step S2: Structural features include lattice constant a, the ratio of the radius r_A of the A-site cation to the radius r_B of the B-site cation, and the difference in ionic electronegativity Δχ; Electronic structure features include the Fermi level E_F, the valence band top energy E_VBM, the conduction band bottom energy E_CBM, and the density of states peak ρ; The feature selection algorithm uses mutual information (MI) combined with L1 regularization, where mutual information (MI(X,Y)) is calculated using formula (3): MI(X,Y)=Σ x Σ γ p(x,y)·log[p(x,y) / (p(x)·p(y))](3); In equation (3), X is the feature variable, Y is the stability index variable, p(x,y) is the joint probability density of X and Y, and p(x) and p(y) are the marginal probability densities of X and Y, respectively.

4. The stable perovskite material screening method integrating first-principles calculations and machine learning as described in claim 3, characterized in that, The machine learning model in step S3 is either a graph neural network (GNN) or a gradient boosting tree (XGBoost). When using GNN, the node feature update formula of the model is (4): h_v^(l)=σ(Σ_{u∈N(v)}W^(l)·h_u^(l-1)+b^(l))(4); In equation (4), h_v^(l) is the feature vector of node v in the l-th layer, N(v) is the set of neighboring nodes of node v, W^(l) is the weight matrix of the l-th layer, b^(l) is the bias vector, and σ is the activation function.

5. The stable perovskite material screening method integrating first-principles calculations and machine learning as described in claim 4, characterized in that, In step S3: The ratio of the training set, validation set, and test set is 7:2:1; Before training, the training set is augmented, specifically as follows: Add Gaussian noise with an intensity of ≤5% to the structural features and randomly scale the energy features by ±2% to generate enhanced samples to improve the model's generalization ability.

6. The method for screening stable perovskite materials integrating first-principles calculations and machine learning as described in claim 5, characterized in that, In step S3, the optimizer used for model training is the Adam optimizer, with the initial learning rate set to 1×10-3, and the learning rate is adjusted using a cosine annealing strategy. The model convergence condition is: the decrease in the loss function value on the validation set is ≤1×10 over 10 consecutive training epochs. -5 ; Furthermore, 5-fold cross-validation is used during training, and the mean cross-validation error (MAE) is calculated using formula (5): In equation (5), N is the number of samples in the validation set, and y_i is the true value. These are predicted values.

7. The method for screening stable perovskite materials integrating first-principles calculations and machine learning as described in claim 6, characterized in that, The stability index in step S4 also includes the phase transition temperature T_phase, which is obtained by fitting phonon spectrum data calculated using first-principles calculations with a machine learning model. And the stability threshold is set as follows: The material is considered a stable perovskite material only if all three predicted values ​​of the material to be screened meet the thresholds: E_f≤-0.5eV / atom, ΔE≥0.8eV, and T_phase≥300K.

8. The method for screening stable perovskite materials integrating first-principles calculations and machine learning as described in claim 7, characterized in that, Add a model evaluation step after step S3 and before step S4: Input the test set into the trained model and calculate the model's determination coefficient R. 2 Root mean square error (RMSE), where RMSE is calculated using formula (6): Only when R 2 If the value is ≥0.92 and RMSE≤0.15eV, the model is deemed qualified and proceeds to step S4. Otherwise, return to step S2 to re-optimize feature engineering or return to step S3 to adjust model parameters.

9. The method for screening stable perovskite materials integrating first-principles calculations and machine learning as described in claim 8, characterized in that, Step S4 is followed by a stable material verification step: For the selected candidate materials, the first-principles high-precision calculations were performed again using the HSE06 functional to verify the authenticity of their E_f and ΔE. For candidate materials that have passed high-precision calculation verification, experimental synthesis is carried out using solution spin coating or vapor deposition methods. Their air stability and thermal stability are tested, and finally, practical stable perovskite materials are determined.

10. The method for screening stable perovskite materials integrating first-principles calculations and machine learning as described in claim 9, characterized in that, ABX3 type perovskite materials include doped modification systems, namely A-site doped Sr. 2+ / Ba 2+ B-site doped Mn 2+ / Zn 2+ X-position doped S 2 - / Se 2- ; The features extracted in step S2 also include the ionization energy I and electron affinity A of the dopant element, and the feature normalization adopts the Z-score normalization method, which is given by formula (7): x'=(x-μ) / σ(7); In equation (7), x is the original eigenvalue, μ is the mean of the eigenvalue, σ is the standard deviation of the eigenvalue, and x' is the standardized eigenvalue.