Machine learning and physical priori knowledge constraint-based double-perovskite halide material layer-by-layer screening method, equipment and medium
By constructing a cascaded screening framework based on SLME and combining machine learning with prior physical knowledge, the problems of low efficiency and insufficient balance of multiple performance in traditional methods are solved, and high-performance double perovskite halide materials are screened efficiently, significantly improving computational efficiency and screening accuracy.
Patent Information
- Application Number
- CN202511200818.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies are inefficient in screening high-power conversion efficiency double perovskite halide materials, have a large candidate space, lack physical prior constraints, and are insufficient in balancing multiple properties, making it difficult to simultaneously ensure photoelectric conversion efficiency and structural stability.
We employ a layer-by-layer screening method based on machine learning and physical prior knowledge constraints. Using SLME as the core indicator, we combine spatial group, bandgap characteristics and thermodynamic stability to construct a cascaded screening framework, which gradually narrows down the candidate space and achieves multi-dimensional and efficient screening.
The candidate space was significantly narrowed, the physical reliability of the screening results was improved, and a variety of high-performance double perovskite halide materials were efficiently screened by combining machine learning prediction with DFT verification, with a computational speedup of approximately 109 times.
Smart Images

Figure CN121034498A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of solar cell material design, and in particular to a method, device and medium for layer-by-layer screening of double perovskite halide materials based on machine learning and physical prior knowledge constraints. BACKGROUND
[0002] Double perovskite halides have excellent application potential in the field of optoelectronic devices, but their reported photovoltaic performance is still unsatisfactory. Therefore, developing an effective screening method that can effectively screen double perovskite halides with high power conversion efficiency has become a primary goal.
[0003] Due to the diversity of the elemental composition of materials, it is time-consuming and costly to discover new materials with multiple target properties through traditional trial-and-error methods. Although high-throughput density functional theory (DFT) calculations can accelerate material screening, the complexity of materials and the vast chemical space greatly limit their efficiency.
[0004] The rapid development of machine learning (ML) technology provides a new solution to this problem. At the present stage, researchers attempt to use ML combined with DFT to perform high-throughput prediction, but existing methods still have many technical bottlenecks in practical applications, such as limited screening indicators, which are often focused on band gaps or formation energies, making it difficult to evaluate material performance from multiple angles; multi-objective optimization algorithms face challenges such as difficulty in setting target weights, difficulty in interpreting Pareto solution sets, and difficulty in introducing physical prior constraints in material design.
[0005] In the prior art, CN115579089A proposes to use elemental intrinsic characteristics as intermediate input, perform Pearson screening, GBRT feature sorting, and symbolic regression modeling to realize fast prediction of the band gap of ABX3-type perovskite, but this method only targets a single band gap, does not consider multiple performance synergies such as structural symmetry and thermodynamic stability, and the component space is still limited to ABX3, making it difficult to be directly extended to double perovskite systems. CN112132185A targets A2B'B"O6-type double perovskite oxides and predicts the band gap through an mRMR-SVM process, but its feature screening is only based on maximum correlation and minimum redundancy, lacks physical prior constraints, resulting in a large candidate space, and also does not establish a progressive screening mechanism for multiple indicators such as band gap, stability, and crystal symmetry, which cannot guarantee both photoelectric conversion efficiency and structural stability.
[0006] Therefore, the existing data-driven screening framework mainly has the following defects: low efficiency, single prediction only focusing on a single attribute, repeated calculation bringing huge time cost; large candidate space, lack of physical priori constraints, exponential expansion of element combination; insufficient performance balance, unable to achieve synergistic optimization among space group, SLME, band gap and thermodynamic stability, and difficult to directly lock high-performance double perovskite halide.
[0007] In view of the above-mentioned defects, it is urgent to propose a layer-by-layer screening method taking SLME as the core and being constrained by physical priori knowledge, to realize multi-dimensional efficient screening from structure, optics, electricity to stability. SUMMARY
[0008] The present application aims at the problems of low efficiency, large candidate space and insufficient performance balance of the traditional data-driven screening framework, and constructs a cascade screening method taking spectral limit maximum efficiency (SLME) as the guide, to realize the discovery of high-precision photovoltaic materials by strictly following the progressive decision chain of space group symmetry, SLME, band gap characteristics and thermodynamic stability.
[0009] The object of the present application can be realized by the following technical solutions:
[0010] The present application provides a layer-by-layer screening method for double perovskite halide materials based on machine learning and physical priori knowledge constraints in the first aspect, comprising the following steps:
[0011] S1, collecting original data of space group, SLME, band gap, Ehull and Ef of double perovskite halide characteristics in public database;
[0012] S2, first eliminating high correlation redundant features by Pearson coefficient, and then optimizing and retaining by Spearman coefficient, and locking the optimal feature subset for double perovskite by genetic algorithm-cross validation on the original data obtained in S1;
[0013] S3, training multiple machine learning models in parallel based on the optimal feature subset in S2, optimizing and integrating, introducing SHAP explanation, and obtaining a prediction engine for multiple target attributes of double perovskite;
[0014] S4, generating all chemical formulas by element substitution strategy under the constraint of charge neutrality according to the coordination environment and oxidation state of each site of double perovskite, and constructing the candidate space of double perovskite halide;
[0015] S5, based on the prediction engine in S3, gradually narrowing the double perovskite halide candidate space in S4 through the progressive chain of new tolerance factor, space group, SLME, band gap type, band gap value and Ehull;
[0016] S6, the materials in the double perovskite halide candidate space obtained in S5 are sequentially subjected to DFT structure optimization, band gap correction, SLME calculation, and verification of dynamic thermal stability, and finally a high-performance double perovskite halide is selected.
[0017] Further, in S1, the specific process of collecting the space group, SLME, band gap, Ehull, and Ef raw data of the double perovskite halide features in the public database includes:
[0018] Extracting double perovskite halide entries containing chemical formula, space group, SLME, band gap type, band gap value, Ehull, and Ef from the Materials Project and JARVIS databases;
[0019] Divide the samples into high photoelectric conversion efficiency positive class and low efficiency negative class with SLME threshold, take space group number 225 as positive class and the rest as negative class, and divide stable positive class and unstable negative class with Ehull threshold;
[0020] For the resulting positive and negative sample imbalance classification task, use the synthetic minority over-sampling technique to balance the data set and obtain the original data.
[0021] Further, in S2, the specific process of removing highly correlated redundant features using Pearson coefficient and retaining the preferred features using Spearman coefficient for the original data obtained in S1 includes:
[0022] Calculate the basic physical descriptors and calculated descriptors to construct the initial feature set, and calculate the Pearson correlation coefficient matrix, and delete the redundant feature pairs with absolute value of correlation coefficient exceeding the threshold.
[0023] Calculate the Spearman correlation coefficient in the remaining features and retain the features with stronger correlation with the corresponding target variables.
[0024] Further, in S2, the specific process of locking the double perovskite-specific optimal feature subset using genetic algorithm-cross validation includes:
[0025] Use the retained features as the initial population, and each chromosome corresponds to a candidate feature combination;
[0026] Set the F1 value of five-fold cross-validation as the fitness function, and optimize through selection, crossover, and mutation iteration;
[0027] Stop when the average F1 value of cross-validation improves by less than the set threshold for three consecutive generations, output the feature combination corresponding to the highest fitness of the chromosome, and the feature combination is the double perovskite-specific optimal feature subset.
[0028] Further, in S3, multiple machine learning models are trained in parallel based on the optimal feature subset in S2, the best is integrated and introduced SHAP explanation, and the specific process of obtaining the prediction engine for the double perovskite multi-objective attribute includes:
[0029] For the five target attributes of space group, SLME, band gap type, band gap value, and Ehull, seven algorithms of decision tree, random forest, k-nearest neighbor, multi-layer perceptron, extreme gradient boosting, extreme random tree, and LightGBM are trained in parallel on the double perovite-specific optimal feature subset obtained in S2;
[0030] The average F1 or R 2 is taken as the benchmark, and the multiple single models with the best performance for each target attribute are integrated into an integrated model through majority voting or mean strategy;
[0031] Then, the SHAP value is called to analyze the contribution of each feature to the prediction result, and after explanation, the integrated model is used as a unified prediction engine.
[0032] Further, in S4, according to the coordination environment and oxidation state of each site of double perovskite, the periodic table of elements is traversed, and all chemical formulas are generated using element substitution strategy under the constraint of charge neutrality, and the specific process of constructing the double perovskite halide candidate space includes:
[0033] According to the general formula A2BB'X6, the coordination number and oxidation state requirements of A site, B site, B' site, and X site are split, and the candidate elements that meet the coordination environment are screened by traversing the periodic table of elements;
[0034] According to the principle of charge neutrality, the elements of each site are combined and replaced to automatically generate all possible chemical formulas;
[0035] After removing the redundant and uniform format, the obtained chemical formula set is obtained.
[0036] Further, in S5, based on the prediction engine in S3, the double perovskite halide candidate space in S4 is gradually reduced in turn through the new tolerance factor, space group, SLME, band gap type, band gap value, and Ehull progressive chain, and the specific process includes:
[0037] Calculate the new tolerance factor for the S4 candidate space, and screen the double perovskite halide with stable structure;
[0038] Call the S3 space group prediction engine, and only keep the cubic phase double perovskite halide with space group number 225;
[0039] Call the SLME prediction engine of S3, and screen out the double perovskite halide of the high photoelectric conversion efficiency category;
[0040] Call S3 band gap type prediction engine, eliminate indirect band gap, keep direct band gap double perovskite halide;
[0041] Call S3 band gap value regression engine, screen double perovskite halide with band gap value in the target interval;
[0042] Call S3 Ehull prediction engine, only keep double perovskite halide with thermodynamic stability, and the obtained set is the reduced double perovskite halide candidate space.
[0043] Further, in S6, the materials in the double perovskite halide candidate space obtained in S5 are sequentially subjected to DFT structure optimization, band gap correction, SLME calculation and dynamic thermal stability verification, and the specific process of finally selecting high-performance double perovskite halide includes:
[0044] DFT structure optimization is performed on the reduced double perovskite halide candidate space material to obtain double perovskite halide with ground state configuration;
[0045] Band gap is calculated by PBE functional, and double perovskite halide with band gap in the target interval is kept;
[0046] Band gap correction is performed on the retained material by HSE06 functional, and double perovskite halide with direct band gap characteristics is confirmed;
[0047] After calculating the absorption spectrum based on the independent particle approximation method, the SLME value is evaluated, and only double perovskite halide with high SLME is kept;
[0048] Finally, the thermal stability is verified at a set temperature through ab initio molecular dynamics simulation, and the material that passes all verifications is the high-performance double perovskite halide.
[0049] The second aspect of the present application provides an electronic device comprising a memory and a processor, wherein the processor is configured to execute a program in the memory to implement the above-mentioned double perovskite halide material layer-by-layer screening method based on machine learning and physical prior knowledge constraints.
[0050] The third aspect of the present application provides a storage medium comprising computer executable instructions, wherein the computer executable instructions are configured to execute the above-mentioned double perovskite halide material layer-by-layer screening method based on machine learning and physical prior knowledge constraints when executed by a computer processor.
[0051] Compared with the prior art, the present application has the following beneficial effects:
[0052] 1) The application proposes a cascade screening strategy with SLME as the core index, combining multiple key performance indicators such as space group, band gap characteristics and thermodynamic stability, to establish a hierarchical screening logic driven by physical priori and a machine learning prediction model, to realize multi-dimensional screening from structure, optics, electricity to stability, significantly reduce the candidate space, and improve the physical credibility of the screening results.
[0053] 2) The application realizes a "prediction-verification" closed-loop integration framework. Compared with the traditional DFT calculation method, the method realizes about 10 9 times of calculation acceleration while maintaining accuracy. After predicting multiple physical properties by machine learning, high-potential materials are selected for DFT verification, and multiple excellent candidate materials with high SLME, direct band gap, good absorption performance and thermal stability are successfully screened out. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 is the method flowchart in application example 1 of the application;
[0055] Figure 2 is the process flowchart of combining genetic algorithm and cross-validation to select the optimal feature set in application example 1 of the application;
[0056] Figure 3 is the process flowchart of constructing a progressive learning model in application example 1 of the application;
[0057] Figure 4 is a diagram of each site element of a double perovskite halide in application example 1 of the application;
[0058] Figure 5 is a multi-attribute screening schematic diagram in application example 1 of the application. DETAILED DESCRIPTION
[0059] The application will be described in detail below in combination with the drawings and specific embodiments. In this technical solution, if the component model, material name, connection structure, circuit structure, control method, algorithm and other features are not explicitly stated, they are considered as common technical features disclosed in the prior art.
[0060] Example 1
[0061] The double perovskite halide material layer-by-layer screening method based on machine learning and physical priori knowledge constraints in this embodiment includes the following steps:
[0062] Step 1: Data collection and preprocessing. Extract relevant data from the Materials Project (MP) and JARVIS databases, focus on space group, SLME, band gap type, band gap value, convex hull energy (Ehull), and form key attributes such as formation energy (Ef).
[0063] Step 2: Feature engineering. Based on the data obtained in Step 1, calculate the basic physical descriptors and calculated descriptors, calculate based on the Pearson correlation coefficient matrix, remove the strong correlation feature pairs with absolute value greater than 0.8, and retain the features with stronger correlation with the target variable according to the Spearman correlation coefficient, to reduce the risk of overfitting caused by feature redundancy and multicollinearity. The optimal feature subset is obtained by using the feature selection strategy combining genetic algorithm and cross-validation.
[0064] Step 3: Model training and selection. Based on the feature dataset obtained after completing the feature engineering in Step 2, train and evaluate multiple machine learning models for each target property, and use ensemble algorithms to improve model prediction ability to obtain machine learning prediction model. Introduce SHAP (SHapley Additive exPlanations) method to analyze the explainability of machine learning prediction model.
[0065] Step 4: Construction of candidate space. According to the coordination environment and oxidation state of each site of the desired material, elements are screened from the periodic table. Based on the element substitution strategy and charge neutrality principle, a large chemical space of candidate compounds is constructed.
[0066] Step 5: Multi-attribute screening. The candidate materials constructed in Step 4 are first screened for stability according to the new tolerance factor, and then the properties of the candidate materials are predicted by the multiple different target attribute models trained in Step 3. The materials are gradually screened through a progressive decision chain strictly following the order of space group symmetry, SLME, band gap characteristics and thermodynamic stability.
[0067] Step 6: First-principles calculation (DFT) verification. A series of DFT calculations are performed on a small number of candidate materials screened in Step 5 to further verify the performance of the materials.
[0068] In specific implementation, the specific process of data collection and preprocessing in Step 1 is as follows:
[0069] Step 1.1: Extract relevant data from MP and JARVIS databases, including chemical formula, space group, SLME, band gap type, band gap value, Ehull and Ef, etc.
[0070] Step 1.2: For the class labels of the classification task, the value of SLME is divided into two categories according to whether it reaches 20%: high photoelectric conversion efficiency materials (positive class) and low photoelectric conversion efficiency materials (negative class), as the label output by the SLME classification prediction model.
[0071] Step 1.3: The space group target attribute is classified as positive class according to space group number 225, and non-225 space group is classified as negative class as the class label of the machine learning model classification task.
[0072] Step 1.4: For the Ehull classification prediction model, Ehull < 0.02 eV / atom is taken as the positive class label, and Ehull >= 0.02 eV / atom is taken as the negative class label.
[0073] Step 1.5: For the data in the initial data set of the classification task, the distribution is unbalanced between positive and negative materials, and the synthetic minority over-sampling technique is used to balance the data set by over-sampling the minority class to avoid imbalance in model prediction.
[0074] In specific implementation, the specific steps of feature engineering in step 2 are as follows:
[0075] Step 2.1: Based on the data obtained in step 1, a feature set is obtained from the calculation of basic physical descriptors and calculated descriptors based on basic physical descriptors from each crystal site.
[0076] Step 2.2: Calculate the Pearson correlation coefficient matrix, and remove the feature pairs with a correlation coefficient greater than 0.8. In the highly correlated feature pairs, the feature with stronger correlation with the target variable is retained based on the Spearman correlation coefficient.
[0077] Step 2.3: To further determine the optimal feature subset corresponding to each target attribute (space group classification, SLME classification, band gap type classification, band gap value regression, Ehull classification), a feature selection strategy combining genetic algorithm and cross-validation is used.
[0078] In specific implementation, the specific process of model training and selection in step 3 is as follows:
[0079] Step 3.1: Based on the feature data set obtained after completing the feature engineering in step 2, for each target attribute to be predicted (space group classification, SLME classification, band gap type classification, band gap value regression, Ehull classification), seven different machine learning algorithms are used for model training: decision tree (DT), random forest (RF), k-nearest neighbor (KNN), multilayer perceptron (MLP), extreme gradient boosting (XGBoost), extreme random tree (EXT), and light gradient boosting machine (LightGBM).
[0080] Step 3.2: For each model trained by each algorithm, use five-fold cross-validation for evaluation, calculate and record the average performance score of cross-validation, and take the average score as the main evaluation index of model performance.
[0081] Step 3.3: Establish a progressive learning model to train the Ef prediction model, and use the predicted Ef as a tool descriptor in the SLME feature set to train the SLME classification model on the optimized feature set.
[0082] Step 3.3: For each target property, select the best model from the seven models and perform ensemble learning on the cross-validated trained models to obtain the final machine learning prediction model. For classification problems (SLME classification, space group classification, band gap type classification, Ehull classification), use majority voting to determine the final prediction class. For regression problems (band gap value prediction), use the average of the model prediction values to determine the final prediction value.
[0083] Step 3.4: Use the SHAP method to perform explainability analysis on the features of the machine learning prediction model corresponding to each property.
[0084] In specific implementation, the specific steps of the candidate space construction process in step 4 are as follows:
[0085] Step 4.1: Traverse the periodic table of elements, and screen elements that meet the conditions according to the different site coordination environments and different oxidation states of elements in double perovskites.
[0086] Step 4.2: Based on the element substitution strategy and strictly following the charge neutrality principle, combine all possible chemical formulas to construct a large chemical space of candidate compounds.
[0087] In specific implementation, the specific content of multi-attribute screening in step 5 is as follows:
[0088] Step 5.1: Calculate the new tolerance factor of the large candidate space constructed in step 4, and screen out materials with a value less than 4.18, i.e., structurally stable materials.
[0089] Step 5.2: Use the space group classification model trained in step 3 to predict whether the structurally stable material has a space group of 225, and screen out cubic phase materials that meet the space group of 225.
[0090] Step 5.3: Use the SLME classification model trained in step 3 to predict the SLME category of the material screened in step 5.2, and screen out SLME >= 25% materials.
[0091] Step 5.4: Use the band gap type classification model obtained in step 3 to predict the band gap type of the material screened in step 5.3, and remove indirect band gap materials to obtain direct band gap materials.
[0092] Step 5.5: Further use the band gap regression model obtained in step 3 to predict the band gap value of the direct band gap material screened above, and select candidate materials with a value in the range of 0.5-1.5 eV for further research.
[0093] Step 5.6: Use the Ehull classification model trained in Step 3 to predict the thermodynamic stability of the material, and finally screen out cubic phase double perovskite materials that meet the requirements of stability, suitable direct band gap and high photoelectric conversion efficiency.
[0094] In practice, the specific content of the DFT verification in step 6 is as follows:
[0095] Step 6.1: Perform DFT verification on the small number of double perovskite materials obtained through step 5, optimize the structure to obtain a stable configuration for each material, further calculate the PBE (Perdew-Burke-Ernzerhof) band gap, and select candidate materials with values in the range of 0.5-1.5 eV for the next verification step.
[0096] Step 6.2: Further calculate the more accurate band gap for the candidate materials screened by PBE using the HSE06 (Heyd-Scuseria-Ernzerhof) hybrid functional, and select direct band gap materials with values in the range of 1-2 eV.
[0097] Step 6.3: To calculate the photovoltaic efficiency of the candidate material as a solar cell absorber layer material, the absorption spectrum of the candidate material is first calculated based on the independent particle approximation method, and then the SLME values of the material in different thickness ranges at 293.15 K are calculated.
[0098] Step 6.4: Verify the thermal stability of the candidate material at 300K and 600K using ab initio molecular dynamics simulations.
[0099] Application Example 1
[0100] like Figure 1 As shown, this application example demonstrates a cascade screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints. The specific steps are as follows:
[0101] Step 1: Data Collection and Preprocessing
[0102] Step 1.1: Based on the studied double perovskite halide system, 741 compound data entries with space group number 225 were extracted from the JARVIS database to obtain high-precision SLME data. 1291 double perovskite halide data entries with space group number 225 were extracted from the MP database, including information such as chemical formula, space group, band gap type, band gap value, Ehull, and Ef. Simultaneously, to construct a space group classification model, additional data with space group numbers other than 225 were extracted from the MP database as negative samples.
[0103] Step 1.2: For the class label of the classification task, SLME ≥ 20% is marked as high photoelectric conversion efficiency material (positive class), and SLME < 20% is marked as low photoelectric conversion efficiency material (negative class) as the label output by the SLME classification prediction model.
[0104] Step 1.3: The space group target attribute is marked as 225 as the positive class, and the non-225 space group is marked as the negative class as the class label of the machine learning model classification task.
[0105] Step 1.4: For the Ehull classification prediction model, Ehull < 0.02 eV / atom is marked as the positive class label, and Ehull ≥ 0.02 eV / atom is marked as the negative class label.
[0106] Step 1.5: For the data with unbalanced distribution between positive and negative materials in the initial data set of the classification task, the synthetic minority over-sampling technique is used to balance the data set by over-sampling the minority class.
[0107] Step 2: Feature engineering
[0108] Step 2.1: For the data obtained in step 1, the matminer tool is used to extract features from the chemical formula of the material for the data set corresponding to SLME, and 139 features are extracted based on the statistical features of element properties, including atomic mass, electronegativity, atomic radius, and atomic orbital characteristics such as valence electron configuration and ionization energy.
[0109] Step 2.2: For the data set obtained from the MP database, 84 descriptors are generated by calculating 12 basic physical descriptors and calculated descriptors based on basic physical descriptors, such as density, dipole polarizability, covalent radius, atomic radius, first ionization energy, valence electron number, electronegativity, and s, p and d electron number, from each crystal position.
[0110] Step 2.3: Calculate the Pearson correlation coefficient matrix, and remove the feature pairs with a correlation coefficient greater than 0.8. In the highly correlated feature pairs, the features with stronger association with the target variable are retained based on the Spearman correlation coefficient. After feature pruning, the number of features in the SLME data set is reduced to 70, and the number of features in the MP data set is reduced to 27.
[0111] Step 2.4: As shown in Figure 2 , a feature selection strategy combining genetic algorithm and cross-validation is used to further determine the optimal feature subset for each target attribute.
[0112] Step 3: Model training and selection
[0113] Step 3.1: Based on the dataset obtained after feature selection in Step 2, seven different machine learning algorithms are used for model training for each target attribute to be predicted (SLME classification, space group classification, band gap type classification, band gap value regression, Ehull classification): Decision Tree (DT), Random Forest (RF), k-Nearest Neighbors (KNN), Multilayer Perceptron (MLP), Extreme Gradient Boosting (XGBoost), Extremely Randomized Trees (EXT), and Light Gradient Boosting Machine (LightGBM).
[0114] Step 3.2: For each model trained by each algorithm, five-fold cross-validation is used for evaluation, and the average performance score of cross-validation is calculated and recorded as the main evaluation indicator of model performance.
[0115] Step 3.3: For each target attribute, the best-performing model is selected from the seven models. For the space group classification model, the XGBoost model performs best, and multiple models trained by five-fold cross-validation are integrated to form the final ensemble space group prediction model.
[0116] Step 3.3: For the SLME classification model, a progressive learning model (as shown in Figure 3 ) is established, and the Ef regression model is trained, with the predicted Ef as a tool descriptor included in the SLME feature set. The SLME classification model is trained on the optimized feature set, with the EXT classifier having the best overall performance. Multiple models trained by five-fold cross-validation are integrated to form the final ensemble SLME prediction model.
[0117] Step 3.4: For the band gap type classification and band gap value regression models, XGBoost and EXT perform best, respectively, and the trained models are integrated.
[0118] Step 3.5: For the Ehull classification model, XGBoost performs best on average, and the model is integrated to form the final ensemble Ehull prediction model.
[0119] Step 3.6: The SHAP method is used for explainability analysis of the features of the machine learning prediction model corresponding to each attribute.
[0120] Step 4: Construction of candidate space
[0121] Step 4.1: Traverse the periodic table, and according to the coordination environment of different sites and different oxidation states of elements in double perovskite (general formula A2BBX'6), screen elements that meet the conditions (as shown in Figure 4 ), determine 7 kinds of A + and 15 kinds of A 2+ cations meet the 12-coordination standard, and 10 kinds of B+ / B' + , 34 kinds of B 2+ / B' 2+ and 45 kinds of B 3+ / B' 3+ Cations meet the 6-coordination requirement.
[0122] Step 4.2: Based on the element substitution strategy and strictly adhering to the principle of charge neutrality, all possible chemical formulas were combined to generate a large chemical space containing 29,364 candidate compounds according to the elements obtained at each site.
[0123] Step 5: Multi-attribute screening (as shown in Figure 5 )
[0124] Step 5.1: Calculate the new tolerance factor of the large candidate space constructed in step 4, and screen out materials with a value less than 4.18, i.e., 8818 structurally stable materials.
[0125] Step 5.2: For structurally stable materials, use the space group classification model trained in step 3 to predict whether their space group is 225, and screen out cubic phase materials that meet the space group 225, reducing the candidate space to 7037.
[0126] Step 5.3: Use the SLME classification model trained in step 3 to predict the SLME category of the materials screened in step 5.2, and obtain 2460 materials with SLME >= 25%.
[0127] Step 5.4: Use the band gap type classification model obtained in step 3 to predict the band gap type of the materials screened in step 5.3, and eliminate indirect band gap materials to obtain 863 direct band gap candidate materials.
[0128] Step 5.5: Further use the band gap regression model obtained in step 3 to predict the band gap value of the direct band gap materials screened above, and select 163 candidate materials with values in the range of 0.5-1.5 eV for further research.
[0129] Step 5.6: Use the Ehull classification model trained in step 3 to predict the thermodynamic stability of the materials, and finally screen out 34 cubic phase double perovskite materials that meet the requirements of stability, appropriate direct band gap, and high photoelectric conversion efficiency.
[0130] Step 6: DFT verification
[0131] Step 6.1: DFT verification was performed on the 34 double perovskite materials screened in Step 5. First, structural optimization was performed to obtain the stable configuration of each material, and further calculation of the PBE (Perdew-Burke-Ernzerhof) band gap was performed to select candidate materials with values in the range of 0.5-1.5 eV.
[0132] Step 6.2: Further calculation of the more accurate band gap using HSE06 (Heyd-Scuseria-Ernzerhof) was performed on the candidate materials screened above using PBE, resulting in 3 materials with direct band gaps in the range of 1-2 eV.
[0133] Step 6.3: To calculate the photovoltaic efficiency of the candidate materials as solar cell absorption layer materials, first, the absorption spectrum of the candidate materials was calculated based on the independent particle approximation method, and further calculation of the SLME value of the materials in the range of different thicknesses at 293.15 K was performed, with the SLME values of the 3 candidate materials all greater than 26%.
[0134] Step 6.4: Verification of the thermal stability of the candidate materials at 300 K and 600 K was performed using ab initio molecular dynamics simulation, and finally the thermal stability of the 3 materials at 300 K and 600 K was confirmed. Thus, the entire screening and verification process of high-performance candidate materials of double perovskite halide was completed, and finally 3 stable candidate materials with excellent comprehensive performance were obtained.
[0135] The double perovskite halide material layer-by-layer screening method based on machine learning and physical prior knowledge constraints in this application example addresses the problems of low efficiency, large candidate space, and insufficient multi-performance balance in traditional data-driven screening frameworks. A cascade screening method guided by the spectral limit maximum efficiency (SLME) is constructed, and through strict adherence to the progressive decision chain of space group symmetry, SLME, band gap characteristics, and thermodynamic stability, high-precision photovoltaic material discovery is achieved. Based on the element coordination environment and charge neutrality principle, the candidate material space is constructed, first screened for structural stability with a new tolerance factor, then screened layer by layer through a machine model, and finally 3 high-performance materials are screened out, providing a new solution for the development of high-efficiency solar cells.
[0136] The above description of the embodiments is for the convenience of those of ordinary skill in the art to understand and use the invention. Those skilled in the art can easily make various modifications to these embodiments, and apply the general principles described herein to other embodiments without having to go through creative labor. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made by those skilled in the art without departing from the scope of the invention should be within the scope of protection of the invention.
[0137] Example 2
[0138] The embodiment provides an electronic device, including a memory and a processor, the processor is used for executing the program in the memory, so as to realize the double perovskite halide material layer-by-layer screening method based on machine learning and physical prior knowledge constraint. The processor can be a general processor, including a central processing unit (CPU), a network processor (NP) and the like; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components; the memory can contain a random access memory (RAM), and can also include a non-volatile memory (Non-Volatile Memory), for example, at least one disk memory. The memory can be an internal memory of a random access memory (RAM) type, and the processor and the memory can be integrated into one or more independent circuits or hardware, such as an application specific integrated circuit (ASIC). It should be noted that the computer program in the memory can be realized in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application.
[0139] Embodiment 3
[0140] The embodiment provides a storage medium containing computer executable instructions, which, when executed by a computer processor, are used to execute the method for layer-by-layer screening of double perovskite halide materials based on machine learning and physical prior knowledge constraints as described above. The storage medium can be an electronic medium, a magnetic medium, an optical medium, an electromagnetic medium, an infrared medium or a semiconductor system or a propagation medium. The storage medium can also include a semiconductor or solid state memory, a magnetic tape, a removable computer disk, a random access memory (RAM), a read-only memory (ROM), a hard disk and an optical disk. The optical disk can include a compact disk-read only memory (CD-ROM), a compact disk-read / write (CD-RW) and a DVD.
Claims
1. A layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints, characterized in that, Includes the following steps: S1. Collect raw data from publicly available databases, including spatial groups, SLMEs, band gaps, Ehull, and Ef data on double perovskite halide characteristics. S2. The original data obtained from S1 is first removed by using the Pearson coefficient to eliminate highly correlated redundant features, and then the Spearman coefficient is used to select the best features to retain. The optimal feature subset for double perovskite is locked by genetic algorithm-cross-validation. S3. Based on the optimal feature subset in S2, train multiple machine learning models in parallel, select the best to integrate and introduce SHAP interpretation to obtain a prediction engine for the multi-objective properties of double perovskite. S4. Based on the coordination environment and oxidation state of the double perovskite traversing the periodic table, all chemical formulas are generated by element substitution strategy under charge neutrality constraints, thus constructing a candidate space for double perovskite halides. S5. Based on the prediction engine in S3, the candidate space of double perovskite halides in S4 is gradually narrowed down through the new tolerance factor, space group, SLME, bandgap type, bandgap value, and Ehull progressive chain. S6. For the materials in the candidate space of double perovskite halides obtained in S5, DFT structural optimization, band gap correction, SLME calculation and kinetic thermal stability verification are carried out in sequence to finally select high-performance double perovskite halides.
2. The layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints as described in claim 1, characterized in that, In S1, the specific process of collecting raw data on space groups, SLMEs, band gaps, Ehull, and Ef features of double perovskite halides from publicly available databases includes: Entries for double perovskite halides containing chemical formula, space group, SLME, band gap type, band gap value, Ehull, and Ef were extracted from the Materials Project and JARVIS databases. The samples were divided into a positive class with high photoelectric conversion efficiency and a negative class with low efficiency using the SLME threshold, with spatial group number 225 as the positive class and the rest as the negative class, and the stable positive class and unstable negative class using the Ehull threshold. For classification tasks with imbalanced positive and negative samples, a synthetic minority class oversampling technique is used to balance the dataset and obtain the original data.
3. The layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints as described in claim 1, characterized in that, In S2, the specific process of first removing highly correlated redundant features from the original data obtained in S1 using the Pearson coefficient, and then selectively retaining them using the Spearman coefficient, includes: The basic physical descriptor and computational descriptor are calculated on the raw data obtained from S1 to construct the initial feature set, and the Pearson correlation coefficient matrix is calculated. Redundant feature pairs with absolute values of correlation coefficients exceeding the threshold are removed. Calculate the Spearman correlation coefficient among the remaining features and retain the features that are more strongly associated with the corresponding target variable.
4. The layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints as described in claim 3, characterized in that, In S2, the specific process of locking the optimal feature subset for dual perovskite using genetic algorithm-cross-validation includes: The retained features are used as the initial population, and each chromosome corresponds to a candidate feature combination; The F1 score of the five-fold cross-validation was set as the fitness function, and iterative optimization was performed through selection, crossover, and mutation. When the average F1 score of cross-validation is lower than the set threshold for three consecutive generations, the process stops and the feature combination corresponding to the chromosome with the highest fitness is output. This feature combination is the optimal feature subset for dual perovskite applications.
5. The layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints as described in claim 1, characterized in that, In S3, the specific process of training multiple machine learning models in parallel based on the optimal feature subset from S2, selectively integrating them, and introducing SHAP interpretation to obtain a prediction engine for the multi-objective properties of dual perovskites includes: For the five target attributes of spatial group, SLME, bandgap type, bandgap value, and Ehull, seven algorithms, namely decision tree, random forest, k-nearest neighbor, multilayer perceptron, extreme gradient boosting, extreme random tree, and LightGBM, were trained in parallel on the dual perovskite-specific optimal feature subset obtained by S2. Five-fold cross-validation average F1 or R 2 Based on this, multiple single models that perform best in each objective attribute are integrated into an ensemble model through majority voting or mean strategy. Then, the SHAP value is called to analyze the contribution of each feature to the prediction result. After the interpretation is completed, the ensemble model is used as a unified prediction engine.
6. The layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints as described in claim 1, characterized in that, In S4, based on the coordination environment and oxidation state of the double perovskite traversing the periodic table, and under charge-neutral constraints, an element substitution strategy is used to generate all chemical formulas. The specific process for constructing the candidate space for double perovskite halides includes: Based on the general formula A2BB′X6, the coordination number and oxidation state requirements of the A, B, B′ and X positions are split, and the periodic table is traversed to screen candidate elements that meet the coordination environment requirements. Based on the principle of charge neutrality, the elements at each point are combined and replaced to automatically generate all possible chemical formulas; After deduplication and format unification, the resulting set of chemical formulas yields a candidate space for double perovskite halides.
7. The layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints as described in claim 1, characterized in that, In S5, based on the prediction engine in S3, the specific process of gradually narrowing down the candidate space of double perovskite halides in S4 through the new tolerance factor, space group, SLME, bandgap type, bandgap value, and Ehull progressive chain includes the following steps: A new tolerance factor was calculated for the S4 candidate space to screen out structurally stable double perovskite halides. The S3 space group prediction engine was invoked, and only the cubic phase double perovskite halide with space group number 225 was retained. The SLME prediction engine of S3 is invoked to screen out double perovskite halides with high photoelectric conversion efficiency. The S3 bandgap type prediction engine is invoked to remove indirect bandgap double perovskite halides and retain direct bandgap double perovskite halides. The S3 bandgap regression engine is invoked to screen out double perovskite halides whose bandgap values are within the target range. The Ehull prediction engine of S3 is invoked, and only thermodynamically stable bisperovskite halides are retained. The resulting set is the reduced bisperovskite halide candidate space.
8. The layer-by-layer screening method for dual perovskite halide materials based on machine learning and physical prior knowledge constraints as described in claim 1, characterized in that, In S6, the specific process of sequentially performing DFT structural optimization, bandgap correction, SLME calculation, and kinetic thermal stability verification on the materials in the candidate space of double perovskite halides obtained in S5, and finally selecting high-performance double perovskite halides, includes: DFT structural optimization was performed on the scaled-down candidate space materials of double perovskite halides to obtain the ground state configuration of the double perovskite halides. The band gap was calculated using the PBE functional, and double perovskite halides with band gaps in the target range were retained. Bandgap correction of the retained material was performed using the HSE06 functional to confirm the direct bandgap characteristics of the double perovskite halide. After calculating the absorption spectrum based on the independent particle approximation method, the SLME value was evaluated, and only double perovskite halides with high SLME were retained. Finally, the thermal stability was verified at a set temperature through ab initio molecular dynamics simulations. The material that passed all the verifications is the high-performance double perovskite halide.
9. An electronic device, comprising a memory and a processor, characterized in that, The processor is used to execute the program in the memory to implement the layer-by-layer screening method for double perovskite halide materials based on machine learning and physical prior knowledge constraints as described in any one of claims 1 to 8.
10. A storage medium containing computer-executable instructions, characterized in that, When executed by a computer processor, the storage medium of the computer-executable instructions is used to perform the layer-by-layer screening method for double perovskite halide materials based on machine learning and physical prior knowledge constraints as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method for quickly predicting double perovskite oxide band gap based on data mining
CN112132185A
Cited By
Shaft sand migration monitoring experiment system based on distributed optical fiber sound wave sensing and migration speed prediction method
CN121298192A