Method and system for optimizing preparation process of perovskite solar cell
By building databases, screening key features and optimizing machine learning models, the problems of insufficient machine learning prediction accuracy and inefficiency of traditional experimental methods in the field of perovskite solar cells are solved, and the photoelectric conversion efficiency of perovskite solar cells is improved and the research cycle is shortened.
Patent Information
- Application Number
- CN202510265979.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
In the field of perovskite solar cells, machine learning methods have insufficient accuracy in predicting photoelectric conversion efficiency, and traditional experimental trial and error methods have problems of long research cycles and inefficiency.
A perovskite solar cell preparation process optimization method is adopted to optimize the preparation process by constructing databases, screening key features, building and optimizing machine learning models, and designing experiments. Specific steps include data cleaning, feature importance sorting, model hyperparameter tuning, experimental design and process optimization.
It improves data prediction accuracy, shortens the research cycle, improves the photoelectric conversion efficiency of perovskite solar cells, and provides a new perspective and scientific basis for the rapid development of high-efficiency solar cells.
Smart Images

Figure CN120197477A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of the preparation of perovskite solar cells, and in particular, to a method and system for optimizing the preparation process of perovskite solar cells. Background Art
[0002] Since perovskite solar cells (PSCs) were first reported in 2009, they have become an important research direction for new solar cells due to their high power conversion efficiency (PCE), low cost, and simple manufacturing process. Perovskite materials, especially lead-based perovskite compounds, have a high PCE, and the bandgap and light absorption characteristics of perovskite materials make them very suitable for use in photovoltaic devices.
[0003] Due to the above characteristics of perovskite solar cells, various domestic and foreign enterprises and teams have invested in the research and development of new technologies. However, the traditional experimental trial-and-error method has the deficiencies of a long research cycle and low efficiency in improving the PCE of PSCs.
[0004] In 2019, Li and his team used machine learning (ML) technology to study the bandgaps of various perovskite materials and their effects on the efficiency of PSCs. The research team constructed an ML model to predict the bandgap and energy levels based on the material composition, thus successfully predicting the PCE of PSCs, which was consistent with the theoretical prediction of the Shockley-Queisser limit, providing ideas for the synthesis of new perovskite materials.
[0005] Although there has been some progress in the application of ML in the field of PSCs, there are still the following deficiencies: there are various ML methods, but the matching degrees with experimental data are different, resulting in insufficient prediction accuracy. Therefore, a new technical solution is proposed. Summary of the Invention
[0006] In order to improve the accuracy of data prediction, the present application provides a method and system for optimizing the preparation process of perovskite solar cells.
[0007] In a first aspect, the present application provides a method for optimizing the preparation process of a perovskite solar cell, adopting the following technical solution:
[0008] A method for optimizing the preparation process of a perovskite solar cell, characterized by comprising:
[0009] Step 1: Construct a database, which includes:
[0010] Using experimental data as the basic data, clean the data and perform standardization processing. Set the PCE threshold to classify the batteries into high-efficiency batteries and low-efficiency batteries, and obtain a data set;
[0011] Step 2: Screen key features, which include:
[0012] Apply the random forest algorithm to rank the feature importance of the data set to obtain the ranking of feature importance order;
[0013] Based on the feature importance order ranking and Pearson correlation coefficient, screen out the feature data set one with weak correlation;
[0014] Use the recursive feature elimination method to reduce the data dimension of the feature data set one again to obtain the feature data set two;
[0015] Step 3: Model construction and optimization, which include:
[0016] Build and train the screened feature data set two, apply the Optuna method to tune the hyperparameters of all models, and screen out the optimal model among logistic regression, random forest, support vector machine, K-nearest neighbor, gradient boosting, and extreme gradient boosting according to the weighted score of the preset evaluation index for prediction tasks;
[0017] Step 4: Design experiments and optimize the preparation process, which include:
[0018] Based on the optimal model, analyze the possible reasons for high photoelectric conversion efficiency for the screened feature data set two, and design experiments to obtain experimental conclusions;
[0019] Based on the experimental conclusions, obtain the best experimental conditions and optimize the preparation process.
[0020] Optionally, the data cleaning includes:
[0021] Preliminarily screen the basic data according to a preset standard, and retain the data that passes the screening;
[0022] Check for outliers in the retained data and delete the outliers.
[0023] Optionally, group the data into continuous data and discrete data based on the data type and add a prefix for identification; among them, the prefix represents the data type;
[0024] Identify the prefix of the data grouping, and use the preset corresponding processing methods to process the data of different groups respectively; among them, perform maximum-minimum normalization on the continuous data and one-hot encoding processing on the discrete data;
[0025] A data balancing method that combines synthetic minority over-sampling technique and Tomek links is used to balance the data.
[0026] Optionally, based on the selected optimal model, perform performance metric analysis on the original training set of unbalanced data;
[0027] Analyze the impact of whether the training set is balanced on the model's prediction ability based on the performance metric analysis results.
[0028] Optionally, the balancing process includes:
[0029] Divide the data in the dataset into a training set and a test set at a preset ratio based on the method of stratified sampling;
[0030] Perform balancing processing on the training set and the test set;
[0031] Perform sample classification based on a preset photovoltaic conversion efficiency threshold to obtain high photovoltaic conversion efficiency samples and low photovoltaic conversion efficiency samples.
[0032] Optionally, optimize the optimal model, which includes:
[0033] Based on the selected optimal model, use the Optuna hyperparameter optimization method to optimize the model performance, and evaluate the feature importance to find the features that have the highest impact on high photovoltaic conversion efficiency samples;
[0034] Identify the features whose importance impact on distinguishing high photovoltaic conversion efficiency samples and low photovoltaic conversion efficiency samples reaches a preset standard.
[0035] Optionally, the selection of the optimal model according to the preset evaluation metrics includes:
[0036] Evaluate the model based on one or more metrics such as accuracy, recall, precision, the harmonic mean of precision and recall, and the area under the curve;
[0037] Compare the performance of multiple models based on each evaluation metric;
[0038] Score the model based on each evaluation metric and determine the optimal model; where the weights of the evaluation metrics are area under the curve > harmonic mean of precision and recall > precision > recall > accuracy.
[0039] Optionally, the method for identifying the importance impact on distinguishing high photovoltaic conversion efficiency samples and low photovoltaic conversion efficiency samples includes:
[0040] Based on the optimal model and using the SHAPley Additive exPlanations method to analyze the contribution values of the features in the second set of feature data to the model prediction, and generate the analysis results of the model prediction contribution values;
[0041] The described designed experiment includes: based on the analysis results of the model prediction contribution value, using the control variable method to regulate each feature, and designing the experimental conditions for multiple experiments;
[0042] The obtaining of the optimal experimental conditions based on the experimental conclusions and the optimization of the preparation process include:
[0043] Obtaining the experimental data of multiple experiments under different experimental conditions and analyzing them using the optimal model to generate experimental analysis conclusions and obtain the optimal experimental conditions;
[0044] Based on the optimal experimental conditions, optimizing the preparation process.
[0045] In a second aspect, the present application provides a preparation process optimization system for perovskite solar cells, adopting the following technical solution:
[0046] A preparation process optimization system for perovskite solar cells includes a memory and a processor, and the memory stores a computer program that can be loaded and executed by the processor for the preparation process optimization method of the perovskite solar cells as described in claims 1 - 8.
[0047] In summary, the present application includes the following beneficial technical effects: overcoming the problem of low prediction accuracy in the application of ML (i.e., machine learning) in the field of PSCs, effectively solving the limitations of the traditional experimental trial - and - error method, providing a new perspective and scientific basis for the rapid development of high - efficiency PSCs, and also providing a reference for the development of other new solar cell technologies. Brief Description of the Drawings
[0048] Figure 1 is the working flow chart of the optimization of the preparation process of this method;
[0049] Figure 2 is the flow chart of this method;
[0050] Figure 3 is the schematic diagram of the relationship between the AUC value and the PCE threshold in this method;
[0051] Figure 4 is the sample distribution diagram when the PCE threshold is 18.5 in this method;
[0052] Figure 5 is the PCC matrix diagram between features and between features and PCE in this method;
[0053] Figure 6 is the diagram of the optimal number of features selected by the RFE method in this method;
[0054] Figure 7 is the PCE key feature diagram after sorting the feature importance in this method;
[0055] Figure 8 is the performance evaluation index diagram of the ML model in this method;
[0056] Figure 9 is the bar chart of the weighted scores of the evaluation indexes in this method;
[0057] Figure 10 is the bar chart of the feature importance based on the optimal model in this application;
[0058] Figure 11 is the distribution diagram of the feature data to be optimized in this application;
[0059] Figure 12 is the current-voltage curve diagram of the battery prepared based on Condition 2 in this application;
[0060] Figure 13 is the external quantum efficiency spectrum and the corresponding current density response diagram of the battery prepared based on Condition 2 in this application. Embodiment
[0062] The following combines the attached Figures 1-13 to further elaborate on this application in detail.
[0063] The embodiment of this application discloses a preparation process optimization method for perovskite solar cells.
[0064] Referring to Figure 1 and Figure 2 , the preparation process optimization method for perovskite solar cells (hereinafter referred to as PSCs) mainly constructs a power conversion efficiency (hereinafter referred to as PCE) prediction model by applying a variety of machine learning (ML) algorithms, and finally selects an optimal model for visual analysis and experimental verification.
[0065] This method first constructs a detailed synthesis process database by identifying potential optimization points in the perovskite material synthesis process, including discrete features (such as the materials used in specific layers) and continuous features (such as the amount of solvent or additive added in specific layers). Subsequently, key features are screened out from the collected PSCs data through feature engineering techniques, and a variety of ML algorithms are further applied to construct a model for predicting PCE, and the optimal ML model is selected. Finally, the preparation process and method are optimized according to the prediction of the model by interpreting the model output and designing experiments to verify the prediction of the model.
[0066] This method solves the key problems such as the insufficient prediction accuracy of ML in the field of PSCs and the long R & D cycle of traditional experimental methods. It not only provides a new perspective and scientific basis for the development of high-efficiency PSCs, but also brings innovative ideas to the research and development work in this field.
[0067] Referring to Figure 1 and Figure 2 , this application mainly includes the following four steps:
[0068] Step 1: Construct a database, which includes:
[0069] 1. Using experimental data as the basic data, cleaning the data, and performing standardization processing. Setting a "PCE threshold" to classify the batteries into high-efficiency batteries and low-efficiency batteries to obtain a dataset.
[0070] Among them, the experimental data comes from the laboratory and can include multiple experimental records of perovskite materials. The experimental records can be as many as possible, such as about 800; for most of the adjustable material synthesis processes specifically involved in the experiment, a small dataset is established. Each record consists of a set of feature vectors and the corresponding PCE value.
[0071] Subsequently, analyze the features included in the dataset and clean the data. The data cleaning in Step 1 includes:
[0072] 1) Preliminarily screen the data according to a preset standard (example: excluding the optoelectronic performance parameters "open circuit voltage", "short circuit current", "fill factor", "light intensity", etc. that are known to significantly affect PCE), and retain the data that passes the screening;
[0073] 2) Check for outliers in the retained data and delete the outliers;
[0074] Example: By setting conditions such as open circuit voltage > 1.1V, fill factor > 75%, 15mA / cm 2 < short circuit current < 23mA / cm 2 , PCE > 15% to screen out the data that meets the requirements to improve the quality and reliability of the data.
[0075] At the same time, in order to ensure that different features have the same dimension in model training, the data can be normalized. Through these data preprocessing steps, the improvement of data quality and the effectiveness of subsequent model training are ensured.
[0076] The above standardization processing mainly includes:
[0077] 1) Group the data into continuous data and discrete data based on the data type and add a prefix for identification; among them, the prefix represents the data type;
[0078] Classify data according to its characteristics. For example, divide it into discrete data and continuous data, and add corresponding prefixes. For example, cat represents discrete data and num represents continuous data. Use different methods to balance data for different classifications.
[0079] 2) Identify the prefix of the data group and process the data of different groups using the preset corresponding processing method; among which, the continuous data is normalized to the maximum and minimum, and the discrete data is processed by one-hot encoding;
[0080] The data type can be easily identified and analyzed based on the prefix, and then processed specifically according to the data type.
[0081] Example: For discrete data, each categorical variable is encoded using the one-hot encoding method.
[0082] Among them, one-hot encoding is a technique for converting categorical variables into binary vectors. The principle is to create a new feature for each category and represent it as a binary vector, in which only one position is 1 and the other positions are 0. Each category corresponds to a unique vector. This method can avoid the introduction of false order relationships in the model by categorical variables while retaining category information.
[0083] For continuous data, the maximum and minimum normalization method is used to map the value of each feature to the interval [0, 1]. The normalization process can be expressed as the following formula 1:
[0084]
[0085] Among them, X is the original data, X min and X max are the minimum and maximum values of the features, respectively. This method can eliminate the differences in dimensions and ranges of different features, making them have equal weight in analysis and modeling, and is particularly suitable for algorithms that are sensitive to distance metrics (such as the K nearest neighbor algorithm and support vector machine algorithm, etc.).
[0086] After one-hot encoding and normalization, the data structure becomes clearer and more consistent, laying a solid foundation for subsequent model construction and feature engineering.
[0087] 3) The data is balanced using a data balancing method that combines synthetic minority oversampling technology (ie, SMOTE) and Tomek link.
[0088] To solve the problem of data imbalance, the SMOTE-Tomek method can be used to balance the dataset. SMOTE-Tomek is a data balancing method that combines SMOTE (Synthetic Minority Over-sampling Technique) and Tomek links, aiming to solve the class imbalance problem by generating new samples and cleaning noise samples simultaneously.
[0089] SMOTE expands the minority class representation by synthesizing new data points between minority class samples instead of simply replicating existing samples. It calculates several neighbors of each minority class sample and randomly generates synthetic samples on the line connecting the selected sample and its neighbors. Tomek links identify sample pairs near the class boundary and remove these noise samples that may cause model confusion. SMOTE-Tomek cleans redundant and noise points in the dataset while generating minority class samples, which is very suitable for scenarios that require enhancing the minority class and improving data quality.
[0090] Therefore, to handle the binary classification problem of imbalanced data, the balancing process includes:
[0091] 3-1) Divide the data in the dataset into a training set and a test set at a preset ratio based on the stratified sampling method;
[0092] The dataset is divided into an 80% training set and a 20% test set through the stratified sampling method to ensure that the class distribution in the training set and the test set is consistent with the original data.
[0093] 3-2) Perform balancing processing on the training set and the test set;
[0094] The SMOTE-Tomek method is used to balance the training set and the test set, making the ratio of positive and negative class samples close to 1:1 to ensure that the class distribution in the training set and the test set is consistent with the original dataset, thereby eliminating the bias caused by uneven data division.
[0095] 3-3) Classify the samples based on a preset PCE threshold to obtain high-PCE samples and low-PCE samples.
[0096] In the method of this application, the samples are binary-classified according to the threshold of the PCE value. For example, PCE > 18.5 is classified into the high-efficiency category (High Effciency), that is, high-PCE samples, while PCE ≤ 18.5 is classified into the low-efficiency category (LowEffciency), that is, low-PCE samples.
[0097] By setting different PCE thresholds, it is possible to adjust which samples are classified into the high-efficiency category, thereby affecting the classification boundary of the model. In particular, if the threshold is too high, a large number of samples with high PCE values may be misclassified as the low-efficiency category, resulting in a decrease in recall rate; if the threshold is too low, samples with low PCE values may be misclassified as the high-efficiency category, leading to an increase in false positives and a decrease in precision.
[0098] Therefore, the selection of an appropriate threshold can find the best balance between precision and recall and optimize the model's ability to recognize different categories.
[0099] To avoid the problem of `n_neighbors > n_samples` (i.e., the number of neighbors is greater than the number of samples) during the data balancing process, a PCE threshold range starting from 17.4 and ending at 20.1 with a step size of 0.1 was selected. This range ensures that each category has a sufficient number of samples after threshold division, thus avoiding the problem of being unable to find enough neighbors during SOMTE oversampling and ensuring the effectiveness of the balancing process.
[0100] Refer to Figure 3 and Figure 4 , the selected threshold can not only cover samples with significant PCE differences but also ensure the feasibility of the data balancing strategy during model training. As can be observed from Figure 3 , in the case of extremely imbalanced classes, the ML model often exhibits overfitting, that is, the model over-learns the features of a certain class, resulting in a falsely high performance evaluation result (AUC value), which is not desirable.
[0101] Therefore, this method selects 18.5 as the binary classification threshold for PCE. At this threshold, the difference in the number of samples between class 0 and class 1 (i.e., high PCE and low PCE) is small, and the AUC value remains within an acceptable range. Based on this, when 18.5 is used as the binary classification threshold, the model can balance its prediction ability when dealing with different classes, thereby improving the generalization ability and practical application value of the model.
[0102] The above steps can solve the data quality problems caused by the uncertainty of the experiment, help improve the data quality, ensure the reliability and representativeness of the analyzed data, exclude the influence of human factors and environmental interference on the experimental data, and improve the data quality.
[0103] Step 2: Screen key features, which include:
[0104] 1. Apply the random forest algorithm to rank the feature importance of the data set to obtain the feature importance order ranking (i.e., Feature Ranking One).
[0105] The random forest algorithm refers to an ML algorithm that repeatedly samples from the original dataset through the bootstrap sampling method to generate multiple sub-datasets. Each sub-dataset is used to train a decision tree, and finally a forest composed of multiple decision trees is obtained.
[0106] This algorithm can quantify the impact of each feature on the model's decision-making by calculating the importance of each feature during the decision tree splitting process and using the Gini impurity (i.e., the Gini impurity, which is used to measure the impurity of a node, that is, the degree of class mixing in the sample set). This method helps to identify the most predictive features, thus providing a scientific basis for subsequent model optimization and process improvement.
[0107] 2. Based on the ranking of feature importance order and the Pearson product-moment correlation coefficient (PCC), a set of feature data with relatively weak correlations is selected.
[0108] To further optimize the model performance, the correlation between features is evaluated based on PCC. The calculation formula is shown in Formula 2:
[0109]
[0110] X i and Y i represent the i-th observations of two variables respectively, and represent the means of the two variables respectively, and n is the number of observations.
[0111] PCC is a statistical index that measures the strength of the linear correlation between two variables, and its value range is [-1, 1]. Among them, a positive value indicates a positive correlation, a negative value indicates a negative correlation, and the closer the absolute value is to 1, the stronger the correlation. Considering that highly correlated features may lead to the problem of multicollinearity, which in turn affects the interpretability and generalization ability of the model, and at the same time redundant features may increase the model complexity and cause the risk of overfitting, by removing highly correlated features to avoid the interference of redundant information. Through this process, those features that are highly correlated with other features and have low importance are preferentially deleted.
[0112] Using PCC analysis, features with a correlation with PCE greater than 0.65 (the set threshold) are removed, and features with lower rankings in Feature Ranking 1 are preferentially deleted to simplify the model and improve the prediction effect, and to avoid the negative impact of redundant information on the model performance.
[0113] Refer to Figure 5, this figure presents the PCC matrix between various features and between each feature and PCE, indicating that after PCC screening, the correlation between all features is lower than 0.65, effectively avoiding the problem of multicollinearity and ensuring that features with high independence and predictive ability are retained.
[0114] As shown in Table 1 (below), the features in the output have been abbreviated. Features with the prefix "num" represent continuous features that have been processed by min-max normalization; while features with the prefix "cat" are discrete features that have been processed using one-hot encoding. In addition, the correlation between each feature and PCE has been calculated and visualized, and the results show that the correlation between these features and PCE is weak, indicating that they do not exhibit significant linear or non-linear relationships statistically, thus avoiding the interference of redundant information and improving the model's ability to capture the independent effects of each feature.
[0115] Table 1: Key features selected based on the PCC heatmap and their abbreviations
[0116]
[0117] 3. Use the recursive feature elimination method (recursive feature elimination, RFE, hereinafter referred to as RFE) to further reduce the data dimension of Feature Data Set 1 (i.e., further screen the number of features in Feature Data Set 1 that contribute best to the model performance), obtaining Feature Data Set 2;
[0118] Among them, RFE is an iterative feature selection method that gradually deletes features with less impact on the prediction performance by recursively training the model and evaluating the contribution of features to the model performance. In each iteration, RFE determines which features are most important for predicting the target variable by evaluating the performance of the model (such as mean squared error, MSE). Through this process, the preparation conditions that have the most influence on PCE can be accurately identified, so as to more effectively adjust the preparation process parameters, optimize the experimental design and improve the performance of PSCs.
[0119] Refer to Figure 6 , through the RFE method, the best number of features is screened based on MSE. Through this process, the number of features that contribute best to the model prediction performance is determined, thus optimizing the model and improving the accuracy of PCE prediction.
[0120] Refer to Figure 7, the feature importance was calculated by the random forest algorithm, showing nine key preparation conditions that have a significant impact on PCE sorted by feature contribution degree. Combining the results of RFE, two features with smaller contributions were further removed. Through this series of steps, redundant features were successfully reduced, obtaining the second set of feature data, improving the accuracy of the model, and providing strong support for optimizing the preparation process of PSCs.
[0121] Step 3: Model construction and optimization, which includes:
[0122] 1. Model and train the second set of feature data selected. Apply the Optuna method to tune the hyperparameters of all models. According to the weighted scores of the preset evaluation metrics, select the optimal model among logistic regression (LR), random forest (RF), support vector machine (SVM), k-nearest neighbors (KNN), gradient boosting (GB), and extreme gradient boosting (XGBoost) to perform the prediction task;
[0123] In this step, the Scikit-learn toolkit was used to model and train the selected features, and the above six ML models were used to perform the prediction task to evaluate the performance and accuracy of different models. Among them, the Scikit-learn toolkit refers to an open-source ML toolkit based on Python (a high-level programming language), which contains numerous ML algorithms and provides various data preprocessing methods, such as feature scaling, feature selection, data cleaning, etc., to help users prepare the dataset for training.
[0124] The preset evaluation metrics mainly include five metrics: accuracy, recall, precision, the harmonic mean of precision and recall (F1 Score, hereinafter referred to as F1 value), and area under the curve (AUC). The above metrics and evaluation criteria will be explained in detail later; the performance of the model is comprehensively evaluated by multiple metrics such as accuracy, recall, precision, F1 value, and AUC to comprehensively measure its classification effect and stability.
[0125] Step 4: Design experiments and optimize the preparation process, which includes:
[0126] 1. Analyze the possible reasons for high PCE based on the optimal model for the second set of feature data selected, and design experiments to obtain experimental conclusions;
[0127] Train the algorithm training model based on the optimal model and use feature importance evaluation to measure the impact of individual features on the recognition and prediction accuracy of high PCE samples. Through feature importance evaluation, the role of each feature in the model can be quantified, and which features are crucial for distinguishing the types of PCE samples can be identified. This not only helps to optimize the model performance but also provides a basis for deeply understanding the key factors affecting PCE. By analyzing the importance of features, it is possible to focus on those features that have a significant impact on PCE, thus making more accurate decisions in the material synthesis process.
[0128] 2. Obtain the best experimental conditions based on the experimental conclusions and optimize the preparation process.
[0129] The Shapley explanation method can be used to evaluate the impact of each feature on improving PCE. Then, experiments under different conditions can be designed according to the evaluation results, and experimental data and conclusions can be obtained. Based on the experimental conclusions, the preparation process of PSCs can be optimized.
[0130] Through the above method, the limitations of the traditional experimental trial-and-error method are effectively solved, and the problem of low prediction accuracy in the application of ML in the field of PSCs is overcome. It provides a new perspective and scientific basis for the rapid development of high-efficiency PSCs and also provides a reference for the development of other new solar cell technologies.
[0131] Refer to Figure 8 and Figure 9 , the screening of the optimal model according to the preset evaluation indicators described above includes:
[0132] 1) Evaluate the model based on multiple indicators such as accuracy, recall, precision, the harmonic mean of precision and recall, and the area under the curve.
[0133] Among them, accuracy is the proportion of samples correctly predicted by the model in the total samples, which is calculated by formula three (shown below):
[0134]
[0135] Among them, TP (true positives, the number of samples correctly predicted as positive by the model), TN (true negatives, the number of samples that the model wrongly predicts the negative class as positive), FP (false positives, the number of samples that the model wrongly predicts the negative class as positive), and FN (false negatives, the number of samples that the model wrongly predicts the positive class as negative); accuracy is an intuitive and easy-to-calculate evaluation indicator, which is suitable for datasets with balanced classes because it can directly reflect the overall proportion of correct classifications of the model.
[0136] Recall, also known as true positive rate, is the proportion of all positive class samples identified by the model, calculated by formula four (shown below):
[0137]
[0138] Recall can effectively evaluate the model's ability to identify positive classes, which is particularly important when the cost of missing positive class samples is high (such as in disease diagnosis). However, when the recall is too high, it often leads to an increase in false positives, reducing the precision of the model, and may even mispredict a large number of negative class samples as positive classes, thus affecting the stability and practical application effect of the model.
[0139] Precision is the proportion of samples that are truly positive among those predicted as positive by the model, calculated by formula five (shown below):
[0140]
[0141] Precision emphasizes the accuracy of the model when predicting positive classes and is applicable to application scenarios where the cost of false positives is high, such as spam filtering and fraud detection. However, when used alone, precision does not reflect the model's comprehensive ability to identify positive classes, and over-optimizing precision may lead to missing positive class samples, thus affecting the recall.
[0142] The harmonic mean of precision and recall is the F1 value (F1 Score), which takes into account the trade-off between the two and is calculated by formula six (shown below):
[0143]
[0144] While considering both, the F1 value avoids the problem of over-optimizing a single metric. In binary classification problems, it helps to find a model that does not overly favor precision or recall.
[0145] The area under the curve (AUC) refers to the area under the receiver operating characteristic (ROC) curve and represents the model's ability to distinguish between positive and negative class samples. The closer the AUC is to 1, the better the model; the closer the AUC is to 0.5, the closer the model's prediction ability is to random guessing. The AUC calculation implemented by Sklearn is based on numerical integration of the ROC curve using the trapezoidal rule. Calculating the area under the ROC curve can be obtained by integration, usually using numerical methods, and is calculated by formula seven (shown below):
[0146]
[0147] Among them, N pos and N neg represent the numbers of positive and negative class samples respectively, y i and y j refer to the predicted scores of the i-th positive class sample and the j-th negative class sample respectively. Φ(y i >y j ) is an indicator function. If y i >y j , that is, the score of the positive class sample is greater than that of the negative class sample, the function value is 1; otherwise, it is 0. AUC provides a more comprehensive evaluation and can reflect the performance of the model under different classification thresholds. It measures the ability of the model to distinguish between positive and negative classes and avoids relying on a single threshold.
[0148] In addition, to optimize the model performance, the Optuna tool can also be used for hyperparameter tuning. Optuna is an automated hyperparameter optimization framework based on the Bayesian optimization algorithm. It defines the hyperparameter space and gradually explores different hyperparameter combinations, adjusts the search strategy based on the results of each experiment, and thus efficiently improves the performance of the model. Optuna builds a probability model to predict which hyperparameters may bring better results and guides the subsequent search process based on these predictions, making the search strategy more efficient and avoiding ineffective exploration.
[0149] Furthermore, to reduce the possible performance fluctuations of the model caused by different train set partitions, 5-fold cross-validation is performed on each model (the dataset is divided into 5 subsets, and each time 4 subsets are used for training and the remaining 1 subset is used for validation. This process is repeated 5 times, and each subset serves as a validation set once. Finally, the evaluation result of the model is the average of all 5 validation results) to further ensure the generalization ability of the model.
[0150] 2) Compare the performances of multiple models based on various evaluation metrics;
[0151] Referring to Figure 8 , after further optimizing the model performance by the above method, comprehensive evaluation is carried out through multiple metrics such as accuracy, recall rate, precision, F1 value, and AUC to comprehensively measure its classification effect and stability. The evaluation metrics are shown in Table 2:
[0152] Table 2: Performance comparison of 6 ML models on 5 evaluation criteria respectively
[0153]
[0154]
[0155] 3) Score the model based on various evaluation metrics and determine the optimal model;
[0156] Refer to Figure 9 , the weight of the evaluation metrics is Area Under the Curve (AUC) > F1-score (harmonic mean of precision and recall) > Precision > Recall > Accuracy.
[0157] The evaluation criteria for selecting a model are in the order of AUC > F1 > Precision > Recall > Accuracy. This is mainly because AUC can comprehensively measure the ability of the model to distinguish positive and negative class samples at different thresholds, which is crucial for identifying high-PCE samples. The F1-score follows closely because it combines precision and recall, balancing these two metrics to ensure that high-PCE samples are not missed while reducing the occurrence of false positives. Precision and recall are also important in evaluating the performance of the model, especially precision, which ensures that the samples predicted by the model as high-PCE actually have high PCE. Accuracy can be misleading in tasks with class imbalance, so it is used as the last evaluation metric. In this study, to select the most suitable model, multiple evaluation metrics were comprehensively considered and different weights were assigned to each metric.
[0158] Specifically, AUC is given the highest weight (0.3) because it can comprehensively evaluate the ability of the model to distinguish positive and negative class samples at different thresholds, which is particularly important for identifying high-PCE samples. The F1-score is the second most important evaluation criterion (weight 0.25), which takes into account both precision and recall, ensuring the accurate identification of high-PCE samples while reducing false positives. The weight of precision is 0.2, highlighting its role in ensuring the reliability of samples predicted as high-PCE. The weight of recall is 0.15 to ensure that true high-PCE samples are not missed, and the weight of accuracy is 0.1, mainly considering the potential misleading nature of accuracy in the case of class imbalance.
[0159] Refer to Figure 9 , through this weighted method, the performance of the model can be comprehensively evaluated, and the optimal model selected is GB (total score 0.717, higher than other models).
[0160] This method also includes:
[0161] 1) Based on the selected optimal model, analyze the performance metrics of the original training set of unbalanced processed data;
[0162] To further explore the impact of data balancing on model performance, this study investigated whether the SMOTE-Tomek sampling method affects model performance based on the original training set. The optimal model, the GB algorithm, selected above was used for the research, and the performance of the model under different data balancing conditions was analyzed.
[0163] 2) Analyze the impact of the balance of the training set and the test set on the model's prediction ability based on the analysis results of the performance metrics.
[0164] Examples are as follows:
[0165] Table 3 (shown below) presents the class distributions of the training set and the test set under balanced scenarios for different datasets.
[0166] Table 3: Number of samples in the training set and the test set under different balancing conditions
[0167]
[0168]
[0169] According to the results of cross-validation in Table 4 (shown below), the performance of Condition 1 (unbalanced training set, unbalanced test set) and Condition 2 (unbalanced training set, balanced test set) is better than that of Condition 3 (balanced training set, unbalanced test set) and Condition 4 (balanced training set, balanced test set).
[0170] Table 4: Cross-validation performance metrics under different data balancing conditions based on GB
[0171]
[0172]
[0173] The reason is that the unbalanced training set can retain the original class distribution of the dataset, enabling the model to better adapt to the unbalanced data situation in the real world, thereby improving the model's prediction ability. For Conditions 3 and 4, by balancing the training set, although it can improve the model's learning of the minority class, it may lead to overfitting of the minority class samples, affecting the overall prediction performance, especially a decrease in precision and F1 score.
[0174] Therefore, the unbalanced training set can more effectively simulate the actual scenario during the training process, resulting in better performance than the balanced training set setting. On the test set (Table 5, shown below), the performance of Condition 1 (unbalanced training set, unbalanced test set) is better than that of Condition 2 (unbalanced training set, balanced test set).
[0175] Table 5: Performance metrics of the test set under different data balancing conditions based on GB
[0176]
[0177]
[0178] As shown in the above table, under Condition 1, the model maintained a high recall rate, and the accuracy, precision, and AUC were also at a good level. This indicates that the model can better identify minority class samples while maintaining good classification performance for the majority class. In Condition 2, although the recall rate was high, the precision and F1-score decreased, indicating that the model was overly biased towards the minority class, resulting in an increase in false positives and affecting the performance in practical applications. The advantage of Condition 1 is that both its training set and test set maintained the true distribution of the imbalanced data, making the model's performance in practical applications more robust.
[0179] In the study of this method, the combination of imbalanced training set and imbalanced test set (Condition 1) performed best on the cross-validation and test sets because it could better simulate the actual situation, avoiding overfitting and false positive problems caused by balanced data. This study shows that training a model on an imbalanced dataset achieves better results. This is because an imbalanced dataset can better reflect the data distribution in practical applications and can help the model better adapt to the class imbalance problem in the real world. Using an imbalanced training set and test set, the model can avoid overfitting minority class samples while maintaining good prediction ability for the majority class.
[0180] In addition, algorithms that can handle imbalanced data (such as GB, RF, etc.) effectively address class imbalance through built-in mechanisms. By adjusting class weights or optimizing the loss function, the model can achieve better prediction performance on imbalanced data, avoiding the overfitting problem that may be caused by a balanced dataset, thereby improving the accuracy and stability of the model. Since the difference between Class 1 and Class 0 is not significant (the quantity ratio is only slightly greater than 2:1), in this case, data balancing may instead introduce unnecessary complexity, and overly adjusting the class distribution may instead affect the generalization ability of the model.
[0181] Refer to Figure 10 , this method also includes optimizing the optimal model, which includes:
[0182] 1) Based on the selected optimal model, use the Optuna hyperparameter optimization method to optimize the model performance, and conduct feature importance evaluation to find the features that have the highest impact on high PCE samples;
[0183] Due to the excellent comprehensive performance of the GB algorithm, a model is trained based on the GB algorithm and feature importance evaluation is used to measure the impact of individual features on the recognition and prediction accuracy of high-PCE samples. Through feature importance evaluation, the role played by each feature in the model can be quantified, and it can be identified which features are crucial for distinguishing high-PCE and low-PCE samples. This not only helps to optimize the model performance but also provides a basis for deeply understanding the key factors affecting PCE. By analyzing the importance of features, it is possible to focus on those features that have a significant impact on PCE, thus making more accurate decisions during the material synthesis process.
[0184] 2) Identify the features whose importance in distinguishing high-PCE samples and low-PCE samples reaches the preset standard.
[0185] The importance analysis of features by the GB model is as Figure 10 shown. It can be observed that nP_DV is the most important feature, with an importance contribution of 0.41, indicating that the dilution volume of perovskite in the research of this method has a decisive impact on predicting high-PCE samples.
[0186] The study found that the dilution volume of perovskite directly affects its concentration. When the dilution volume is large, the concentration of the perovskite solution decreases, which may affect its crystallization quality, film uniformity, and PCE. If the concentration is too high, it may lead to an increase in defects in the perovskite film, affecting electron conductivity and light absorption ability, thus reducing PCE. A reasonable dilution volume helps to optimize the stability and film-forming quality of the perovskite solution, further improving PCE.
[0187] Secondly, the perovskite rotation speed (nP_CS, contribution 0.24) in the spin-coating method also has a crucial impact on PCE. A higher rotation speed usually can form a thinner and more uniform perovskite film, which helps to improve light absorption ability and carrier migration efficiency, thereby enhancing PCE. However, too high a rotation speed may result in an overly thin film layer, affecting light absorption or triggering an unstable crystal structure, reducing charge transport efficiency. Therefore, an appropriate rotation speed can not only optimize the film uniformity and thickness but also promote good crystallization of the perovskite material, thus improving PCE. Finding the optimal rotation speed is also one of the keys to improving the efficiency of PSCs in this study.
[0188] Subsequently, nS_CAT(0.16) represents the annealing time of the self-assembled monolayers (SAM) cleaner. An appropriate annealing time is crucial for removing organic residues on the substrate surface. By controlling the annealing time, the cleanliness and stability of the thin film can be significantly improved. The annealing process not only helps remove surface contaminants but also promotes the rearrangement of surface atoms or molecules, thereby enhancing the orderliness of the SAM layer and ensuring the required cleanliness and surface quality of the substrate surface during the self-assembly process. Good surface treatment can reduce surface defects and non-uniformities, thus optimizing the self-assembly process and improving the efficiency and stability of the final device.
[0189] Next, nP_IPC, namely the additive concentration, with a contribution of 0.13, is also a key factor. The influence of this feature also stems from the crucial role of concentration in the synthesis of perovskite materials. Additives in perovskite materials, especially passivators and interface modifiers, can effectively improve the surface properties and crystallization quality of the materials, thereby increasing the PCE. An appropriate passivation concentration helps reduce surface defects in perovskite materials, which often serve as recombination centers for charge carriers, leading to a decrease in PCE. Additives passivate surface defects, reduce the recombination rate, and thereby increase the charge carrier mobility and material stability.
[0190] However, too low a passivator concentration may result in ineffective passivation of surface defects, thus affecting the PCE; while too high a concentration may lead to non-uniform distribution of additives in the material, forming new defects and instead reducing the PCE of the material. Reasonable optimization of the additive concentration can not only increase the photoconductivity and PCE but also enhance the material stability, which is the key to improving the PCE of perovskite materials.
[0191] Finally, nP_AV(0.03) represents the volume of the antisolvent used in this task. The antisolvent helps control the crystallization rate of the perovskite thin film and improve the film quality. Although it is generally believed that the dosage of the perovskite antisolvent added during the spin-coating process has a relatively small direct impact on the device efficiency, its indirect effects on crystal nucleation, growth uniformity, and interface quality cannot be ignored. Too little antisolvent dosage may lead to uneven nucleation of the thin film, generating defects, while too much dosage may carry away too much precursor solution, affecting the film thickness and composition distribution.
[0192] The antisolvent dosage may also have potential effects on the experimental repeatability and device performance stability. Therefore, in the preparation of high-efficiency devices, it is still of great significance to systematically evaluate the effects of the antisolvent dosage on the film morphology, interface characteristics, and defect state density.
[0193] In addition, cP_I Is PM_MI(0.02) and cP_I IsPM_MBr(0.01) indicates that MDAI2 and MDABr2 (a certain specific organic molecule) are added to the perovskite solution as additives. The model predicts that these additives have a certain impact on improving the PCE of perovskite materials. These molecular additives may enhance the light absorption ability, carrier mobility, and device stability of the material by regulating the crystallinity, energy band structure, or interface characteristics of the perovskite material, thereby improving the PCE of the final device. These characteristics play an important role in the preparation of the batteries in this study. By optimizing the parameters of these characteristics, the PCE and stability of the material should be further improved.
[0194] The method for identifying the important influence on distinguishing high-PCE samples and low-PCE samples in step four of this method includes:
[0195] 1) Based on the optimal model and using the SHapley Additive exPlanations (SHAP) method (a method for explaining the prediction results of ML models), analyze the contribution values of the features in the second set of feature data to the model prediction, and generate the analysis results of the model prediction contribution values;
[0196] Example: Rank the features according to their importance in contributing to the model prediction in the SHAP plot and analyze the SHAP values. The features ranked higher have the greatest impact on the prediction results. The SHAP value of each feature represents the specific contribution of the feature to a single prediction result. A positive value indicates that the feature promotes an increase in the prediction result, while a negative value indicates that it promotes a decrease in the prediction result.
[0197] The designed experiments include:
[0198] 1) Based on the analysis results of the model prediction contribution values, use the method of controlling variables to regulate each feature and design the experimental conditions for multiple experiments;
[0199] By analyzing the color and distribution in the SHAP plot, all features can be analyzed, and experiments can be designed based on the analysis of the SHAP plot to determine the experimental conditions that may improve the PCE.
[0200] According to the research of this method, it is finally concluded that the feature nP_IPC is the primary factor affecting the PCE; the additive concentration has a significant impact on improving the PCE, and the low concentration is more obvious for improving the PCE. Therefore, the concentration of the additive can be continuously regulated to find low concentrations that may not have been explored in the current experiments to improve the PCE.
[0201] In step four of this method, obtaining the best experimental conditions based on the experimental conclusions and optimizing the preparation process includes:
[0202] 1) Obtain the experimental data of multiple experiments under different experimental conditions and analyze them using the optimal model (i.e., GB) to generate the experimental analysis conclusions and obtain the best experimental conditions;
[0203] Reference Figure 11 , example: Based on the research and analysis of this method, it was finally decided to temporarily not regulate nS_CAT and cP_I Is PM_MBr, which are set to 0 (i.e., no annealing after cleaning) and 0 (i.e., not using this additive) respectively. The specific experimental design can refer to Table 6 (shown below). By systematically optimizing the key features (such as nP_IPC, cP_I Is PM_MI, nP_CS, nP_DV, and nP_AV), and controlling other features, to find the best experimental conditions for improving PCE.
[0204] Table 6:
[0205] The experimental conditions of PSCs optimized based on SHAP analysis are used for experimental verification.
[0206]
[0207] 2) Based on the best experimental conditions, optimize the preparation process.
[0208] Reference Figure 12 and Figure 13 , through the implementation of these optimized conditions, finally through experimental research, it is obtained that the preparation process based on Condition 2 can achieve a greater improvement in the PCE of the battery than other conditions, and the preparation process is optimized based on this condition.
[0209] In summary, this method introduces ML technology and deeply explores its application potential in the R & D of PSCs materials. Based on the analysis of PSCs experimental data, an innovative data-driven analysis framework is proposed, which successfully identifies the key features affecting PCE and optimizes the preparation process in the material R & D process.
[0210] This application also discloses a preparation process optimization system for perovskite solar cells, including a memory and a processor. The memory stores a computer program that can be loaded and executed by the processor to perform the preparation process optimization method for perovskite solar cells described above.
[0211] The above are all preferred embodiments of this application. The protection scope of this application is not limited by this. Therefore, all equivalent changes made according to the structure, shape, and principle of this application should be covered within the protection scope of this application.
Claims
1. A method for optimizing the preparation process of a perovskite solar cell, characterized in that: include: Step 1: Build a database, which includes: Using the experimental data as the basic data, the data is cleaned and standardized, and the PCE threshold is set to classify the batteries into high-efficiency batteries and low-efficiency batteries to obtain the data set; Step 2: Screen key features, including: Apply the random forest algorithm to sort the feature importance of the data set and obtain the feature importance order ranking; Based on the feature importance ranking and Pearson correlation coefficient, the feature data set with weak correlation is selected; The recursive feature elimination method is used to further reduce the data dimension of feature data set 1 to obtain feature data set 2; Step 3: Model construction and optimization, which includes: Model and train the selected feature data set 2, apply the Optuna method to tune the hyperparameters of all models, and select the best model among logistic regression, random forest, support vector machine, K nearest neighbor, gradient boosting, and extreme gradient boosting according to the weighted scores of the preset evaluation indicators for prediction tasks; Step 4: Design experiments and optimize the preparation process, which includes: Based on the optimal model, the two pairs of selected characteristic data sets are analyzed for possible reasons leading to high photoelectric conversion efficiency, and experiments are designed to obtain experimental conclusions; Based on the experimental conclusions, the optimal experimental conditions are obtained and the preparation process is optimized.
2. The method for optimizing the preparation process of a perovskite solar cell according to claim 1, characterized in that: The data cleaning includes: Conduct a preliminary screening of basic data using preset standards and retain the data that passes the screening; The retained data is checked for outliers and outliers are deleted.
3. The method for optimizing the preparation process of a perovskite solar cell according to claim 1, characterized in that: The standardization process comprises: The data are grouped into continuous data and discrete data based on the data type, and prefixes are added for identification; wherein the prefix indicates the data type; Identify the prefix of the data group and process the data of different groups respectively using the preset corresponding processing methods; among which, the continuous data is normalized to the maximum and minimum, and the discrete data is processed by one-hot encoding; The data is balanced by combining the synthetic minority class oversampling technique and Tomek link data balancing method.
4. The method for optimizing the preparation process of a perovskite solar cell according to claim 3, characterized in that: include: Based on the selected optimal model, the performance index analysis is performed on the original training set of unbalanced data; Based on the performance indicator analysis results, the impact of whether the original training set is balanced on the model's prediction ability is analyzed.
5. The method for optimizing the preparation process of a perovskite solar cell according to claim 3, characterized in that: The balancing process includes: The data in the dataset is divided into a training set and a test set in a preset ratio based on the stratified sampling method; Balance the training set and the test set; The samples are classified based on a preset photoelectric conversion efficiency threshold to obtain samples with high photoelectric conversion efficiency and samples with low photoelectric conversion efficiency.
6. The method for optimizing the preparation process of a perovskite solar cell according to claim 4, characterized in that: Optimizing the optimal model includes: Based on the selected optimal model, the Optuna hyperparameter optimization method is used to optimize the model performance and perform feature importance evaluation to find out the features that have the highest impact on samples with high photoelectric conversion efficiency; Identify features that distinguish high from low photoelectric conversion efficiency samples and have an impact on meeting preset criteria.
7. The method for optimizing the preparation process of a perovskite solar cell according to claim 1, characterized in that: The step of selecting the optimal model according to the preset evaluation index comprises: Evaluate the model based on one or more metrics including accuracy, recall, precision, harmonic mean of precision and recall, and area under the curve; Compare the performance of multiple models based on various evaluation indicators; The models are weighted based on various evaluation indicators and the optimal model is determined. The weights of the evaluation indicators are as follows: area under the curve > harmonic mean of precision and recall > precision > recall > accuracy.
8. The method for optimizing the preparation process of a perovskite solar cell according to claim 6, characterized in that: The method for identifying the importance of distinguishing samples with high photoelectric conversion efficiency from samples with low photoelectric conversion efficiency comprises: Based on the optimal model and using the Sharp-Galli interpretation method, the contribution values of the features in the feature data set 2 to the model prediction are analyzed to generate the model prediction contribution value analysis results; The design experiment includes: based on the model prediction contribution value analysis results, using the control variable method to adjust various characteristics, and designing experimental conditions for multiple experiments; The method of obtaining the best experimental conditions based on the experimental conclusions and optimizing the preparation process includes: Obtain experimental data from multiple experiments under different experimental conditions and use the optimal model analysis to generate experimental analysis conclusions and obtain the best experimental conditions; Based on the best experimental conditions, the preparation process was optimized.
9. A system for optimizing the preparation process of a perovskite solar cell, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executes the method for optimizing the preparation process of a perovskite solar cell as described in claims 1 to 8.