Mesoscopic perovskite solar cell conductive carbon electrode component design method and system based on machine learning model

By collecting data in the same laboratory and optimizing model hyperparameters using automatic encoder and grid search, a high-quality perovskite solar cell conductive carbon electrode database was constructed, and data integrity and accuracy problems in the industrialization of perovskite solar cells were solved, and efficient design of conductive carbon electrode components and device performance optimization were achieved.

CN120340683APending Publication Date: 2025-07-18SUNRISE (XIAMEN) PHOTOVOLTAIC IND CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510483007.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing perovskite solar cell industrialization process is limited by the lack of high-quality device production process-related information and formula databases, and machine learning models face data integrity and accuracy problems in perovskite research.

Method used

A mesoscopic perovskite solar cell conductive carbon electrode component design method is constructed based on machine learning models. By collecting data in the same laboratory, a high-quality perovskite solar cell conductive carbon electrode database is constructed, using an automatic encoder for feature dimensionality reduction, and the model hyperparameters are optimized through grid search, and the fit prediction model is trained to design conductive carbon electrode components.

Benefits of technology

Ensure the reliability and uniformity of data sources, reduce the risk of model overfitting, clearly compare the impact of preparation processes and electrode components on device efficiency, and promote the industrial application of perovskite solar cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340683A_ABST
    Figure CN120340683A_ABST
Patent Text Reader

Abstract

The invention discloses a mesoscopic perovskite solar cell conductive carbon electrode component design method and system based on a machine learning model, and the method comprises the steps: constructing a sample data set based on the basic data of a perovskite cell device, and obtaining a perovskite solar cell conductive carbon electrode database based on the sample data set; extracting conductive carbon electrode features based on data in a perovskite solar cell conductive carbon electrode database, and then encoding and performing dimension reduction to obtain dimension reduction features; constructing an initial fitting prediction model based on a machine learning model, optimizing hyper-parameters of the fitting prediction model by using grid search, and training the optimized fitting prediction model by using dimension reduction features to obtain a trained fitting prediction model; and inputting data of the conductive carbon electrode of the perovskite solar cell to be predicted into the trained fitting prediction model to obtain a component formula of the conductive carbon electrode of the mesoscopic perovskite solar cell, and completing component design of the conductive carbon electrode of the mesoscopic perovskite solar cell based on the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of perovskite solar cells, machine learning, and prediction of new material properties, and particularly relates to a method and system for designing the conductive carbon electrode components of a mesoscopic perovskite solar cell based on a machine learning model. Background Art

[0002] In recent years, significant progress has been made in the field of regression models in machine learning, mainly focusing on two core themes: the extension of linear regression and the deepening of non-linear regression models. The linear regression model is the most basic regression model, which widely used for predictive analysis by fitting the linear relationship between input features and target variables. However, linear regression often has difficulty capturing complex non-linear relationships. To address this issue, researchers have developed various variants of linear regression, such as ridge regression and lasso regression, which handle multicollinearity by adding regularization terms and enhance the robustness of the model. At the same time, non-linear regression models such as support vector regression (SVR) and decision tree regression have also emerged rapidly. These models can handle non-linear associations in data and provide more flexible fitting capabilities. Support vector regression maps input features to a high-dimensional space by introducing a kernel function to handle complex relationships, while decision tree regression generates tree-shaped decision nodes for prediction by recursively splitting data. Also, ensemble learning methods such as random forest and gradient boosting machine have performed well in regression tasks. Random forest achieves enhanced prediction accuracy and robustness by constructing multiple decision trees and integrating them, while gradient boosting machine improves the accuracy of the regression model by gradually optimizing the loss function.

[0003] Despite many advances, regression models still face challenges in practice, such as overfitting, selecting appropriate features, and their complexity. However, through continuous theoretical research and practical innovation, the regression models in machine learning are becoming increasingly powerful and can solve practical problems more accurately and effectively.

[0004] With the continuous maturity of machine learning technology, especially its excellent performance in processing complex data sets and extracting multi-dimensional data information, the research and development of perovskite materials are undergoing an unprecedented transformation. Traditionally, the research and development of perovskite materials relied on repeated trial and error through experiments and theoretical calculations, which usually required a large amount of manpower and time investment. However, machine learning has injected new vitality into this field. In recent years, researchers have started to use machine learning models for modeling structure-property relationships, predicting material properties, and optimizing synthesis processes. In perovskite research, machine learning algorithms are used to identify key factors affecting material performance and predict parameters such as the photoelectric conversion efficiency, stability, and manufacturability of materials. This enables scientists to quickly screen out new perovskite compounds with potential high performance, thus effectively shortening the research and development cycle.

[0005] Despite significant progress, the application of machine learning in perovskite research still faces challenges. For example, issues such as the accuracy and generalization ability of models, the integrity and quality of data, and the interpretability of physical explanations still need to be further explored. At the same time, the data in the current perovskite database mainly comes from the publication of journal papers, lacking unified and standardized data, which seriously affects the quality of the database and the prediction effect of the model. In addition, the industrialization process of perovskite solar cells is still restricted. The main problem is that the industry lacks relevant information in the device production process and high-quality database resources for formulations, which has become a bottleneck for the wide application of perovskite solar cells. Summary of the Invention

[0006] To solve the above problems, the present invention proposes a method and system for designing the conductive carbon electrode components of a mesoscopic perovskite solar cell based on a machine learning model. The method specifically includes:

[0007] Step S1: Construct a sample data set based on the basic data of perovskite battery devices, preprocess the sample data set and save it in a database to obtain a conductive carbon electrode database for perovskite solar cells;

[0008] Step S2: Extract the conductive carbon electrode features from the data in the conductive carbon electrode database for perovskite solar cells, encode and reduce the dimension of the conductive carbon electrode features to obtain reduced-dimensional features;

[0009] Step S3: Construct an initial fitting prediction model based on a machine learning model, optimize the hyperparameters of the fitting prediction model using grid search, and train the optimized fitting prediction model using the reduced-dimensional features to obtain a trained fitting prediction model;

[0010] Step S4: Input the data of the conductive carbon electrode of the perovskite solar cell to be predicted into the trained fitting prediction model to obtain the formulation of the conductive carbon electrode components of the mesoscopic perovskite solar cell, and complete the design of the conductive carbon electrode components of the mesoscopic perovskite solar cell based on the machine learning model.

[0011] Optionally, the basic data includes the preparation process of perovskite battery devices, the formulation of conductive carbon electrodes, and the performance indicators of the corresponding prepared perovskite solar cells.

[0012] Optionally, the performance indicators of the perovskite solar cell include energy conversion efficiency, fill factor, short-circuit current, and open-circuit voltage.

[0013] Optionally, the preprocessing process specifically includes:

[0014] Perform data cleaning based on the performance indicators of the perovskite solar cells in the sample data set and then perform outlier detection.

[0015] Optionally, the specific method of step S2 includes:

[0016] Perform feature encoding on the preprocessed data to obtain feature-encoded data;

[0017] Use Keras to construct an autoencoder model, and input the feature-encoded data into the autoencoder to obtain the dimension-reduced feature scalar.

[0018] The present invention also discloses a system for designing the conductive carbon electrode components of a mesoscopic perovskite solar cell based on a machine learning model. The system includes:

[0019] A dataset construction module for constructing a sample dataset based on the basic data of perovskite battery devices, preprocessing the sample dataset and saving it in a database to obtain a conductive carbon electrode database for perovskite solar cells;

[0020] A feature dimension reduction module for extracting the conductive carbon electrode features based on the data in the conductive carbon electrode database for perovskite solar cells, encoding the conductive carbon electrode features and then reducing the dimension to obtain dimension-reduced features;

[0021] A model training module for constructing an initial fitting prediction model based on a machine learning model, optimizing the hyperparameters of the fitting prediction model using grid search, and training the optimized fitting prediction model using the dimension-reduced features to obtain a trained fitting prediction model;

[0022] A component design module for inputting the data of the conductive carbon electrode of the perovskite solar cell to be predicted into the trained fitting prediction model to obtain the formulation of the conductive carbon electrode components of the mesoscopic perovskite solar cell, and completing the design of the conductive carbon electrode components of the mesoscopic perovskite solar cell based on the machine learning model.

[0023] Optionally, the basic data includes the preparation process of perovskite battery devices, the formulation of conductive carbon electrodes, and the performance indicators of the corresponding perovskite solar cells prepared.

[0024] Optionally, the performance indicators of the perovskite solar cell include energy conversion efficiency, fill factor, short-circuit current, and open-circuit voltage.

[0025] Compared with the prior art, the beneficial effects of the present invention are:

[0026] Different from other perovskite databases with data sources from different laboratories, the data in this project was collected through experiments in the same laboratory. This ensures the reliability and unity of the data source, constructs a high-quality dataset of conductive carbon electrodes for perovskite solar cells, and finds the optimal conductive carbon electrode formulation and preparation process according to different environments, which is of great significance for accelerating the industrial application of perovskite solar cells.

[0027] Through the autoencoder, the features related to the conductive carbon electrode are reduced to a single feature, which enriches the means of storing the dataset and effectively reduces the risk of model overfitting. At the same time, it is possible to more clearly compare the influence degree of the preparation process and the composition of the conductive carbon electrode on the device efficiency during the device preparation process. Brief Description of the Drawings

[0028] To more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0029] Figure 1 It is a method flowchart of the method for designing the conductive carbon electrode components of the mesoscopic perovskite solar cell based on the machine learning model in the embodiment of the present invention;

[0030] Figure 2 It is a dataset (excerpt) of the PSC samples in the embodiment of the present invention;

[0031] Figure 3 It is a schematic diagram of the importance ranking of features in the embodiment of the present invention;

[0032] Figure 4 It is a schematic diagram for verifying the accuracy of the model in the embodiment of the present invention;

[0033] Figure 5 It is a schematic diagram of the SHAP value dependence graph of the carbon content in the embodiment of the present invention. Detailed Embodiments

[0034] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0035] Autoencoder is a neural network structure used to learn effective representations and features of high-dimensional data. By compressing the data into a low-dimensional space and reconstructing the original input as losslessly as possible, data dimensionality reduction and feature extraction are achieved. Autoencoders usually consist of three parts: an encoder, a bottleneck layer, and a decoder.

[0036] Specifically, the encoder part of the autoencoder maps the input data to a low-dimensional latent space. This latent space is usually achieved through a linear or non-linear transformation, compressing the data and retaining its key information. The bottleneck layer is the narrowest layer in the autoencoder structure, responsible for forcing the learning of the most effective feature representations in order to reconstruct the original information as much as possible during decoding.

[0037] The decoder then reconstructs the input data from the latent space. This is the opposite of the encoder, mapping the low-dimensional representation back to a high-dimensional one through a process similar to the reverse transformation. During the training process, the autoencoder is optimized by minimizing the error between the input data and the reconstructed data, with the goal of retaining effective information compression as much as possible.

[0038] In terms of feature extraction, the advantage of the autoencoder lies in its ability to automatically learn the hidden patterns and structural information in the data, being particularly suitable for processing non-linear and high-dimensional data. Through the bottleneck layer, the autoencoder can filter out noise and retain only the most important features for input to subsequent machine learning tasks.

[0039] In addition, autoencoders can also be used for anomaly detection and denoising applications. In anomaly detection, since autoencoders have a smaller reconstruction error for normal data and a larger error for abnormal data, abnormal points can be identified by examining the reconstruction error. In denoising tasks, autoencoders can extract and reconstruct clear data from noisy data.

[0040] Embodiment 1

[0041] A method for designing the composition of a mesoscopic perovskite solar cell's conductive carbon electrode based on a machine learning model, as Figure 1 shown, the method includes:

[0042] Step S1, constructing a sample data set based on the basic data of perovskite battery devices, preprocessing the sample data set and saving it to a database to obtain a conductive carbon electrode database for perovskite solar cells.

[0043] Based on the hole-free mesoscopic carbon-based perovskite solar cell structure proposed by the Han Hongwei research group, the formulation and preparation process of the conductive carbon electrode were optimized and adjusted. Through systematic experiments, a detailed data set containing 1478 PSC samples was established for use in the examples. This data set details the preparation process of perovskite battery devices, the formulation of the conductive carbon electrode, and the performance indicators of the corresponding prepared perovskite solar cells, including power conversion efficiency (PCE), fill factor (FF), short-circuit current (J), and open-circuit voltage (V). An excerpt of the data is as Figure 2 shown.

[0044] Clean the dataset item by item based on prior material knowledge and mathematical statistics. For example, the photoelectric conversion efficiency of PCE should be between 0 and 27. Due to testing reasons, for some data, the PCE will exceed this range. In this embodiment, such abnormal data will be screened and deleted. Other examples are that the fill factor should be < 90, the current density should be < 26 mA / cm 2 , the voltage < 1.2 V, etc. are all the bases for screening abnormal data based on prior material knowledge in this embodiment. By the above method, abnormal "dirty data" are removed to ensure the correctness of the data.

[0045] First, three types of samples not within the scope of this study need to be excluded from the data, specifically including:

[0046] (1) Samples with abnormal tests, machine abnormalities, and curve abnormalities. As shown in the answer to the previous question, there are thresholds for judging corresponding data. For example, samples with severely jittery curves.

[0047] (2) Data samples not within the measurement standard. The data included in the object of this study should be the best photoelectric conversion efficiency measured under standard test conditions (STC) (AM1.5, 1000 W / m 2 , 25°C). Data measured under non-STC conditions will be deleted.

[0048] Exclude data with PCE = 0 from the data. Then, through the visual display of some data, identify controversial abnormal points and remove them from the sample set. The abnormal point detection method is as follows:

[0049] Abnormal points are called outliers, defined as "one or more observations in a sample that are far from other observations, indicating that they may come from different populations". Usually, the appearance of abnormalities may be caused by human recording and measurement errors, or they may be real special samples. In this embodiment, methods such as material theory knowledge, mathematical statistics, and model prediction can be used to detect abnormal values:

[0050] (1) Import methods such as describe() and value_counts() in the Pandas dependency to view the descriptive statistical information of the data to quickly understand the overall distribution of the data;

[0051] (2) Observe the scatter plot between variable v1 and variable v2 or variable v i and the label target;

[0052] (3) According to the 3σ principle, when the data set is close to a normal distribution, if the observed value x exceeds 3 times the standard deviation σ from the mean μ, it can be regarded as an outlier; because under the standard normal distribution X~N(0,1), almost all observed values x are concentrated in the interval [μ - 3σ, μ + 3σ], and the theoretical probability (calculated by probability) of exceeding this range is less than 0.3%;

[0053] (4) Box plot, by measuring the interquartile range (IQR) of the sample data to find out the outliers. In this embodiment, the number equal to 25% after sorting the samples by value is taken as the first quartile Q1, the number equal to 75% after all arrangements is taken as the third quartile Q3, and the median (50%) is taken as Q2, then IQR = Q3 - Q1. Then, 1.5 times of IQR is specified as the standard, and the points exceeding Q3 + 1.5IQR or Q1 - 1.5IQR are outliers;

[0054] (5) Model detection method, by establishing a probability distribution model of a random forest to determine the probability that the data conforms to the model, and regarding the samples with low probability as outliers.

[0055] Outlier detection is a relatively mature technology in the industry, and the above methods are more applicable to the outlier detection method of this embodiment. In a specific implementation case, two or three methods are selected as appropriate, and the intersection of the detected outliers is taken to determine the outliers. For example, two models, RF and LGBM, are used, and the find_outliers function is used to detect the outliers. Among them, the RF model gives 49 outliers, and the LGBM model gives 53 outliers. The intersection of these outliers is taken, and finally 28 outliers are determined and processed such as deletion. After finding the possible outliers, it is necessary to trace the source of this piece of data, and if there is an error, it should be modified or deleted.

[0056] For supervised machine learning with labels, generally LGBM and tree models are used for prediction. For small and medium-sized databases with a data volume of 1k - 10k, the RF and LGBM models perform better and have stronger anti-noise ability. In terms of model interpretability, for high interpretability requirements such as SHAP analysis of the model, models such as logistic regression, decision tree, and random forest are better choices.

[0057] Step S2: Extract the conductive carbon electrode features based on the data in the perovskite solar cell conductive carbon electrode database, encode the conductive carbon electrode features, and then perform dimensionality reduction to obtain the dimensionality-reduced features.

[0058] According to different requirements, in the regression task of the model, the label encoding method is generally used because in this embodiment, we are more interested in seeing which aspect has a greater impact on the device PCE. In an interpretable model, the one-hot encoding method is generally used because in this embodiment, we are more interested in knowing which material has a good or bad impact on the PCE of the device. For example, in the regression task, this embodiment knows that the conductive agent in the conductive paste has an important impact on the PCE. In the interpretable model, this embodiment knows that carbon nanotubes in the conductive agent have a positive impact on the PCE, while graphene has a negative impact. That is, label focuses more on the overall, while one-hot focuses more on more detailed features within a single feature.

[0059] Encode the input features using one-hot or label encoder. In particular, reduce the dimensionality of all features related to the conductive carbon electrode into a single feature through an autoencoder, which can effectively reduce the risk of overfitting and accelerate the calculation speed.

[0060] Use an autoencoder to reduce the dimensionality of the features. Build an autoencoder model using Keras. Take the feature descriptor of the conductive carbon electrode formulation as the input layer. In the encoding layer, map several features of the input layer into one feature. In the decoding layer, restore one feature into several features of the input layer to build and compile the model. Use the existing data to train the encoder model. After training, the corresponding features of the conductive carbon electrode components can be reduced to one feature. This method can not only reduce the dimensionality but also retain the correlation between the input features and control the quality of dimensionality reduction through the reconstruction error of the autoencoder. The dimensionality-reduced feature is a scalar, and the output layer of the encoder is a layer with only one neuron.

[0061] Step S3: Build an initial fitting prediction model based on the machine learning model, optimize the hyperparameters of the fitting prediction model using grid search, and train the optimized fitting prediction model using the dimensionality-reduced features to obtain a trained fitting prediction model.

[0062] Divide the dataset into a training set and a test set in a ratio of 4:1. In the selection of the regression model, mainstream machine learning algorithms including Random Forest, XGBoost, Catboost, and LGBM are selected for prediction. Optimize the hyperparameters of the machine learning model through grid search and save the test results. GridSearchCV is grid search cross-validation for tuning parameters. It traverses all permutations and combinations of the input parameters and returns the evaluation metric scores under all parameter combinations through cross-validation.

[0063] Specifically, grid search is an exhaustive search method that searches for the optimal hyperparameters by traversing all possible combinations of hyperparameters. Grid search first sets a set of candidate values for each hyperparameter, then generates the Cartesian product of these candidate values to form a combination grid of hyperparameters. Next, grid search trains and evaluates the model for each hyperparameter combination to find the hyperparameter combination with the best performance.

[0064] By optimizing the R2 and RMSE of the machine learning model, the machine learning model with the best prediction performance is selected and used for subsequent calculations.

[0065] Different models are used to fit and predict the training data. At the same time, based on the characteristics of different models, targeted adjustments are made to features and parameters.

[0066] The number of hidden layers, the number of nodes and the activation function were adjusted, and then a variety of mainstream algorithms (random forest, XGBoost, AdaBoost, LGBM, CatBoost, ANN) were used to establish a forward prediction model for PCE. The model was repeatedly fine-tuned based on the results of cross-validation, and finally the effect of the test set was used as the basis for model selection.

[0067] In machine learning regression problems, feature importance is a quantitative indicator that analyzes the contribution or influence of each input feature (feature variable, features) on the model's prediction results. This is especially critical for black box models (such as random forests, XGBoost, neural networks, etc.) because their internal mechanisms are complex and it is difficult to directly explain the model's decision-making process. By analyzing feature importance, you can better understand how the model makes predictions. By visualizing feature importance, you can more clearly understand the importance of each influencing factor.

[0068] Step S4, inputting the data of the conductive carbon electrode of the perovskite solar cell to be predicted into the trained fitting prediction model, obtaining the component formula of the conductive carbon electrode of the mesoscopic perovskite solar cell, and completing the design of the conductive carbon electrode component of the mesoscopic perovskite solar cell based on the machine learning model.

[0069] The high-efficiency strategy given by the trained machine learning model was verified through experiments. The conductive carbon electrode slurry was obtained by weighing the drugs, ball milling, rotary evaporation and other steps, and the perovskite solar cell device was prepared by screen printing. Finally, the photoelectric conversion efficiency, fill factor, open circuit voltage and short circuit current of the device were measured by the solar simulator K2400. The data obtained from the experiment will be used to verify the accuracy of the model and continue to be included in the conductive carbon electrode database, starting a new round of optimization process.

[0070] Specifically, it includes: randomly generating an array within the corresponding range through AI, putting these data into a machine learning model for PCE prediction, finding the array with a high PCE prediction value, finding the corresponding conductive carbon electrode components for experimental verification, and finally achieving the purpose of optimizing the conductive carbon electrode components of the mesoscopic perovskite solar cell. 2. Through the method of SHAP value dependence graph, according to the positive and negative distribution of SHAP values, the characteristic distribution interval of high-efficiency PCE is given, as shown in the figure: the optimal carbon content of the conductive carbon electrode is about 0.75. The SHAP value dependence graph of carbon content is as Figure 5 shown.

[0071] Example 2

[0072] A design system for the conductive carbon electrode components of a mesoscopic perovskite solar cell based on a machine learning model, the system includes:

[0073] A data set construction module, used to construct a sample data set based on the basic data of perovskite battery devices, preprocess the sample data set and save it in a database to obtain a conductive carbon electrode database for perovskite solar cells.

[0074] Based on the hole-free mesoscopic carbon-based perovskite solar cell structure proposed by the Han Hongwei research group, the formula and preparation process of the conductive carbon electrode have been optimized and adjusted. Through systematic experiments, a detailed data set containing 1478 PSC samples has been established for use in the examples. This data set details the preparation process of perovskite battery devices, the formula of the conductive carbon electrode, and the performance indicators of the corresponding perovskite solar cells prepared, including power conversion efficiency (PCE), fill factor (FF), short-circuit current (J), and open-circuit voltage (V). An excerpt of the data is as Figure 2 shown.

[0075] Clean the data set item by item based on prior material knowledge and mathematical statistics. For example, the photoelectric conversion efficiency of PCE should be between 0 and 27. Due to testing reasons, some data may have PCE values outside this range, and this example will screen and delete such abnormal data. Other examples such as the fill factor should be <90, the current density should be <26 mA / cm 2 , and the voltage <1.2 V, etc. are all the bases for screening abnormal data based on prior material knowledge in this example. By the above method, abnormal "dirty data" is eliminated to ensure the correctness of the data.

[0076] First, three types of samples not within the scope of this study need to be excluded from the data, specifically including:

[0077] (1) Samples with abnormal testing, machine anomalies, and curve anomalies. As shown in the answer to the previous question, there are thresholds for judging corresponding data. For example, samples with severely jittery curves.

[0078] (2) Data samples outside the measurement standards. The data included in the research object of this study should be the best photoelectric conversion efficiency measured under standard test conditions (STC) (AM1.5, 1000 W / m 2 , 25°C), and the data under non-STC should be deleted.

[0079] Data with PCE = 0 was excluded from the data. Then, through the visualization display of some data, controversial outliers were identified and removed from the sample set. The outlier detection method is as follows:

[0080] Outliers are called extreme points and are defined as "one or more observations in a sample that are far from other observations, indicating that they may come from different populations". Usually, the appearance of outliers may be caused by human recording or measurement errors, or they may be real special samples. In this embodiment, methods such as material theory knowledge, mathematical statistics, and model prediction can be used to detect outliers:

[0081] (1) Import methods such as describe() and value_counts() in the Pandas dependency to view the descriptive statistical information of the data to quickly understand the overall distribution of the data;

[0082] (2) Observe the scatter plot between variable v1 and variable v2 or variable v i and the label target;

[0083] (3) According to the 3σ principle, when the data set is close to a normal distribution, if the observed value x exceeds 3 times the standard deviation σ from the mean μ, it can be regarded as an outlier; because under the standard normal distribution X~N(0,1), almost all observed values x are concentrated in the interval [μ - 3σ, μ + 3σ], and the theoretical possibility (calculated by probability) of exceeding this range is less than 0.3%;

[0084] (4) Box plot, by measuring the interquartile range (IQR) of the sample data to find outliers. In this embodiment, the number equal to 25% after sorting the sample by value is used as the first quartile Q1, and all the numbers equal to 75% after sorting are used as the third quartile Q3, and the median (50%) is used as Q2, then IQR = Q3 - Q1. Then, 1.5 times of IQR is specified as the standard, and the points exceeding Q3 + 1.5IQR or Q1 - 1.5IQR are outliers;

[0085] (5) Model detection method, by establishing a probability distribution model of a random forest to determine the probability that the data conforms to the model, and regarding the samples with low probability as outliers.

[0086] Outlier detection is a relatively mature technology in the industry. The above-mentioned methods are outlier detection methods that are more applicable to this embodiment. In a specific implementation case, two or three methods are selected as appropriate, and the intersection of the detected outliers is taken to determine the outliers. For example, two models, RF and LGBM, are used, and the find_outliers function is used to detect outliers. Among them, the RF model gives 49 outliers, and the LGBM model gives 53 outliers. The intersection of these outliers is taken, and finally 28 outliers are determined and processed such as deletion. After finding the possible outliers, it is necessary to trace the source of this piece of data, and if there is an error, it should be modified or deleted.

[0087] For supervised machine learning with labels, LGBM and tree models are generally used for prediction. For small and medium-sized databases with a data volume of 1k - 10k, the RF and LGBM models perform better and have stronger anti-noise capabilities. In terms of model interpretability, for high interpretability requirements such as SHAP analysis of the model, models such as logistic regression, decision trees, and random forests are better choices.

[0088] The feature dimensionality reduction module is used to extract conductive carbon electrode features based on the data in the perovskite solar cell conductive carbon electrode database, encode the conductive carbon electrode features, and then reduce the dimensionality to obtain reduced-dimensional features.

[0089] According to different requirements, one-hot or label encoder is used to encode the input features. In particular, all features related to the conductive carbon electrode are reduced to a single feature through an autoencoder, which can effectively reduce the risk of overfitting and speed up the calculation rate.

[0090] In the regression task of the model, this embodiment generally uses the label method for encoding because this embodiment wants to see more which aspect has a greater impact on the device PCE. In an interpretable model, this embodiment generally uses the one-hot encoding method because this embodiment wants to know which material has a good or bad impact on the PCE of the device. For example, in the regression task, this embodiment knows that the conductive agent in the conductive paste has an important impact on the PCE. In the interpretable model, this embodiment knows that the carbon nanotubes in the conductive agent have a positive impact on the PCE, while graphene has a negative impact. That is, label pays more attention to the overall, and onehot pays more attention to more detailed features within a single feature.

[0091] Use an autoencoder to reduce the dimensionality of features. Build an autoencoder model using Keras. Take the feature descriptors of the conductive carbon electrode formulation as the input layer. In the encoding layer, map several features of the input layer into one feature. In the decoding layer, restore one feature into several features of the input layer to construct a compiled model. Use existing data to train the encoder model. After training, the corresponding features of the conductive carbon electrode components can be reduced to one feature. This method can not only reduce the dimensionality but also retain the correlation between input features and control the quality of dimensionality reduction through the reconstruction error of the autoencoder. The reduced-dimensional feature is a scalar, and the output layer of the encoder is a layer with only one neuron.

[0092] A model training module, used to build an initial fitting prediction model based on a machine learning model, optimize the hyperparameters of the fitting prediction model using grid search, and train the optimized fitting prediction model using the reduced-dimensional features to obtain a trained fitting prediction model.

[0093] Divide the dataset into a training set and a test set in a ratio of 4:1. In the selection of the regression model, mainstream machine learning algorithms including Random Forest, XGBoost, Catboost, and LGBM were selected for prediction. Optimize the hyperparameters of the machine learning model through grid search and save the test results. GridSearchCV is grid search cross-validation parameter tuning. It traverses all permutations and combinations of the input parameters and returns the evaluation metric scores under all parameter combinations through cross-validation.

[0094] Specifically, grid search is an exhaustive search method that finds the optimal hyperparameters by traversing all possible combinations of hyperparameters. Grid search first sets a set of candidate values for each hyperparameter, then generates the Cartesian product of these candidate values to form a combination grid of hyperparameters. Then, grid search will train and evaluate the model for each hyperparameter combination to find the hyperparameter combination with the best performance.

[0095] By optimizing the R2 and RMSE of the machine learning model, select the machine learning model with the best prediction performance for subsequent calculations.

[0096] Use different models to fit and predict the training data, and at the same time, make targeted adjustments to features and parameters based on the characteristics of different models.

[0097] The parameters of the number of hidden layers, the number of nodes, and the activation function were tuned. Subsequently, multiple mainstream algorithms (Random Forest, XGBoost, AdaBoost, LGBM, CatBoost, ANN) were used to establish a forward prediction model for PCE. Based on the results of cross-validation, fine-tuning was performed repeatedly. Finally, the performance on the test set was used as the basis for model selection.

[0098] In machine learning regression problems, feature importance is a quantitative metric for analyzing the contribution or influence of each input feature (feature variable, features) on the model's prediction results. This is particularly crucial for black-box models (such as Random Forest, XGBoost, neural networks, etc.) because their internal mechanisms are complex and it is difficult to directly explain the model's decision-making process. By analyzing feature importance, one can better understand how the model makes predictions. As Figure 3 shown, through the visualization of feature importance, one can more clearly understand the importance of each influencing factor.

[0099] The component design module is used to input the data of the conductive carbon electrode of the perovskite solar cell to be predicted into the trained fitting prediction model to obtain the formulation of the conductive carbon electrode components of the mesoscopic perovskite solar cell, thereby completing the design of the conductive carbon electrode components of the mesoscopic perovskite solar cell based on the machine learning model.

[0100] Through the high-efficiency strategy given by the trained machine learning model, the verification of machine learning prediction is carried out through experiments, as Figure 4 shown. The conductive carbon electrode slurry was obtained through steps such as weighing the medicine, ball milling, and rotary evaporation. The perovskite solar cell device was prepared by screen printing. Finally, the photoelectric conversion efficiency, fill factor, open-circuit voltage, and short-circuit current of the device were measured using the solar simulator K2400. The data obtained from the experiment will be used to verify the accuracy of the model and will continue to be incorporated into the conductive carbon electrode database to start a new round of optimization process.

[0101] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for designing the composition of a conductive carbon electrode of a mesoscopic perovskite solar cell based on a machine learning model, characterized in that, The method includes: Step S1: Construct a sample data set based on the basic data of perovskite solar cell devices, preprocess the sample data set and save it in a database to obtain a database of conductive carbon electrodes for perovskite solar cells; Step S2: Extract the characteristics of the conductive carbon electrode based on the data in the database of the conductive carbon electrode for the perovskite solar cell, encode the characteristics of the conductive carbon electrode and then reduce the dimension to obtain the reduced-dimensional characteristics; Step S3: Construct an initial fitting prediction model based on a machine learning model, optimize the hyperparameters of the fitting prediction model using grid search, and train the optimized fitting prediction model with the reduced-dimensional characteristics to obtain a trained fitting prediction model; Step S4: Input the data of the conductive carbon electrode of the perovskite solar cell to be predicted into the trained fitting prediction model to obtain the component formula of the conductive carbon electrode of the mesoscopic perovskite solar cell, and complete the design of the component of the conductive carbon electrode of the mesoscopic perovskite solar cell based on the machine learning model.

2. The method for designing the conductive carbon electrode component of the mesoscopic perovskite solar cell based on the machine learning model according to claim 1, wherein The basic data includes the preparation process of perovskite solar cell devices, the formula of the conductive carbon electrode, and the performance indicators of the corresponding perovskite solar cells prepared.

3. The method for designing the conductive carbon electrode composition of the mesoscopic perovskite solar cell based on the machine learning model according to claim 2, wherein The performance indicators of the perovskite solar cell include energy conversion efficiency, fill factor, short-circuit current, and open-circuit voltage.

4. The method for designing the conductive carbon electrode component of the mesoscopic perovskite solar cell based on the machine learning model according to claim 3, wherein The specific preprocessing process includes: Perform data cleaning on the performance indicators of the perovskite solar cell in the sample data set and then perform outlier detection.

5. The method for designing the conductive carbon electrode component of the mesoscopic perovskite solar cell based on the machine learning model according to claim 1, wherein The specific method of Step S2 includes: Perform feature encoding on the preprocessed data to obtain feature-encoded data; Use Keras to construct an autoencoder model, and input the feature-encoded data into the autoencoder to obtain the reduced-dimensional feature scalar.

6. A system for designing the composition of a conductive carbon electrode of a mesoscopic perovskite solar cell based on a machine learning model, the system being used to implement the method according to any one of claims 1-5, characterized in that, The system includes: A data set construction module for constructing a sample data set based on the basic data of perovskite solar cell devices, preprocessing the sample data set and saving it in a database to obtain a database of conductive carbon electrodes for perovskite solar cells; A feature dimension reduction module for extracting the characteristics of the conductive carbon electrode based on the data in the database of the conductive carbon electrode for the perovskite solar cell, encoding the characteristics of the conductive carbon electrode and then reducing the dimension to obtain the reduced-dimensional characteristics; A model training module for constructing an initial fitting prediction model based on a machine learning model, optimizing the hyperparameters of the fitting prediction model using grid search, and training the optimized fitting prediction model with the reduced-dimensional characteristics to obtain a trained fitting prediction model; A component design module for inputting the data of the conductive carbon electrode of the perovskite solar cell to be predicted into the trained fitting prediction model to obtain the component formula of the conductive carbon electrode of the mesoscopic perovskite solar cell, and completing the design of the component of the conductive carbon electrode of the mesoscopic perovskite solar cell based on the machine learning model.

7. The system for designing the component of the conductive carbon electrode of the mesoscopic perovskite solar cell based on the machine learning model according to claim 6, wherein The basic data includes the preparation process of perovskite solar cell devices, the formula of the conductive carbon electrode, and the performance indicators of the corresponding perovskite solar cells prepared.

8. The system for designing the conductive carbon electrode composition of the mesoscopic perovskite solar cell based on the machine learning model according to claim 7, wherein The performance indicators of the perovskite solar cell include energy conversion efficiency, fill factor, short-circuit current, and open-circuit voltage.

Citation Information

Cited By

  • Photocurrent prediction method and system for perovskite / crystalline silicon two-end laminated cell

    CN120636622A

  • A method and system for predicting the photocurrent of a perovskite / crystalline silicon tandem cell

    CN120636622B