A method and system for band gap regulation of perovskite materials based on machine learning

By constructing a perovskite material database using machine learning methods and performing multi-dimensional feature extraction and gradient boosting regression model prediction, the problem of insufficient bandgap control accuracy in multi-component systems in existing technologies is solved. This enables efficient and accurate material screening and design, improving the development efficiency and adaptability of photovoltaic materials.

CN122637985APending Publication Date: 2026-08-25SICHUAN EVERSEY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610535123.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-22
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing perovskite material design methods rely on empirical screening or high-cost calculations, making it difficult to achieve precise bandgap control in multi-component systems. This results in low material screening efficiency, long design cycles, and difficulty in obtaining the optimal material combination that meets the requirements of photovoltaic applications.

Method used

A perovskite material composition database was constructed using a machine learning-based approach. Multi-dimensional features were extracted and a comprehensive material feature vector was generated. A gradient boosting regression model was used to predict the bandgap. The bandgap was then screened by combining structural stability constraints and performance indicators. Through reverse optimization and iterative adjustment of the material composition, high-throughput and high-precision bandgap control was achieved.

Benefits of technology

It improves the accuracy and screening efficiency of bandgap prediction for multi-component perovskite materials, ensures that the screened materials have good structural stability within the target bandgap range, reduces the cost of high-throughput computing and experimental trial and error, and enhances the self-learning and evolution capabilities of material design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637985A_ABST
    Figure CN122637985A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of perovskite photovoltaic material design, and discloses a perovskite material band gap regulation method and system based on machine learning, which comprises the following steps: S1, constructing a perovskite material composition database to form a multi-component material combination space; S2, performing feature extraction and fusion processing on the material composition data to generate a comprehensive material feature vector; S3, inputting the comprehensive feature vector into a pre-trained band gap prediction model to obtain a band gap prediction value; S4, based on the deviation of the band gap prediction result and a target range, completing material screening through multi-constraint optimization; and S5, performing performance evaluation on candidate materials and executing iterative optimization. The perovskite material band gap regulation method and system based on machine learning improve the band gap prediction accuracy and material screening efficiency of the multi-component system, effectively represent the complex coupling relationship between components, realize the collaborative optimization of structural stability and component feasibility in the target band gap interval, and greatly shorten the new material development cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of perovskite photovoltaic material design technology, specifically to a method and system for predicting the bandgap and performing high-throughput screening of perovskite photovoltaic materials based on machine learning. Background Technology

[0002] With the global energy structure shifting towards cleaner and lower-carbon energy, the photovoltaic new energy industry has experienced rapid development and widespread application. The development of high-efficiency, low-cost photovoltaic materials has become a core element in improving solar energy utilization efficiency and promoting the upgrading of the photovoltaic industry. Perovskite materials, as a new generation of high-performance optoelectronic materials, have become a research hotspot due to their excellent photoelectric conversion characteristics. Their photoelectric performance and application effects largely depend on the material composition and structure and the control of key parameters. Among these, the precise control of the material's bandgap directly determines the light absorption range, photoelectric conversion efficiency, and device practicality of perovskite materials.

[0003] Existing perovskite material design and bandgap control techniques generally rely on empirical trial and error, first-principles calculations, or simple statistical models for material screening and composition optimization. Bandgap adjustment is achieved by adjusting the proportions of components at the A, B, and X sites. Most existing techniques employ analytical methods that depend on experimental experience, fixed computational logic, or single linear models. This trial-and-error, costly, or weakly fitted approach is based on the core assumption that there is a simple linear relationship between material composition and performance, and that computational resources and exploration space are mutually compatible. However, in actual material development, factors such as strong coupling between A, B, and X sites in multi-component hybrid systems, crystal structure stability constraints, and nonlinear effects caused by disorder are unavoidable. Fixed empirical rules and single models cannot perceive or characterize these factors. These complex dynamic coupling relationships lead to insufficient accuracy in material bandgap prediction and a limited screening range, making it difficult to achieve efficient optimization within the target bandgap range (1.2-1.6 eV), and consistency and accuracy are hard to guarantee. As the material composition space expands, traditional methods inevitably experience a surge in computational costs and a decrease in exploration efficiency. Or, when new cations or halogen components are introduced, the material interaction laws change. At this point, the original empirical parameters and calculation models may no longer be applicable, requiring researchers to conduct tedious experimental trial and error and parameter adjustments. This limitation will greatly reduce the development efficiency of perovskite materials, making it difficult to guarantee the stability of bandgap control and performance screening in multi-component systems, and affecting the rapid iteration and large-scale application of perovskite photovoltaic materials.

[0004] To address the aforementioned issues, there is an urgent need for a perovskite material design method that can achieve high-throughput, high-precision prediction and screening, in order to improve material development efficiency and optimize photovoltaic performance. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for bandgap control of perovskite materials based on machine learning, in order to solve the problem that the design of existing perovskite materials mainly relies on empirical screening or high-cost calculation methods, which makes it difficult to achieve precise bandgap control in multi-component systems. This results in low material screening efficiency, long design cycles, and difficulty in obtaining the optimal material combination that meets the requirements of photovoltaic applications.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A machine learning-based method for bandgap modulation of perovskite materials, the method comprising the following steps:

[0008] S1: Construct a database of perovskite material composition, collect information on the types and molar ratios of A-site cations, B-site metal ions and X-site halide anions, and form a multi-component material combination space.

[0009] S2: Feature extraction and fusion of the material composition data to generate a comprehensive material feature vector specifically includes: extracting the ionization energy, electron affinity, electronegativity, and ionic radius of each component as intrinsic physical properties; calculating the structure factor and interaction features characterizing the stability tendency of the material's crystal structure and the coupling relationship between multiple components, wherein the structure factor includes at least the tolerance factor, octahedral factor, and parameters; the interaction features include at least the BX ionization energy difference, BX electron affinity difference, and BX electronegativity difference; and assigning fusion weights to different feature groups containing the above features and weighting them together based on the mean feature importance output by the pre-trained bandgap prediction model to generate the comprehensive material feature vector.

[0010] S3: Input the comprehensive material feature vector into a pre-trained gradient boosting regression model to predict the band gap index of perovskite materials. The model is used to capture the direct correlation features between material composition and band gap, so as to capture the nonlinear relationship between multi-component material composition, structural parameters and band gap.

[0011] S4: Based on the deviation between the bandgap prediction result and the target range, material screening is completed through structural stability conditions, and the optimal material composition ratio is dynamically determined. The structural stability constraint limits the tolerance factor and octahedral factor to fall within the preset stability range.

[0012] S5: Real-time evaluation of the performance indicators of the selected candidate materials. If the deviation of the performance indicators from the target range exceeds a preset threshold, the gradient boosting regression model is used as a surrogate evaluator, and a constrained optimization search algorithm (such as genetic algorithm, particle swarm optimization, or Bayesian optimization) is used to perform reverse optimization in the material combination space. During this process, the stoichiometric normalization constraint of step S1 and the structural stability constraint of step S4 must be satisfied. During the optimization process, the system automatically adjusts the proportion of components at sites A, B, and X, and returns the updated candidate scheme to step S1 to reconstruct the material combination space. Step S2 re-extracts material features and executes a new round of optimization loop. This iterative process continues until all performance indicators meet the preset targets, and finally, a perovskite material system that meets the requirements is output.

[0013] Preferably, in the S1 stage, the perovskite multi-component system is subjected to stoichiometric normalization and proportion coding to form a standardized material combination space. The specific material composition data includes the following: A-site cation data, including the types and proportions of organic / inorganic cations; B-site metal ion data, including the types and proportions of metal cations; X-site anion data, including the types and proportions of halide anions; and the ternary / quaternary mixed components are subjected to proportion alignment processing to ensure the stoichiometric integrity of the material system.

[0014] The types and proportions of the organic / inorganic cations include one or more of Cs, Rb, K, Na, Li, MA, FA, EA, BA, PEA, and NH4, and their proportions.

[0015] The types and proportions of the metal cations include one or more of Sn, Ge, Pb, Bi, Sb, Ti, Mn, Fe, Co, Ni, Cu, and Zn, and their proportions.

[0016] The anion data includes one or more of F, Cl, Br, I, and SCN, and their proportions.

[0017] Preferably, step S2 includes: S21: extracting intrinsic physical property characteristics of elements, including electronegativity, ionization energy, electron affinity, and ionic radius, and constructing a basic feature set; S22: calculating structural factors and interaction characteristics, including tolerance factor Tf, octahedral factor Of, τ parameter, BX ionization energy difference, BX electron affinity difference, BX electronegativity difference, AB size matching degree, and mixing entropy characteristics; S23: normalizing the basic features, structural features, interaction characteristics, and mixing entropy characteristics, and assigning fusion weights to different feature groups according to feature importance, and generating a unified comprehensive material feature vector by weighted splicing.

[0018] Preferably, the bandgap prediction model in S3 adopts a gradient boosting regression model to capture the direct correlation between element composition and bandgap.

[0019] Preferably, the bandgap prediction model in S3 is further configured with a model version management scheme, including the following steps: S301: Create an independently stored model version each time the model is updated, and record the version metadata; S302: When the mean prediction error within the sliding window continuously exceeds a preset threshold, calculate the distribution difference between the current material feature data and the training sample feature data of each historical model version; S303: Select the historical model with the smallest distribution difference as the optimal historical model, and perform update training based on the feature engineering configuration and hyperparameter combination corresponding to the historical model to generate a new model version, or roll back to a stable version when the preset distribution threshold is met, so as to ensure the stability of prediction and screening.

[0020] Preferably, in step S4, candidate materials are screened using a combination of hard constraint cascade filtering and comprehensive scoring ranking. The hard constraints include at least bandgap range constraints, Tf / Of / τ stability interval constraints, and component fabrication constraints. The comprehensive scoring ranks candidate materials based at least on bandgap deviation, structural stability, and component complexity.

[0021] Preferably, the monitoring and optimization loop steps for candidate material performance indicators in S5 are as follows: S51: Evaluate the performance of the screened candidate materials, including indicators such as band gap value, structural stability, and composition rationality; S52: Establish a mapping relationship model of material composition-characteristics-band gap-stability to form the computational basis for reverse optimization; S53: Perform reverse optimization based on the mapping relationship model and the constraint optimization search algorithm, using the band gap prediction model as the candidate material evaluator, and under the premise of satisfying the constraints of normalization of component ratios at positions A, B, and X, structural stability constraints, and fabrication constraints, search for a material composition scheme that best approximates the target performance, and automatically correct the composition space and re-predict when the performance does not meet the standard.

[0022] This invention also provides a perovskite material bandgap tuning system, which includes the following:

[0023] The material data construction module is used to collect and organize the components and proportions of positions A, B, and X, and construct a standardized combination space;

[0024] The multi-feature fusion module is used to extract physical, structural, interactive, and entropy features and fuse them to generate a comprehensive feature vector.

[0025] The bandgap prediction module is used to load the trained GBR model to achieve fast and accurate bandgap prediction.

[0026] The multi-constraint screening module is used to perform candidate screening based on bandgap targets and stability constraints;

[0027] The closed-loop iterative module is used for performance evaluation and feedback, driving continuous optimization of the component space and model parameters.

[0028] Furthermore, the system also includes a model management module for the bandgap prediction model, which includes the following:

[0029] Model version storage unit, used to independently store prediction models at different training stages;

[0030] The version matching analysis unit is used to perform matching analysis between the current material feature data and the training sample feature data of the historical model version based on the feature distribution difference index, and output the optimal historical model version.

[0031] The model update and rollback unit is used to call the feature engineering configuration and hyperparameter combination corresponding to the best historical model to perform update training when the prediction error continues to exceed the standard, or to switch to a stable version that meets the distribution threshold requirements.

[0032] Compared with the prior art, the beneficial effects of the present invention are: the bandgap control method and system for perovskite materials improve the bandgap prediction accuracy and screening efficiency of multi-component perovskite materials, especially in the material design of complex systems with multi-component coupling at A / B / X sites. It integrates multi-dimensional features in real time and accurately predicts the bandgap, ensuring that the screened candidate materials have good structural stability within the target bandgap range.

[0033] By employing this systematic approach and employing multi-constraint optimization, the system proactively seeks the optimal balance between bandgap accuracy and structural stability, achieving efficient screening while meeting photovoltaic application requirements. This directly reduces the costs of high-throughput computation and experimental trial and error. Simultaneously, it integrates the design experience of materials experts into the intelligent algorithm. Even when introducing new components or expanding the material combination space, the system can automatically adjust the prediction model to adapt to the new system through model self-learning and version management, maintaining optimal screening results. This reduces reliance on specialized experts in specific fields and alleviates the burden on manual labor. Furthermore, it constructs a closed loop from composition design to performance feedback, continuously improving the screening level and forming a continuous improvement cycle of "data—features—prediction—screening—iteration." The system feeds back the performance evaluation results of the final candidate materials to the data construction module, using this data to continuously revise and optimize the prediction model and screening strategy. This enables the entire materials design process to possess self-learning and evolutionary capabilities, improving the efficiency of materials screening and the adaptability of the design. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the operation flow of the method of the present invention;

[0035] Figure 2 This is a schematic diagram of the model management process of the present invention (optional enhancement module).

[0036] Figure 3 This is a schematic diagram illustrating the closed-loop optimization of the performance indicators of the candidate materials of this invention.

[0037] Figure 4 This is a schematic diagram of the system modules of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] The inventive concept of this invention is:

[0040] Existing perovskite bandgap tuning techniques mostly rely on fixed computational logic or single linear models, with the core assumption being a simple linear relationship between material composition and performance. However, in multi-component mixed systems, strong nonlinear dynamic coupling exists between the A-site, B-site, and X-site. When new cations or halogen components are introduced, the material interaction laws fundamentally change, and traditional pre-trained models experience a precipitous degradation in predictive performance due to sample distribution drift, necessitating tedious manual experimental trial and error and parameter adjustment. This invention abandons the conventional approach of treating machine learning merely as a "black box" mapping tool, instead deeply integrating the physicochemical constraints of perovskite crystals with a dynamic model version management mechanism. White-box extraction of physical features: This invention does not stop at simple element splicing, but innovatively introduces interactive features characterizing the strong coupling of multiple components (such as the electronegativity difference and ionization energy difference between B and X) and the mixing entropy feature reflecting the disorder of the system. Adaptive Dynamic Evolution of Distribution Drift: Addressing prediction inaccuracies caused by the introduction of new material systems, this invention innovatively introduces a model matching analysis mechanism based on distribution difference indices such as Maximum Mean Difference (MMD). The system can proactively perceive the difference between the current feature data and the feature distribution of historical model training samples, using this as a trigger condition to select the optimal historical model for feature engineering reuse or version rollback. Existing technologies do not offer technical insights into dynamically switching or updating material prediction models using data distribution difference. This concept, anchoring "algorithm distribution drift monitoring" to "material composition spatial expansion," constitutes the substantial distinguishing technical feature of this application.

[0041] Based on the above concepts, this application achieves technical effects that exceed conventional expectations: breaking the zero-sum game between "accuracy and computing power": through optimal historical model matching and feature configuration retrieval, the system achieves rapid response and convergence when facing a completely new multi-component material space, completely avoiding the enormous time and computing power overhead caused by retraining the model from scratch under a new system. Achieving synergistic evolution of "structure-performance-process": This invention not only uses structural stability parameters such as tolerance factor and octahedral factor, as well as component fabrication feasibility, as pre-filtering conditions, but also embeds them uniformly into the reverse closed-loop optimization process. The bandgap prediction model is transformed into an intelligent evaluator in this process, automatically correcting unreasonable component variations. This ensures that the system outputs not only a theoretically acceptable combination of bandgap values, but also a truly "living" material solution that meets photovoltaic application requirements, is structurally stable, and possesses high fabrication feasibility. Giving the system the "vitality" of self-iteration: Through the closed-loop iteration module, the system continuously feeds back the evaluation feedback of new materials to the prediction model, forming a continuous improvement cycle of "data-feature-prediction-screening-iteration", which transforms the material screening process from static "experience output" into an intelligent closed loop with self-learning and evolution capabilities.

[0042] This invention provides an embodiment of a method for bandgap control of perovskite materials, the method comprising the following steps:

[0043] S1: Construct a perovskite material composition database, collect multi-component material ratio data, including the types and molar ratios of A-site cations, B-site metal ions, and X-site halide anions; perform stoichiometric normalization and ratio coding to form a standardized combination space; clean the raw data, including null value removal, outlier removal, and bandgap data analysis.

[0044] S2: Extract and fuse features from material composition data, and uniformly represent elemental properties, structural factors, interactions and mixing entropy features to generate a comprehensive material feature vector.

[0045] S21: Extract intrinsic physical properties of elements, including electronegativity, ionization energy, electron affinity, and ionic radius, and construct a basic feature set;

[0046] S22: Calculate the structural factors and interaction characteristics, including the tolerance factor Tf, octahedral factor Of, τ parameter, BX ionization energy difference, BX electron affinity difference, BX electronegativity difference, AB size matching degree, and mixing entropy characteristics.

[0047] S23: Normalize the basic features, structural features, interaction features, and mixed entropy features, and assign fusion weights to different feature groups according to feature importance, and generate a unified comprehensive material feature vector by weighted splicing.

[0048] S3: Input the S2 integrated feature vector into the gradient boosting regression (GBR) model for training and bandgap prediction; optimize the model parameters through random hyperparameter search and K-fold cross-validation.

[0049] Preferably, the bandgap prediction model in S3 further includes a model version management scheme, as follows:

[0050] S301: Creates an independently stored model version for each model update and records version metadata, where the initial model consists of experimental and computational data in a multi-component perovskite material system;

[0051] S302: When the prediction error continues to exceed the preset threshold, the system performs a difference analysis on the feature distribution of the current material feature data and the corresponding training samples of each historical model version, and selects the historical model version that best matches the current data based on the degree of distribution difference.

[0052] Specifically, the system uses a sliding window approach to monitor prediction errors on newly added validation samples or external test samples. With a window length of M, the average absolute error of samples within the window is statistically analyzed. When a condition is met, the system determines that the current model's prediction error continues to exceed the limit, triggering the model matching analysis process; where is a preset error threshold.

[0053] During model matching analysis, the feature set of the current material to be processed is denoted as Xcurrent, and the feature set of the training data corresponding to the v-th historical model version is denoted as Xv. The distribution dissimilarity index is used to measure the closeness between the two. Preferably, the distribution dissimilarity index is one or more of the following: maximum mean difference (MMD), Kullback-Leibler divergence, Wasserstein distance, or Mahalanobis distance.

[0054] Preferably, the distribution deviation between the current data and historical data is calculated using the maximum mean difference (MMD).

[0055] The system selects the historical model version that satisfies the minimum requirement as the optimal historical model. That is, when the minimum distribution difference is less than the preset distribution threshold, the historical model is determined to have a high degree of adaptability to the current material system.

[0056] S303: Select the best historical model as the basis for updating training or rolling back the version to improve the prediction stability and convergence efficiency of the model under the current material distribution.

[0057] Specifically, the system retrieves the feature engineering configuration, feature selection order, and optimal hyperparameter combination corresponding to the best historical model from the model version storage unit, including learning rate, number of weak learners, tree depth, and minimum sample splitting parameter; then, it merges the newly added material samples with the training samples corresponding to the historical model according to a preset ratio, retrains the gradient boosting regression model, and obtains a new model version adapted to the current data distribution.

[0058] When there is a historical model version with a distribution difference that meets the preset requirements and has stable historical performance indicators, the system can also directly switch to the stable version to perform bandgap prediction in order to improve the robustness of system operation.

[0059] By adopting the above methods, we can achieve rapid response to situations such as the introduction of new material systems, sample distribution drift, and prediction performance degradation, avoiding the time and computing power overhead of retraining the model from scratch.

[0060] S4: Based on the deviation between the predicted bandgap value and the target bandgap range, and combined with the structural stability constraints of Tf, Of, and τ, as well as the preparability of the components, candidate materials are screened through multi-constraint optimization.

[0061] Preferably, in S4:

[0062] The screening criteria are set with band gap close to the target range, good structural stability and high component prepareability as the screening objectives. The band gap range, Tf, Of, τ stability range and component feasibility are uniformly embedded into the screening process to achieve synergistic screening of structural stability and composition feasibility under the premise of meeting the band gap requirements. The candidate materials are screened and ranked according to the band gap deviation and the constraint satisfaction to determine the final candidate material scheme.

[0063] Among them, the target bandgap interval, Tf / Of / τ stability interval, error threshold, distribution difference threshold, and component complexity threshold can be set based on historical experimental sample statistical results, cross-validation performance, and domain experience, and can be adaptively adjusted as the model is updated.

[0064] S5: Real-time evaluation of candidate material performance indicators. If the performance indicators deviate from the target range by more than a preset threshold, return to S1 to reconstruct the combination space and perform a new round of iterative optimization until the indicators meet the requirements and output the optimal candidate material system.

[0065] S51: Evaluate the performance of the candidate materials output from the screening, including indicators such as band gap value, structural stability, and composition rationality;

[0066] S52: Establish a mapping model of material composition, characteristics, band gap, and stability to form the computational basis for inverse optimization;

[0067] S53: Based on the mapping relationship model and the constraint optimization search algorithm, perform reverse optimization. Using the bandgap prediction model as the candidate material evaluator, under the premise of satisfying the normalization constraints of the component ratios of A-position, B-position and X-position, structural stability constraints and fabrication constraints, search for the material composition scheme that best approximates the target performance, and automatically correct the composition space and re-predict when the performance does not meet the standard.

[0068] Example 1: This invention provides a technical solution: a method for bandgap control of perovskite materials based on machine learning, comprising the following steps:

[0069] S1: Construct a database of perovskite material composition to form a multi-component material combination space;

[0070] The construction of a perovskite material composition database is fundamental for subsequent feature extraction, bandgap prediction, and candidate material screening. The preferred perovskite material is an ABX3 type structure, where the A-site is an organic or inorganic cation, the B-site is a metal cation, and the X-site is a halide anion or pseudohalogen anion.

[0071] In this embodiment, the A-site component includes one or more of Cs, Rb, K, Na, Li, MA, FA, EA, BA, PEA, and NH4; the B-site component includes one or more of Sn, Ge, Pb, Bi, Sb, Ti, Mn, Fe, Co, Ni, Cu, and Zn; and the X-site component includes one or more of F, Cl, Br, I, and SCN.

[0072] Each record in the material composition database includes the following fields: A-position component name, A-position component coefficient, B-position component name, B-position component coefficient, X-position component name, X-position component coefficient, and material band gap data.

[0073] To ensure consistency of data from different sources, the original composition records need to be preprocessed, including: normalizing different naming conventions and mapping different aliases to standard component names; normalizing the component proportions at positions A, B, and X; removing null values, outliers, and erroneous records; and uniformly parsing the bandgap data, converting bandgap data in text, interval, or mixed formats into calculable values.

[0074] In the above scheme, stage S1 involves cleaning, standardizing, and uniformly encoding the compositional data of perovskite materials to form a basic database for subsequent modeling. The specific compositional data includes: A-site compositional data, used to characterize the component types and corresponding proportions of cation sites in the perovskite material; B-site compositional data, used to characterize the component types and corresponding proportions of metal cations in the perovskite framework; X-site compositional data, used to characterize the component types and corresponding proportions of halogen or pseudohalogen anions; and bandgap data, used to characterize the experimental bandgap values ​​of the corresponding materials.

[0075] The preprocessing and normalization of this composition database: After obtaining the raw composition data, cleaning and standardized format conversion are necessary to form effective modeling data. This involves data preprocessing, including null value identification, invalid marker removal, outlier determination, and bandgap numerical analysis. Because material names, component notation, and quantitative expressions differ across literature and databases, they need to be standardized, and the component proportions at each point need to be normalized to ensure that all material records meet a uniform comparability standard. After completing these processes, a perovskite material database with a unified structure and complete fields is generated, providing a guarantee for subsequent feature extraction and model training.

[0076] S2: Extract and fuse features from the material composition data to generate a comprehensive material feature vector;

[0077] The above scheme converts the compositional information of the A-site, B-site, and X-site in perovskite materials into a comprehensive feature vector that can characterize the electronic properties, geometric structure, and degree of mixed disorder of the material, thereby establishing a calculable mapping relationship between the material composition and the band gap.

[0078] S21: Extract the elemental physical property characteristics from the material composition data and process the differences in site properties in the perovskite multi-component system;

[0079] The elemental property feature extraction step is mainly used to capture the basic atomic property information carried by each component of the material. Since the perovskite band gap is affected by the intrinsic properties of each element at each point, this step extracts the basic elemental property parameters corresponding to each component of the material. These basic elemental property parameters include atomic number, group number, period number, atomic weight, electronegativity, ionization energy, electron affinity, ionic radius, and the number of valence shell s, p, d, and f electrons.

[0080] In this embodiment, for the case where sites A, B, and X contain multiple components, the coefficients corresponding to each component are used as weights to perform weighted statistics on the physical properties of the basic elements, resulting in the mean, minimum, maximum, and standard deviation. This forms site A features, site B features, and site X features, such as A_en_mean, B_IE_mean, X_r_ion_mean, A_Z_max, B_mass_std, and X_val_p_mean, which are used to describe the comprehensive properties of the multi-component sites.

[0081] Furthermore, in this embodiment, the mixing degree at positions A, B, and X is preferably characterized by calculating the mixing entropy at position A, position B, and position X, respectively, to reflect the degree of disorder in the multi-component system. Preferably, the entropy value is calculated using the component ratio at each position, thereby obtaining features such as H_mix_A, H_mix_B, and H_mix_X. Through the above steps, discrete material component information can be converted into continuous elemental attribute statistical features, enabling different materials to be compared and learned within a unified feature space.

[0082] S22: Extract the structural and interaction features from the material composition data and output a feature vector representing the material coupling relationship;

[0083] In this step, structural features and interaction features are constructed to further characterize the structural formation tendency of perovskite materials and the interaction relationships between different sites.

[0084] Among them, structural features include tolerance factor T f Octahedral factor O f And the τ parameter. Specifically, it is calculated using the effective radius at site A, the average ionic radius at site B, and the average ionic radius at site X. Where T... f Used to characterize the size matching relationship between position A and the BX3 skeleton; O f The τ parameter is used to characterize the geometric fit between the B-site and the X-site to form an octahedral structure; the τ parameter is used to further evaluate whether the material has a tendency to form a stable perovskite structure.

[0085] Furthermore, to more accurately describe the coupling effects between different sites during bandgap formation, interactive features were also constructed. These interactive features include: BX ionization energy difference features, BX electron affinity difference features, BX electronegativity difference features, B-site valence orbital correlation features, structure factor coupling features, and mixing entropy coupling features.

[0086] In this embodiment, the BX ionization energy difference characteristic is defined as the difference between the average ionization energy at the X site and the average ionization energy at the B site; the BX electron affinity difference characteristic is defined as the difference between the average electron affinity at the X site and the average electron affinity at the B site; the BX electronegativity difference characteristic is defined as the absolute difference between the average electronegativity at the X site and the average electronegativity at the B site; and the structure factor coupling characteristic is defined as T f With O f The product of the total mixed entropy S_mix_total and the product of the mixed entropy at different sites.

[0087] Through the above processing, features such as BX_IE_diff, BX_EA_diff, BX_en_absdiff, Tf_Of, S_mix_total, S_mix_AX, S_mix_BX, S_mix_AB, BX_valence_p_sum, BX_valence_s_sum, and B_d_fraction_proxy are obtained, which can be used to more precisely characterize the differences in electronic structure, geometric matching, and disorder effects in perovskite materials.

[0088] S23: Normalize and weightedly fuse elemental physical properties, structural features, interaction features, and hybrid features to generate a comprehensive material feature vector;

[0089] In this step, the elemental property statistical features and mixed entropy features obtained in S21, as well as the structural features and interaction features obtained in S22, are first subjected to missing value imputation and numerical normalization. It is preferable to use Z-score normalization or min-max normalization so that features of different dimensions can participate in subsequent modeling at a unified scale.

[0090] After normalization, the features are divided into elemental property feature group, structural feature group, interaction feature group, and mixed entropy feature group, and fusion weights are assigned according to the contribution of each feature group to the bandgap prediction task. Preferably, the feature importance output by the pre-trained bandgap prediction model is used as the basis to calculate the mean importance of features within each feature group, and the mean importance of each group is normalized to obtain the corresponding group weights.

[0091] Let the average importance of the g-th feature group be... Then its fusion weight is defined as: ,in, The g-th feature group represents the average feature importance; G is the total number of feature groups; h is the index of the feature groups being traversed. This represents the average feature importance of the h-th feature group.

[0092] Then, the feature groups are weighted and concatenated according to their corresponding weights to obtain the comprehensive material feature vector:

[0093] Where F represents the overall score of the candidate materials; Scoring of the physical properties of basic elements; Score the structural stability; Scoring of interactions between different sites; Scoring the entropy of a multi-component mixture; , , , The weighting coefficients for each scoring indicator can be set based on feature importance or experience, and can be normalized. .

[0094] Through the above processing, the comprehensive feature vector not only includes the basic elemental properties of the material components, but also information such as structural stability, site coupling relationship and mixing disorder effect, thereby improving the model's ability to characterize the bandgap variation law of multi-component perovskites.

[0095] During the model update phase, the weights of each feature group can be recalculated based on the latest training round, so that the feature fusion strategy can be adaptively adjusted as the material distribution changes.

[0096] The technical effects of feature fusion and white-box extraction (S1-S2): By extracting multi-dimensional features such as elemental properties, structural factors, interactions, and mixing entropy, and then performing weighted fusion, this step overcomes the deficiency of traditional methods where a single linear model cannot perceive the strong coupling effects between multiple components. It transforms complex physicochemical constraints into a computable comprehensive feature vector, significantly improving the model's ability to characterize and predict nonlinear effects in multi-component mixed systems.

[0097] S3: Input the comprehensive material feature vector generated in S2 into the gradient boosting regression (GBR) model for training and bandgap prediction; optimize the model parameters through stochastic hyperparameter search and K-fold cross-validation.

[0098] In this embodiment, the bandgap prediction model is a gradient boosting regression model. This model can effectively fit the complex nonlinear relationship between the composition, structural parameters, and bandgap of multi-component perovskite materials through multiple rounds of weak learner iteration. To obtain optimal model parameters, the training and test sets are partitioned, random hyperparameter search is performed, and K-fold cross-validation is used to optimize the model during training.

[0099] The bandgap prediction model outputs predicted bandgap values ​​for the material and can further output model evaluation metrics on the training and test sets to evaluate model performance. These evaluation metrics include the coefficient of determination (R²), root mean square error (RMSE), and mean absolute error (MAE).

[0100] Preferably, the bandgap prediction model in this embodiment further includes a model version management scheme, as follows:

[0101] S301: First, create an independently stored model version for each model training iteration and record version metadata. The initial model is trained based on historical experimental data of perovskite materials. After each model training iteration is completed, save the new model as an independent file and assign it a unique version number. At the same time, record key metadata, including the training data range, number of samples, number of features, hyperparameter combinations, test set metrics, and cross-validation metrics.

[0102] S302: When the prediction error continues to exceed the preset threshold, the system performs a difference analysis on the feature distribution of the current material feature data and the corresponding training samples of each historical model version, and selects the historical model version that best matches the current data based on the degree of distribution difference.

[0103] Specifically, the system uses a sliding window approach to monitor prediction errors on newly added validation samples or external test samples. Let the window length be M, and the mean absolute error of the samples within the window is statistically analyzed. When the following conditions are met... When the current model prediction error continues to exceed the limit, the model matching analysis process is triggered; where δ is the preset error threshold.

[0104] During model matching analysis, the feature set of the current material to be processed is denoted as Xcurrent, and the feature set of the training data corresponding to the v-th historical model version is denoted as Xv. The distribution dissimilarity index is used to measure the closeness between the two. Preferably, the distribution dissimilarity index is one or more of the following: maximum mean difference (MMD), Kullback-Leibler divergence, Wasserstein distance, or Mahalanobis distance.

[0105] Preferably, the distribution deviation between the current data and historical data is calculated using the maximum mean difference (MMD): Dv = MMD * current, Xv.

[0106] The system selects the historical model version that satisfies the minimum Dv as the optimal historical model, that is: When the minimum distribution difference is less than the preset distribution threshold δd, it is determined that the historical model has a high degree of adaptability to the current material system.

[0107] S303: Select the best historical model as the basis for updating training or rolling back the version to improve the prediction stability and convergence efficiency of the model under the current material distribution.

[0108] Specifically, the system retrieves the feature engineering configuration, feature selection order, and optimal hyperparameter combination corresponding to the best historical model from the model version storage unit, including learning rate, number of weak learners, tree depth, and minimum sample splitting parameter; then, it merges the newly added material samples with the training samples corresponding to the historical model according to a preset ratio, retrains the gradient boosting regression model, and obtains a new model version adapted to the current data distribution.

[0109] When there is a historical model version with a distribution difference that meets the preset requirements and has stable historical performance indicators, the system can also directly switch to the stable version to perform bandgap prediction in order to improve the robustness of system operation.

[0110] By adopting the above methods, we can achieve rapid response to situations such as the introduction of new material systems, sample distribution drift, and prediction performance degradation, avoiding the time and computing power overhead of retraining the model from scratch.

[0111] The technical benefits of Dynamic Model Version Management (S3): It introduces a model matching analysis mechanism based on distribution difference indices such as the maximum mean difference. When the prediction error exceeds a preset threshold, the system can automatically match and call upon the best historical model for feature engineering reuse or version rollback. This mechanism overcomes the distribution drift problem that easily occurs in traditional pre-trained models when introducing new material systems, completely avoiding the enormous time and computational overhead of retraining the model from scratch.

[0112] S4: Based on the deviation between the predicted results and the target values, and in combination with structural stability constraints and component prepareability conditions, candidate materials are screened.

[0113] In this step, a two-layer screening mechanism of "hard constraint cascade filtering + comprehensive scoring and ranking" is adopted to conduct multi-objective collaborative screening of candidate materials. The candidate materials must simultaneously meet the requirements of target bandgap, structural stability, and practical fabrication.

[0114] In this embodiment, screening rules are constructed based on the criteria of band gap close to the target range, good structural stability, and high compositional feasibility. Specifically, the target band gap range is set to 1.34. 1.50 eV; structural constraints are set to T f Located at 0.82 Within the range of 1.08, O f Located at 0.40 Within the range of 1.00, τ 4.18; At the same time, the number of mixed components at sites A, B, and X is limited to a preset value to improve the feasibility of candidate materials in actual preparation.

[0115] Specifically, a large number of candidate ABX3 combinations are first generated, and material features are extracted for each candidate combination. These features are then input into the bandgap prediction model to obtain the predicted bandgap value. .

[0116] The first stage involves executing hard-constraint cascaded filtering:

[0117] (1) Stoichiometric validity filtering: The proportions of components at positions A, B and X are required to meet the normalization constraints, and the sum of the proportions at each point is 1;

[0118] (2) Structural stability filtering: requires tolerance factor T f Octahedral factor O f The parameters T and τ fall within a preset stable range; preferably, T f At 0.82 Within the range of 1.08, O f At 0.40 Within the range of 1.00, τ 4.18;

[0119] (3) Preparability filtration: Limit the number of mixed components at site A, site B and site X to no more than a preset value, and limit the proportion of a single component to no less than a preset lower limit, in order to avoid combinations that are too complex or difficult to prepare in practice;

[0120] (4) Bandgap window filtering: Candidate materials whose predicted bandgap values ​​are within the target range are preferred.

[0121] In the second stage, a comprehensive scoring and ranking process is performed on the candidate materials that have passed the hard constraint filtering. Let the comprehensive scoring function be: ,in, This indicates a score indicating that the band gap is close to the target value. Indicates the structural stability score. Indicates the preparability score. These are the corresponding weighting coefficients.

[0122] Preferably, the bandgap score is defined as: ,in R represents the target bandgap center value, and R is the bandgap normalization range.

[0123] For structural stability scoring, it can be based on T f O f The degree to which τ deviates from the center value of the stable interval is normalized; for the prepareability score, a value can be assigned based on the number of components, the degree of dispersion of proportions, and empirical feasibility.

[0124] By using the above two-layer screening method, it is possible to preferentially screen out candidate material schemes that simultaneously meet the target band gap, structural stability and actual preparation requirements from a large-scale candidate material space.

[0125] In another preferred embodiment, a unified constraint objective function can also be constructed:

[0126]

[0127] in, This is a structural stability penalty term. This is a prepareability penalty term; the system optimizes and screens candidate materials by minimizing the objective function J.

[0128] The technical effect of multi-constraint collaborative screening (S4): A multi-layered cascade filtering mechanism is employed, considering bandgap range, structural stability (tolerance factor, octahedral factor, etc.), and component fabrication feasibility. This step ensures that the selected materials are not merely combinations of theoretically acceptable bandgap values, but rather high-quality candidate material solutions that truly meet photovoltaic application requirements, possess structural stability, and demonstrate high fabrication feasibility. S5: Real-time evaluation of candidate material performance indicators. If the performance indicators deviate from the target range beyond a preset threshold, the process returns to S1 to reconstruct the combination space and execute a new round of iterative optimization until the indicators meet the requirements, outputting the optimal candidate material system.

[0129] S51: Evaluate the performance of the candidate materials output from the screening, including indicators such as band gap value, structural stability, and composition rationality;

[0130] S52: Establish a mapping model of material composition, characteristics, band gap, and stability to form the computational basis for inverse optimization;

[0131] S53: Based on the mapping relationship model and the constraint optimization search algorithm, perform reverse optimization, using the bandgap prediction model as the candidate material evaluator, and under the premise of satisfying the normalization constraints of the component ratios of A-site, B-site and X-site, the structural stability constraints and the fabrication constraints, search for the material composition scheme that best approximates the target performance.

[0132] The specific method for achieving closed-loop optimization is as follows: In the process of achieving closed-loop optimization, the reverse optimization does not directly perform analytical inversion of the bandgap prediction model, but uses the bandgap prediction model as a surrogate evaluation model and combines it with a constrained optimization search algorithm to find the optimal solution in the material composition space.

[0133] Specifically, the proportion vector of component A is denoted as Let m be the total number of possible components at position A; the component proportion vector at position B is denoted as... n is the total number of possible components at position B; the component proportion vector at position X is denoted as Let k be the total number of possible components at position X, and satisfy the following stoichiometric constraints:

[0134]

[0135] For any candidate component scheme, the system first generates a corresponding material feature vector based on the component ratio, and then inputs it into the bandgap prediction model to obtain the predicted bandgap value. Then, combining structural stability parameters and fabrication conditions, the comprehensive objective function is calculated:

[0136]

[0137] in, For the target bandgap value, This is a structural stability penalty term. As a preparability penalty term, This is a component complexity penalty term.

[0138] Preferably, the structural stability penalty term is based on T f O f The parameters τ and τ are determined based on whether they deviate from the preset stable range; the prepareability penalty term is determined based on the number of components in the candidate material, the lower limit of the proportion of a single component, and the rules for prohibited components; the component complexity penalty term is determined based on whether the number of non-zero components at positions A, B, and X exceeds the preset threshold.

[0139] Preferably, the constrained optimization search algorithm is any one of genetic algorithm, particle swarm optimization algorithm, or Bayesian optimization algorithm. Taking genetic algorithm as an example, the system first generates an initial candidate population that satisfies the above stoichiometric constraints, and then performs selection, crossover, and mutation operations on the candidate components; in each iteration, the bandgap prediction model is called to calculate the objective function value of each candidate material, and the candidate scheme with the better objective function value is retained to enter the next generation; when the objective function value is lower than a preset threshold or the maximum number of iterations is reached, the optimal material composition scheme is output.

[0140] When the band gap or structural parameters of the current candidate material do not meet the target requirements, the system triggers the above-mentioned constraint optimization search process, automatically adjusts the components and their proportions at positions A, B, and X, and returns the updated component scheme to S1-S4 to re-execute feature extraction, band gap prediction, and multi-constraint screening until the performance of the candidate material gradually converges to the target band gap range and stability requirements.

[0141] Preferably, the structural stability penalty term Defined as T f O f The sum of the deviations of τ from the preset stable interval; when T f O f When both τ and τ fall within the preset stable interval, .

[0142] The technical benefits of inverse closed-loop optimization (S5): By using the bandgap prediction model as a surrogate evaluation model and combining it with a constrained optimization search algorithm, the optimal solution is automatically found within the material composition space. This transforms static material screening into a closed-loop system with adaptive correction capabilities, automatically correcting unreasonable composition variations and re-predicting when performance is substandard.

[0143] Example 2: The present invention further provides a machine learning-based perovskite material bandgap control system, as detailed below:

[0144] The material data construction module is used to construct a perovskite material composition database, forming a multi-component combination space; the multi-feature fusion module is used to extract, statistically process, and fuse material composition data to generate a comprehensive material feature vector; the bandgap prediction module is used to call a pre-trained bandgap prediction model to predict the bandgap of candidate materials; the multi-constraint screening module is used to complete the screening of candidate materials based on bandgap targets, structural stability constraints, and component fabrication conditions; the closed-loop iteration module is used to continuously optimize material composition parameters based on candidate material performance feedback.

[0145] The system also includes a model management module for the bandgap prediction model, which includes:

[0146] Model version storage unit, used to independently store different versions of the prediction model;

[0147] The version matching analysis unit is used to perform matching analysis between the current material feature data and the training sample feature data of the historical model version based on the feature distribution difference index, and output the optimal historical model version.

[0148] The model update and rollback unit is used to call the feature engineering configuration and hyperparameter combination corresponding to the best historical model to perform update training when the prediction error continues to exceed the standard, or to switch to a stable version that meets the distribution threshold requirements.

[0149] The various technical features of this invention can produce the following coordinated technical effects:

[0150] The synergistic effect of "comprehensive feature representation" and "gradient boosting regression model": A unified comprehensive material feature vector, including interactive features (such as the electronegativity difference of BX) and mixed entropy features, is extracted and input into the gradient boosting regression model (GBR) for training. The "strong coupling effect representation" at the feature level and the "multiple weak learner nonlinear fitting" at the algorithm level work together to break the limitations of traditional high-cost first-principles calculations, achieving high-throughput, high-precision bandgap prediction within a vast material composition space. The synergistic effect of "distribution difference monitoring" and "dynamic model evolution": The system measures distribution deviation by calculating the maximum mean difference between current data and historical training data. When the prediction error continues to exceed the limit, the optimal historical model is selected for updating training based on this difference. This closed-loop combination of "algorithm distribution drift monitoring" and "material composition space expansion" gives the system strong adaptability and robustness when facing entirely new multi-component material spaces. The synergistic effect of the "surrogate evaluator" and "objective function optimization with penalty terms": In the reverse optimization stage, the system not only calls the prediction model to calculate the bandgap value, but also introduces a comprehensive objective function that includes penalties for structural stability and compositional complexity for optimization. Minimizing this objective function using search techniques such as genetic algorithms can actively guide the material composition to evolve towards higher fabrication feasibility and higher stability. This combination effectively avoids the algorithm falling into the trap of "pure mathematical extrema," achieving synergistic evolution of "structure-performance-process."

[0151] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for bandgap modulation of perovskite materials based on machine learning, characterized in that, The method includes the following steps: S1: Construct a database of perovskite material composition, collect information on the types and molar ratios of A-site cations, B-site metal ions and X-site halide anions, and form a multi-component material combination space. S2: Feature extraction and fusion of the material composition data to generate a comprehensive material feature vector specifically includes: extracting the ionization energy, electron affinity, electronegativity, and ionic radius of each component as intrinsic physical properties; calculating the structure factor and interaction features characterizing the stability tendency of the material's crystal structure and the coupling relationship between multiple components, wherein the structure factor includes at least the tolerance factor, octahedral factor, and parameters; the interaction features include at least the BX ionization energy difference, BX electron affinity difference, and BX electronegativity difference; and assigning fusion weights to different feature groups containing the above features and weighting them together based on the mean feature importance output by the pre-trained bandgap prediction model to generate the comprehensive material feature vector. S3: Input the comprehensive material feature vector into a pre-trained gradient boosting regression model to predict the band gap index of perovskite materials. The model is used to capture the direct correlation features between material composition and band gap, so as to capture the nonlinear relationship between multi-component material composition, structural parameters and band gap. S4: Based on the deviation between the bandgap prediction result and the target range, material screening is completed through structural stability conditions, and the optimal material composition ratio is dynamically determined. The structural stability constraint limits the tolerance factor and octahedral factor to fall within the preset stability range. S5: Real-time evaluation of the performance indicators of the selected candidate materials. If the deviation of the performance indicators from the target range exceeds a preset threshold, the gradient boosting regression model is used as a surrogate evaluator, and a constrained optimization search algorithm (such as genetic algorithm, particle swarm optimization, or Bayesian optimization) is used to perform reverse optimization in the material combination space. During this process, the stoichiometric normalization constraint of step S1 and the structural stability constraint of step S4 must be satisfied. During the optimization process, the system automatically adjusts the proportion of components at sites A, B, and X, and returns the updated candidate scheme to step S1 to reconstruct the material combination space. Step S2 re-extracts material features and executes a new round of optimization loop. This iterative process continues until all performance indicators meet the preset targets, and finally, a perovskite material system that meets the requirements is output.

2. The method for bandgap control of perovskite materials based on machine learning according to claim 1, characterized in that: In the S1 stage, the perovskite multi-component system is stoichiometrically normalized and coded to form a standardized material combination space. The specific material composition data includes the following: A-site cation data, including the types and proportions of organic / inorganic cations; B-site metal ion data, including the types and proportions of metal cations; X-site anion data, including the types and proportions of halogen anions; and the ternary / quaternary mixed components are calibrated to ensure the stoichiometric integrity of the material system. The types and proportions of the organic / inorganic cations include one or more of Cs, Rb, K, Na, Li, MA, FA, EA, BA, PEA, and NH4, and their proportions. The types and proportions of the metal cations include one or more of Sn, Ge, Pb, Bi, Sb, Ti, Mn, Fe, Co, Ni, Cu, and Zn, and their proportions. The anion data includes one or more of F, Cl, Br, I, and SCN, and their proportions.

3. The method for bandgap control of perovskite materials based on machine learning according to claim 1, characterized in that: Step S2 includes: S21: Extracting intrinsic physical properties of elements, including electronegativity, ionization energy, electron affinity, and ionic radius, and constructing a basic feature set; S22: Calculating structure factors and interaction features, including tolerance factor T. f Octahedral factor O f S23: Normalize the basic features, structural features, interaction features and mixed entropy features, and assign fusion weights to different feature groups according to feature importance, and generate a unified comprehensive material feature vector by weighted splicing.

4. The method for bandgap control of perovskite materials based on machine learning according to claim 1, characterized in that: The bandgap prediction model in S3 adopts a gradient boosting regression model to capture the direct correlation between elemental composition and bandgap.

5. The method for bandgap control of perovskite materials based on machine learning according to claim 1, characterized in that: The bandgap prediction model in S3 is also configured with a model version management scheme, including the following steps: S301: Create an independently stored model version each time the model is updated, and record the version metadata; S302: When the mean prediction error within the sliding window continuously exceeds a preset threshold, calculate the distribution difference between the current material feature data and the training sample feature data of each historical model version; S303: Select the historical model with the smallest distribution difference as the optimal historical model, and perform update training based on the feature engineering configuration and hyperparameter combination corresponding to the historical model to generate a new model version, or roll back to a stable version when the preset distribution threshold is met, so as to ensure the stability of prediction and screening.

6. The method for bandgap control of perovskite materials based on machine learning according to claim 1, characterized in that: In step S4, a combination of hard-constraint cascade filtering and comprehensive scoring ranking is used to screen candidate materials. The hard constraints include at least bandgap range constraints and T... f / O f / τ stability interval constraints and component fabrication constraints; the comprehensive score is based on at least bandgap deviation, structural stability and component complexity to rank candidate materials.

7. The method for bandgap control of perovskite materials based on machine learning according to claim 1, characterized in that: The monitoring and optimization loop steps for candidate material performance indicators in S5 are as follows: S51: Evaluate the performance of the screened candidate materials, including indicators such as band gap value, structural stability, and composition rationality; S52: Establish a mapping relationship model of material composition-characteristics-band gap-stability to form the computational basis for reverse optimization; S53: Perform reverse optimization based on the mapping relationship model and the constraint optimization search algorithm, using the band gap prediction model as the candidate material evaluator, and under the premise of satisfying the normalization constraints of the component ratios of positions A, B, and X, structural stability constraints, and fabrication constraints, search for a material composition scheme that best approximates the target performance, and automatically correct the composition space and re-predict when the performance does not meet the standard.

8. A machine learning-based perovskite material bandgap control system, applied to the machine learning-based perovskite material bandgap control method according to any one of claims 1-7, characterized in that, The system includes: a material data construction module for collecting and organizing the types and proportions of components A, B, and X, and constructing a standardized multi-component material combination space; a multi-feature fusion module for extracting property, structure, interaction, and entropy features and fusing them to generate a comprehensive material feature vector; a bandgap prediction module for loading a trained gradient boosting regression model to accurately predict the bandgap index of perovskite materials; a multi-constraint screening module for performing candidate material screening based on bandgap targets and stability constraints; and a closed-loop iteration module for evaluating and providing feedback on candidate material performance, driving continuous optimization of the component space and model parameters.

9. The perovskite material bandgap control system based on machine learning according to claim 8, characterized in that: The system also includes a model management module for bandgap prediction models, comprising: a model version storage unit for independently storing prediction models and corresponding metadata for different training stages; a version matching analysis unit for performing model matching analysis based on the distribution difference between the current material feature data and the feature data of historical model training samples when the prediction error continues to exceed the standard, and outputting the optimal historical model version; and a model update and rollback unit for calling the feature engineering configuration and hyperparameter combination corresponding to the optimal historical model to perform update training, or switching to a stable version that meets the preset distribution threshold requirements to ensure the stability of system prediction and screening.