Global search method for stable adsorption configurations based on similarity distribution and active learning

Through similarity distribution and active learning methods, a high-throughput model of knowledge embedding and a small sample data set were constructed. Combined with the Gaussian process regression model, the problems of high cost and high resource consumption in heterogeneous catalysis research were solved, and efficient and stable adsorption configuration search was achieved.

CN119557676BActive Publication Date: 2025-09-26ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411535308.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-09-26
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

In the existing technology of heterogeneous catalysis research, the high-throughput calculation of stable adsorption configuration search is costly and resource-intensive, making it difficult to search for stable adsorption configurations efficiently.

Method used

A method based on similarity distribution and active learning was adopted to construct a high-throughput adsorption structure model with knowledge embedding, cluster and hierarchically construct a small sample data set, combine it with the Gaussian process regression model, establish an agent model, and use a hybrid convergence criterion to search for stable adsorption configurations.

Benefits of technology

It achieves efficient search for stable adsorption configurations with less computing resources and iterations, reduces computing costs and resource consumption, and improves search efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557676B_ABST
    Figure CN119557676B_ABST
Patent Text Reader

Abstract

The present invention discloses a global search method for stable adsorption configurations based on similarity distribution and active learning, comprising the following steps: S1, constructing a knowledge-embedded high-throughput adsorption structure model; S2, constructing a small-sample adsorption structure dataset based on clustered stratified sampling; S3, establishing a proxy model for adsorption energy prediction by combining the small-sample adsorption structure dataset with a Gaussian process regression model; and S4, training the proxy model to search for stable adsorption configurations. The global search method for stable adsorption configurations based on similarity distribution and active learning, disclosed in the present invention, achieves efficient search for stable adsorption configurations from a large structure space by combining high-throughput modeling of catalytic configurations with knowledge embedding, constructing a small-sample dataset based on clustered stratified sampling, a small-sample recommendation model based on Gaussian process regression, and a hybrid convergence criterion that considers data distribution and proxy model prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computational chemistry, and in particular relates to a global search method for stable adsorption configurations based on similarity distribution and active learning. Background Art

[0002] Computational catalysis research based on knowledge and first-principles methods holds enormous potential for cost-effectively exploring a vast material space and accelerating expensive, long-term trial-and-error experimental research. It plays a leading role in catalyst development within the third scientific paradigm, driven by computational simulations, and has become a cornerstone of the fourth scientific paradigm, data-driven catalysis research. In theoretical research on heterogeneous catalysis, the calculation of absorption energies of key species and the identification of stable adsorption structures are fundamental for elucidating catalytic reaction mechanisms and structure-activity relationships, as well as for the rational design and screening of new catalytic systems. In practical research, high-throughput computational methods are often used to determine the optimal binding configurations of adsorbates, such as molecules or atoms, on specific crystal faces under complex surface chemistry. However, this in turn leads to significant computational resource and cost demands, dominating the computational resources required for high-throughput studies of catalytic mechanisms and machine learning research. Therefore, efficiently and cost-effectively searching for stable adsorption configurations within the vast structural space of all possible surface-bound configurations of adsorbates is crucial for reducing the cost of theoretical catalyst design and development.

[0003] Combining the potential energy surface fitting of first-principle calculations or machine learning with the search scheme of basin jumping or minimum jumping algorithms can effectively reduce the number of calculated configurations and reduce research costs compared to enumeration-based calculation studies. However, the potential energy surface fitting data and multiple calculation iterations still inevitably require a large amount of computing resources. Summary of the Invention

[0004] In view of this, an embodiment of the present invention discloses a global search method for stable adsorption configurations based on similarity distribution and active learning, comprising the steps of:

[0005] S1. Constructing a high-throughput adsorption structure model with knowledge embedding; specifically including:

[0006] Enumerate all potential surface sites or site combinations in the adsorption system;

[0007] Combine domain knowledge of the adsorbate on the specific catalytic surface of the adsorption system to remove clearly unstable sites or site combinations;

[0008] Automatically construct a surface adsorption model of the adsorbate binding to the catalytic crystal surface site based on the selected site or site combination;

[0009] S2. Constructing a small sample dataset of adsorption structure based on clustering and stratified sampling; specifically including:

[0010] Combined with the physical property information of the surface atoms of the catalytic configuration, candidate features describing the adsorption structure are constructed;

[0011] Cluster analysis was performed on all candidate adsorption structures using clustering method;

[0012] According to the clustering, the proportion of adsorption structures in different clusters is obtained, and the farthest point sampling is used to extract the corresponding number of adsorption structures from each cluster;

[0013] High-throughput calculation of the binding energy of adsorbates on selected adsorption structures to construct a small sample data set of adsorption structures with similar data distribution to the large candidate space;

[0014] S3. Combining a small sample data set of adsorption structures with a Gaussian process regression model, a proxy model for adsorption energy prediction is established.

[0015] S4. Training the agent model to search for stable adsorption configurations; specifically including:

[0016] Based on the predicted value and uncertainty output by the surrogate model, the acquisition function value of all adsorption structure samples is calculated, that is, the probability of improvement relative to the optimal value of the small sample of adsorption structure, or the expected improvement value;

[0017] Select the candidate adsorption structure with the largest improvement value as the search result, calculate the binding energy of the adsorbate on the candidate adsorption structure, and update the small sample data set of the adsorption structure;

[0018] Determine whether the hybrid convergence criterion is met. If so, stop the iteration, and the optimal structure in the small sample data set of adsorption structures is the global optimal adsorption structure, which is determined as the stable adsorption configuration; if not, retrain the proxy model and conduct the next search.

[0019] Furthermore, in some embodiments of the global search method for stable adsorption configurations based on similarity distribution and active learning, in step S4, the hybrid convergence criterion is the union of the first criterion and the second criterion; wherein the first criterion is a criterion based on similarity distribution, and the second criterion is a criterion based on the prediction results of the agent model.

[0020] In some embodiments, the global search method for stable adsorption configurations based on similarity distribution and active learning is disclosed, wherein the first criterion is a criterion for directly determining the global optimal value based on Gaussian distribution; specifically, the method includes:

[0021] Fitting a Gaussian distribution to the binding energy data of a small sample data set of adsorption structures;

[0022] Determine whether the optimal adsorption energy of a small sample data set of adsorption structures is in the area to the left of the mean value of the Gaussian distribution minus two times the standard deviation;

[0023] If it is satisfied, the optimal adsorption structure of the small sample data set is determined to be the global most stable adsorption structure of the large candidate space; if it is not satisfied, the optimal adsorption structure of the small sample data set is determined not to be the global most stable adsorption structure of the large candidate space.

[0024] In some embodiments of the global search method for stable adsorption configurations based on similarity distribution and active learning, the second criterion is a criterion for determining the uncertainty change in the binding energy of the global optimal structure predicted by the surrogate model; specifically, it includes:

[0025] It is determined whether the uncertainty change of the binding energy of the global optimal structure predicted by the proxy model is less than the set threshold. If it is satisfied, the proxy model stops iterating and determines that the current predicted global optimal structure is the globally most stable adsorption structure; if it is not satisfied, the proxy model continues to iterate.

[0026] In the global search method for stable adsorption configurations based on similarity distribution and active learning disclosed in some embodiments, the threshold for the uncertainty change of the binding energy is set according to the accuracy of the density functional calculation.

[0027] In the global search method for stable adsorption configurations based on similarity distribution and active learning disclosed in some embodiments, the threshold for uncertainty change in binding energy is set to 0.025 eV.

[0028] Some embodiments disclose a global search method for stable adsorption configurations based on similarity distribution and active learning. For an alloy system, sites or site combinations containing at least two elements among the surface atoms bound by the adsorbate are selected as potential surface sites or site combinations.

[0029] Some embodiments disclose a global search method for stable adsorption configurations based on similarity distribution and active learning. For co-adsorption configurations, site combinations whose distances between adsorbates satisfy a set range are selected as potential site combinations.

[0030] In the global search method for stable adsorption configurations based on similarity distribution and active learning disclosed in some embodiments, in step S4, multiple candidate adsorption structures with the largest improvement values ​​are selected as search results.

[0031] In the global search method for stable adsorption configurations based on similarity distribution and active learning disclosed in some embodiments, in step S4, two candidate adsorption structures with the largest improvement values ​​are selected as search results.

[0032] The global search method for stable adsorption configurations based on similarity distribution and active learning disclosed in an embodiment of the present invention realizes efficient search for stable adsorption configurations from a large structure space through high-throughput modeling of catalytic configurations including knowledge embedding, construction of a small sample data set based on clustering and stratified sampling, a small sample recommendation model based on Gaussian process regression, and a hybrid convergence criterion considering data distribution and proxy model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 , a flow chart of the global search method for stable adsorption configurations based on similarity distribution and active learning in Example 1;

[0034] Figure 2 , Schematic diagram of the co-adsorption model of NH and OH on the Cu diatomic configuration on the Pt(111) crystal surface in Example 1;

[0035] Figure 3 , the performance of the NH_OH adsorption configuration binding energy prediction model in Example 1;

[0036] Figure 4 , the search results based on the initial small sample derived from different random numbers in Example 1;

[0037] Figure 5 , Active learning search results based on small sample data sets of different sizes (20, 30 and 40) and different sampling strategies in Example 1. DETAILED DESCRIPTION

[0038] The term "embodiment" is used herein specifically to describe any embodiment as "exemplary," and should not be construed as superior or preferable to other embodiments. Performance indicators in the embodiments of the present invention were tested using conventional testing methods in the art, unless otherwise specified. It should be understood that the terms used in the embodiments of the present invention are intended solely to describe specific implementations and are not intended to limit the disclosure of the embodiments of the present invention.

[0039] Unless otherwise specified, the technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which the embodiments of the present invention pertain; any experimental methods and technical means not otherwise specified in the embodiments of the present invention refer to experimental methods and technical means commonly used by those skilled in the art.

[0040] As used herein, the terms "substantially" and "approximately" are used to describe small fluctuations. For example, they can refer to less than or equal to ±5%, such as less than or equal to ±2%, such as less than or equal to ±1%, such as less than or equal to ±0.5%, such as less than or equal to ±0.2%, such as less than or equal to ±0.1%, such as less than or equal to ±0.05%. Numerical data expressed or presented in range format herein are used for convenience and brevity only and should therefore be interpreted flexibly to include not only the values ​​explicitly listed as the limits of the range, but also all independent values ​​or subranges contained within the range. For example, a numerical range of "1-5%" should be interpreted to include not only the explicitly listed values ​​of 1% to 5%, but also the independent values ​​and subranges within the indicated range. Thus, included in this numerical range are independent values ​​such as 2%, 3.5%, and 4%, and subranges such as 1% to 3%, 2% to 4%, and 3% to 5%, etc. This principle also applies to ranges that only list a single value. Furthermore, this interpretation applies regardless of the width of the range or the characteristics described.

[0041] Throughout this document, including in the claims, transitional terms such as "comprises," "includes," "with," "having," "contains," "involving," and "accommodating" are understood to be open-ended, meaning "including but not limited to." Only the transitional terms "consisting of" and "composed of" are closed transitional terms.

[0042] In order to better illustrate the present invention, numerous specific details are provided in the following specific examples. It should be understood by those skilled in the art that the present invention can be practiced without certain specific details. In the examples, some methods, means, instruments, equipment, etc. well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present invention.

[0043] Under the premise of no conflict, the technical features disclosed in the embodiments of the present invention can be arbitrarily combined, and the resulting technical solutions belong to the contents disclosed in the embodiments of the present invention.

[0044] Generally, a large structure space refers to a structure space consisting of all possible structures that contain the optimal structure. Global search refers to a method that attempts to consider the entire search space of a problem to ensure that the global optimal solution is found. It is a professional term in contrast to local search; local search starts from an initial solution and gradually finds the optimal solution through continuous small steps of improvement. It is prone to getting stuck in local optimality, that is, the optimal solution obtained may only be within the current neighborhood, not necessarily within the entire search space.

[0045] In some embodiments, a global search method for stable adsorption configurations based on similarity distribution and active learning comprises the steps of:

[0046] S1. Constructing a high-throughput adsorption structure model with knowledge embedding; specifically including:

[0047] Enumerate all potential surface sites or site combinations in the adsorption system;

[0048] Combine domain knowledge of the adsorbate on the specific catalytic surface of the adsorption system to remove clearly unstable sites or site combinations;

[0049] Automatically construct a surface adsorption model of the adsorbate binding to the catalytic crystal surface site based on the selected site or site combination;

[0050] S2. Construct a small sample data set of adsorption structure based on clustering and stratification; specifically including:

[0051] Combined with the electronegativity, atomic mass, valence electrons and other physical properties of the surface atoms of the catalytic configuration, candidate features describing the adsorption structure are constructed, such as the average electronegativity, atomic mass, and valence electrons of surface atoms, the average electronegativity, atomic mass, and valence electrons of subsurface atoms, and the average electronegativity, atomic mass, and valence electrons of the atoms of the adsorbed species;

[0052] Cluster analysis is performed on all candidate adsorption structures using clustering methods; clustering methods usually include K-means clustering algorithm, etc.

[0053] According to the proportion of adsorption structures in different clusters obtained by clustering, the farthest point sampling method is used to extract the corresponding number of adsorption structures from each cluster;

[0054] High-throughput calculation of the binding energy of adsorbates on selected adsorption structures to construct a small sample data set of adsorption structures with similar data distribution to the large candidate space;

[0055] S3. Combining a small sample data set of adsorption structures with a Gaussian process regression model, a proxy model for adsorption energy prediction is established.

[0056] S4. Training the agent model to search for stable adsorption configurations; specifically including:

[0057] Based on the predicted value and uncertainty output by the surrogate model, the acquisition function value of all adsorption structure samples is calculated, that is, the probability of improvement (PI) or expected improvement (EI) relative to the optimal value of the small sample of adsorption structure.

[0058] The candidate adsorption structure with the largest lift value is selected as the search result, the binding energy of the adsorbate on the candidate adsorption structure is calculated, and the small sample data set of the adsorption structure is updated; usually, one or more candidate adsorption structures with the largest lift value can be selected; multiple candidate structures with the largest lift value usually refer to sorting by the expected lift value size, and the required number of candidate adsorption structures are selected in the order of ranking position.

[0059] Determine whether the hybrid convergence criterion is met. If so, stop the iteration, and the optimal structure in the small sample data set of adsorption structures is the global optimal adsorption structure, which is determined as the stable adsorption configuration; if not, retrain the proxy model and conduct the next search.

[0060] In some embodiments of the disclosed global search method for stable adsorption configurations based on similarity distribution and active learning, in step S4, the hybrid convergence criterion is the union of a first criterion and a second criterion; wherein the first criterion is based on the similarity distribution, and the second criterion is based on the prediction results of the surrogate model. Typically, the union of the first and second criteria serves as the hybrid convergence criterion, typically meaning that if both the first and second criteria are satisfied, the optimal structure found is determined to be the global optimal adsorption structure, i.e., the most stable adsorption configuration.

[0061] Typically, in a hybrid convergence criterion, the first criterion based on similarity distribution can determine the optimal structure with no or few iterations. The second criterion, based on the prediction results of the surrogate model, can stop model iterations promptly when the search model reaches a certain accuracy. This prevents excessive active learning in potential optimal regions, which can lead to deviations in the distribution of small samples from the candidate space data and ultimately lead to search failure. The combination of these two criteria ensures that the model searches for the global optimal structure with fewer iterations. This hybrid convergence criterion based on similarity distribution and surrogate models effectively balances efficiency and accuracy while determining the global optimal value and suppressing excessive model iterations and distribution deviations.

[0062] In some embodiments, a stable adsorption configuration search method based on similarity distribution and active learning achieves efficient search for the global optimum in a large space based on a small initial sample of less than 10% and a small number of iterations of less than 2.

[0063] In some embodiments, the first criterion is a criterion for directly determining the global optimal value based on a Gaussian distribution; specifically, the first criterion includes: fitting a Gaussian distribution to the binding energy data of a small sample data set of adsorption structures; determining whether the optimal adsorption energy of the small sample data set of adsorption structures is in the area to the left of the mean value of the Gaussian distribution minus two standard deviations; that is, determining whether the optimal adsorption energy satisfies the following conditions:

[0064] E min ≤μ-2σ;

[0065] Among them, E min is the optimal adsorption energy of the small sample data set, μ is the mean of the Gaussian distribution, and σ is the standard deviation;

[0066] If it is satisfied, the optimal adsorption structure of the small sample data set is determined to be the global most stable adsorption structure of the large candidate space; if it is not satisfied, the optimal adsorption structure of the small sample data set is determined not to be the global most stable adsorption structure of the large candidate space.

[0067] In some embodiments, the second criterion is a criterion for determining the change in uncertainty of the binding energy of the global optimal structure predicted by the surrogate model; specifically, it includes:

[0068] The surrogate model determines whether the uncertainty change in the binding energy of the global optimal structure predicted by the surrogate model is less than a set threshold. If so, the surrogate model stops iterating and determines that the current predicted global optimal structure is the most stable adsorption structure. If not, the surrogate model continues iterating. The threshold for the uncertainty change in the binding energy is set to 0.025 eV based on the density functional calculation accuracy.

[0069] The technical details are further illustrated below with reference to embodiments.

[0070] Example 1

[0071] In Example 1, the search for a stable adsorption configuration of NH and OH groups co-adsorbed on a Cu diatomic configuration on a Pt(111) crystal surface is taken as an example to illustrate the global search method for a stable adsorption configuration based on similarity distribution and active learning; the Cu diatomic configuration on a Pt(111) crystal surface is as follows: Figure 2 As shown; Figure 2 Where 1 represents Cu atom, 2 represents N and 3 represents OH;

[0072] A global search method for the stable adsorption configuration of NH and OH groups co-adsorbed (NH_OH) on the Cu diatomic configuration on the Pt(111) crystal surface is used, such as Figure 1 Shown, including:

[0073] S1. Construct a knowledge-embedded high-throughput adsorption structure model; specifically, enumerate all potential surface sites or site combinations on the Pt(111) crystal surface; remove clearly unstable sites or site combinations by combining domain knowledge of adsorbate NH and OH digroups on the Pt(111) crystal surface; and automatically construct an adsorption surface model for the adsorbate bound to the catalytic crystal surface site based on the selected sites or site combinations;

[0074] S2. Construct a small sample dataset of adsorption structures based on clustering and stratification. This includes: combining the physical property information of the surface atoms of the catalytic configuration to construct candidate features describing the adsorption structure; using clustering methods to perform cluster analysis on all candidate adsorption structures; using farthest point sampling to extract a corresponding number of adsorption structures from each cluster based on the number of adsorption structures in different clusters; and high-throughput calculation of the binding energy of the adsorbate on the selected adsorption structure to construct a small sample dataset of adsorption structures with a data distribution similar to that of the large candidate space.

[0075] S3. Combining a small sample data set of adsorption structures with a Gaussian process regression model, a proxy model for adsorption energy prediction is established.

[0076] S4. Training the agent model to search for stable adsorption configurations; specifically including:

[0077] According to the predicted value and uncertainty output by the surrogate model, the improvement probability or expected improvement value of all adsorption structure samples relative to the optimal value of the adsorption structure small sample is calculated;

[0078] Select the candidate adsorption structure with the largest improvement value as the search result, calculate the binding energy of the adsorbate on the candidate adsorption structure, and update the small sample data set of the adsorption structure;

[0079] Determine whether the hybrid convergence criterion is met. If so, stop the iteration, and the optimal structure in the small sample data set of adsorption structures is the global optimal adsorption structure, which is determined as the stable adsorption configuration; if not, retrain the proxy model and conduct the next search.

[0080] In Example 1, a high-throughput adsorption structure model based on knowledge embedding was constructed, and a large dataset of 405 NH_OH co-adsorption structure models was constructed;

[0081] Combining K-means clustering, farthest point sampling, and high-throughput computing, a small sample data set of adsorption configurations containing 30 samples was constructed. A Gaussian distribution was fitted to the small sample data set of adsorption configurations, and the boundaries of the 95% confidence interval of the distribution were determined.

[0082] Based on Gaussian process regression and a small sample data set of adsorption configurations, a prediction model for the binding energy of NH_OH adsorption configurations was trained.

[0083] Predict the binding energy of the system in a large material space and the uncertainty of the minimum value in the predicted value, and calculate the mean absolute error (MAE) and root mean square error (RMSE) of the training set and the test set using 5-fold cross validation; Figure 3 As shown, it shows that the prediction model based on small samples has better accuracy;

[0084] The sampling effect is evaluated using a hybrid convergence criterion. If the hybrid convergence criterion is met, iteration is terminated. Otherwise, active learning based on Bayesian optimization is used to search for new data points from the large material space, adjust the distribution of the small sample dataset, and optimize the prediction model. The prediction model is used to calculate the binding energies corresponding to all adsorption structure models in the large dataset. The expected improvement (EI) acquisition function is used to evaluate the expected improvement of each adsorption configuration. The two adsorption configurations with the largest improvement are recommended and their corresponding binding energies are calculated. The small sample dataset of adsorption configurations is then added, and the hybrid criterion is again used to evaluate whether the model has recommended the globally optimal adsorption configuration. This process is repeated until the model converges.

[0085] The search results based on the initial small sample obtained using different random numbers are as follows Figure 4 As shown, considering the calculation error, the binding energy results of -6.942 and -6.944 eV correspond to the same NH_OH adsorption structure. It can be seen that in Example 1, the global optimal value, i.e., the most stable NH_OH adsorption structure, can be determined in less than two iterations, demonstrating the excellent robustness of the global search method based on active learning and similarity distribution.

[0086] The search performance of the global search method is further evaluated by combining small sample datasets of adsorption configurations of different sizes and different sampling strategies, e.g. Figure 5 As shown. The results show that compared with random sampling and farthest point sampling, the active learning search based on the composite sampling strategy of clustering and farthest point sampling can achieve the determination of the global optimal configuration with few active learning iterations or no iterations, and has the highest sampling efficiency. On the other hand, Figure 5 As shown in Figure c, the active learning framework can achieve iteration-free search to the global optimum in small sample data sets containing 20, 30, and 40 data points, indicating that the active learning framework of this global search method has strong robustness.

[0087] The global search method for stable adsorption configurations based on similarity distribution and active learning disclosed in an embodiment of the present invention realizes efficient search for stable adsorption configurations from a large structure space through high-throughput modeling of catalytic configurations including knowledge embedding, construction of a small sample data set based on clustering and stratified sampling, a small sample recommendation model based on Gaussian process regression, and a hybrid convergence criterion considering data distribution and proxy model prediction.

[0088] The technical solutions and technical details disclosed in the embodiments of the present invention are merely illustrative of the inventive concept of the present invention and do not constitute a limitation on the technical solutions of the embodiments of the present invention. Any conventional changes, replacements or combinations of the technical details disclosed in the embodiments of the present invention have the same inventive concept as the present invention and are within the scope of protection of the claims of the present invention.

Claims

1. A global search method for stable adsorption configurations based on similarity distribution and active learning, characterized by: Including steps: S1. Constructing a high-throughput adsorption structure model with knowledge embedding; specifically including: Enumerate all potential surface sites or site combinations in the adsorption system; Combine domain knowledge of the adsorbate on the specific catalytic surface of the adsorption system to remove clearly unstable sites or site combinations; Automatically construct a surface adsorption model of the adsorbate binding to the catalytic crystal surface site based on the selected site or site combination; S2. Constructing a small sample dataset of adsorption structure based on clustering and stratified sampling; specifically including: Combined with the physical property information of the surface atoms of the catalytic configuration, candidate features describing the adsorption structure are constructed; Cluster analysis was performed on all candidate adsorption structures using clustering method; According to the clustering, the proportion of adsorption structures in different clusters is obtained, and the farthest point sampling is used to extract the corresponding number of adsorption structures from each cluster; High-throughput calculation of the binding energy of adsorbates on selected adsorption structures to construct a small sample data set of adsorption structures with similar data distribution to the large candidate space; S3. Combining a small sample data set of adsorption structures with a Gaussian process regression model, a proxy model for adsorption energy prediction is established. S4. Training the agent model to search for stable adsorption configurations; specifically including: Based on the predicted value and uncertainty output by the proxy model, the acquisition function value of all adsorption structure samples is calculated, that is, the probability of improvement or expected improvement relative to the optimal value of the small sample of adsorption structure; Select the candidate adsorption structure with the largest improvement value as the search result, calculate the binding energy of the adsorbate on the candidate adsorption structure, and update the small sample data set of the adsorption structure; Determine whether the hybrid convergence criterion is met. If so, stop the iteration, and the optimal structure in the small sample data set of adsorption structures is the global optimal adsorption structure, which is determined to be the stable adsorption configuration; if not, retrain the proxy model and perform the next search result.

2. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 1, characterized in that: In step S4, the hybrid convergence criterion is the union of the first criterion and the second criterion; wherein the first criterion is a criterion based on similarity distribution, and the second criterion is a criterion based on the prediction result of the surrogate model.

3. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 2, characterized in that: The first criterion is a criterion for directly determining the global optimal value based on Gaussian distribution; specifically, it includes: Fitting a Gaussian distribution to the binding energy data of a small sample data set of adsorption structures; Determine whether the optimal adsorption energy of a small sample data set of adsorption structures is in the area to the left of the mean value of the Gaussian distribution minus two times the standard deviation; If it is satisfied, the optimal adsorption structure of the small sample data set is determined to be the global most stable adsorption structure of the large candidate space; if it is not satisfied, the optimal adsorption structure of the small sample data set is determined not to be the global most stable adsorption structure of the large candidate space.

4. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 2, characterized in that: The second criterion is a criterion for determining the uncertainty change of the binding energy of the global optimal structure predicted by the surrogate model; specifically, it includes: It is determined whether the uncertainty change of the binding energy of the global optimal structure predicted by the proxy model is less than the set threshold. If it is satisfied, the proxy model stops iterating and determines that the current predicted global optimal structure is the globally most stable adsorption structure; if it is not satisfied, the proxy model continues to iterate.

5. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 1, characterized in that: For alloy systems, sites or site combinations containing at least two elements among the surface atoms bound by the adsorbate are selected as potential surface sites or site combinations.

6. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 1, characterized in that: For the co-adsorption configuration, the site combinations whose distances between the adsorbates meet the set range are selected as potential site combinations.

7. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 1, characterized in that: In step S4, a plurality of candidate adsorption structures with the largest improvement values ​​are selected as the push search results.

8. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 1, characterized in that: In step S4, two candidate adsorption structures with the largest improvement values ​​are selected as search results.

9. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 4, characterized in that: The threshold for the uncertainty change in the binding energy is set according to the accuracy of the density functional calculation.

10. The global search method for stable adsorption configuration based on similarity distribution and active learning according to claim 9, characterized in that: The uncertainty threshold of the binding energy change was set to 0.025 eV.