Catalyst screening method, device and system and electronic equipment
By combining Bayesian optimization models and large language models, highly efficient catalyst combinations were screened, solving the problem of low screening efficiency caused by the large number of Lewis acid-base combinations, and improving the efficiency of PET glycolysis reaction.
Patent Information
- Application Number
- CN202511134660.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, the number of Lewis acid and base combinations is enormous. Exhaustive testing of all combinations is inefficient and costly, making it difficult to efficiently screen for high-efficiency catalyst combinations to improve the efficiency of PET glycolysis.
By combining Bayesian optimization models and large language models, we screened and predicted Lewis acid-base pairs, used Bayesian optimization models for initial screening, and combined large language models for data analysis and hypothesis generation to find highly efficient catalyst combinations.
This significantly improved the efficiency of catalyst screening, leading to the discovery of highly efficient catalyst combinations such as zinc neopentanoate and N,N'-diethylethylenediamine, which increased the BHET yield from 86% to 95%, reducing the workload for researchers.
Smart Images

Figure CN121034474A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of plastic degradation technology, and in particular to a catalyst screening method, apparatus, system and electronic equipment. Background Technology
[0002] PET (especially polyethylene terephthalate) is a polymer widely used in packaging and textiles. PET has extremely poor natural degradation capabilities, and its long-term presence in ecosystems will pollute soil and water systems. Although mechanical recycling remains the most economical and sustainable solution for treating uncontaminated PET waste, its limitations in handling contaminated materials highlight the need for alternative solutions. Against this backdrop, chemical recycling technology, by depolymerizing the polymer into monomer components and achieving a circular life cycle, has become a breakthrough solution to address the PET waste crisis. Currently, the depolymerization of PET polymers into monomer components through glycolysis for recycling, thereby reducing environmental pollution, has proven to be a feasible technical approach.
[0003] To improve the efficiency of PET glycolysis, various catalysts have been explored to accelerate the reaction. Lewis acids and bases, due to their wide availability and significant catalytic activity, show promising application prospects. However, the sheer number of Lewis acids and bases, resulting in an overwhelming number of combinations, makes exhaustive testing of every possible combination inefficient and costly—clearly impractical. Therefore, finding efficient ways to explore these combinations to reduce the workload of researchers and improve testing efficiency has become a pressing issue. Summary of the Invention
[0004] In view of the above problems, this application provides a catalyst screening method, apparatus, system, and electronic equipment to improve testing efficiency and reduce workload. The specific solution is as follows:
[0005] The first aspect of this application provides a catalyst screening method, comprising:
[0006] The first number of Lewis acid-base pairs were subjected to PET glycolysis reaction experiments to obtain the first number of yield data;
[0007] The second number of Lewis acid-base pairs were screened based on the Bayesian optimization model to obtain the third number of Lewis acid-base pairs. The third number of Lewis acid-base pairs were then subjected to the PET glycolysis reaction experiment to obtain the third number of yield data. The second number of Lewis acid-base pairs were obtained by expanding the first number of Lewis acid-base pairs.
[0008] The experimental results data and prompt words are input into a large language model, which then uses the experimental results data to predict new Lewis acid-base pairs according to the prediction rules indicated by the prompt words, resulting in a fourth number of Lewis acid-base pairs and their associated hypothesis types and logical bases. The fourth number of Lewis acid-base pairs includes at least zinc neopentanoate and N,N'-diethylethylenediamine. The experimental results data includes the first number of Lewis acid-base pairs and yield data, as well as the third number of Lewis acid-base pairs and yield data.
[0009] In one possible implementation, the cue words include: a first content characterizing the role of the catalyst in the PET glycolysis reaction; a second content characterizing the role of Lewis acids and Lewis bases in the PET glycolysis reaction; a third content characterizing the screening of Lewis acid-base pairs synergistically; and a fourth content characterizing the generated content. The step of inputting experimental result data and cue words into a large language model, enabling the large language model to predict new Lewis acid-base pairs based on the prediction rules indicated by the cue words and referencing the experimental result data, includes:
[0010] The large language model refers to the experimental results data and predicts Lewis acid-base pairs based on the first content and the second content to obtain the first prediction result data;
[0011] The large language model filters the first prediction result data based on the third content to obtain the second prediction result data;
[0012] The large language model filters the second prediction result data based on the fourth content, and outputs the third prediction result data, the hypothesis type, and the logical basis.
[0013] In one possible implementation, the third content includes: a definition of positive synergy, a definition of negative synergy, and a Lewis acid-base pair recommendation based on positive synergy. The large language model filters the first prediction result data according to the third content to obtain second prediction result data, including:
[0014] The large language model selects a fifth number of Lewis acid-base pairs from the first prediction result data based on the definitions of positive and negative synergistic effects.
[0015] The large language model expands the fifth number of Lewis acid-base pairs based on the recommended content of the Lewis acid-base pairs to obtain the second prediction result data.
[0016] In one possible implementation, before filtering the second number of Lewis acid-base pairs based on a Bayesian optimization model to obtain the third number of Lewis acid-base pairs, the following steps are also included:
[0017] Based on the large language model, generate attribute description text for each Lewis acid-base pair;
[0018] Convert each attribute description text into an embedding vector;
[0019] Each of the embedded vectors is transformed into a subspace vector based on the projection matrix.
[0020] In one possible implementation, the process of screening a second number of Lewis acid-base pairs based on a Bayesian optimization model to obtain a third number of Lewis acid-base pairs includes:
[0021] Based on the Bayesian optimization model, the second number of Lewis acid-base pairs are screened in a preset number of rounds, and after each round of screening, the top preset number of Lewis acid-base pairs with the highest collected function values are retained.
[0022] In one possible implementation, the Bayesian optimization model uses a Gaussian process model as a surrogate model, the similarity between Lewis acid-base pairs is modeled by a kernel function, and the upper confidence bound (UCB) is used as the acquisition function. After each screening, the surrogate model is updated by the obtained yield data.
[0023] In one possible implementation, the process of the PET glycolysis reaction experiment includes:
[0024] Use a solid sample grinder to grind PET resin particles into PET powder of a preset size;
[0025] A mobile robot is used to transfer the target weight of PET powder to the reaction platform and react it with a Lewis acid-base pair to obtain a reaction mixture.
[0026] After a portion of the reaction mixture was aspirated using an automated pipetting system and an NMR reference was added, the yield was determined using an NMR spectrometer.
[0027] A second aspect of this application provides a catalyst screening device, comprising:
[0028] The literature data experimental module is used to conduct PET glycolysis reaction experiments on a first number of Lewis acid-base pairs to obtain a first number of yield data.
[0029] The experimental data expansion module is used to screen the second number of Lewis acid-base pairs based on the Bayesian optimization model to obtain the third number of Lewis acid-base pairs, and to conduct the PET glycolysis reaction experiment on the third number of Lewis acid-base pairs to obtain the third number of yield data. The second number of Lewis acid-base pairs are obtained by expanding the first number of Lewis acid-base pairs.
[0030] The catalyst prediction module is used to input experimental result data and prompt words into a large language model. The large language model then uses the experimental result data and the prediction rules indicated by the prompt words to predict new Lewis acid-base pairs, resulting in a fourth number of Lewis acid-base pairs and their associated hypothesis types and logical bases. The fourth number of Lewis acid-base pairs includes at least zinc neopentanoate and N,N'-diethylethylenediamine. The experimental result data includes the first number of Lewis acid-base pairs and yield data, as well as the third number of Lewis acid-base pairs and yield data.
[0031] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the catalyst screening method of the first aspect or any implementation thereof.
[0032] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0033] The memory is used to store computer programs;
[0034] The processor is used to execute the computer program so that the electronic device can implement the catalyst screening method of the first aspect or any implementation thereof.
[0035] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs that, when executed by an electronic device, enable the electronic device to perform the catalyst screening method described in the first aspect or any implementation thereof.
[0036] The sixth aspect of this application provides a catalyst screening system, comprising: the electronic device described in the fourth aspect above, and a PET glycolysis reaction experimental platform connected to the electronic device.
[0037] By employing the above technical solution, the catalyst screening method provided in this application first conducts PET glycolysis reaction experiments on a first number of Lewis acid-base pairs to obtain a first number of yield data. Then, by using a combination of Lewis acids and Lewis bases, the first number of Lewis acid-base pairs is expanded to obtain a second number of Lewis acid-base pairs. Based on a Bayesian optimization model, the second number of Lewis acid-base pairs is screened to obtain a third number of Lewis acid-base pairs. The experimental results and prompt words are then input into a large language model, which, referring to the experimental results and based on the prediction rules indicated by the prompt words, predicts new Lewis acid-base pairs, obtaining a fourth number of Lewis acid-base pairs and their associated hypothesis types and logical basis. This allows researchers to verify the predicted catalysts, accelerating the process of finding effective catalysts, overcoming the limitations of traditional catalyst methods that rely on trial and error and limited literature data, effectively reducing the workload of subsequent researchers and improving testing efficiency. Attached Figure Description
[0038] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0039] Figure 1 A flowchart of a catalyst screening method provided in this application;
[0040] Figure 2 An NMR spectrum provided for this application;
[0041] Figure 3 A flowchart of a Lewis acid-base pair topology provided in this application;
[0042] Figure 4 Another flowchart of a catalyst screening method provided in this application;
[0043] Figure 5 A structural diagram of an automated catalyst screening system provided in this application;
[0044] Figure 6 A structural diagram of a catalyst screening device provided in this application;
[0045] Figure 7 This is a structural diagram of an electronic device provided in this application. Detailed Implementation
[0046] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0047] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0048] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0049] To improve the efficiency of PET glycolysis, researchers have explored various catalytic systems, including homogeneous catalysts such as metal salts, chlorides, and metal alkoxides, as well as organic base catalysts such as urea and ionic liquids. Heterogeneous catalysts such as metal oxides, supported noble metal nanoparticles, and molecular sieves are also being explored. Among these, Lewis acids and bases show promising potential due to their wide availability and significant catalytic activity. While single-component pairings have achieved some success, research on pairing Lewis acids and bases to leverage potential synergistic effects remains largely unexplored. Such synergistic effects hold the promise of significantly reducing the activation energy of PET depolymerization.
[0050] Although there have been a few reports on Lewis acid-base combinations, systematic research remains very limited. This insufficiency is mainly due to the vast number of existing acids and bases, resulting in an excessively large combination space, making it neither practical nor economical to exhaustively test all possible combinations. Furthermore, traditional trial-and-error methods lack the scalability and speed required to handle such a complex search space.
[0051] To address the aforementioned problems, this application provides a catalyst screening method. The catalyst screening method of this application embodiment will be described in detail below with reference to the accompanying drawings.
[0052] Reference Figure 1 , Figure 1 This is a flowchart illustrating a catalyst screening method provided in an embodiment of this application, as shown below. Figure 1 As shown in the embodiments of this application, a catalyst screening method may include steps 101 to 103, which are described in detail below.
[0053] 101. The first number of Lewis acid-base pairs were subjected to PET glycolysis reaction experiments to obtain the first number of yield data.
[0054] Specifically, we systematically collected reported Lewis acid-base catalyst combinations from publicly available literature, including metal salt Lewis acids (such as zinc acetate, zinc trifluoromethanesulfonate, etc.) and nitrogen- or oxygen-containing Lewis bases (such as DMAP, DBU, etc.).
[0055] Under uniform reaction conditions (instrument set temperature 185℃, stirring at 1000 rpm, reaction time 20 minutes, catalyst dosage 0.1 equivalents, ethylene glycol 20 equivalents), verification tests were conducted using an automated experimental platform. Each experiment was repeated three times, and the BHET yield was quantitatively analyzed using 1H NMR spectroscopy, referring to... Figure 2 As shown:
[0056] Using DMF as an internal standard, the calculation was performed using the characteristic peak area ratios of 8.15 ppm and 7.95 ppm. Specifically, this was achieved through the equation... and Calculate the yield, where n BHET and n DMF S represents the amount of substance, respectively. BHET and S DMF Corresponding NMR peak area. Based on known DMF dosage (n) DMF Peak area ratio, molecular weight (=192 g / mol) and mass (m) of PET structural units PET The final yield (W) was calculated using the formula. Considering that the pyridine ring of DMAP may play a key catalytic role, the range of organic bases was further expanded to include more nitrogen heterocyclic compounds (including pyridine, imidazole, quinoline, and acridine derivatives) and the superbase tetramethylguanidine (TMG), and combined with acetate or zinc Lewis acids for testing. The test results showed that the combination of zinc acetate and DMAP performed best, with an average yield of 86%, and this result was established as the performance benchmark for subsequent optimization. By analogy, 78 sets of Lewis acid-base pair data were obtained.
[0057] 102. Based on the Bayesian optimization model, the second number of Lewis acid-base pairs were screened to obtain the third number of Lewis acid-base pairs. The third number of Lewis acid-base pairs were then subjected to PET glycolysis reaction experiments to obtain the yield data for the third number of pairs. The second number of Lewis acid-base pairs were obtained by expanding the first number of Lewis acid-base pairs.
[0058] Specifically, based on the multiple initial benchmark data obtained above, the catalyst chemical space was expanded to 60 Lewis acids (including organic acid salts, inorganic metal salts, etc.) and 186 Lewis bases (covering amines, heterocyclic compounds, etc.), resulting in a total of 11,160 possible combinations. A second set of Lewis acid-base pairs was then screened in a predetermined number of rounds based on this Bayesian optimization model. After each round of screening, the top predetermined number of Lewis acid-base pairs with the highest collected function values were retained.
[0059] For example, in each round of Bayesian optimization, the acquisition function values of all candidate acid-base pairs are calculated, and the six combinations with the highest scores are selected for experimental verification. To ensure fair comparison with the baseline (the Lewis acid-base pair data obtained in step 101), the number of rounds of Bayesian optimization is fixed at 13 (consistent with the total number of baseline experiments), ultimately completing 78 sets of experimental evaluations (including one repeated sampling point), which is comparable in size to the dataset obtained in step 101.
[0060] It is understood that those skilled in the art can adjust the number of the Lewis acid-base pairs as needed, and no restrictions are imposed here.
[0061] 103. Input the experimental results data and prompt words into the large language model, so that the large language model refers to the experimental results data and makes new Lewis acid-base pair predictions according to the prediction rules indicated by the prompt words, to obtain the fourth number of Lewis acid-base pairs and the associated hypothesis type and logical basis. The fourth number of Lewis acid-base pairs includes at least: zinc neopentanoate and N,N'-diethylethylenediamine. The experimental results data include: the first number of Lewis acid-base pairs and yield data and the third number of Lewis acid-base pairs and yield data.
[0062] Specifically, based on the 155 sets of experimental data obtained in the first two steps (one set being duplicate data), a large language model was used to perform multi-dimensional analysis of the dataset and generate chemical hypotheses to guide subsequent catalyst development. All Lewis acid-base pairs and their corresponding BHET yields were integrated into a unified input prompt and submitted to the large language model. The model ultimately output five categories of hypotheses, each including: a description of the potential structure-activity relationship or catalytic mechanism, an explanation of the corresponding chemical principles, and targeted experimental verification suggestions. To ensure the hypotheses are verifiable and actionable, the model was also required to identify specific supporting evidence from existing data and recommend three new acid-base combinations for each hypothesis for testing, ultimately generating 15 novel catalyst combinations. These suggestions include both exploratory candidate systems designed based on new mechanisms and optimized schemes for high-activity regions. The content generated by the large language model is shown in the table below:
[0063]
[0064] As can be seen above, this paradigm shift from Bayesian optimization-guided exploration to large language model-driven mechanism discovery expanded the total number of tested catalysts to 170 and increased the highest observed BHET yield from 86% to 90%. More significantly, this stage successfully positioned the large language model as a hypothesis generation engine. By bridging experimental data with mechanistic reasoning, the model demonstrated the ability to identify potential patterns from massive amounts of data and propose chemically plausible research directions. This means that artificial intelligence can enhance the experimental research process through data-driven insights.
[0065] By organically integrating Bayesian optimization for efficient sampling of known chemical spaces with mechanistic interpretations extracted from large language models, this framework enables researchers to perform inductive reasoning and propose groundbreaking "out-of-sample" catalyst candidates. This collaborative model led to a key discovery: when zinc neopentanoate (rather than the previously optimal zinc acetate) was used as a Lewis acid in combination with N,N'-diethylethylenediamine as a Lewis base, the BHET yield was further increased to 95%. This achievement vividly illustrates the synergistic effect of human-machine collaboration—machine-generated knowledge not only accelerates the exploration of known spaces but also empowers researchers to break through traditional cognitive boundaries.
[0066] The large language model also generates interpretable and verifiable hypotheses by analyzing the continuously accumulating experimental data. These hypotheses not only guided multiple sets of experimental verifications (with some catalyst systems outperforming the original baseline), but more importantly, they initiated the inductive reasoning process, propelling us to explore regions beyond the original chemical space. Ultimately, a high-performance catalyst composed of zinc neopentanoate and N,N'-diethylethylenediamine was discovered, achieving a BHET yield of 95%. In addition, 18 catalyst combinations with yields exceeding 80% were identified, such as anhydrous zinc acetate + 4-dimethylaminopyridine, anhydrous zinc acetate + 2,2'-bipyridine, anhydrous zinc acetate + 1-propylimidazolium, anhydrous zinc acetate + 8-hydroxyquinoline, zinc acetate dihydrate + N,N'-dimethyl-1,3-propanediamine, copper acetate + tetramethylguanidine, zinc acetate dihydrate + 4-acetylpyridine, zinc neopentanoate + 4-dimethylaminopyridine, and anhydrous zinc acetate + N,N-dimethyldipropylenetriamine, etc. By combining automated experiments, artificial intelligence, and collaborative efforts with human experts, this framework transcends traditional data-driven optimization, enabling true chemical discovery. It not only utilizes experimental data but also integrates coded chemical knowledge and human intuition, establishing a scalable and interpretable paradigm for intelligent, autonomous research. Its innovation lies both in its algorithms and its deep roots in domain-specific knowledge. This allows researchers to validate predicted catalysts, accelerating the search for effective catalysts and overcoming the limitations of traditional catalyst research that relies on trial and error and limited literature data. This effectively reduces the workload for subsequent researchers and improves testing efficiency.
[0067] As a feasible implementation of the above embodiments, the prompt words include: a first content characterizing the role of the catalyst in the PET glycolysis reaction; a second content characterizing the role of Lewis acids and Lewis bases in the PET glycolysis reaction; a third content characterizing the screening of Lewis acid-base pairs; and a fourth content characterizing the generated content. The experimental results data and the prompt words are input into a large language model, which then uses the experimental results data to predict new Lewis acid-base pairs according to the prediction rules indicated by the prompt words, including:
[0068] The large language model references the experimental results data and makes predictions about Lewis acid-base pairs based on the first and second content, thus obtaining the first prediction result data.
[0069] The large language model filters the first prediction result data based on the third content to obtain the second prediction result data.
[0070] The large language model filters the second prediction result data based on the fourth content, and outputs the third prediction result data, as well as the hypothesis type and logical basis.
[0071] The third part includes: definitions of positive synergy, definitions of negative synergy, and Lewis acid-base pair recommendations based on positive synergy. The large language model filters the first prediction result data based on the third part to obtain the second prediction result data, which includes:
[0072] The large language model defines the content based on positive and negative synergistic effects, and selects the fifth number of Lewis acid-base pairs from the first prediction result data.
[0073] The large language model expands the fifth number of Lewis acid-base pairs based on the recommended content of Lewis acid-base pairs to obtain the second prediction result data.
[0074] Specifically, the text content of the prompt words may include the following:
[0075] First content:
[0076] Polyethylene terephthalate (PET) is a widely used thermoplastic known for its strength, stability, and safety; however, its ever-increasing demand presents significant environmental challenges and relies on non-renewable resources. The efficient recycling of PET via ethylene glycol (EG) depolymerization to produce bis(hydroxyethyl) terephthalate (BHET) offers a solution for sustainable development. This process, operating under mild conditions, has been commercialized, with the catalyst playing a crucial role in enhancing reaction efficiency. This study aims to develop high-performance Lewis acid-base catalysts using machine learning and to establish an automated platform for plastic degradation. Several acid-base pairs have been tested. All acids and bases and their corresponding catalytic yields are listed in the table below:
[0077]
[0078] Second content:
[0079] Here is some additional chemical information:
[0080] 1. The role of Lewis acids in the catalytic degradation of plastics
[0081] 1.1 Effects on anions
[0082] The optimal effect is achieved in metal salts with moderate basicity to counteract the anion. The anion plays two key roles: coordinating with the metal ion and acting as a base. The binding constant between the anion and the metal ion should be maintained at a moderate level to balance these two effects.
[0083] (1) It reduces the binding between anion and metal ion, which is beneficial for the coordination of metal ion with ester carbonyl group. This coordination will activate carbonyl group and enhance the nucleophilic attack of oxygen atom of alcohol hydroxyl group.
[0084] (2) It exhibits strong anionic basicity, thereby accelerating the deprotonation of alcohol hydroxyl groups.
[0085] The binding ability of metal ions to counter anions and the basicity of anions are inherently conflicting; therefore, an optimal middle ground must be determined.
[0086] 1.2 The Influence of Metal Ions
[0087] The coordination of metal ions with the oxygen atom of the carbonyl group promotes the degradation process through two key mechanisms:
[0088] (1) Increase the solubility of the polymer in ethylene glycol, thereby promoting the reaction. (2) Activate the carbonyl group by enhancing the electrophilicity of the carbonyl carbon, making it more susceptible to nucleophilic attack by ethylene glycol.
[0089] Current experimental results indicate that zinc ions have the most significant activation effect on carbonyl groups.
[0090] Third content:
[0091] Now we need to consider the synergistic effect of Lewis acid-base pairs in the catalytic degradation of plastics.
[0092] Definition of positive synergistic effect: Assuming the degradation yield of Lewis acid alone is x and the yield of base alone is y, if the combined yield is significantly higher than the larger of x and y, then the pair is considered to exhibit a positive synergistic effect.
[0093] The definition of negative synergy is as follows: Conversely, if the combined yield is significantly lower than the smaller of x and y, then they are considered to exhibit negative synergy.
[0094] Lewis acid-base pair recommendations: Our goal is to identify the common properties of acids or bases that exhibit positive concerted effects. For example, if acid A and base B exhibit a positive concerted effect, our goal is to determine the characteristics that acid C might share with acid A, thereby predicting that it might also exhibit a positive concerted effect with base B. By utilizing these characteristics, we can systematically recommend Lewis acid-base pairs with positive concerted effects.
[0095] Fourthly, please analyze the data based on the Lewis acid-base theory described above and generate five hypotheses regarding preferences and data trends for further research. Please also provide your reasoning and the data points that support these hypotheses.
[0096] Based on the above prompts, the large language model uses existing experimental data to make predictions, thereby obtaining the above results. It is understood that those skilled in the art can adjust the content of the above prompts as needed, which will not be elaborated here.
[0097] In one embodiment, to improve the execution efficiency of Bayesian optimization, before screening the second number of Lewis acid-base pairs based on the Bayesian optimization model to obtain the third number of Lewis acid-base pairs, the following steps are also included:
[0098] Generate attribute description text for each Lewis acid-base pair based on a large language model.
[0099] Convert the text describing each attribute into an embedding vector.
[0100] Each embedding vector is transformed into a subspace vector based on the projection matrix.
[0101] Specifically, refer to Figure 3As shown, we use an enhanced Bayesian optimization (e-BO) model inspired by embedded subspaces. In the embedding generation stage, we input standardized prompts into a large language model such as GPT-4o: "Please describe the chemical properties and characteristics of {molecule name}", to obtain a structured text description containing molecular formula, structural features, physicochemical properties, chemical characteristics and typical applications. Then, we use the OpenAI API to convert these texts into 1536-dimensional high-dimensional embedding vectors. Since molecules with similar chemical properties will produce similar embedding vectors, these vector representations can be used as descriptors with chemical meaning in the BO process. (3) To reduce dimensionality, we use random embedding technology: using the projection matrix generated by Gaussian sampling, each 1536-dimensional vector is projected into 20-dimensional subspace vectors of acid components and base components, and then concatenated into a 40-dimensional feature vector. This strategy significantly reduces computational complexity while preserving semantic similarity.
[0102] The Bayesian optimization model uses a Gaussian process model as a surrogate model, the similarity between Lewis acid-base pairs is modeled by a kernel function, and the upper confidence bound (UCB) is used as the acquisition function. After each screening, the surrogate model is updated by the obtained yield data.
[0103] Specifically, in each round of Bayesian optimization, the acquisition function values for all candidate acid-base pairs are calculated, and the six highest-scoring combinations are selected for experimental verification. The optimization process needs to consider both exploration strategies—selecting unexplored regions—and development strategies—selecting high-yield regions already explored. The significance of the acquisition function is to consider both strategies simultaneously during sampling, avoiding falling into local optima or overly random global exploration. A Gaussian process model is used as a surrogate model to predict experimental results based on historical data. The similarity between data points is modeled using the Matérn kernel function, and an upper confidence bound (UCB) is used as the acquisition strategy, defined as: ,in Let μ represent the posterior mean of the predicted values, σ represent the posterior variance of the uncertainty, and β be the hyperparameter controlling the exploration-exploitation tradeoff. After each round of experiments, the Gaussian process model is updated with the newly acquired yield data to improve the accuracy of subsequent predictions. To ensure fair comparison with the baseline (Phase 0), the number of BO rounds is fixed at 13 (consistent with the total number of baseline experiments), ultimately completing the evaluation of 78 sets of experiments (including one repeated sampling point), which is comparable in size to the Phase 0 dataset.
[0104] In other embodiments, the PET glycolysis reaction experiment process includes:
[0105] Use a solid sample grinder to grind PET resin particles into PET powder of a preset size.
[0106] A mobile robot is used to transfer the target weight of PET powder to the reaction platform and react it with a Lewis acid-base pair to obtain a reaction mixture.
[0107] After a portion of the reaction mixture was aspirated using an automated pipetting system and an NMR reference was added, the yield was determined using an NMR spectrometer.
[0108] Specifically, the PET glycolysis reaction experimental platform mainly consists of the following modules: (1) Solid processing unit: equipped with a grinder and vibrating screen to crush PET particles to 60-120 mesh; (2) Reaction workstation: including a reactor, solid weighing station and pipetting device; (3) Liquid processing system: using a high-precision liquid workstation to realize automatic distribution of liquid after reaction; (4) Analysis and detection module: an automatic sample introduction system coupled with 400MHz NMR, equipped with vortex mixing and sample dilution functions. Each module realizes material transfer through a mobile robot, and the entire system is coordinated by central control software.
[0109] For example, you can refer to Figure 5 As shown, the entire automatic catalyst screening system, under the coordinated control of a computer (a central system that can run and control various parts of the experimental platform), controls the movement of each robot module between the various components of the experimental platform and performs corresponding experimental actions to carry out the above-mentioned PET glycolysis reaction experiment. It also automatically performs each step of the catalyst screening method in the above embodiment, realizing full automation of the entire catalyst screening process and effectively improving screening efficiency.
[0110] As a specific application of the above catalyst screening method, refer to Figure 4 As shown, the specific processing steps may include the following:
[0111] All experiments were conducted as follows. Commercially available PET resin granules were first pulverized using an XA-1 solid sample grinder and sieved to obtain a 60-120 mesh powder. All reactions were carried out in 8 mL sample vials with pre-installed magnetic stirrers. The reaction rack contained reagent vials containing Lewis acid-base catalysts, ethylene glycol, and DMF. 0.46 g of PET powder was accurately weighed using a Mettler balance and placed into the reaction vials, which were then transferred to the ChemRoboX platform by a mobile robot. The platform automatically added 0.1 equivalents of catalyst (approximately 0.2 mmol of anhydrous zinc acetate and DMAP in a 1:1 molar ratio) and 20 equivalents of ethylene glycol (approximately 2.7 mL). Each experiment was repeated three times to ensure statistical reliability. The reaction conditions were set at 185 °C and 1000 rpm for 20 minutes. After the reaction, the system was cooled to 120 °C, and 5 equivalents of DMF (approximately 0.92 mL) were added as an internal standard and mixed for 10 minutes. A mobile robot transfers the reaction mixture to an automated pipetting system, and 100 μL of sample is aspirated for NMR analysis. The entire NMR sample preparation process is automated: 1 mL of deuterated DMSO is added and vortexed until completely dissolved to visual inspection. Then, 0.5 mL of the solution is transferred into an NMR tube and detected using a 400 MHz NMR spectrometer.
[0112] Phase 0: First, the baseline performance was established using the literature data according to the above method, and a total of 78 data points containing catalysts and yields were obtained. Based on the initial baseline data, the chemical space of the catalyst was expanded to 60 Lewis acids (including organic acid salts, inorganic metal salts, etc.) and 186 Lewis bases (covering amines, heterocyclic compounds, etc.), forming a total of 11,160 possible combinations. We used the developed embedded subspace-inspired enhanced Bayesian optimization (e-BO) framework to characterize these catalysts: (1) In the embedding generation step, we input the standardized prompt words into GPT-4o: "Please describe the chemical characteristics and properties of {anhydrous zinc acetate}{2,2'-bipyridine}", and obtained a structured text description containing molecular formula, structural features, physicochemical properties, chemical characteristics and typical applications. (2) These texts were converted into 1536-dimensional high-dimensional embedding vectors through the OpenAI API. (3) Using the projection matrix generated by Gaussian sampling, the 1536-dimensional vector was projected into a 20-dimensional subspace vector, and then the acid-base combination was spliced into a 40-dimensional feature vector.
[0113] Phase 1: In each round of Bayesian optimization, the acquisition function values for acid and base were calculated, and the six highest-scoring combinations were selected for experimental verification. The first round selected vanadium chloride + dibutylamine, vanadium chloride + N-methylbutylamine, vanadium chloride + diisobutylamine, sodium trifluoromethanesulfonate + N-methyl-n-propylamine, zinc trifluoromethanesulfonate + N-methyl-n-propylamine, and magnesium trifluoromethanesulfonate + N-methyl-n-propylamine. After each round of experiments, the Gaussian process model was updated with the newly acquired yield data to improve subsequent prediction accuracy. A total of 78 experimental sets were evaluated, comparable in size to the Phase 0 dataset. Three catalyst combinations produced yields greater than 80%.
[0114] Phase 2: Based on the 155 sets of experimental data obtained in the first two phases, all Lewis acid-base pairs and their corresponding BHET yields were integrated into a unified input prompt and submitted to the large language model. The model ultimately output five hypotheses and specific supporting evidence, and recommended three new acid-base combinations for each hypothesis for testing—ultimately producing 15 new catalyst combinations, namely: copper acetate + tetramethylguanidine, cobalt acetate + 4-methylmorpholine, calcium acetate + 4-methylmorpholine, anhydrous zinc acetate + 1,8-diazabicyclo[5.4.0] eleven The catalyst combinations were: C7-ene, zinc acetate dihydrate + N,N'-dimethyl-1,3-propanediamine, magnesium acetate + N,N,N,N-tetramethylethylenediamine, copper acetate + 2,6-diisopropylpyridine, magnesium acetate + N-tert-butylisopropyl, copper acetate + N,N-diisopropylethylamine, magnesium trifluoromethanesulfonate + 2,4-dimethylpyrrole, calcium trifluoromethanesulfonate + triethylenediamine, zinc trifluoromethanesulfonate + 4-methylmorpholine, anhydrous zinc acetate + 8-hydroxyquinoline, zinc acetate dihydrate + 4-acetylpyridine, and zinc acetate dihydrate + 4-pyrrolidinylpyridine. Six catalyst combinations yielded yields greater than 80%. The combination of anhydrous zinc acetate + 8-hydroxyquinoline achieved a yield of 90%.
[0115] Phase 3: In the human-AI collaborative design phase, by organically integrating the efficient sampling capabilities of Bayesian optimization for known chemical spaces and the mechanistic interpretation extracted by large language models, this framework enables researchers to perform inductive reasoning and propose groundbreaking "out-of-sample" candidate catalysts. This collaborative model led to a key discovery: when zinc neopentanoate was combined as a Lewis acid with N,N'-diethylethylenediamine as a Lewis base, the BHET yield was further increased to 95%. Simultaneously, a total of 7 catalyst combinations produced yields greater than 80%, and 3 catalyst combinations produced yields greater than 90%. This achievement vividly illustrates the synergistic effect of human-machine collaboration—machine-generated knowledge not only accelerates the exploration of known spaces but also empowers researchers to break through traditional cognitive boundaries.
[0116] The above describes a catalyst screening method provided by the embodiments of this application. The following describes the apparatus for performing the above catalyst screening method.
[0117] Please see Figure 6 , Figure 6 This is a schematic diagram of a catalyst screening device provided in an embodiment of this application. Figure 6 As shown, the catalyst screening device includes:
[0118] The literature data experimental module 601 is used to conduct a PET glycolysis reaction experiment on a first number of Lewis acid-base pairs to obtain a first number of yield data.
[0119] The experimental data expansion module 602 is used to screen the second number of Lewis acid-base pairs based on the Bayesian optimization model to obtain the third number of Lewis acid-base pairs, and to conduct PET glycolysis reaction experiments on the third number of Lewis acid-base pairs to obtain the third number of yield data. The second number of Lewis acid-base pairs are obtained by expanding the first number of Lewis acid-base pairs.
[0120] The catalyst prediction module 603 is used to input experimental result data and prompt words into the large language model, so that the large language model refers to the experimental result data and makes new Lewis acid-base pair predictions according to the prediction rules indicated by the prompt words, to obtain a fourth number of Lewis acid-base pairs and the associated hypothesis type and logical basis. The fourth number of Lewis acid-base pairs includes at least: zinc neopentanoate and N,N'-diethylethylenediamine. The experimental result data includes: a first number of Lewis acid-base pairs and yield data and a third number of Lewis acid-base pairs and yield data.
[0121] In one possible implementation, the cue words include: a first element characterizing the role of the catalyst in the PET glycolysis reaction; a second element characterizing the role of Lewis acids and Lewis bases in the PET glycolysis reaction; a third element characterizing the screening of Lewis acid-base pair synergies; and a fourth element characterizing the generated content. The catalyst prediction module 603 inputs the experimental results data and the cue words into the large language model, enabling the large language model to predict new Lewis acid-base pairs based on the prediction rules indicated by the cue words, referencing the experimental results data. This process includes:
[0122] The large language model refers to the experimental results data and makes predictions on Lewis acid-base pairs based on the first and second contents to obtain the first prediction result data;
[0123] The large language model filters the first prediction result data based on the third content to obtain the second prediction result data;
[0124] The large language model filters the second prediction result data based on the fourth content, and outputs the third prediction result data, as well as the hypothesis type and logical basis.
[0125] In one possible implementation, the third content includes: definitions of positive synergistic effects, definitions of negative synergistic effects, and Lewis acid-base pair recommendations based on positive synergistic effects. The process by which the catalyst prediction module 603 large language model filters the first prediction result data based on the third content to obtain the second prediction result data includes:
[0126] The large language model defines content based on positive and negative synergistic effects, and selects the fifth number of Lewis acid-base pairs from the first prediction result data;
[0127] The large language model expands the fifth number of Lewis acid-base pairs based on the recommended content of Lewis acid-base pairs to obtain the second prediction result data.
[0128] In one possible implementation, before the experimental data augmentation module 602 filters the second number of Lewis acid-base pairs based on the Bayesian optimization model to obtain the third number of Lewis acid-base pairs, it is also used for:
[0129] Generate attribute description text for each Lewis acid-base pair based on a large language model;
[0130] Convert the text describing each attribute into an embedding vector;
[0131] Each embedding vector is transformed into a subspace vector based on the projection matrix.
[0132] In one possible implementation, the experimental data augmentation module 602, based on a Bayesian optimization model, filters the second number of Lewis acid-base pairs to obtain the third number of Lewis acid-base pairs, including:
[0133] The second number of Lewis acid-base pairs are screened in a preset number of rounds based on a Bayesian optimization model, and after each round of screening, the top preset number of Lewis acid-base pairs with the highest collected function values are retained.
[0134] In one possible implementation, the Bayesian optimization model in the experimental data augmentation module 602 uses a Gaussian process model as a surrogate model, the similarity between Lewis acid-base pairs is modeled by a kernel function, and the upper confidence bound (UCB) is used as the acquisition function. After each screening, the surrogate model is updated by the obtained yield data.
[0135] In one possible implementation, the process of the PET glycolysis reaction experiment in the literature data experiment module 601 and the experimental data expansion module 602 includes:
[0136] Use a solid sample grinder to grind PET resin particles into PET powder of a preset size;
[0137] A mobile robot is used to transfer the target weight of PET powder to the reaction platform and react it with a Lewis acid-base pair to obtain a reaction mixture.
[0138] After a portion of the reaction mixture was aspirated using an automated pipetting system and an NMR reference was added, the yield was determined using an NMR spectrometer.
[0139] This application also provides an electronic device in its embodiments. (See reference...) Figure 7 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as laptops, desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0140] like Figure 7 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. When the electronic device is powered on, the RAM 703 also stores various programs and data required for the operation of the electronic device. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0141] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, memory cards, hard drives, etc.; and communication devices 709. Communication device 709 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.
[0142] This application also provides a catalyst screening system, including: the electronic device described in the above embodiments and a PET glycolysis reaction experimental platform connected to the electronic device.
[0143] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the catalyst screening methods provided in this application.
[0144] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the catalyst screening methods provided in this application.
[0145] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0147] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0148] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A catalyst screening method, characterized in that, include: The first number of Lewis acid-base pairs were subjected to PET glycolysis reaction experiments to obtain the first number of yield data; The second number of Lewis acid-base pairs were screened based on the Bayesian optimization model to obtain the third number of Lewis acid-base pairs. The third number of Lewis acid-base pairs were then subjected to the PET glycolysis reaction experiment to obtain the third number of yield data. The second number of Lewis acid-base pairs were obtained by expanding the first number of Lewis acid-base pairs. The experimental results data and prompt words are input into a large language model, which then uses the experimental results data to predict new Lewis acid-base pairs according to the prediction rules indicated by the prompt words, resulting in a fourth number of Lewis acid-base pairs and their associated hypothesis types and logical bases. The fourth number of Lewis acid-base pairs includes at least zinc neopentanoate and N,N'-diethylethylenediamine. The experimental results data includes the first number of Lewis acid-base pairs and yield data, as well as the third number of Lewis acid-base pairs and yield data.
2. The catalyst screening method according to claim 1, characterized in that, The cue words include: a first content characterizing the role of the catalyst in the PET glycolysis reaction; a second content characterizing the role of Lewis acids and Lewis bases in the PET glycolysis reaction; a third content characterizing the screening of Lewis acid-base pairs synergistically; and a fourth content characterizing the generated content. The experimental results data and cue words are input into a large language model, which then uses the experimental results data to predict new Lewis acid-base pairs according to the prediction rules indicated by the cue words. This includes: The large language model refers to the experimental results data and predicts Lewis acid-base pairs based on the first content and the second content to obtain the first prediction result data; The large language model filters the first prediction result data based on the third content to obtain the second prediction result data; The large language model filters the second prediction result data based on the fourth content, and outputs the third prediction result data, the hypothesis type, and the logical basis.
3. The catalyst screening method according to claim 2, characterized in that, The third content includes: definitions of positive synergy, definitions of negative synergy, and Lewis acid-base pair recommendations based on positive synergy. The large language model filters the first prediction result data based on the third content to obtain second prediction result data, including: The large language model selects a fifth number of Lewis acid-base pairs from the first prediction result data based on the definitions of positive and negative synergistic effects. The large language model expands the fifth number of Lewis acid-base pairs based on the recommended content of the Lewis acid-base pairs to obtain the second prediction result data.
4. The catalyst screening method according to claim 1, characterized in that, Before selecting the third number of Lewis acid-base pairs based on a Bayesian optimization model, the process includes: Based on the large language model, generate attribute description text for each Lewis acid-base pair; Convert each attribute description text into an embedding vector; Each of the embedded vectors is transformed into a subspace vector based on the projection matrix.
5. The catalyst screening method according to claim 1 or 4, characterized in that, The second number of Lewis acid-base pairs is screened using a Bayesian optimization model to obtain a third number of Lewis acid-base pairs, including: Based on the Bayesian optimization model, the second number of Lewis acid-base pairs are screened in a preset number of rounds, and after each round of screening, the top preset number of Lewis acid-base pairs with the highest collected function values are retained.
6. The catalyst screening method according to claim 5, characterized in that, The Bayesian optimization model uses a Gaussian process model as a surrogate model, models the similarity between Lewis acid-base pairs using a kernel function, and uses the upper confidence bound (UCB) as the acquisition function. After each screening, the surrogate model is updated with the obtained yield data.
7. The catalyst screening method according to claim 1, characterized in that, The process of the PET glycolysis reaction experiment includes: Use a solid sample grinder to grind PET resin particles into PET powder of a preset size; A mobile robot is used to transfer the target weight of PET powder to the reaction platform and react it with a Lewis acid-base pair to obtain a reaction mixture. After a portion of the reaction mixture was aspirated using an automated pipetting system and an NMR reference was added, the yield was determined using an NMR spectrometer.
8. A catalyst screening device, characterized in that, include: The literature data experimental module is used to conduct PET glycolysis reaction experiments on a first number of Lewis acid-base pairs to obtain a first number of yield data. The experimental data expansion module is used to screen the second number of Lewis acid-base pairs based on the Bayesian optimization model to obtain the third number of Lewis acid-base pairs, and to conduct the PET glycolysis reaction experiment on the third number of Lewis acid-base pairs to obtain the third number of yield data. The second number of Lewis acid-base pairs are obtained by expanding the first number of Lewis acid-base pairs. The catalyst prediction module is used to input experimental result data and prompt words into a large language model, so that the large language model refers to the experimental result data and makes new Lewis acid-base pair predictions according to the prediction rules indicated by the prompt words, to obtain a fourth number of Lewis acid-base pairs and their associated hypothesis types and logical basis. The fourth number of Lewis acid-base pairs includes at least: zinc neopentanoate and N,N'-diethylethylenediamine. The experimental result data includes: the first number of Lewis acid-base pairs and yield data, and the third number of Lewis acid-base pairs and yield data.
9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the catalyst screening method as described in any one of claims 1 to 7.
10. A catalyst screening system, characterized in that, include: The electronic device as described in claim 9 and the PET glycolysis reaction experimental platform connected to the electronic device.