Catalyst parameter determination method, apparatus, device, and storage medium
By acquiring the experimental parameters and rule set of the catalyst and using an intelligent agent to recommend parameters, the problem of rapidly determining the parameters of high-quality catalysts is solved, and the discovery of high-performance catalysts is achieved efficiently.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-05-29
AI Technical Summary
How to quickly determine the optimal combination of catalyst parameters from multiple parameter combinations in order to improve the discovery efficiency of catalytic performance.
By acquiring the experimental parameter set and rule set of the catalyst, and comprehensively utilizing knowledge mining and intelligent agents to recommend parameters, target experimental parameters are selected to prepare catalysts with high catalytic performance.
This improved the efficiency of catalyst discovery, ensured that the recommended parameters were based on both practical evidence and theoretical feasibility, and enhanced the efficiency of catalytic performance discovery.
Smart Images

Figure CN122117147A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chemical technology, and in particular to a method, apparatus, device, and storage medium for determining catalyst parameters. Background Technology
[0002] Catalysts play an indispensable role in the chemical industry. For the same material, different preparation parameters result in different catalytic performances. Therefore, it is of great significance to quickly determine the optimal parameter combination from a large number of parameter combinations for the efficient discovery of catalysts with high catalytic performance. Summary of the Invention
[0003] The main technical problem addressed by this application is to provide a method, apparatus, device, and storage medium for determining catalyst parameters, which can improve the discovery efficiency of catalysts with high catalytic performance.
[0004] To address the aforementioned technical problems, this application provides a method for determining catalyst parameters. This method includes: acquiring an experimental parameter set and a rule set for a first catalyst; the experimental parameter set includes multiple sets of original experimental data obtained from preparing the first catalyst, and each rule in the rule set is obtained through knowledge mining; combining the experimental parameter set and the rule set to recommend parameters, resulting in several sets of recommended preparation parameters; selecting at least one set of target experimental parameters from the several sets of recommended preparation parameters; wherein each set of target experimental parameters is used to prepare a second catalyst.
[0005] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a catalyst parameter determination device, which includes: an acquisition module, a parameter recommendation module, and a parameter determination module; the acquisition module is used to acquire an experimental parameter set and a rule set for a first catalyst; the experimental parameter set includes multiple sets of original experimental data obtained from the preparation of the first catalyst, and each rule in the rule set is obtained through knowledge mining; the parameter recommendation module is used to comprehensively analyze the experimental parameter set and the rule set to recommend parameters, obtaining several sets of recommended preparation parameters; the parameter determination module is used to select at least one set of target experimental parameters from the several sets of recommended preparation parameters; wherein, each set of target experimental parameters is used to prepare a second catalyst.
[0006] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide an electronic device, including a memory and a processor coupled to each other, wherein the memory stores program instructions; and the processor is used to execute the program instructions stored in the memory to implement the above-mentioned method.
[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium for storing program instructions that can be executed to implement the above-mentioned method.
[0008] The above-described scheme recommends preparation parameters by integrating the original experimental data obtained from the preparation of the first catalyst with relevant knowledge rules. In this process, parameters conforming to the knowledge rules can be recommended as preparation parameters, thereby obtaining the target experimental parameters for the preparation of the second catalyst. Compared to methods that only utilize the original experimental data obtained from the preparation process to determine parameters, this application's approach, which introduces knowledge rules, ensures that the recommended preparation parameters are not only based on practical experience but also theoretically feasible, thus improving the discovery efficiency of catalysts with high catalytic performance. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating an embodiment of the catalyst parameter determination method provided in this application; Figure 2 yes Figure 1 The flowchart of step S13 shown is a schematic diagram of one embodiment. Figure 3 This is a flowchart illustrating an embodiment of the first rule provided in this application; Figure 4 This is a schematic diagram of a framework of an embodiment of the catalyst parameter determination device provided in this application; Figure 5 This is a schematic diagram of the framework of an embodiment of the electronic device provided in this application; Figure 6 This is a schematic diagram of the framework of the computer-readable storage medium provided in this application. Detailed Implementation
[0010] To make the purpose, technical solution and effects of this application clearer and more explicit, the following describes this application in further detail with reference to the accompanying drawings and embodiments.
[0011] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0012] It should be noted that the catalyst parameter determination method provided in this application is applicable to catalysts of any material, such as covalent organic frameworks (COFs), metal oxides, metal-organic frameworks, etc.
[0013] Different preparation parameters for catalyst materials result in different structures and catalytic performances. Taking covalent organic frameworks (COFs) with photocatalytic hydrogen evolution performance as an example, by carefully selecting building blocks, connection methods, and synthesis routes, the band structure, photogenerated carrier separation efficiency, and surface catalytic activity of COFs can be precisely controlled, thus significantly affecting their photocatalytic hydrogen evolution performance. Since the performance of COFs is determined by multiple dimensions of parameters, including precursor selection, connection methods, synthesis methods, and post-modification strategies, and these parameters have complex nonlinear interactions, a huge synthetic design space is formed. Determining the preparation parameters for preparing high-performance catalysts from this vast synthetic design space is extremely difficult. Therefore, it is of great significance to quickly determine the optimal preparation parameters for catalysts.
[0014] To address this problem, this application proposes the following method and related apparatus for determining catalyst parameters: Please see Figure 1 , Figure 1 This is a schematic flowchart of an embodiment of the catalyst parameter determination method provided in this application. It should be noted that if substantially the same result is obtained, this embodiment is not necessarily identical. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, this embodiment includes: S11: Obtain the experimental parameter set and rule set for the first catalyst; the experimental parameter set includes multiple sets of original experimental data obtained from the preparation of the first catalyst, and each rule in the rule set is obtained through knowledge mining.
[0015] This embodiment is used to determine the parameters of the second catalyst by utilizing the original experimental data obtained from the preparation of the first catalyst and the rules obtained from knowledge mining; wherein the materials of the first catalyst and the second catalyst can be the same or different; the determined parameters of the second catalyst can be the optimized parameters of the first catalyst or the preparation parameters of a completely new catalyst.
[0016] Each rule in the rule set includes at least one of the first rule, the second rule, and the third rule.
[0017] The first rule was obtained through knowledge mining based on the original experimental data of each group (see below for specific acquisition steps). Figure 3 (Description of the illustrated embodiment), the second rule is a catalyst design rule obtained through knowledge mining based on general chemical knowledge; the third rule is based on... The rules were derived from knowledge mining of the key retrieval data for the first catalyst. To ensure the dynamic evolution of these rules, the latest research abstracts retrieved regarding the first catalyst were used as key retrieval data.
[0018] In one implementation, knowledge mining can be performed using relevant intelligent agents to obtain corresponding knowledge rules. For example, agent 1 can be used to mine the first rule, agent 2 can be used to mine the second rule, and agent 3 can be used to mine the third rule.
[0019] It should be noted that the catalyst parameter determination method provided in this embodiment is essentially a method for determining new parameters of the first catalyst or discovering new catalyst materials different from the first catalyst by using the original experimental data (including existing preparation data of the first catalyst) and knowledge rules about the first catalyst obtained from the preparation of the first catalyst (i.e., preparing a new catalyst material using the new parameters).
[0020] In one embodiment, multiple sets of original experimental parameters of the first catalyst can be directly used as multiple sets of original experimental data related to the first catalyst. These multiple sets of original experimental parameters are obtained from relevant literature / patents and / or from experimental data obtained during the preparation process; for example, by searching based on keywords to obtain relevant literature / patents, and then using relevant models (e.g., LLM) or tools to extract relevant experimental parameters from the relevant literature / patents.
[0021] In another embodiment, the acquired original experimental parameters can be further processed by mechanism annotation to obtain mechanism annotation data. Then, the original experimental parameters and the corresponding mechanism annotation data are combined to obtain combined data as the original experimental related data. That is, in this embodiment, the original experimental related data includes: original experimental parameters and corresponding mechanism annotation data.
[0022] In some embodiments, to obtain high-quality raw experimental parameters, after extracting the experimental parameters using relevant models, each extracted experimental parameter can be quality checked and revised to obtain high-quality raw experimental parameters. This quality checking and revision can be performed manually or by machine.
[0023] The original experimental parameters for each group include: the preparation parameters of the first catalyst, and the performance-related parameters of the first catalyst under the preparation parameters.
[0024] The preparation parameters include at least one of the following: synthesis parameters of the first catalyst, structural characterization parameters, and reaction condition parameters / reaction parameters. Specifically, synthesis parameters include parameters related to the preparation process, such as monomer type, solvent selection, synthesis temperature, and reaction time; structural characterization parameters include parameters reflecting the microstructure of the material, such as specific surface area, pore size distribution, and crystallinity; reaction condition parameters are the external environment and operating parameters of the catalyst during actual application / performance testing; taking photocatalytic hydrogen evolution as an example, reaction condition parameters include light source parameters such as light source type, light intensity, and light wavelength range, as well as parameters such as reaction liquid volume, sacrificial agent type and concentration; performance-related parameters include parameters directly related to catalyst performance, such as the hydrogen evolution reaction (HER) rate.
[0025] For example, the preparation parameters of the first catalyst include: synthesis parameters, structural characterization parameters, and reaction condition parameters. Its mechanism annotation data includes: mechanism annotations for the selection of synthesis parameters, analysis data on the influence of structural characterization parameters on catalytic performance, and mechanism annotations for reaction parameters.
[0026] In one specific embodiment, a relevant intelligent agent can be used to obtain mechanistic annotation data of the original experimental parameters based on a relevant model (e.g., the Gemini-2.5-pro model) and convert it into structured data. Taking the first catalyst as a covalent organic framework material as an example, the converted structured data is organized into three modules: Monomer and Synthesis Module: Records the molecular structure, synthesis route, and mechanistic annotation of the parameter selection of the chemical substance of the first catalyst; Structure Characterization Module: Records structural parameters and analysis of their impact on photocatalytic performance; Photocatalytic Hydrogen Evolution Module: Records reaction parameters and provides an explanation of their mechanism of action.
[0027] S12: Combine the experimental parameter set and rule set to recommend parameters and obtain several sets of recommended preparation parameters.
[0028] Among them, the recommended preparation parameters for each group are the preferred experimental parameters generated in order to obtain highly active catalysts, that is, the recommended preparation parameters for each group are highly active experimental parameters.
[0029] Furthermore, the recommended preparation parameters for each group should conform to the rules in the rule set.
[0030] In one embodiment, the intelligent agent 4 can be used to combine the experimental parameter set and rule set to recommend parameters and obtain several sets of recommended preparation parameters.
[0031] It should be noted that the experimental parameter set includes the existing historical experimental parameters (original experimental parameters) obtained from the preparation of the first catalyst. It contains multiple complete sets of "synthesis parameters (such as monomer SMILES structure, hydrothermal temperature), structural characterization parameters (such as specific surface area, pore size distribution, crystallinity), reaction condition parameters (such as electrolyte concentration, reaction temperature), and performance labels (such as HER rate, stability)", which serve as the "practical data basis" for the recommended experimental parameters.
[0032] The rule set includes at least one of the first, second, and third rules mentioned above; wherein, the first rule is an empirical rule mined from the original experimental parameters, the second rule is a theoretical rule derived from general chemical knowledge, and the third rule is a specific rule mined from the key retrieval data of the first catalyst. For example: the first rule: "Ni-based catalysts with hydrothermal temperature of 180-200℃ and specific surface area ≥120m² / g have an HER rate ≥15mA / cm²"; the second rule: "Monomer structures containing amino (-NH2) functional groups can improve the separation efficiency of photogenerated carriers"; the third rule: "The catalytic stability of the first catalyst (CoS2 / TiO2) is optimal when the CoS2 loading is 5%-8%".
[0033] Agent 4 can provide several sets of recommended preparation parameters that conform to the patterns of the experimental parameter set and the constraints of the rule set by using the original experimental parameters and the various knowledge rules in the rule set.
[0034] In another embodiment, the agent 4 can be used to recommend parameters by integrating the experimental parameter set, rule set, and catalyst-related database to obtain several sets of recommended preparation parameters; wherein, the catalyst-related database is used to provide structural and computational data in the field of catalysis (such as numerical data such as electronic structure descriptors).
[0035] Among them, the catalyst-related database provides "supplementary basis at the microscopic mechanism level" for parameter recommendations. For example, it verifies the rationality of "amino functional groups enhance electron transfer efficiency" in the rule set by verifying electronic structure descriptors, or screens parameter combinations that are more conducive to hydrogen evolution reaction by using adsorption energy data. This helps the agent 4 output several sets of recommended preparation parameters that take into account practical basis, theoretical feasibility, and microscopic mechanism support.
[0036] In one specific embodiment, agent 4 recommends parameters based on the following method: First, feature extraction is performed on the original experimental parameters in the experimental parameter set: unstructured / semi-structured data are transformed into computable features, and knowledge rules in the rule set are transformed into logical expressions, which facilitates subsequent computer parsing and matching.
[0037] Then, agent 4 traverses each set of original experimental data in the experimental parameter set, removes original experimental parameters that violate the rule set, and filters out parameters that fully or partially conform to the rule set. Where the rule set includes at least two types of rules (e.g., including a first rule and a second rule), the priority of different rule types can be pre-set (e.g., prioritizing the second rule, then the first rule) for filtering.
[0038] Furthermore, the selected experimental parameters are optimized and adjusted. For example, the experimental parameters of existing catalysts with high catalytic performance can be optimized and adjusted to ensure that the recommended preparation parameters can produce high-quality catalysts.
[0039] It should be noted that the above-mentioned method of recommending parameters by combining relevant data from the original experiment (such as the original experimental parameters) and knowledge rules can avoid the blindness of purely data-driven approaches (such as simply replicating the original experimental parameters, which cannot break through existing knowledge) and the abstractness of purely knowledge rules (such as relying solely on theory and being divorced from actual experience and laws), ensuring that the final recommended parameters are both based on evidence and have a logical basis.
[0040] S13: Select at least one set of target experimental parameters from several sets of recommended preparation parameters; wherein each set of target experimental parameters is used to prepare the second catalyst.
[0041] Among them, the selected target experimental parameters are the experimental parameters corresponding to the predicted second catalyst with high catalytic activity.
[0042] In one embodiment, the rationality of each group of recommended preparation parameters can be evaluated, and then each group of recommended preparation parameters that pass the evaluation can be used as the target experimental parameters, or at least one group of target experimental parameters can be randomly selected from the group of recommended preparation parameters that pass the evaluation.
[0043] In another embodiment, a performance prediction model can be pre-trained, and then the performance prediction model can be used to predict the performance prediction results corresponding to each set of recommended preparation parameters. Then, based on each performance prediction result, at least one set of recommended preparation parameters whose performance prediction results meet the requirements (which can be set according to actual needs) can be selected from several sets of recommended preparation parameters as candidate experimental parameters. Then, at least one set of target experimental parameters can be selected from each set of candidate experimental parameters.
[0044] Specifically, please refer to Figure 2 , Figure 2 yes Figure 1 The flowchart shown in step S13 is a schematic diagram of an embodiment. This embodiment includes: S21: Use the performance prediction model to process the recommended preparation parameters for each group and obtain the performance prediction results corresponding to the recommended preparation parameters for each group.
[0045] The performance prediction results are used to characterize the predicted catalytic activity of the catalysts prepared under the corresponding recommended preparation parameters; among them, the performance prediction results of each group of recommended preparation parameters include one of high activity, medium activity and low activity.
[0046] In one embodiment, training samples can be obtained first, and then sorted based on the values of performance-related parameters (e.g., HER rate) corresponding to all training samples. Based on the sorting results, performance labels are assigned to each training sample. Then, a performance prediction model is trained using the training samples with assigned performance labels, specifically including the following steps: First, obtain the original experimental data for each group and the relationship representation data between the rules in the rule set; among them, the relationship representation data is used to characterize whether the original experimental data conforms to each rule.
[0047] In one embodiment, if the original experimental data satisfies the corresponding rule, the relational representation data is determined to be a first constant (e.g., 1); if the corresponding rule is not satisfied, the relational representation data is determined to be a second constant (e.g., 0).
[0048] Second, the data representing the fusion relationship and the corresponding original experimental data are used as a training sample.
[0049] In this embodiment, for each set of original experimental data (e.g., data A), data A and the relational representation data corresponding to each rule can be concatenated to obtain a training sample.
[0050] For example, the original experimental data and the relationship representation data between the rules in the rule set are vectorized to obtain a first vector, and the original experimental data are vectorized to obtain a second vector. Then, the first and second vectors are concatenated to obtain the corresponding training samples. For a detailed understanding of the original experimental data and rule set, please refer to the previous description, which will not be repeated here.
[0051] Third, the performance prediction model is trained using each training sample.
[0052] In this embodiment, after obtaining each training sample, the training samples are sorted based on the performance-related parameter values (e.g., HER rate) corresponding to all training samples. Based on the sorting results, a performance label is set for each training sample. Then, a performance prediction model is trained using the training samples with set performance labels.
[0053] Assuming the total number of training samples is N, and the HER rate sequence is... ,in Then, based on the total number of training samples, the classification thresholds are determined, and performance labels are assigned to each training sample using these thresholds. The classification thresholds are determined as follows: High activity label: ,when
[0054] Active label: ,when
[0055] Low-activity label: ,when
[0056] in This represents the floor function. This classification method based on relative ranking ensures a balanced label distribution while avoiding the dependence of absolute thresholds on dataset preferences.
[0057] Understandably, the trained performance prediction model is a classification model (such as a gradient boosting tree or a neural network). This classification model can classify based on the input experimental parameters, predicting the category (high activity, medium activity, or low activity) corresponding to the performance-related parameters of the second catalyst prepared under those experimental parameters, and can be trained based on the following objective function:
[0058] in, This represents the true label of sample i in category c (e.g., in one-hot form). This represents the probability output predicted by the model.
[0059] S22: Select at least one set of recommended preparation parameters whose performance prediction results meet the requirements as candidate experimental parameters.
[0060] In one embodiment, the recommended preparation parameters with high performance prediction results can be used as candidate experimental parameters to ensure that the catalyst activity corresponding to each candidate experimental parameter is higher than that of the recommended preparation parameters that were not selected as candidate experimental parameters.
[0061] S23: Select at least one set of target experimental parameters from each set of candidate experimental parameters.
[0062] In this embodiment, the intelligent agent 5 can be used to perform a secondary evaluation on each set of candidate experimental parameters selected to obtain the candidate experimental parameters that pass the evaluation, and the candidate experimental parameters that pass the secondary evaluation can be used as at least one set of target experimental parameters.
[0063] In one embodiment, the intelligent agent 5 combines its expertise in materials chemistry to evaluate the rationality of candidate experimental parameters from multiple dimensions such as synthesis feasibility, cost-effectiveness, and safety, and finally determines high-performance experimental parameters with practical application value.
[0064] In another implementation, after obtaining the candidate experimental parameters that have passed the secondary evaluation, density functional theory (DFT) calculations are performed on them to quickly screen out the target experimental parameters that meet the target performance.
[0065] In a specific implementation scenario, both the first catalyst and the second catalyst are photocatalysts.
[0066] Taking the first and second catalysts as covalent organic framework materials as an example, after obtaining the candidate experimental parameters of the covalent organic framework materials that have passed the secondary evaluation, density functional theory (DFT) calculations are performed on them to predict the electronic structure, band structure and adsorption performance of the covalent organic framework materials. Then, based on the prediction results, the target experimental parameters that meet the high photocatalytic hydrogen evolution potential are quickly screened out.
[0067] The above-described scheme recommends preparation parameters by integrating the original experimental data obtained from the preparation of the first catalyst with relevant knowledge rules. In this process, parameters conforming to the knowledge rules can be recommended as preparation parameters, thereby obtaining the target experimental parameters for the preparation of the second catalyst. Compared to methods that only utilize the original experimental data obtained from the preparation process to determine parameters, this application's approach, which introduces knowledge rules, ensures that the recommended preparation parameters are not only based on practical experience but also theoretically feasible, thus improving the discovery efficiency of catalysts with high catalytic performance.
[0068] In some embodiments, after selecting at least one set of target experimental parameters, the method further includes: preparing a catalyst based on each set of target experimental parameters to obtain a corresponding second catalyst; and optimizing the rule set and / or performance prediction model based on the performance verification results of each second catalyst.
[0069] For example, for each set of target experimental parameters, actual catalyst materials are prepared according to the target experimental parameters, and then the performance of the second catalyst is verified by testing and analyzing its catalytic performance.
[0070] If the performance verification results indicate that the performance of the second catalyst meets the performance requirements, then it is used as a new training sample to enrich the performance prediction model and / or the training sample set of each agent; otherwise, the rules in the rule set are optimized using the target experimental parameters (e.g., deleting or modifying the rules) and / or the performance prediction model.
[0071] Please see Figure 3 , Figure 3This is a flowchart illustrating an embodiment of obtaining the first rule provided in this application. In this embodiment, obtaining the first rule includes: S31: Based on the preparation parameters, classify the relevant data of multiple sets of original experiments into categories to obtain the data of each category.
[0072] In this embodiment, the original experimental data includes: preparation parameters and performance-related parameters of the first catalyst under the preparation parameters.
[0073] It should be noted that in this embodiment, the original experimental data are classified only based on the preparation parameters of the first catalyst, without introducing the performance-related parameters of the first catalyst. This approach is to ensure that the classified data only reflects the structural characteristics of the first catalyst itself, so as to divide multiple sets of original experimental data into multiple categories with clear structural features.
[0074] In one specific implementation, a feature space can be constructed from the preparation parameters in each set of original experimental data. Then, a relevant clustering algorithm (such as K-means clustering algorithm) can be used to divide the multiple sets of original experimental data into multiple different structural feature categories to obtain the divided data of each category.
[0075] The clustering objective function is expressed as follows:
[0076] Where k represents the number of clusters, C i This represents the i-th cluster (category). Indicates cluster C i The centroid of x is the sample feature vector (the constructed feature space).
[0077] It should be noted that before constructing the feature space of the preparation parameters in the original experimental data of each group, feature vectors are generated for each dimension of the preparation parameters (such as synthesis parameters, structural characterization parameters, and reaction condition parameters). Then, mean pooling is used to obtain the comprehensive vector representation of the preparation parameters. Principal component analysis is then used to reduce the dimensionality of the comprehensive vector representation by calculating the covariance matrix and eigenvalue decomposition, so as to retain the main information while reducing the computational complexity.
[0078] S32: Based on the data of various classifications, knowledge mining is performed to obtain the first rule.
[0079] In this embodiment, the original experimental data in the various types of data classification include: the original experimental parameters mentioned above, and the mechanism annotation data of the original experimental parameters.
[0080] In one embodiment, intra-class knowledge mining is performed on various types of segmented data based on a first preset prompt word to obtain unique design rules corresponding to each type of segmented data; and / or, inter-class knowledge mining is performed on at least two types of segmented data based on a second preset prompt word to obtain general design rules between at least two types of segmented data.
[0081] Preferably, intra-class knowledge mining and Leijian knowledge mining can be performed separately for various types of data to fully uncover the material (parameter) design rules for various structural characteristics.
[0082] In one specific embodiment, before performing intra-class and inter-class knowledge mining, a first preset number of original experimental data sets can be selected from each class of data, and a second preset number of original experimental data sets can be selected from different class of data sets (e.g., every two class sets of data).
[0083] The first preset quantity and the second preset quantity may be the same or different; for example, the first preset quantity and the second preset quantity are both 10 groups; in addition, data can be selected by random selection, and the original experimental data of each selected group are different.
[0084] In one implementation scenario, intra-class knowledge mining can be performed on various types of data based on the first preset prompt words to obtain unique design rules corresponding to each type of data; and inter-class knowledge mining can be performed on at least two types of data based on the second preset prompt words to obtain general design rules between at least two types of data.
[0085] Among them, the first preset prompt word and the second prompt word are pre-designed prompt word frameworks used to guide the corresponding agents (agents 6 and 7) to perform knowledge mining.
[0086] For example, the design of the first preset prompt and / or the second preset prompt is as follows: "You are a research expert who focuses on the design and performance regulation of covalent organic frameworks (COFs) materials."
[0087] There is now a set of experimental data on COFs materials, which records their synthesis conditions (monomer, temperature, time, atmosphere, etc.), electronic structure information, photocatalytic reaction conditions (light source wavelength, co-catalyst, etc.), as well as their photocatalytic performance parameters and rates.
[0088] Your task is: Based on this data, we propose **40 important design rules related to the excellent photocatalytic performance of COFs materials**.
[0089] ### **Content Focus (The rules must be closely related to the following content)** The rules must clearly define the associated molecular structure, pore size, and surface chemistry, and can be based on, but are not limited to, any of the following dimensions: * **Monomer Design**: Element type, donor-acceptor structure (D–A), π-conjugated system, heteroatom type (e.g., N, S, O); * **Skeleton structure**; * **Physicochemical properties and electronic structure characteristics** (the subject to which the characteristics belong must be clearly identified); **Synthesis process conditions:** reaction type, temperature, time, catalyst, solvent, atmosphere, etc. * **Photocatalytic reaction conditions**: co-catalyst, sacrificial agent, light source wavelength, etc.
[0090] ### **Output Format and Requirements** Please output **40 rules**, using **Chinese characters**, and number them sequentially (1–40). - You should avoid simply repeating data and instead extract general trends and causal relationships across multiple samples.
[0091] * Each rule must be a **quantifiable or categorizable criterion**, such as: * Numerical type (e.g., "Calculate the LUMO level of the monomer, which needs to be higher than... to meet the requirements of electron reduction"); * Classification (e.g., "Check if there are sacrificial agents (such as triethanolamine, Na2S-Na2SO3) that can consume photogenerated holes and improve electron utilization"). * Combinations of conditions (e.g., "high specific surface area + use of Pt co-catalyst → improved catalytic performance"); * Each rule should **not exceed two sentences**, be semantically clear, specify the object being described, and use accurate terminology; * Do not include introductions, background information, or summaries; **only output the rule itself.** In one embodiment, for each type of data segmentation, the unique design rules obtained by intra-class knowledge mining include: a first unique design rule obtained by knowledge mining based on each original experimental parameter of the category, and a second unique design rule obtained by knowledge mining based on each original experimental related data of the category.
[0092] Similarly, for data with at least two categories, the general design rules include: a first general design rule obtained by knowledge mining of the original experimental parameters belonging to different categories, and a second general design rule obtained by knowledge mining of the relevant data of the original experiments belonging to different categories.
[0093] It should be noted that the first unique design rule is the first association rule between the original experimental parameters of the first catalyst of the corresponding category and the corresponding catalytic performance; the second unique design rule is the first causal rule between the original experimental parameters, mechanism of action and catalytic performance of the first catalyst of the corresponding category; the first association rule is the unique association rule of the corresponding category, and the first causal rule is the unique causal rule of the corresponding category. The first general design rule is a second association rule concerning the original experimental parameters and catalytic performance of different categories. The second general design rule is a second causal rule concerning the original experimental parameters, mechanism of action, and catalytic performance of different categories of the first catalyst. The second association rule is a general association rule for different categories, and the second causal rule is a general causal rule for different categories.
[0094] It should also be noted that the first unique design rule and the first general design rule are obtained by knowledge mining of the original experimental parameters (i.e., historical experimental parameters), for example, by using agent 6 to perform knowledge mining on the original experimental parameters; the second unique design rule and the second general design rule are obtained by knowledge mining of the corresponding original experimental related data (i.e., historical experimental parameters + mechanism annotation data), for example, by using agent 7 to perform knowledge mining on the original experimental related data (i.e., historical experimental parameters + mechanism annotation data).
[0095] Among them, compared with the rules mined by agent 6, the rules mined by agent 7 not only describe "what" but also explain "why", greatly improving the interpretability and practicality of the rules.
[0096] In one specific embodiment, the aforementioned knowledge rules are obtained by using different intelligent agents to mine knowledge from different data.
[0097] Specifically, agent 1 performs knowledge mining on multiple sets of original experimental data related to the first catalyst to obtain the first rule; agent 2 performs knowledge mining on general chemical knowledge to obtain the second rule; agent 3 performs knowledge mining on the latest retrieved research abstracts to obtain the third rule. Further, agent 6 performs knowledge mining on the original experimental parameters to obtain the first unique design rule and the first general design rule; agent 7 performs knowledge mining on the original experimental data (i.e., original experimental parameters + mechanism annotation data) to obtain the second unique design rule and the second general design rule. Agent 1 can be considered as an agent that includes agents 6 and 7.
[0098] The workflow and interaction logic of each agent are implemented through specific prompt words; the prompt words for agents 6 and 7 can be referenced from the design of the first preset prompt words and / or the second preset prompt words described above.
[0099] The prompts for Agent 2 can be found as follows: "You are a research expert specializing in the design and performance regulation of photocatalytic covalent organic frameworks (COFs). Based on your expertise, please propose 40 important design rules related to the excellent photocatalytic performance of organic polymers."
[0100] ### **Content Focus (The rules must be closely related to the following content)** The rules must be explicitly related and based on any of the following dimensions: * **Monomer Design**: Element type, donor-acceptor structure (D–A), π-conjugated system, heteroatom type (e.g., N, S, O); **Skeleton structure:** Topology, crystal order, pore size, interlayer spacing, etc. * **Physicochemical properties and electronic structure characteristics** (the subject to which the characteristics belong must be clearly identified); **Synthesis process conditions:** reaction type, temperature, time, catalyst, solvent, atmosphere, etc. * **Photocatalytic reaction conditions**: co-catalyst, sacrificial agent, light source wavelength, etc.
[0101] ### **Output Format and Requirements** Please output **40 rules**, using **Chinese characters**, and number them sequentially (1–40). * **Each rule must begin with "Calculate..." or "Check..."**; * Each rule must be a **quantifiable or categorizable criterion**, such as: * Numerical type (e.g., "Calculate the LUMO level of the monomer, which needs to be higher than... to meet the requirements of electron reduction"); * Classification (e.g., "Check if there are sacrificial agents (such as triethanolamine, Na2S-Na2SO3) that can consume photogenerated holes and improve electron utilization"). * Combinations of conditions (e.g., "high specific surface area + use of Pt co-catalyst → improved catalytic performance"); * Each rule should **not exceed two sentences**, be semantically clear, specify the object being described, and use accurate terminology; * Do not include introductions, background information, or summaries; **only output the rule itself.** In one specific embodiment, agent 2 acts as a knowledge integrator. Based on the knowledge system of the large language model itself, it extracts universal material design principles from the model's pre-training data without relying on a specific dataset. Although these comprehensive rules do not originate directly from experimental data, they provide higher-level design concepts and directional guidance, compensating for the limitations of data-driven methods.
[0102] The prompts for Agent 3 can be found as follows: "You are a model expert in extracting rules for COF material design and photocatalytic regulation. Determine whether it involves **covalent organic framework (COF) material design and photocatalytic performance regulation**. If so, summarize **1-2 rules**."
[0103] The rules must clearly define the associated molecular structure, pore size, and surface chemistry, and can be based on, but are not limited to, any of the following dimensions: * **Monomer Design**: Element type, donor-acceptor structure (D–A), π-conjugated system, heteroatom type (e.g., N, S, O); * **Skeleton structure**; * **Physicochemical properties and electronic structure characteristics** (the entity to which the characteristics belong must be clearly identified); **Synthesis process conditions:** reaction type, temperature, time, catalyst, solvent, atmosphere, etc. * **Photocatalytic reaction conditions**: co-catalyst, sacrificial agent, light source wavelength, etc.
[0104] ### Output Requirements: * **Each rule must begin with "Calculate..." or "Check..."**; * Each rule must be a **quantifiable or categorizable criterion**, such as: * Numerical type (e.g., "Calculate the LUMO level of the monomer, which needs to be higher than... to meet the requirements of electron reduction"); * Classification (e.g., "Check if there are sacrificial agents (such as triethanolamine, Na2S-Na2SO3) that can consume photogenerated holes and improve electron utilization"). * Combinations of conditions (e.g., "high specific surface area + use of Pt co-catalyst → improved catalytic performance"); * Each rule should **not exceed two sentences**, be semantically clear, specify the object being described, and use accurate terminology; * Do not include introductions, background information, or summaries; **only output the rule itself**.
[0105] - If this study **does not involve COF materials**, please output: "**No photocatalytic COF materials involved.**" It should be noted that, through this prompt, the third rule mined by agent 3 can be strictly limited to Boolean or numerical expression, such as quantifiable judgment conditions like "specific surface area greater than 1000m² / g" or "using platinum-based co-catalyst", ensuring the clarity and operability of the rule.
[0106] It should also be noted that the rules mined by the various intelligent agents are integrated into a unified rule set after rigorous deduplication and consistency checks. The advantage of this multi-agent collaborative mining mechanism lies in maintaining the objectivity of data-driven rules while incorporating the depth of mechanistic annotations. It also supplements the universality of domain knowledge and the latest research progress, forming a complete rule system that is multi-layered, multi-faceted, and mutually verifying, providing a solid theoretical foundation for subsequent parameter recommendation and material design.
[0107] It should be noted that, apart from knowledge mining, other steps can also be implemented using relevant intelligent agents. For example, agent 4 can be used to perform step S12 above: combining the experimental parameter set and rule set to recommend parameters and obtain several sets of recommended preparation parameters; agent 5 can also be used to perform a secondary evaluation on each set of candidate experimental parameters to obtain candidate experimental parameters that pass the evaluation.
[0108] The workflow and interaction logic of Agent 4 and Agent 5 are also implemented through specific prompt words, which will not be described in detail here.
[0109] Please see Figure 4 , Figure 4 This is a schematic diagram of an embodiment of the catalyst parameter determination device provided in this application. In this embodiment, the catalyst parameter determination device 40 includes: an acquisition module 41, a parameter recommendation module 42, and a parameter determination module 43. The acquisition module 41 is used to acquire the experimental parameter set and rule set of the first catalyst; the experimental parameter set includes multiple sets of original experimental data obtained from the preparation of the first catalyst, and each rule in the rule set is obtained through knowledge mining; the parameter recommendation module 42 is used to comprehensively analyze the experimental parameter set and the rule set to recommend parameters, obtaining several sets of recommended preparation parameters; the parameter determination module 43 is used to select at least one set of target experimental parameters from the several sets of recommended preparation parameters; wherein each set of target experimental parameters is used to prepare the second catalyst.
[0110] In some embodiments, the parameter determination module 43 selects at least one set of target experimental parameters from several sets of recommended preparation parameters, including: processing each set of recommended preparation parameters using a performance prediction model to obtain the performance prediction results corresponding to each set of recommended preparation parameters; using at least one set of recommended preparation parameters whose performance prediction results meet the requirements as candidate experimental parameters; and selecting at least one set of target experimental parameters from each set of candidate experimental parameters.
[0111] In some embodiments, before the parameter determination module 43 selects at least one set of target experimental parameters from several sets of recommended preparation parameters, the catalyst parameter determination device 40 is further configured to: acquire each set of original experimental data and relationship characterization data between each rule in the rule set; use the relationship characterization data to characterize whether the original experimental data conforms to each rule; fuse the relationship characterization data and the corresponding original experimental data as a training sample; and use each training sample to train a performance prediction model.
[0112] In some embodiments, the first catalyst and the second catalyst are catalysts made of the same or different materials; after selecting at least one set of target experimental parameters from each set of candidate experimental parameters, the catalyst parameter determination device 40 is further configured to: prepare a catalyst based on each set of target experimental parameters to obtain a corresponding second catalyst; and optimize the rule set and / or performance prediction model based on the performance verification results of each second catalyst.
[0113] In some embodiments, the original experimental data acquired by the acquisition module 41 includes: original experimental parameters, which include preparation parameters and performance-related parameters of the first catalyst under the preparation parameters. The preparation parameters include at least one of the synthesis parameters, structural characterization parameters, and reaction condition parameters of the first catalyst. The rule set includes: a first rule obtained by knowledge mining based on each set of original experimental data. The steps for acquiring the first rule include: classifying multiple sets of original experimental data based on the preparation parameters to obtain class classification data; and performing knowledge mining based on the class classification data to obtain the mined first rule.
[0114] In some embodiments, the rule set acquired by the acquisition module 41 further includes at least one of a second rule and a third rule: the second rule is a catalyst design rule obtained by knowledge mining based on general chemical knowledge, and the third rule is obtained by knowledge mining based on key retrieval data of the first catalyst.
[0115] In some embodiments, knowledge mining is performed based on various types of segmented data to obtain the first rule mined, including: performing intra-class knowledge mining on various types of segmented data based on a first preset prompt word to obtain unique design rules corresponding to various types of segmented data; and / or, performing inter-class knowledge mining on at least two types of segmented data based on a second preset prompt word to obtain general design rules between at least two types of segmented data.
[0116] In some embodiments, the original experimental data acquired by the acquisition module 41 further includes mechanistic annotation data of the original experimental parameters, wherein the mechanistic annotation data is obtained by performing mechanistic annotation processing on the original experimental parameters; for various types of partitioned data, the unique design rules for partitioned data include: a first unique design rule obtained by knowledge mining based on each original experimental parameter belonging to a category, and a second unique design rule obtained by knowledge mining based on each original experimental data belonging to a category; the general design rules include: a first general design rule obtained by knowledge mining based on each original experimental parameter belonging to different categories, and a second general design rule obtained by knowledge mining based on each original experimental data belonging to different categories.
[0117] In some embodiments, the original experimental parameters include: preparation parameters, and performance correlation parameters corresponding to the first catalyst having structural characterization data, wherein the preparation parameters include at least one of the synthesis parameters, structural characterization parameters, and reaction condition parameters of the first catalyst; and / or, a first unique design rule is a first correlation rule between the original experimental parameters and corresponding catalytic performance of the first catalyst of a corresponding category; a second unique design rule is a first causal rule between the original experimental parameters, mechanism of action, and corresponding catalytic performance of the first catalyst of a corresponding category; the first correlation rule is a unique correlation rule for the corresponding category, and the first causal rule is a unique causal rule for the corresponding category; a first general design rule is a second correlation rule between the original experimental parameters and catalytic performance of different categories, and the second general design rule is a second causal rule between the original experimental parameters, mechanism of action, and catalytic performance of the first catalyst of different categories; the second correlation rule is a general correlation rule for different categories, and the second causal rule is a general causal rule for different categories.
[0118] In some embodiments, the first catalyst and the second catalyst are photocatalysts, and / or both the first catalyst and the second catalyst are covalent organic framework materials.
[0119] Please see Figure 5 , Figure 5 This is a schematic diagram of an embodiment of the electronic device provided in this application. In this embodiment, the electronic device 50 includes a memory 51 and a processor 52 coupled to each other.
[0120] The memory 51 stores program instructions, and the processor 52 executes the program instructions stored in the memory 51 to implement the steps of any of the above-described method implementations. In a specific implementation scenario, the electronic device 50 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 50 may also include mobile devices such as laptops and tablets, which are not limited here.
[0121] Specifically, processor 52 controls itself and memory 51 to implement the steps of any of the above embodiments. Processor 52 may also be referred to as a CPU (Central Processing Unit). Processor 52 may be an integrated circuit chip with signal processing capabilities. Processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 52 may be implemented using integrated circuit chips.
[0122] Please see Figure 6 , Figure 6 This is a schematic diagram of the framework of the computer-readable storage medium provided in this application. The computer-readable storage medium 60 of this application embodiment stores program instructions 61, which, when executed, implement the methods provided in any embodiment or any non-conflicting combination of the above-described methods. The program instructions 61 can form a program file and be stored in the computer-readable storage medium 60 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) can execute all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 60 includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.
[0123] The above-described scheme recommends preparation parameters by integrating the original experimental data obtained from the preparation of the first catalyst with relevant knowledge rules. In this process, parameters conforming to the knowledge rules can be recommended as preparation parameters, thereby obtaining the target experimental parameters for the preparation of the second catalyst. Compared to methods that only utilize the original experimental data obtained from the preparation process to determine parameters, this application's approach, which introduces knowledge rules, ensures that the recommended preparation parameters are not only based on practical experience but also theoretically feasible, thus improving the discovery efficiency of catalysts with high catalytic performance.
[0124] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0125] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0126] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0127] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0128] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0129] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0130] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for determining catalyst parameters, characterized in that, The method includes: Obtain the experimental parameter set and rule set for the first catalyst; the experimental parameter set includes multiple sets of original experimental data obtained from the preparation of the first catalyst, and each rule in the rule set is obtained through knowledge mining; By combining the experimental parameter set and rule set, several sets of recommended preparation parameters are obtained; From a number of recommended preparation parameters, at least one set of target experimental parameters is selected; wherein each set of target experimental parameters is used to prepare the second catalyst.
2. The method according to claim 1, characterized in that, The selection of at least one set of target experimental parameters from several sets of recommended preparation parameters includes: The recommended preparation parameters for each group are processed using a performance prediction model to obtain the performance prediction results corresponding to the recommended preparation parameters for each group. At least one set of recommended preparation parameters whose performance prediction results meet the requirements are used as candidate experimental parameters. Select at least one set of target experimental parameters from the candidate experimental parameters described in each group.
3. The method according to claim 2, characterized in that, Before selecting at least one set of target experimental parameters from a plurality of recommended preparation parameters, the method further includes: Obtain the original experimental data for each group and the relationship representation data between the rules in the rule set; the relationship representation data is used to characterize whether the original experimental data conforms to each rule. The relational representation data and the corresponding original experimental data are combined to form a training sample. The performance prediction model is trained using each training sample.
4. The method according to claim 2, characterized in that, The first catalyst and the second catalyst may be catalysts made of the same or different materials; After selecting at least one set of target experimental parameters from each set of candidate experimental parameters, the method further includes: Based on the target experimental parameters of each group, the catalyst was prepared to obtain the corresponding second catalyst; Based on the performance verification results of each second catalyst, the rule set and / or the performance prediction model are optimized.
5. The method according to claim 1, characterized in that, Each set of original experimental data includes: original experimental parameters, which include preparation parameters and performance-related parameters of the first catalyst under the preparation parameters. The preparation parameters include at least one of the synthesis parameters, structural characterization parameters, and reaction condition parameters of the first catalyst. The rule set includes: first rules obtained by knowledge mining based on the first set of original experimental parameters. The steps for obtaining the first rule include: Based on the preparation parameters, multiple sets of original experimental data were categorized to obtain data for each category. Knowledge mining is performed based on various data categories to obtain the first rule.
6. The method according to claim 1 or 5, characterized in that, The rule set also includes at least one of a second rule and a third rule: the second rule is a catalyst design rule obtained by knowledge mining based on general chemical knowledge, and the third rule is obtained by knowledge mining based on key retrieval data of the first catalyst.
7. The method according to claim 5, characterized in that, The knowledge mining based on various data segments yields the first rule, which includes: Based on the first preset prompt words, intra-class knowledge mining is performed on each type of data segmentation to obtain unique design rules corresponding to each type of data segmentation; and / or, Based on the second preset prompt words, inter-class knowledge mining is performed on the data divided into at least two categories to obtain general design rules between the data divided into at least two categories.
8. The method according to claim 7, characterized in that, The original experimental data also includes mechanistic annotation data for the original experimental parameters; the mechanistic annotation data is obtained by performing mechanistic annotation processing on the original experimental parameters. For each type of partitioned data, the unique design rules for the partitioned data include: a first unique design rule obtained by knowledge mining based on each original experimental parameter belonging to the category, and a second unique design rule obtained by knowledge mining based on each original experimental related data of the category. The general design rules include: a first general design rule obtained by knowledge mining based on the original experimental parameters belonging to different categories, and a second general design rule obtained by knowledge mining based on the relevant data of the original experiments belonging to different categories.
9. The method according to claim 8, characterized in that, The first unique design rule is a first association rule relating the original experimental parameters and corresponding catalytic performance of the first catalyst of the corresponding category; the second unique design rule is a first causal rule relating the original experimental parameters, mechanism of action, and corresponding catalytic performance of the first catalyst of the corresponding category; the first association rule is a unique association rule for the corresponding category, and the first causal rule is a unique causal rule for the corresponding category; The first general design rule is a second association rule for each original experimental parameter and catalytic performance of different categories; the second general design rule is a second causal rule for the original experimental parameters, mechanism of action and catalytic performance of the first catalyst of different categories; the second association rule is a general association rule for different categories; the second causal rule is a general causal rule for different categories.
10. The method according to claim 1, characterized in that, The first catalyst and the second catalyst are photocatalysts, and / or the first catalyst and the second catalyst are both covalent organic framework materials.
11. A catalyst parameter determination device, characterized in that, The device includes: The acquisition module is used to acquire the experimental parameter set and rule set of the first catalyst; the experimental parameter set includes multiple sets of original experimental data obtained from the preparation of the first catalyst, and each rule in the rule set is obtained through knowledge mining; The parameter recommendation module is used to recommend parameters by combining the experimental parameter set and the rule set, and obtain several sets of recommended preparation parameters. The parameter determination module is used to select at least one set of target experimental parameters from several sets of recommended preparation parameters; wherein each set of target experimental parameters is used to prepare the second catalyst.
12. An electronic device, characterized in that, Including interconnected memory and processor, The memory stores program instructions; The processor is used to execute program instructions stored in the memory to implement the method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that can be executed by a processor to implement the method according to any one of claims 1-10.