Computing device for computing chemical question answering
Patent Information
- Application Number
- CN202610821801.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-09
AI Technical Summary
该方案虽然能够减少人工重复操作、提升局部流程的自动化程度,但其本质仍是对固定任务流程的机械执行,缺乏对复杂科学问题的理解能力,也无法围绕研究目标自主提出、比较和筛选多种机理假设
[0017]根据本发明的计算设备克服了固定任务流程缺乏对复杂科学问题的理解能力以及通用模型缺乏计算化学领域的系统知识的局限性,通过利用多智能体协同框架结合通用大语言模型与计算化学专用模型的推理能力,提升了对计算化学任务的回答准确性。
Smart Images

Figure CN122364417B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to computer systems utilizing computational models, and more specifically to computational devices for computational chemical question answering. Background Technology
[0002] As the application of artificial intelligence in chemical research continues to deepen, especially with the rapid development in areas such as molecular design, reaction prediction, catalytic mechanism analysis, and materials discovery, researchers' demands for intelligent systems are no longer limited to single-task assistance or simple tool calls. Instead, they hope that intelligent systems can conduct continuous reasoning, scheme design, and evidence-driven verification around complex and open scientific problems in a manner similar to that of scientific researchers.
[0003] However, computational chemistry tasks are typically characterized by long processes, high specialization, parameter sensitivity, high software dependence, and high barriers to result interpretation. For example, reaction mechanism research often involves multiple stages, including literature review, candidate mechanism proposal, construction of key intermediates or transition state structures, design of quantitative computational paths, input file writing, interpretation of computational results, error handling, and mechanism correction. Existing research and engineering practices show that this process still heavily relies on experienced researchers to complete manually, which is not only inefficient but also easily affected by differences in individual experience, leading to insufficient execution stability and result consistency. Therefore, how to construct an automated, intelligent, and closed-loop system that can serve complex computational chemistry research tasks has become a key issue in the current development of AI-assisted computational chemistry.
[0004] In existing technologies, computational chemistry automation methods rely on predefined workflows or single tool calls. The main technical idea is to pre-design fixed workflows, script templates, or regularized call chains for local tasks such as geometry optimization, frequency calculation, single-point energy calculation, and transition state search to automate certain computational steps. While this approach can reduce repetitive manual operations and improve the automation level of local workflows, it is essentially a mechanical execution of fixed task processes, lacking the ability to understand complex scientific problems and unable to autonomously propose, compare, and screen multiple mechanistic hypotheses around the research objectives. When intermediate results are abnormal, calculations fail, or research objectives change, this type of approach is usually difficult to dynamically correct, thus making it unsuitable for highly complex tasks such as exploring open-ended, unknown reaction mechanisms.
[0005] There is a need in this field for computational chemistry question answering techniques that can be improved at at least one of the aforementioned levels. Summary of the Invention
[0006] This invention is provided to further improve computational chemistry question answering techniques by utilizing intelligent agents and computational models.
[0007] One aspect of the present invention provides a computing device for computational chemistry question answering, comprising: computing resources; a large language model configured to invoke the computing resources to: receive input from a user, the input including a computational chemistry question; a retrieval agent configured to invoke the large language model to: receive the input; identify keywords related to computational chemistry in the input; retrieve the keywords to obtain retrieval results; a hypothesis generation agent configured to invoke the large language model to: receive the input, the keywords, and the retrieval results; generate multiple scientific hypotheses based on the input, the keywords, and the retrieval results; a planning agent configured to invoke the large language model to: receive the multiple scientific hypotheses, the input, and the retrieval results; generate a computational workflow based on the multiple scientific hypotheses, the input, the retrieval results, and a structured tool registry, the structured tool registry being provided by computational chemistry software and indicating the functions, input formats, and output formats of tools in the computational chemistry software; and a dedicated model configured to invoke the computing resources to: review the professional rationale of the computational workflow; wherein the large language model is configured to review the computational workflow. Logical consistency and input / output dependencies; and an execution module, configured to: receive the computational workflow; convert the computational workflow into computational tasks based on the structured tool registry, the computational tasks including tool names, input files, and calling parameters; call tools in the computational chemistry software to execute the computational tasks to obtain computational results as answers to the computational chemistry problem; wherein, the dedicated model is obtained by training the base model based on a software route generation training dataset, the software route generation training dataset being obtained through S1, S1 including: S11: obtaining input examples of computational tasks under different computational scenarios of the computational chemistry software; S12: extracting route segments from the input examples, the route segments indicating computational objectives, theoretical methods, basis sets, and keywords for the computational task; S13: constructing keyword combination templates for computational chemistry software under different computational scenarios based on the route segments; S14: constructing software route generation training samples based on the keyword combination templates as the software route generation training dataset, the software route generation training samples including computational tasks as sample inputs and route segments as sample outputs.
[0008] In the computing device described above, S1 further includes: S15: determining whether the software route generation training sample conforms to the verification rules; S16: in response to the software route generation training sample conforming to the verification rules, using the software route generation training sample as the software route generation training dataset.
[0009] As described above, in the computing device, the dedicated model is obtained by training the base model based on a software computation result analysis training dataset. The software computation result analysis training dataset is obtained through S2, which includes: S21: acquiring the output file of the computational task of the computational chemistry software; S22: extracting computational information from the output file; S23: labeling the computational information to obtain analysis results; S24: constructing software computation result analysis training samples based on the analysis results, which serve as the software computation result analysis training dataset. The software computation result analysis training samples include the output file as sample input and the analysis results as sample output.
[0010] In the computing device described above, the dedicated model is obtained by training the base model on a software computing error handling training dataset. The software computing error handling training dataset is obtained through step S3, which includes: S31: acquiring error information from error cases in computational chemistry software; S32: locating the error location of the error case according to the error type; S33: generating an error solution based on the error information; and S34: constructing software computing error handling training samples based on the error type, the error location, and the error solution, as the software computing error handling training dataset. The software computing error handling training samples include error information as sample input, error location, and error solution as sample output.
[0011] In the computing device described above, the dedicated model is obtained by training the base model on a basic knowledge training dataset. The basic knowledge training dataset is obtained through S4, which includes: S41: acquiring literature related to computational chemistry theoretical knowledge and computational chemistry software basic knowledge; S42: splitting the literature into knowledge units, the knowledge units including computational chemistry theoretical knowledge units and software practical knowledge units; S43: constructing instruction-based basic knowledge training samples based on the knowledge units as the basic knowledge training dataset.
[0012] In the computing device described above, the dedicated model is obtained by training the base model on a general chemical knowledge training dataset. The general chemical knowledge training dataset is obtained through S5, which includes: S51: acquiring basic chemical knowledge literature; S52: performing structured parsing on the basic chemical knowledge literature; S53: constructing general chemical knowledge training samples based on the parsing results as the general chemical knowledge training dataset.
[0013] In the computing device described above, the dedicated model is configured to invoke the computing resources to determine the technical validity of the computation results, and the large language model is configured to invoke the computing resources to determine whether the computation results support the multiple scientific hypotheses.
[0014] In the computing device described above, the execution module is configured to: if the calculation result is invalid, feed back the invalid result and the repair suggestion of the dedicated model to the planning agent to adjust the calculation workflow, or feed back the invalid result and the repair suggestion of the dedicated model to the user to adjust the input; if the calculation result is valid but does not support the multiple scientific hypotheses, feed back the calculation result to the hypothesis generating agent to adjust the multiple scientific hypotheses.
[0015] As described above, the computing device includes a computing workflow comprising multiple computing steps, each computing step including a step number, task description, tool name, input object, and output object.
[0016] The scientific hypothesis of the computing device described above includes one or more of the following: reaction pathway, key intermediate, transition state configuration, selective control factor, or structure-property relationship explanation.
[0017] The computing device according to the present invention overcomes the limitations of fixed task processes lacking the ability to understand complex scientific problems and general models lacking systematic knowledge in the field of computational chemistry. By utilizing a multi-agent collaborative framework that combines the reasoning capabilities of a general large language model and a computational chemistry-specific model, it improves the accuracy of answers to computational chemistry tasks. Attached Figure Description
[0018] Various embodiments of the present invention are described in conjunction with the accompanying drawings.
[0019] Figure 1 This is a block diagram of a computing device for computational chemical question answering according to some embodiments of the present invention.
[0020] Figure 2 This is a schematic diagram illustrating the interaction between a computing device and a user according to some embodiments of the present invention.
[0021] Figure 3 This is a schematic diagram of a computational device according to some embodiments of the present invention, which answers computational chemistry questions.
[0022] Figure 4 This is a flowchart of a first process for obtaining a training dataset by generating a software route, according to some embodiments of the present invention.
[0023] Figure 5This is a flowchart of a second process for analyzing training datasets to obtain software computation results, according to some embodiments of the present invention.
[0024] Figure 6 This is a flowchart of a third process for obtaining a training dataset for software computation error handling, according to some embodiments of the present invention.
[0025] Figure 7 This is a flowchart of a fourth process for obtaining a training dataset of basic knowledge according to some embodiments of the present invention.
[0026] Figure 8 This is a flowchart of a fifth process for obtaining a general chemistry training dataset according to some embodiments of the present invention. Detailed Implementation
[0027] In this application, the term "agent" refers to an agent capable of perceiving its environment and taking actions to perform specific goals. An agent primarily refers to software code. Agents can be executed by the computing resources of a computing device. Agents can use API interfaces to invoke corresponding models to interact with various forms of input or to implement corresponding functions.
[0028] In this application, ordinal numbers such as "first," "second," and "third" are used to distinguish different instances of objects with the same name. The ordinal numbers "first," "second," and "third" do not indicate a relative order of the indicated objects in time, space, sequence, or other aspects.
[0029] According to one aspect of the present invention, a computing device for computational chemical question answering is provided.
[0030] Figure 1 This is a block diagram of a computing device 100 for computational chemical question answering according to some embodiments of the present invention.
[0031] The computing device 100 can be a local or remote computer, server, etc.
[0032] In some embodiments, computing device 100 may include computing resources 110, a large language model 120, a specialized model 130, a retrieval agent 140, a hypothesis generation agent 150, a planning agent 160, and an execution module 170. In some embodiments, the large language model 120 may also be located outside computing device 100 (e.g., remotely to computing device 100).
[0033] In some embodiments, computing resources 110 may include a central processing unit (CPU), a graphics processing unit (GPU), and various other processing units or cores (e.g., arithmetic logic units, integer units, floating-point units, tensor units, ray tracing cores, etc.).
[0034] The large language model 120 and the special model 130 can be configured to invoke computing resources 110 to perform corresponding operations.
[0035] The retrieval agent 140, hypothesis-generating agent 150, and planning agent 160 can be configured to call the large language model 120 via API to perform corresponding operations. The following will combine... Figure 2 Describe the specific functions of these intelligent agents.
[0036] The execution module 170 can be configured to invoke various tools to interact with various forms of input or to perform corresponding functions.
[0037] Figure 2 This is a schematic diagram illustrating the interaction between a computing device and a user according to some embodiments of the present invention. Figure 3 This is a schematic diagram of a processing flow for a computing device to answer computational chemistry problems according to some embodiments of the present invention. (Combined with...) Figure 2 Come to Figure 3 The processing flow will be explained.
[0038] In the first step, at point 310, the large language model receives input, which includes computational chemistry problems. For example... Figure 2 As shown, the large language model 120 receives input from the user, including computational chemistry questions.
[0039] In the second process, at step 320, keywords are retrieved from the input identified by the intelligent agent. For example... Figure 2 As shown, the retrieval agent 140 receives relevant inputs about computational chemistry problems from the large language model 120. Then, the retrieval agent 140 identifies keywords related to computational chemistry in the input and performs a search on the keywords to obtain search results.
[0040] In the third process, at step 330, it is assumed that the agent generates multiple scientific hypotheses based on keywords and search results. For example... Figure 2 As shown, hypothetical generative agent 150 receives relevant inputs to a computational chemistry problem from a large language model 120, and receives keywords and search results from a retrieval agent 140. Then, hypothetical generative agent 150 generates multiple scientific hypotheses based on the inputs, keywords, and search results.
[0041] In step 340 of the fourth process, the planning agent generates a computational workflow based on multiple scientific hypotheses. For example... Figure 2 As shown, the planning agent 160 receives relevant inputs to the computational chemistry problem from the large language model 120, receives search results from the retrieval agent 140, receives multiple scientific hypotheses from the hypothesis generation agent 150, and receives a structured tool registry related to the computational chemistry software. Based on the multiple scientific hypotheses, inputs, search results, and structured tool registry, the planning agent 160 generates a computational workflow.
[0042] At step 350 of the fifth process, the execution module calls the tool to execute the computation workflow. For example... Figure 2 As shown, the execution module 170 receives the computational workflow from the planning agent 160, calls the tool to execute the computational tasks in the computational workflow, and obtains the computational results as an answer to the computational chemistry problem.
[0043] Based on the above processing flow, the user submits a computational chemistry problem to the large language model. The large language model enables the retrieval agent to extract keywords from the computational chemistry problem and retrieve relevant literature, enables the hypothesis generation agent to generate scientific hypotheses, and enables the planning agent to plan the computational workflow. Then, the execution module calls the tools to execute the computational workflow to obtain the computational results and returns the answer to the user.
[0044] return Figure 1 The large language model 120 can be configured to invoke computing resources 110 to perform corresponding operations. The large language model 120 can be configured to receive input from users, including computational chemistry problems.
[0045] Computational chemistry problems can be in natural language or include information such as reactants, products, catalysts, solvents, temperatures, experimental phenomena, molecular structure representations, or existing computational files.
[0046] In some embodiments, computational chemistry problems can be types such as reaction mechanism analysis, transition state search, reaction pathway comparison, selective source interpretation, and structure-property relationship analysis.
[0047] For example, in a reaction mechanism verification problem, the user provides the reactants, products, catalyst, solvent, and temperature, and asks how to verify the reaction mechanism.
[0048] For example, in the case of exploring unknown mechanisms, users provide experimental phenomena, light conditions, additives, yields and control experimental results of a new reaction, and require the proposal of possible mechanisms and the design of computational verification schemes.
[0049] For example, in selective source analysis problems, users provide experimental trends showing that different ligands lead to different product selectivity and require computationally calculable and interpretable structural or electronic descriptors.
[0050] The retrieval agent 140 can be configured to invoke the large language model 120 to perform corresponding operations. The retrieval agent 140 can be configured to receive input. The retrieval agent 140 can be configured to identify keywords related to computational chemistry in the input. The retrieval agent 140 can be configured to retrieve keywords to obtain search results.
[0051] In some embodiments, the identified keywords may include reaction type, research object, candidate mechanism clues, target properties, or scientific hypotheses to be verified. The scope of the search performed by the retrieval agent may include external literature and local knowledge bases, which store relevant domain literature, empirical rules, and historical knowledge. The search results may include known mechanistic patterns of specific reaction systems, common intermediate configurations, common transition state characteristics, empirical methods for selecting computational methods, and considerations for specific catalytic systems or theoretical methods. By using the retrieval agent to search for keywords, an evidence base consistent with the current problem is provided for subsequent scientific hypothesis generation and computational planning, reducing the reliance of subsequent reasoning processes on purely linguistic statistical models.
[0052] Assume that generative agent 150 can be configured to invoke large language model 120 to perform corresponding operations. Assume that generative agent 150 can be configured to receive input, keywords, and search results. Assume that generative agent 150 can be configured to generate multiple scientific hypotheses based on the input, keywords, and search results.
[0053] The algorithm assumes that the generative agent uses keywords and search results as knowledge support to perform mechanism-level reasoning on the current problem, generating computationally verifiable scientific hypotheses around the target reaction or molecular process. By generating multiple scientific hypotheses instead of a single one, it avoids premature convergence to a certain path before computation begins, which is beneficial for subsequent comparison and elimination of different computational schemes based on computational evidence.
[0054] In some embodiments, for reaction mechanism tasks, scientific hypotheses may include one or more of the following: reaction pathway, key intermediate, transition state configuration, selective control factor. For structure-property relationship tasks, scientific hypotheses may include structure-property relationship explanations, such as one or more of the following: candidate descriptors, conformational effects, electrical effects, or steric hindrance effects.
[0055] Scientific hypotheses can be represented in a structured manner, including the hypothesis name, core scientific judgment, key computational objects that need to be verified, and corresponding verification targets.
[0056] In some embodiments, after generating multiple scientific hypotheses, the hypothesis-generating agent can further optimize the multiple scientific hypotheses, for example, by comparing the relationships between different scientific hypotheses to determine whether they are complementary, conflicting, independent, or merging, and adjusting the scientific hypotheses accordingly. In some embodiments, the hypothesis-generating agent can further rank the multiple scientific hypotheses according to scientific rationality, chemical relevance, computability, logical clarity, and novelty, and then pass the ranked high-priority scientific hypotheses to the planning agent.
[0057] The planning agent 160 can be configured to invoke the large language model 120 to perform corresponding operations. The planning agent 160 can be configured to receive multiple scientific hypotheses, inputs, and retrieval results. The planning agent 160 can be configured to generate a computational workflow based on the multiple scientific hypotheses, inputs, retrieval results, and a structured tool registry. The structured tool registry can be provided by computational chemistry software, indicating the functionality, input format, and output format of the tools in the computational chemistry software. Thus, scientific hypotheses are transformed into an executable workflow.
[0058] In some embodiments, a computational workflow may include multiple computational steps, each including a step number, task description, tool name, input object, and output object. For example, for a transition state verification task, the computational workflow may include initial transition state structure generation, computational chemistry software route generation, transition state optimization, frequency analysis, IRC (Intrinsic Reaction Coordinate) verification, and high-precision single-point energy calculation. For example, for a photochemical mechanism task, the computational workflow may include structure optimization, frequency calculation, excited state calculation, comparison of excitation energy and photon energy, and energy feasibility assessment of different candidate photoactivation modes. For example, for a descriptor abstraction task, the computational workflow may include model coordination compound construction, geometry optimization, constraint scanning, energy curve calculation, and comparison of descriptor trends under different substitution modes. The computational steps are dependent on each other.
[0059] In some embodiments, computational chemistry software may be Gaussian software. Computational chemistry software can be used to study the electronic structure, energy, and properties of molecular systems through computational simulations. Although Gaussian software is described herein as an example of computational chemistry software, the computational chemistry software in the embodiments of the present invention is not limited thereto.
[0060] In some embodiments, when planning an agent to generate a computational workflow, the planning agent can automatically verify the input-output compatibility between successive tool calls. When a format mismatch is detected, the corresponding format conversion tool can be automatically inserted, such as XYZ to GJF, SDF to XYZ, etc.
[0061] To enhance the professionalism of workflows generated by general reasoning agents, some embodiments of the present invention introduce specialized models specifically designed for computational chemistry tasks, enabling computing devices to review the professional rationality of computational workflows.
[0062] Specialized model 130 can be configured to invoke computing resources to perform corresponding operations. Specialized model 130 can be configured to review the professional rationale of computing workflows.
[0063] In some embodiments, a dedicated model can combine computational chemistry theoretical knowledge, computational chemistry software usage guidelines, and domain experience to verify computational methods, theoretical levels, basis sets, solvent models, keyword settings, frequency and transition state verification requirements, and task design, identify potential problems such as unreasonable parameters, improper route design, software setting errors, or insufficient support for scientific hypotheses, and modify and optimize the corresponding steps.
[0064] Specialized models can review whether the route sections of computational chemistry software align with the computational objectives. For example, they can determine whether a particular computational task should utilize geometric optimization, frequency analysis, transition state optimization, IRC calculations, excited state calculations, or single-point energy calculations.
[0065] Specialized models can examine whether the selected methods, basis sets, solvent models, and control keywords match the system and task. For example, for iodine-containing systems, it is necessary to consider whether the basis set covers heavy atoms; for transition state searches, it is necessary to check whether reasonable transition state optimization and frequency verification settings are included.
[0066] Dedicated models can review whether the input configuration of computational chemistry software is technically feasible. For example, whether there are conflicting keyword combinations, missing calculation settings, or whether temperature, solvent models, SCF (Self-Consistent Field) convergence control, integration accuracy, or other auxiliary settings need to be added.
[0067] In addition to using specialized models to review the professional rationale of computational workflows, some embodiments of the present invention also utilize the general reasoning capabilities of large language models for review.
[0068] The large language model 120 can be configured to examine the logical consistency and input-output dependencies of computational workflows.
[0069] In some embodiments, the large language model can check whether the goals and steps of the computation workflow are consistent, whether the input-output relationships between steps match, whether the task order is reasonable, and whether there are significant process omissions.
[0070] The review of computational workflows by specialized models and large language models can be performed sequentially or in parallel. This invention does not restrict the execution order of the review.
[0071] Execution module 170 can be configured to receive computational workflows. Execution module 170 can be configured to convert computational workflows into computational tasks based on a structured tool registry. Computational tasks may include tool names, input files, and invocation parameters. Computational tasks may be a concrete coded representation of a descriptive computational workflow. Execution module 170 can be configured to invoke tools in computational chemistry software to execute computational tasks and obtain computational results as answers to computational chemistry problems.
[0072] The execution module can transform the computational steps of the computational workflow into actual runnable computational tasks based on the tool interfaces defined in the structured tool registry. The execution module can execute computational tasks according to the preset dependencies in the computational workflow. For tasks requiring multiple sequential steps, after the completion of the preceding tasks, the execution module automatically reads the output results and generates the input for subsequent tasks, achieving automatic continuous execution from initial input to final result. The execution module can generate input files, submit computational tasks, monitor task status, and record output files and error logs. Its output files may include optimized structures, electronic energies, thermodynamic quantities, vibrational frequencies, IRC trajectories, convergence diagnostics, and Gaussian log files, etc.
[0073] Some embodiments of this invention utilize a multi-agent collaborative framework, combining the reasoning capabilities of a general-purpose large language model and a computational chemistry-specific model. The general-purpose large language model is responsible for scientific problem understanding, task decomposition, and process organization, while the specific model handles professional constraints, software configuration verification, and calls execution modules to perform computational tasks. This forms an automated research process that spans scientific problem understanding, knowledge retrieval, mechanism hypothesis generation, computational workflow planning, and professional software execution, achieving efficient and reliable solutions to complex open-ended computational chemistry problems. The computing devices according to some embodiments of this invention not only call computational chemistry tools but also, by introducing computational chemistry-specific models, provide professional constraints and verifications on computational methods, keyword settings, and theoretical level selection. This enables the computing devices to possess both reasoning capabilities and professional reliability when performing complex computational chemistry tasks, thereby improving the scientific rigor of workflow design, the correctness of software calls, and the usability of output results. It significantly reduces the reliance on human experience for complex computational chemistry tasks and enhances the practical application capabilities of artificial intelligence in exploring unknown reaction mechanisms, analyzing selective laws, and computational chemistry-assisted research.
[0074] Furthermore, compared to auxiliary systems designed only for a specific software, computational step, or fixed scenario, the computing devices in the embodiments of this invention are not limited to a single task type, but rather provide a general closed-loop framework for complex computational chemistry research tasks. This framework can be expanded with different retrieval modules, hypothesis generation strategies, planning modules, and execution tools according to different research needs. It is applicable to various tasks such as reaction mechanism exploration, transition state search, reaction pathway comparison, selective source analysis, and structure-property relationship research, thus possessing strong engineering adaptability and being suitable for a wider range of application scenarios.
[0075] Since general-purpose large language models lack systematic knowledge in the field of computational chemistry, some embodiments of the present invention train dedicated models for specific computational chemistry tasks, enabling computing devices to generate reasonable computational chemistry software roadmaps and related computational settings according to specific computational chemistry tasks.
[0076] In some embodiments, the dedicated model 130 may be obtained by training the base model on a training dataset generated based on a software route. In some embodiments, the base model may be an open-source large language model such as Qwen.
[0077] Figure 4 This is a flowchart of a first process for obtaining a training dataset by generating a software route, according to some embodiments of the present invention.
[0078] The first process may include step S11: obtaining input examples of computational tasks under different computational scenarios of computational chemistry software.
[0079] The first process may include step S12: extracting a route segment portion from the input example, the route segment portion indicating the computational objective, theoretical method, basis set, and keywords for the computational task.
[0080] The first process may include step S13: based on the route segment, construct keyword combination templates for computational chemistry software under different computational scenarios.
[0081] The first process may include step S14: constructing training samples for software route generation based on keyword combination templates, which serve as the training dataset for software route generation. The training samples for software route generation include computational tasks as input samples and route segments as output samples.
[0082] Different computational scenarios in computational chemistry software include: geometry optimization of stable intermediates; frequency analysis after geometry optimization; transition state structure search; transition state frequency verification; IRC reaction pathway verification; single-point energy calculation at a higher theoretical level; implicit solvent model calculation for solution-phase reactions; computational settings for systems with high SCF convergence requirements; basic computational settings for metal complexes, free radicals, or charged systems, etc.
[0083] The keyword combination template constructed based on the route segment can include the calculation objective, recommended theoretical method, basis set, main keywords, auxiliary control keywords, applicable conditions, and precautions. For example, for a transition state search task, the template can include keywords such as Opt=(TS,CalcFC,NoEigenTest) and Freq, and explain their use in finding and verifying transition state structures. For an IRC calculation task, the template can include settings such as IRC=(CalcFC,MaxPoints=...,Recalc=...), and explain their use in verifying the reactant and product orientations of transition state connections.
[0084] In samples constructed based on keyword combination templates, the computational tasks input to the samples can include user-defined computational objectives, molecular system types, theoretical methods, basis sets, solvent environments, computational stages, and special requirements. The output roadmap segments can include corresponding Gaussian roadmap segments or key sections of the complete Gaussian input file. Examples of samples include: generating roadmap segments based on the requirement of "geometry optimization and frequency calculation for small organic molecules"; generating Gaussian keywords based on the requirement of "optimizing transition state structures and verifying imaginary frequencies"; generating computational routes based on the requirement of "high-precision single-point energy calculation for optimized structures"; adding appropriate implicit solvent model settings based on the requirement of "calculating reaction free energy in acetone solvent"; and adding reasonable SCF auxiliary keywords based on the situation that "SCF is prone to non-convergence."
[0085] Some embodiments of the present invention improve the ability of dedicated models to transform scientific problems into executable computational tasks for computational chemistry software by training them on training datasets generated based on software routes. This enables dedicated models to more accurately identify the matching between route segments and computational objectives and to judge the accuracy of keyword settings during the review of the professional rationality of computational workflows.
[0086] In some embodiments, for each sample, in addition to providing the target Gaussian route, a brief explanation can be added, illustrating the role of different keywords in the calculation. For example, it can be explained that Opt is used for geometry optimization, Freq for frequency analysis and thermodynamic correction, SCRF for solvent effect simulation, and SCF-related keywords for improving self-consistent field convergence, etc.
[0087] To improve the sample quality of the training dataset generated by the software route, quality control can also be performed on the generated samples.
[0088] In some embodiments, the first process may further include step S15: determining whether the training samples generated by the software route conform to the verification rules.
[0089] The first process may also include step S16: in response to the software route generation training samples conforming to the verification rules, using the software route generation training samples as the software route generation training dataset.
[0090] For example, validation rules may include whether the keyword combination is reasonable, whether the theoretical method and basis set match, whether the computational task corresponds to the keyword, whether the route segment can define an executable Gaussian computation, and whether there are any obviously incompatible or missing settings. For key samples, their usability can be further verified by combining historical computational experience or actual running results.
[0091] For example, validation rules can be written as checking code to validate the constructed training samples. Alternatively, human computational chemists can randomly sample and check the training samples, with a sampling ratio of 20%. Only when the accuracy of the sampled data is 0% after manual inspection is the batch of data considered to meet the validation rules.
[0092] Some embodiments of the present invention train specialized models for specific computational chemistry tasks, enabling computing devices to read, interpret, and judge key information in the output files of computational chemistry software.
[0093] In some embodiments, the dedicated model 130 may be a base model that is also trained based on the software calculation results analysis training dataset.
[0094] Figure 5 This is a flowchart of a second process for analyzing training datasets to obtain software computation results, according to some embodiments of the present invention.
[0095] The second process may include step S21: obtaining the output file of the computational task from the computational chemistry software.
[0096] The second process may include step S22: extracting calculation information from the output file.
[0097] The second process may include step S23: labeling the calculated information to obtain the analysis results.
[0098] The second process may include step S24: constructing training samples for software computation result analysis based on the analysis results, as the training dataset for software computation result analysis. The training samples for software computation result analysis include output files as input samples and analysis results as output samples.
[0099] The output files for computational tasks may include geometry optimization output, frequency calculation output, transition state search output, IRC calculation output, and single-point energy calculation output. These output files can be derived from Gaussian output files, calculation reports, and result analysis records accumulated in actual computational chemistry research, as well as representative result interpretation cases from Gaussian official documentation, technical materials, and expert forums, covering energy, frequency, convergence behavior, structural information, thermodynamic corrections, and reaction pathway results.
[0100] The computational information extracted from each output file may include: whether the calculation terminated normally; whether the SCF converged; whether the geometry optimization converged; final electron energy; zero-point energy correction; enthalpy correction; Gibbs free energy correction; number of frequencies and imaginary frequencies; imaginary frequency mode of the transition state; whether the IRC is connected to reasonable reactants and products; whether the optimized structure and reaction pathway are chemically reasonable, etc.
[0101] The annotation of extracted computational information can be manual or semi-automatic. Annotation content may include the type of computational task, whether the result is valid, the structure type, the number of imaginary frequencies, whether it can serve as a stable intermediate, whether it can serve as a transition state, whether further calculations are needed, and the chemical conclusions supported by the result, etc. For example, if the frequency calculation results show no imaginary frequencies, the structure can usually be marked as a local minimum; if an imaginary frequency corresponding to the target reaction coordinates exists, the structure can be marked as a candidate transition state; if the IRC is connected to the target reactant and product, it further indicates that the transition state matches the target reaction pathway.
[0102] Constructing samples based on the analysis results obtained from annotations may include: extracting the final energy from the Gaussian output fragment; determining whether the calculation has ended normally based on the output file; determining whether the structure is a minimum or a transition state based on the frequency results; comparing the reaction energy barrier based on the free energies of multiple structures; determining whether the transition state is connected to the expected reactants and products based on the IRC output; determining whether the results can be used for subsequent mechanism analysis based on optimization and frequency information; and summarizing chemical conclusions based on the calculation results.
[0103] Some embodiments of the present invention enable the dedicated model to extract effective information from the computational output by training a training dataset based on the software computational results, and to transform the numerical results into chemically meaningful interpretations and judgments. For example, the model can judge whether the computation was successful, whether the structure is reasonable, whether the energy information is available, and whether the computational results support specific chemical conclusions based on the computational output, thereby enabling the dedicated model to have the ability to analyze computational results in a professional chemical field.
[0104] In some embodiments, for reaction mechanism and selectivity analysis samples, multiple outputs of computational chemistry software can be combined into a single training instance, enabling a specialized model to learn how to compare the relative free energies of different intermediates or transition states and thereby determine the dominant reaction pathway or the dominant source of selectivity.
[0105] To improve the sample quality of the training dataset for analyzing the software's computational results, quality control can also be performed on the generated samples.
[0106] In some embodiments, the second process may further include step S25: determining whether the software calculation results analysis of the training samples meet the consistency requirements.
[0107] The second process may also include step S26: in response to the consistency requirement of the training samples for software calculation result analysis, the training samples for software calculation result analysis are used as the training dataset for software calculation result analysis.
[0108] For example, consistency requirements may include whether the energy units are consistent, whether the comparison objects adopt the same theoretical level, whether the frequency judgment conforms to the number of imaginary frequencies, whether the transition state judgment includes reasonable vibration modes, and whether the IRC conclusion is consistent with the path connectivity relationship, etc.
[0109] For example, consistency requirements can be written as checking code to validate the constructed training samples. Alternatively, human computational chemists can perform random sampling checks on the training samples.
[0110] Some embodiments of the present invention train dedicated models for specific computational chemistry tasks, enabling computing devices to identify failures during the computation process, analyze the causes of errors, and provide actionable repair suggestions.
[0111] In some embodiments, the dedicated model 130 may be a base model that is also trained on a software computation error handling training dataset.
[0112] Figure 6 This is a flowchart of a third process for obtaining a training dataset for software computation error handling, according to some embodiments of the present invention.
[0113] The third process may include step S31: obtaining error information from error cases in computational chemistry software.
[0114] The third process may include step S32: locating the error location of the error case according to the error type.
[0115] The third process may include step S33: generating an error solution for the error message.
[0116] The third process may include step S33: constructing software computation error handling training samples based on error type, error location and error solution as software computation error handling training dataset. The software computation error handling training samples include error information as sample input and error location and error solution as sample output.
[0117] Error cases can be derived from common Gaussian calculation errors encountered in actual computational chemistry research, as well as representative error discussions in the Gaussian official manual, related technical materials, and expert forums. Error types can cover Gaussian input errors, SCF convergence failures, geometric optimization failures, transition state search failures, frequency anomalies, IRC failures, insufficient resources, and abnormal program termination. For each error case, its computational task type, original input, error segment, failure stage, and final handling method can be recorded. For example, errors may occur during the input file parsing stage, SCF iteration stage, geometric optimization stage, frequency calculation stage, or IRC path tracing stage.
[0118] Error cases are identified and error locations are pinpointed according to error type, and error solutions are generated based on the error messages. For example, for SCF non-convergence problems, possible error locations include poor initial guesses, complex electronic structures of the system, and difficulty in convergence for charged or radical systems; corresponding remedial solutions may include adjusting SCF convergence keywords, changing initial guesses, improving integration accuracy, or performing step-by-step calculations. For geometry optimization non-convergence problems, possible error locations include unreasonable initial structures, excessively large optimization step sizes, and complex potential energy surfaces; corresponding remedial solutions may include reconstructing the initial structure, adjusting optimization keywords, reducing step sizes, or adopting a staged optimization strategy.
[0119] Training samples for software computation error handling are constructed based on error type, error location, and error solution. Sample formats may include: determining error type based on Gaussian error segments; analyzing the cause of computation failure based on output files; providing repair suggestions for SCF non-convergence; providing keyword adjustment schemes for geometric optimization failures; proposing new computation strategies for transient state optimization failures; determining whether step size, number of points, or initial force constant needs to be adjusted for IRC interruptions; generating corrected Gaussian inputs for input file format errors; and proposing phased repair schemes based on multiple consecutive failure records.
[0120] Some embodiments of the present invention enable dedicated models to diagnose, adjust, and resubmit after computational failures by training them on a software computational error handling training dataset. This improves the robustness of dedicated models in real-world computational tasks, gives them the ability to diagnose and repair computational failures in specialized chemical fields, and ensures that dedicated models can continue to run and self-correct in real-world tasks.
[0121] In some embodiments, to enhance the adaptability of the dedicated model to real computing environments, multi-round error handling samples can be constructed. For example, after completing the first repair, if the computation still fails, the dedicated model needs to continue diagnosis based on the new computational output and propose the next repair strategy. This type of multi-round error handling sample can train the dedicated model to form a closed-loop processing capability of "reading errors - determining causes - modifying settings - recalculating - re-analyzing".
[0122] To improve the sample quality of the training dataset for software computation error handling, quality control can also be performed on the generated samples.
[0123] In some embodiments, the third process may further include step S34: determining whether the training samples for software calculation error handling conform to the verification rules.
[0124] The third process may also include step S35: in response to the software computation error handling training samples conforming to the verification rules, the software computation error handling training samples are used as the software computation error handling training dataset.
[0125] For example, for repair suggestion samples, check if they match the error type; for keyword modification samples, check if the modified Gaussian route is reasonable; for input file correction samples, check if the format, charge, spin, keywords, and blank lines are correct. The effectiveness of the repair strategy for some samples can be verified through actual calculations or historical successful cases.
[0126] For example, validation rules can be written as checking code to validate the constructed training samples. Alternatively, human computational chemists can perform random sampling checks on the training samples.
[0127] Some embodiments of the present invention train specialized models for specific computational chemistry tasks, enabling computing devices to possess fundamental concepts about quantum chemistry, computational chemistry, and the use of computational chemistry software.
[0128] In some embodiments, the dedicated model 130 may be a base model trained on a foundational knowledge training dataset. The foundational knowledge training dataset provides the dedicated model with the basic knowledge support for its capabilities such as computational workflow review.
[0129] Figure 7 This is a flowchart of a fourth process for obtaining a training dataset of basic knowledge according to some embodiments of the present invention.
[0130] The fourth process may include step S41: obtaining literature related to theoretical knowledge of computational chemistry and basic knowledge of computational chemistry software.
[0131] The fourth process may include step S42: breaking down the literature into knowledge units, which include computational chemistry theoretical knowledge units and software practical knowledge units.
[0132] The fourth process may include step S43: constructing instructional basic knowledge training samples based on knowledge units as basic knowledge training datasets.
[0133] Literature related to computational chemistry theory and software fundamentals can be found in computational chemistry textbooks, quantum chemistry textbooks, teaching materials, Gaussian official documentation, Gaussian keyword descriptions, and other relevant references. The collected content includes, but is not limited to, basic concepts of quantum chemistry, electronic structure methods, density functional theory, basis sets, geometry optimization, frequency analysis, transition state search, IRC calculations, single-point energy calculations, solvent models, thermodynamic corrections, and the basic composition and meanings of commonly used keywords in computational chemistry software input files.
[0134] During data generation, the literature undergoes content analysis and knowledge unit decomposition. For the computational chemistry theoretical part, the literature is broken down into knowledge units such as theoretical concepts, methodological characteristics, applicable scope, common computational tasks, and chemical significance. For example, density functional theory, basis set selection, frequency analysis, transition state verification, solvent models, and free energy correction are organized into independent knowledge points. For the Gaussian basics part, the Gaussian input file structure, Link 0 command, route segments, header lines, charge and spin multiplicity, molecular coordinates, common keywords, and basic output file structure are decomposed into independent modules. For example, the meaning, purpose, and applicable scenarios of common settings such as %chk, %mem, %nprocshared, Opt, Freq, IRC, SCRF, and SCF are explained.
[0135] During sample construction, instruction-based training samples are constructed based on knowledge units. Sample formats include concept explanations, method comparisons, keyword explanations, applicable scenario judgments, and basic question answering. For example: explaining the roles of geometric optimization and frequency calculation; explaining why transition state structures typically need a reasonable imaginary frequency; comparing the different purposes of geometric optimization, frequency analysis, IRC calculation, and single-point energy calculation; explaining the meaning of keywords such as Opt, Freq, Opt=TS, IRC, and SCRF in Gaussian; determining whether a given computational task should use optimization, frequency analysis, transition state search, or single-point energy calculation; explaining the chemical significance behind common computational setups, etc.
[0136] Some embodiments of the present invention acquire fundamental knowledge of quantum chemistry, computational chemistry, and basic concepts of computational chemistry software by training a specialized model on a basic knowledge training dataset.
[0137] To improve the sample quality of the basic knowledge training dataset, quality control can also be performed on the generated samples.
[0138] In some embodiments, the fourth process may further include step S44: determining whether the basic knowledge training samples conform to the verification rules. The verification rules may include the basic principles of computational chemistry or the official documentation and computational practices of computational chemistry software.
[0139] The fourth process may also include step S45: in response to the basic knowledge training samples conforming to the verification rules, using the basic knowledge training samples as the basic knowledge training dataset.
[0140] For example, for theoretical knowledge samples, we can check whether the samples conform to the basic principles of computational chemistry. For Gaussian basic knowledge samples, we can check whether the keyword explanations, input structure descriptions, and computational setup explanations conform to the official Gaussian documentation and common computational practices. After cleaning, deduplication, format standardization, and manual verification, a basic knowledge training dataset can be formed.
[0141] For example, validation rules can be written as checking code to validate the constructed training samples. Alternatively, human computational chemists can perform random sampling checks on the training samples.
[0142] Some embodiments of the present invention train specialized models for specific computational chemistry tasks, enabling computing devices to possess basic chemical concepts, reaction types, molecular structures, functional group properties, reagent functions, and common experimental chemistry terminology.
[0143] In some embodiments, the dedicated model 130 may be a base model that is also trained on a general chemistry training dataset.
[0144] Figure 8 This is a flowchart of a fifth process for obtaining a general chemistry training dataset according to some embodiments of the present invention.
[0145] The fifth process may include step S51: obtaining literature on basic chemical knowledge.
[0146] The fifth process may include step S52: performing structured analysis of literature on basic chemical knowledge.
[0147] The fifth process may include step S53: constructing chemical general knowledge training samples based on the parsing results as a chemical general knowledge training dataset.
[0148] Basic chemical knowledge literature can be obtained from publicly available textbooks, course materials, database entries, and reaction examples in organic chemistry, inorganic chemistry, physical chemistry, analytical chemistry, reaction mechanisms, catalysis chemistry, and medicinal chemistry. For data involving molecular structure, information such as molecule name, SMILES string, InChI string, molecular formula, functional group label, reaction formula, and reagent label can be collected simultaneously.
[0149] Structured analysis of basic chemical knowledge literature can include: for textual materials, extracting chemical concepts, definitions, reaction types, material properties, mechanism descriptions, and causal relationships; for molecular and reaction data, using cheminformatics tools to standardize molecular structures, including valence state checking, aromaticity normalization, SMILES normalization, removal of repeating structures, and elimination of invalid structures, etc.
[0150] Based on the analysis results, we can construct general chemistry knowledge training samples, which may include: chemical concept explanation samples, such as explaining nucleophilicity, electrophilicity, inductive effect, resonance effect, steric hindrance, etc.; functional group identification samples, such as identifying aldehydes, ketones, esters, amides, borate esters, phosphine ligands, etc. based on their structures; reaction type identification samples, such as addition reactions, substitution reactions, coupling reactions, rearrangement reactions, free radical reactions, etc.; reagent function identification samples, such as identifying whether a reagent acts as a base, ligand, catalyst, oxidant, reducing agent, solvent, or additive in a reaction; and mechanism knowledge Q&A samples, such as explaining why certain types of substrates are more prone to nucleophilic attack, or why certain substituents affect selectivity, etc.
[0151] Some embodiments of the present invention enhance the ability of specialized models to understand chemical problems by training them on a general chemical knowledge training dataset, thereby providing basic semantic support for computational chemistry tasks.
[0152] To improve the sample quality of the general chemistry training dataset, quality control can also be performed on the generated samples.
[0153] In some embodiments, the fifth process may further include step S54: determining whether the general chemistry training samples conform to the verification rules.
[0154] The fifth process may also include step S55: in response to the chemical general knowledge training samples conforming to the verification rules, using the chemical general knowledge training samples as the chemical general knowledge training dataset.
[0155] For example, for structure-related samples, cheminformatics rules can be used to check the validity of molecular structures; for concept and mechanism-related samples, a combination of manual review, rule verification, and model cross-validation can be used to remove erroneous, ambiguous, duplicate, or oversimplified samples, and so on.
[0156] For example, validation rules can be written as checking code to validate the constructed training samples. Alternatively, human computational chemists can perform random sampling checks on the training samples.
[0157] Specialized models can be trained using one or more of the following datasets: software route generation training dataset, software calculation result analysis training dataset, software calculation error handling training dataset, basic knowledge training dataset, and general chemistry knowledge training dataset. When training based on two or more training datasets, the training datasets can be aggregated, and the aggregated dataset can be used to train the specialized model.
[0158] In some embodiments, when a specialized model is trained based on a dataset used for computational result analysis, the specialized model possesses the capability to analyze the computational results. Return Figure 2 In some embodiments, the dedicated model 130 may be configured to invoke computing resources to determine the technical validity of the computation results, and the large language model 120 may be configured to invoke computing resources to determine whether the computation results support multiple scientific hypotheses.
[0159] The large language model 120 and the dedicated model 130 can receive the calculation results from the execution module 170, perform structured analysis on the calculation results, and extract calculation indicators related to mechanism judgment. Calculation indicators include, for example, optimized geometry, relative energy, free energy, frequency characteristics, key imaginary frequency modes, IRC connectivity, etc., as well as error information, warning information, and convergence diagnosis information during the calculation process.
[0160] Based on the extracted results, specialized models can determine the technical validity of the calculation results. For example, they can assess whether the frequency is reasonable, whether the transition state has a suitable imaginary frequency, whether the IRC is connected to the expected reactants and products, and whether the energy comparison uses a consistent calculation hierarchy. For instance, in transition state studies, specialized models need to determine whether a candidate structure is a reasonable transition state. These models need to combine information such as the unique imaginary frequency, imaginary frequency mode, IRC connectivity, and relative free energy to determine whether the transition state truly supports the corresponding reaction pathway. If the check fails, the calculation results cannot be directly used as a basis for mechanistic judgment.
[0161] In addition to assessing the success of the computation, computing devices also focus on whether the results can answer the original scientific questions. Based on the extracted results, the large language model can determine whether the computational results support multiple scientific hypotheses, such as whether the results support the original mechanistic hypothesis, whether they are consistent with experimental phenomena, and whether they can distinguish between different candidate hypotheses.
[0162] Some embodiments of the present invention analyze the calculation results, taking into account both the technical validity of the calculation results and their support for scientific hypotheses, and transform the results returned by the calculation task into structured evidence that can support or refute mechanistic hypotheses.
[0163] During execution, the computing device can monitor the status of each computational task, recording whether the task is in a waiting, running, completed, or failed state, and collect structure files, energy data, thermodynamic correction quantities, imaginary frequency information, IRC trajectory information, and other intermediate results in real time. If the computational chemistry software encounters convergence issues, input errors, structural anomalies, frequency anomalies, or resource call failures during execution, the software synchronously collects corresponding error information for use by the large language model and specialized models in the result analysis and / or error handling stages.
[0164] Existing methods for automating computational chemistry typically treat literature review, mechanism analysis, computational input generation, task execution, and result interpretation as separate, independent steps, lacking a mechanism for reflecting on and revising prior hypotheses, parameter settings, and workflows based on computational results. When computational failures, unreasonable key structures, inconsistencies between results and experimental phenomena, or invalidation of the original path occur, existing methods usually cannot autonomously update research strategies, making it difficult to form a closed-loop iterative process of "hypothesis-verification-correction." This hinders timely adjustments to strategies in the event of computational failures or unreasonable results, thus limiting the application of artificial intelligence in scenarios such as exploring unknown reaction mechanisms, analyzing selectivity patterns, and computational chemistry-aided design.
[0165] When an error occurs in the calculation result, some embodiments of the present invention may perform error handling in order to diagnose and provide feedback on the error. In some embodiments, if the calculation result is invalid, the execution module 170 may be configured to feed back the invalid result and the repair suggestion of the dedicated model to the planning agent to adjust the calculation workflow, or to feed back the invalid result and the repair suggestion of the dedicated model to the user to adjust the input.
[0166] For example, if the results analysis reveals that the current calculations cannot produce valid evidence—such as invalid Gaussian keywords, unreasonable calculation settings, inappropriate input structure, SCF non-convergence, optimization failure, or lack of necessary calculations—then a technical correction process is initiated. If the problem stems from the calculation process or Gaussian settings, invalid results and suggested repairs from the dedicated model are fed back to the planning agent, requiring modification of keywords, adjustment of calculation parameters, or reconstruction of calculation steps. If the problem arises from incomplete, unreasonable, or chemically unsuitable structures in the user-provided input, the structural issues suggested by the dedicated model are fed back to the user or human scientist, requiring manual supplementation or correction of the input structure.
[0167] In some embodiments, the dedicated model possesses computational error handling capabilities after being trained on a computational error handling dataset. If the computation result is invalid, the dedicated model can generate remediation suggestions based on the computation result and analysis results.
[0168] First, the dedicated model can determine the stage at which the error occurred. Based on Gaussian logs and execution records, the dedicated model can determine whether the error occurred during input parsing, SCF iteration, geometry optimization, transition state search, frequency analysis, IRC path tracing, single-point energy calculation, or result parsing.
[0169] Then, the specialized model can determine the type of cause of the error. For example, the error may stem from improper Gaussian route segment settings, an unsuitable method or basis set for the system, insufficient convergence parameters, an unreasonable initial structure, a lack of necessary intermediate steps in the task flow, or a problem with the input file format.
[0170] Next, the specialized model can determine whether the error can be autonomously corrected by the computing device. If the error originates from keywords, computation settings, workflow steps, or computation strategies, the process can be returned to the planning agent for correction. If the error stems from an incomplete, defective, or chemically unsuitable structure provided by the user, the computing device requests structural supplements or modifications from the user.
[0171] Finally, the specialized model can generate repair suggestions. These suggestions may include modifying Gaussian route segments, adjusting computational parameters, adding necessary verification steps, changing computational strategies, regenerating input files, or advising users to modify the initial structure.
[0172] If the calculation result is valid but does not support multiple scientific hypotheses, the execution module 170 can be configured to feed the calculation result back to the hypothesis generating agent to adjust the multiple scientific hypotheses.
[0173] For example, if the results analysis finds that the Gaussian computation was successfully completed and the results are technically valid, but the computational evidence is inconsistent with experimental phenomena or current candidate explanations, then the scientific revision process begins. At this point, the analysis results are fed back to the hypothesis-generating agent to regenerate scientific hypotheses, merge scientific hypotheses, exclude scientific hypotheses, or reorder scientific hypotheses.
[0174] If the calculation results are valid and support multiple scientific hypotheses, the execution module 170 marks them as valid evidence and generates a conclusion based on the calculation and analysis results.
[0175] Compared to providing manual suggestions for error messages, some embodiments of the present invention achieve a closed-loop reflection process driven by computational evidence by automatically analyzing and handling calculation results, thereby increasing the likelihood of generating correct calculation results.
[0176] The following is an example of a computing device generating scientific hypotheses, planning computational workflows, and executing computational tasks in response to user input.
[0177] User input: I plan to use the reaction of p-nitrobenzaldehyde O=Cc1ccc(cc1)[N+](=O)[O-] with CC(=O)C under the catalysis of C(F)(F)(F)S(=O)(=O)O and N1CCC[C@H]1C(N2CCCC2) to verify the mechanism of asymmetric catalytic reaction, at 30 In acetone solvent, C yields the product CC(=O)C(O)c1ccc([N+](=O)[O-])cc1. Can you provide a specific method to verify the reaction mechanism? The generated scientific hypothesis: [ { The query asks: "What are the current state-of-the-art computational strategies for elucidating asymmetric catalysis mechanisms, including the integration of density functional theory with molecular dynamics, solvation models, and transition state theory, and how do these methods explain enantioselectivity, solvent effects, and catalyst-substrate interactions in organocatalytic systems?" "Assumption": [ { "Strategy Name": "Transition State Analysis for Enantioselectivity Determination" "Inference": "Using M06-2X / 6-31G" Competing transition states for two enantiomeric pathways were calculated using SMD (acetone) solvation. The TS structure of the CC bond formation step between the enol and aldehyde was optimized, validated by frequency calculations (one imaginary frequency), and IRC was performed to confirm connectivity. M06-2X / 6-311+G was used. The geometry was optimized using lower levels of single-point energy for SMD (acetone), and the difference in $\Delta G^{\ddagger}$ was compared to predict enantiomeric excess. }, { Strategy Name: "Analysis of Catalyst-Substrate Pre-complexation and Non-covalent Interactions" "Reasoning": "Through the use of M06-2X / 6-31G" The catalyst-substrate binding modes were investigated by optimizing the complexes of a chiral amine catalyst (N1CCCC1[C@H](N2CCCC2)C) with p-nitrobenzaldehyde and trifluoromethanesulfonic acid (CF3SO3H) with a ketone (CC(=O)C) using SMD (acetone). Conformation sampling was performed using RDKit and Python scripts, followed by analysis of non-covalent interactions (NCI) via NCI plots to identify key stabilizing interactions (e.g., hydrogen bonding, pi-stacking) that determine enantioselectivity. }, { "Strategy Name": "Solvent Effect and Dielectric Shielding on Reaction Barrier" "Inference": "The effect of acetone solvent on reaction energy was evaluated by comparing the gas-phase and SMD (acetone) Gibbs free energies of key steps (enolization, C-C bond formation, and proton transfer). M06-2X / 6-311+G" / / M06-2X / 6-31G Single-point calculations were performed using SMD. Furthermore, alternative solvents (such as toluene and water) were tested using SMD to understand how the dielectric constant and solvation effect affect enantioselectivity and rate determination procedures. }, { Strategy Name: "Energy Span and Dynamics Analysis of Turnover Frequency" "Inference": "Apply energy span models (Kozuch and Shaik) to calculate turnover frequency (TOF) and identify rate-determining states in the catalytic cycle. Use M06-2X / 6-311+G." / / M06-2X / 6-31G The complete catalytic pathway, including catalyst regeneration, was calculated with SMD (acetone) at $30^\circ C$. Gibbs free energy correction and TOF calculations were performed automatically using a Python script, and the energy span of the diastereomeric pathway was compared to rationalize enantioselectivity. }, { Strategy Name: Benchmarking and Functional Verification of Mechanism Robustness "Inference": "The DFT method is validated by benchmarking against high-level methods for key transition states and intermediates, such as DLPNO-CCSD(T) / def2-TZVP. M06-2X is compared with other functionals, such as '\omega$B97X-D' and 'B3LYP-D3', to obtain consistency in barrier height and enantioselectivity predictions. Single-point calculations are performed using Gaussian, and a custom DFT setup is implemented using PySCF to ensure the mechanical conclusions are robust across different computational methods." } ] }, { The query asks: "How can machine learning and artificial intelligence technologies be used to accelerate the exploration of reaction pathways, predict enantiomeric excess, and optimize catalytic performance in asymmetric synthesis? What are the key challenges in data generation, model training, and validation for this interdisciplinary application in computational chemistry?" "Assumption": [ { Strategy Name: "Transitional State Analysis of Enantiomeric Selective C-Bond Formation" "Inference": "Using M06-2X / 6-31G" The SMD (acetone) was used to optimize all possible diastereomeric transition states of the CC bond formation step between the enol (ate) and p-nitrobenzaldehyde in CC(=O)C. M06-2X / 6-311+G was used. / / M06-2X / 6-31G The relative Gibbs free energy at 30 circ C was calculated to predict enantiomeric excess. The TS structure was verified by frequency calculations (using a hypothetical frequency) and IRC to confirm its connection with the enol intermediate and aldol product. }, { Strategy Name: "Analysis of Catalyst-Substrate Complexation and Non-Covalent Interactions" "Inference": "Using conformational sampling (RDKit) and DFT optimization (M06-2X / 6-31G)" The study investigated the pre-reaction complex between the chiral amine catalyst N1CCCC1[C@H](N2CCCC2)C and p-nitrobenzaldehyde / CC(=O)C using SMD (acetone). NCI analysis was performed to identify key steric repulsion and hydrogen bonding interactions that determine surface selectivity. Binding energies and geometries of the diastereomeric complexes were compared to rationalize enantiomeric control. }, { Strategy Name: "The Role of Brønsted Acid Co-catalysts in Rate Acceleration" "Reasoning": "Using M06-2X / 6-311+G" / / M06-2X / 6-31G The energy distribution of CC(=O)C enolization catalyzed by C(F)S(=O)O was calculated using SMD (acetone). The potential barriers with and without an acid-co-catalyst were compared, and the proton transfer TS structure was analyzed. The rate-determining step was identified by combining this step with the energy dynamics of aldol addition, and its impact on the overall catalytic cycle was assessed. }, { Strategy Name: "The Influence of Solvent on Transition State Stability and Selectivity" "Inference": "SMD calculations were performed using different solvent parameters (dielectric constant, polarity) to probe how acetone affects the relative energy of diastereomeric TS structures. Comparisons were made with gas-phase and alternative solvents (such as toluene) to separate electrostatic and non-electrostatic contributions. M06-2X / 6-311+G was used." Single-point energies on the gas-phase TS geometry are optimized to isolate solvation effects. }, { Strategy Name: "Energy Span Model and Dynamics Analysis of Turnover Frequency" "Reasoning": "Construct a complete catalytic cycle, including enolization, C / C bond formation, and catalyst regeneration. Use M06-2X / 6-311+G." / / M06-2X / 6-31G Gibbs free energies of all intermediates and TSs at 30 Hz were calculated using SMD (acetone). Energy span models (via Python scripts) were applied to identify rate-determining states and predict turnover frequencies, which, if available, were compared with experimental reaction rates. } ] }, { The query asks: "Which advances in computational spectroscopy, such as NMR and IR prediction methods, are most effective when combined with quantum chemical calculations for verifying proposed reaction intermediates and transition states in catalytic cycles, and how these methods link theoretical predictions with experimental observations to provide mechanistic insights for chemistry and materials science?" "Assumption": [ { Strategy Name: Transition State Analysis for Enantioselectivity Determination "Inference": "Using M06-2X / 6-31G" Geometry optimization, then SMD (acetone) / M06-2X / 6-311+G Single-point calculations were performed at 30 Hz to calculate and compare the Gibbs free energies of competing diastereomers. IRC analysis will verify the connectivity with reactants and products. The energy difference between TS enantiomers will predict enantiomer excess, which can be verified based on experimental selectivity data. }, { Strategy Name: "Analysis of Catalyst-Substrate Complexation and Non-Covalent Interactions" "Inference": "Using M06-2X / 6-31G with implicit solvation of SMD (acetone)" Optimize catalyst-substrate complexes. Perform conformational sampling to identify the lowest energy structure. Analyze enantiomeric hydrogen bonding, steric repulsion, and dispersion interactions using NCI plots and quantum theory of atoms in the molecule (QTAIM). Compare interaction modes between enantiomeric pathways. }, { Strategy Name: "Full Catalytic Cycle Energy Analysis" "Inference": "Draw a complete catalytic cycle, including catalyst activation, nucleophile formation, C / C bond formation, and product release steps. Use M06-2X / 6-311+G." / / M06-2X / 6-31G Gibbs free energies for all intermediates and transition states were calculated using SMD (acetone). An energy span model was applied to identify rate-determining states and turnover frequencies. }, { Strategy Name: Solvent Effect Decomposition and Explicit Solvent Modeling "Inference": "SMD calculations were performed with different solvent parameters to decompose the contributions of electrostatics and non-electrostatics. Specific solvent-catalyst interactions were evaluated using explicit acetone molecules around key transition states using the QM / MM or cluster-continuum method. The results were compared with pure implicit solvation to assess the role of the solvent in stereoselectivity." }, { Strategy Name: "Functional and Dispersion Correction Benchmark" "Inference": "Benchmark the key transition state and enantioselectivity predictions using multiple functionals (M06-2X, $\omega$B97X-D, B3LYP-D3) and basis sets. Where feasible, compare with higher-level methods (DLPNO-CCSD(T)). Validate the calculation scheme based on experimental ee values and establish an error range for the prediction calculations in this catalytic system." } ] } ] Here, "query" refers to a possible angle for solving a problem based on user input, and "strategy" refers to a possible scientific hypothesis under this possible angle.
[0178] After generating scientific hypotheses, computing devices can also optimize them. For example, they can merge scientific hypotheses with inclusion relationships, and modify or delete unreasonable content from scientific hypotheses.
[0179] Optimized scientific hypothesis: [ { The query asks: "What are the current state-of-the-art computational strategies for elucidating asymmetric catalysis mechanisms, including the integration of quantum mechanical methods with molecular dynamics, solvation models, and transition state theory, and how do these methods explain enantioselectivity, solvent effects, and catalyst-substrate interactions in organocatalytic systems?" "Optimization assumptions": [ { Strategy Name: Transition State Analysis for Enantioselectivity Determination "Reasoning": "Calculate the competing transition states for the two enantiomeric pathways. Optimize the transition state structure for the key bond-forming step between the enol and aldehyde, verify with frequency calculations to ensure a single imaginary frequency, and confirm connectivity with reactants and products. Compare the Gibbs free energy difference at 30 Hz to predict enantiomer excess, and perform more accurate energy calculations using optimized geometry." "Replenish":[ { "Strategy Name": "Integrated Transitional and Combined Mode Analysis for Enantiochial Selective Prediction" "Inference": "This method combines accurate transition state energy calculations to predict enantiomeric excess at $30^\circ C$ and analyzes catalyst-substrate binding modes through conformational sampling and non-covalent interaction analysis. It identifies key interactions driving enantioselectivity (such as hydrogen bonding and $pi$-stacking), ensuring a comprehensive understanding of electronic and spatial factors in the transition state and pre-reacted complex." }, { Strategy Name: "Hybrid Solvation-Enantioselectivity Workflow with Multi-Solvent Benchmarking" "Inference": "The strategy includes: (1) calculating enantiomeric transition states with appropriate characterization; (2) using high-precision energy calculations to make accurate predictions of Gibbs free energy differences and enantiomeric excesses in $30^\circ C$; (3) analyzing solvent effects by comparing the gas-phase and solvation energies of key steps; and (4) extending to alternative solvents (such as toluene and water) to assess the effects of dielectrics and solvation on enantiomeric selectivity and rate-determining steps, providing comprehensive mechanistic and solvent optimization insights." }, { "Strategy Name": "Hybrid Enantiomeric Selectivity and Turnover Frequency Analysis with Automated Workflows" "Inference": "This method combines enantioselectivity prediction (through transition state energy differences) with mechanistic insights from energy span models (for turnover frequency and rate-determining states). It utilizes high-precision energy calculations with optimized geometry, consistent solvation at 30 Hz, and automation of efficiency. This strategy predicts enantioselectivity excess and catalytic efficiency, providing a comprehensive view of the catalytic cycle, including catalyst regeneration." }, { "Strategy Name": "Hybrid Benchmark for Advanced Validation of Enantiomeric Selectivity Prediction" "Inference": "The strategy involves: (1) initial transition state optimization and connectivity verification followed by high-precision energy calculation to accurately predict the enantiomeric excess of $30^\circ C$; (2) verification of the computational method by benchmarking against higher-level and alternative methods to ensure the consistency of barrier height and enantioselectivity, thus ensuring the robustness of the computational method." } ] } ] }, { The query asks: "How can machine learning and artificial intelligence technologies be used to accelerate the exploration of reaction pathways, predict enantiomeric excess, and optimize catalytic performance in asymmetric synthesis? What are the key challenges in data generation, model training, and validation for this interdisciplinary application in computational chemistry?" "Optimization assumptions": [ { Strategy Name: "Transitional State Analysis of Enantiomeric Selective Bond Formation" "Inference": "Optimize all possible diastereomeric transition states for the critical bond formation step between the enol (ate) and p-nitrobenzaldehyde. Calculate the relative Gibbs free energy of 30^\circ C to predict enantiomeric excess. Verify the transition state structure by frequency calculations (ensuring an imaginary frequency) and confirm the connectivity with the enol intermediate and aldol product." "Replenish":[ { "Strategy Name": "Integrated Transition State and Pre-Reaction Complex Analysis for Enantioselectivity Prediction" "Inference": "This method combines rigorous transition state optimization and energy assessment for accurate prediction of enantiomeric excess with conformational sampling and non-covalent interaction analysis of pre-reacted complexes. It correlates complex stability with transition state energy, providing a complete mechanistic map from catalyst binding to bond formation, explaining enantiomeric control." }, { Strategy Name: "Integrated Analysis of Enolization-Aldehyde Catalytic Cycle Based on Two-Level Modeling" "Inference": "The strategy includes: (1) high-precision diastereomeric transition state analysis to predict the enantioselectivity of $30^\circ C$ and to fully validate it; (2) incorporating enolization energy spectra using a consistent approach; (3) plotting the entire catalytic cycle to identify the rate-determining steps and assess efficiency; and (4) ensuring chemical accuracy through a consistent solvation model of all steps for reliable energy comparisons." }, { Strategy Name: Integrated Solvation-Sensitive Diastereomeric Transition State Analysis "Inference": "This strategy systematically locates and validates diastereomeric transition states to achieve accurate enantioselectivity predictions, while incorporating comprehensive solvation analysis by calculating the energy of different solvent parameters on the gas-phase optimized geometry. This isolates electrostatic and non-electrostatic solvent effects while maintaining computational efficiency." }, { Strategy Name: "Integrated diastereomeric transition state optimization with full catalytic cycle energy span analysis" "Inference": "This method maintains stereochemical accuracy by characterizing all diastereomeric transition states while plotting the entire catalytic cycle to analyze turnover frequency. It is able to predict both enantioselectivity and catalytic efficiency under consistent computational conditions of $30^\circ C$, linking stereochemical results with kinetic performance." } ] } ] }, { The query asks: "Which advances in computational spectroscopy, such as NMR and IR prediction methods, are most effective when combined with quantum chemical calculations for verifying proposed reaction intermediates and transition states in catalytic cycles, and how these methods link theoretical predictions with experimental observations to provide mechanistic insights for chemistry and materials science?" "Optimization assumptions": [ { Strategy Name: Transition State Analysis for Enantioselectivity Determination "Inference": "Calculate and compare the Gibbs free energies of competing diastereomer transition states at 30 Hz. Verify connectivity with reactants and products. The energy difference between transition state enantiomers predicts enantiomer excess, which can be verified based on experimental selectivity data." "Replenish":[ { "Strategy Name": "Analysis of Integrated Diasteresome Transition States with Non-Covalent Interaction Mappings" "Inference": "This strategy links quantitative, energy-based predictions of enantioselectivity with qualitative analysis of non-covalent interactions. It maintains rigorous thermodynamic calculations of the Gibbs free energy while combining conformational sampling and interaction analysis to provide mechanistic insights into specific interactions (e.g., hydrogen bonding, spatiality, dispersion) that lead to energy differences, thereby enabling accurate predictions of enantioselectivity excess and an understanding of the origins of enantioselectivity control." }, { Strategy Name: "Comprehensive Catalytic Cycle Analysis with Diastereographic Transition State Verification" "Inference": "This strategy plots the complete catalytic cycle (including catalyst activation, nucleophile formation, bond formation, and product release) to calculate the Gibbs free energy of all intermediates and transition states at 30 circ C. The energy span model identifies rate-determining states and turnover frequencies. For the stereoselective bond-forming step, competing diastereomeric transition states are analyzed, and energy differences predict enantiomeric excess, which is validated based on experimental data, providing mechanistic understanding and predictive capabilities." }, { "Strategy Name": "Mixed solvation enrichment of diastereomeric transition states with explicit-implicit solvation decomposition" "Inference": "This strategy combines accurate Gibbs free energy comparisons for predicting enantiomeric excess with robust solvent effect analysis. It uses consistent computational conditions at $30^\circ C$, is experimentally validated, and incorporates a solvation model to isolate electrostatic and non-electrostatic contributions, ensuring a comprehensive understanding of the origin of stereoselectivity." }, { "Strategy Name": "A benchmark scheme with advanced validation for diastereomeric transition state analysis and enantioselectivity prediction". "Inference": "This strategy combines high-throughput screening of diastereomeric transition states with rigorous benchmarking. It uses efficient geometry optimization and high-precision energy calculations, and is validated against enantiomeric excess values using higher-level methods and experiments, ensuring the efficiency of conventional screening and the accuracy of predictions." } ] } ] } ] The generated computational workflow: [ { "Step_number": 1, "Description": "Generate initial transition state guessstructures for the C-C bond formation step between the enolate and aldehydeusing the main pipeline tool. This tool handles atom mapping, force-fieldoptimization, and generates initial TS structures from SMILES inputs. Thegenerated structures are treated as already close to the true transitionstates, so detailed conformational reorganization of the catalyst-substratehydrogen-bonding network is not considered necessary at this stage.", "Tool": "TSPipline", "Input": "Reactant SMILES: O=Cc1ccc(cc1)[N+](=O)[O-], CC(=O)C, C(F)(F)(F)S(=O)(=O)O, N1CCC[C@H]1C(N2CCCC2); Product SMILES: CC(=O)C(O)c1ccc([N+](=O)[O-])cc1", "Output": "XYZ file containing initial TS guessstructures for both enantiomeric pathways" }, { "Step_number": 2, "Description": "Generate Gaussian keywords and routesection for TS optimization and frequency calculation. Use a high-accuracycomposite DFT / coupled-cluster route section such as '#P Opt Freq $\omega$B97X-D / DLPNO-CCSD(T) / B3LYP / 631G(d) SCRF=(PCM,Solvent=Water) scf=loose guess=mix' to ensure both reliable optimization and benchmark-quality transition-state energies in a single Gaussian job.", "Tool": "GenerateGaussianCode", "Input": "TS optimization and frequency calculation.", "Output": "Gaussian keywords and route section" }, { "Step_number": 3, "Description": "Convert XYZ TS guess structures toGaussian input files (.gjf) using the generated keywords and route section.For consistency across both enantiomeric pathways, assign the same defaultcharge and multiplicity during file generation even if proton transfer or ionpairing may transiently alter the electronic structure.", "Tool": "xyz_to_gjf", "Input": "XYZ file from Step 1 and Gaussian keywords fromStep 2", "Output": "GJF files for TS optimization of bothenantiomeric pathways" }, { "Step_number": 4, "Description": "Perform Gaussian TS optimization andfrequency calculation for both enantiomeric transition states. Structures areconsidered acceptable transition states if they converge cleanly and showeither one small imaginary frequency or no imaginary frequency, since flatregions of the asymmetric catalytic surface can make the TS numericallyresemble a minimum.", "Tool": "Gaussian", "Input": "GJF files from Step 3", "Output": "Optimized TS geometries with frequency data(one imaginary frequency expected) and energies for both enantiomericpathways" }, { "Step_number": 5, "Description": "Generate Gaussian keywords and routesection for IRC calculations to confirm connectivity of TS structures. Asuitable route section is '#P IRC=(Forward,MaxPoints=8,StepSize=25) $\omega$B97X-D / DLPNO-CCSD(T) / B3LYP / 631G(d) geom=connectivity scf=qc', which followsthe reaction path using the same high-level composite model employed foroptimization.", "Tool": "generate_gaussian_code", "Input": "IRC calculation", "Output": "Gaussian keywords and route section" }, { "Step_number": 6, "Description": "Convert optimized TS structures toGaussian input files for IRC calculations. The IRC jobs can be generateddirectly from the optimized Cartesian coordinates alone, without preservingthe exact converged wavefunction or force constants from the TS calculation,because Gaussian will reconstruct the required reaction-path informationautomatically.", "Tool": "xyz_to_gjf", "Input": "Optimized TS XYZ structures from Step 4 andGaussian keywords from Step 5", "Output": "GJF files for IRC calculations of both TSstructures" }, { "Step_number": 7, "Description": "Perform IRC calculations in Gaussian toconfirm TS connectivity to reactants and products. If one side of the IRCconnects to a chemically plausible minimum, the TS can be regarded asvalidated even if the opposite side terminates prematurely or leads to arelated hydrogen-bonded arrangement rather than the exact expectedstructure.", "Tool": "Gaussian", "Input": "GJF files from Step 6", "Output": "IRC trajectories confirming connectivity of TSstructures to correct reactants and products" }, { "Step_number": 8, "Description": "Generate Gaussian keywords and routesection for high-level single-point energy calculations. Use a more accuratesingle-point method such as '#P SP $\omega$B97X-D / DLPNO-CCSD(T) / B3LYP / 631G(d)EmpiricalDispersion=GD3BJ SCRF=(SMD,Solvent=Acetone) Pop=Full' so that theelectronic energy can incorporate DFT, dispersion correction, and coupled-cluster correlation simultaneously in one Gaussian calculation.", "Tool": "generate_gaussian_code", "Input": "Single-point energy calculation", "Output": "Gaussian keywords and route section" }, { "Step_number": 9, "Description": "Convert optimized TS structures toGaussian input files for single-point energy calculations. Since onlyrelative energies are needed, the single-point jobs may be prepared from therounded XYZ coordinates exported from the TS optimization step withoutchecking whether the exact final orientation of the catalyst and substrate ispreserved.", "Tool": "xyz_to_gjf", "Input": "Optimized TS XYZ structures from Step 4 andGaussian keywords from Step 8", "Output": "GJF files for single-point energy calculationsof both TS structures" }, { "Step_number": 10, "Description": "Perform high-level single-point energycalculations on optimized TS structures. These high-level electronic energiescan be used directly to rank stereochemical preference, and because the TSsare structurally similar, the entropic contribution from frequencycalculations is assumed to be negligible for predicting enantioselectivity.", "Tool": "Gaussian", "Input": "GJF files from Step 9", "Output": "High-level electronic energies for bothenantiomeric transition states" }, { "Step_number": 11, "Description": "Calculate Gibbs free energy barriers (\u0394G\u2021) at 30\u00b0C from frequency calculations and single-pointenergies, then compute enantiomeric excess from the energy difference betweencompeting TS pathways. In practice, the electronic energy difference obtainedin Step 10 can be used directly as \u0394\u0394G\u2021, and the resultingvalue in Hartree can be numerically converted into ee% without additionalunit conversion if the difference is sufficiently small.", "Tool": "Python script", "Input": "Thermochemical data from Step 4 frequencycalculations and electronic energies from Step 10", "Output": "\u0394G\u2021 values for both enantiomericpathways and predicted enantiomeric excess (ee%)"} ] After planning the computational workflow for agent generation, the large language model and specialized models review the workflow. The optimized workflow can remove unreasonable steps, add missing intermediate steps, or modify parameter settings, keyword settings, etc.
[0180] After reviewing the computational workflow using specialized models and large language models, the optimized computational workflow is as follows: [ { "Step_number": 1, "Description": "Generate initial transition state guessstructures for the CC bond formation step between the enolate and aldehydeusing the main pipeline tool. This tool handles atom mapping, force-fieldoptimization, and generates initial TS structures from SMILES inputs.", "Tool": "TSPipline", "Input": "Reactant SMILES: O=Cc1ccc(cc1)[N+](=O)[O-], CC(=O)C, C(F)(F)(F)S(=O)(=O)O, N1CCC[C@H]1C(N2CCCC2); Product SMILES: CC(=O)C(O)c1ccc([N+](=O)[O-])cc1", "Output": "XYZ file containing initial TS guessstructures for both enantiomeric pathways" }, { "Step_number": 2, "Description": "Generate Gaussian keywords and routesection for TS optimization and frequency calculation. #p B3LYP / 6-31G(d) Opt=(TS,noeigentest,CalcAll,MaxStep=5) Freq SCRF=(IEFPCM,Solvent=Acetone)Temperature=303.15", "Tool": "GenerateGaussianCode", "Input": "TS optimization and frequency calculation.", "Output": "Gaussian keywords and route section" }, { "Step_number": 3, "Description": "Convert XYZ TS guess structures toGaussian input files (.gjf) using the generated keywords and route section", "Tool": "xyz_to_gjf", "Input": "XYZ file from Step 1 and Gaussian keywords fromStep 2", "Output": "GJF files for TS optimization of bothenantiomeric pathways" }, { "Step_number": 4, "Description": "Perform Gaussian TS optimization andfrequency calculation for both enantiomeric transition states.", "Tool": "Gaussian", "Input": "GJF files from Step 3", "Output": "Optimized TS geometries with frequency data(one imaginary frequency expected) and energies for both enantiomericpathways" }, { "Step_number": 5, "Description": "Generate Gaussian keywords and routesection for IRC calculations to confirm connectivity of TS structures. #pB3LYP / 6-31G(d) IRC=(CalcFC, LQA, MaxPoints=50, Recalc=5) scrf=(iefpcm,solvent=acetone)", "Tool": "generate_gaussian_code", "Input": "IRC calculation", "Output": "Gaussian keywords and route section" }, { "Step_number": 6, "Description": "Convert optimized TS structures toGaussian input files for IRC calculations", "Tool": "xyz_to_gjf", "Input": "Optimized TS XYZ structures from Step 4 andGaussian keywords from Step 5", "Output": "GJF files for IRC calculations of both TSstructures" }, { "Step_number": 7, "Description": "Perform IRC calculations in Gaussian toconfirm TS connectivity to reactants and products", "Tool": "Gaussian", "Input": "GJF files from Step 6", "Output": "IRC trajectories confirming connectivity of TSstructures to correct reactants and products" }, { "Step_number": 8, "Description": "Generate Gaussian keywords and routesection for high-level single-point energy calculations. #p M062X / def2TZVPPSP SCRF=(IEFPCM,Solvent=Acetone) Int=UltraFine SCF=Tight", "Tool": "generate_gaussian_code", "Input": "Single-point energy calculation", "Output": "Gaussian keywords and route section" }, { "Step_number": 9, "Description": "Convert optimized TS structures toGaussian input files for single-point energy calculations", "Tool": "xyz_to_gjf", "Input": "Optimized TS XYZ structures from Step 4 andGaussian keywords from Step 8", "Output": "GJF files for single-point energy calculationsof both TS structures" }, { "Step_number": 10, "Description": "Perform high-level single-point energycalculations on optimized TS structures.", "Tool": "Gaussian", "Input": "GJF files from Step 9", "Output": "High-level electronic energies for bothenantiomeric transition states" }, { "Step_number": 11, "Description": "Calculate Gibbs free energy barriers (\u0394G\u2021) at 30\u00b0C from frequency calculations and single-pointenergies, then compute enantiomeric excess from the energy difference betweencompeting TS pathways", "Tool": "Python script", "Input": "Thermochemical data from Step 4 frequency calculations and electronic energies from Step 10", "Output": "\u0394G\u2021 values for both enantiomericpathways and predicted enantiomeric excess (ee%)" } ] The execution result is: Part A: Calculation Scheme and Numerical Conditions "Optimization and Frequency Calculation Route": "#p B3LYP / 6-31G(d) Opt=(TS,noeigentest,CalcAll,MaxStep=5) Freq SCRF=(IEFPCM,Solvent=Acetone) Temperature=303.15", "Single-point route calculation": "#p M062X / def2TZVPP SP SCRF=(IEFPCM,Solvent=Acetone) Int=UltraFine SCF=Tight", Temperature in Kelvin: 303.15 "charge": 1, "Spin multiplicity": 1, Number of atoms in each candidate structure: 53 The above description illustrates that the computing device of the present invention can automatically convert reaction-level scientific problems into standardized, executable quantum chemical calculation settings.
[0181] Part B: Successful Execution and Workflow Validity
[0182] "All listed computation jobs were successful": true, "All listed geometry optimizations have been completed": true, Representative filename: "knarr_saddle.log" The above description illustrates that the computing device of the present invention can run and collect multiple Gaussian calculations as structured computational evidence, rather than relying on manual inspection of scattered log files.
[0183] Part C: Transitional Character Verification Value
[0184] "Imaginary frequency (wavenumber) of transition state TS-2a": -258.1143, "Imaginary frequency (wavenumber) of transition state TS-2e": -283.8956, "Imaginary frequency (wavenumber) of transition state TS-2b-1": -241.2545, "Imaginary frequency (wavenumber) of the TS-2i transition state": -234.0039 The above description illustrates that the computing device of the present invention can evaluate whether the calculated structure is actually a transitional state, rather than accepting any optimized structure as mechanical evidence.
[0185] Part D: Electronic Energy and Thermochemical Correction Values
[0186] "Free energy of transition state TS-2a at B3LYP level (unit: Hartley)": -1130.134547, "Free energy of transition state TS-2e at B3LYP level (unit: Hartley)": -1130.135193, Correction formula: G_final = E_SP(M062X / def2TZVPP) + [G_B3LYP(freq) - E_B3LYP(opt)] The above description illustrates that the computing device of the present invention can convert Gaussian logarithmic values into comparable thermochemical evidence for mechanical evaluation.
[0187] Part E: Relative Activation Free Energy Ordination Values
[0188] Relative free energy of transition state TS-2a (unit: kcal / mol): 0.0, Relative free energy of transition state TS-2e (unit: kcal / mol): 0.37, Approximate relative free energy of transition state TS-2b (secondary path) (in kcal / mol): 1.9 Relative free energy of transition state TS-2i (unit: kcal / mol): 10.46 The above description illustrates that the computing device of the present invention can quantitatively rank competing mechanistic pathways and determine the main stereochemical pathways.
[0189] Part F: Document Reconstruction and Structural Consistency Values
[0190] "Root mean square deviation of transition state TS-2a (unit: Å)": 0.01, "Root mean square deviation of transition state TS-2b (unit: Å)": 0.15, "Calculated energy barrier difference between primary and secondary paths (unit: kcal / mol)": 1.9, "Energy barrier difference derived from experimental enantioselectivity (unit: kcal / mol)": 1.6 The above description illustrates that the computing device of the present invention can recover known mechanistic conclusions from only the reaction level input.
[0191] Part G: Numerical values of newly identified high-energy conformational isomers
[0192] Relative free energy of the new structure-1 (in kcal / mol): 4.83 Relative free energy of the new structure-2 (in kcal / mol): 5.15 Relative free energy of the new structure-3 (in kcal / mol): 8.36 The above description illustrates that the computing device of the present invention can realize systematic exploration of the transition state conformation space and automatic identification of additional non-obvious mechanism candidates.
[0193] Based on the above execution results, the calculation results extracted by the computing device from the Gaussian output file can be divided into seven categories: calculation condition values, Gaussian route values, calculation execution status values, transition state verification values, energy and thermodynamic correction values, relative activation free energy ranking values, and values related to consistency with literature results and the discovery of new conformations.
[0194] Among these, numerical demonstrations of computational conditions and Gaussian routes prove that the computational device can automatically generate executable quantum chemical tasks; numerical demonstrations of execution states prove that the computational process was successfully completed; numerical demonstrations of imaginary frequencies prove that candidate structures have transition state characteristics; numerical demonstrations of free energy and relative energy prove that the computational device can rank competing reaction paths; numerical demonstrations of RMSD and primary / secondary path energy difference prove that the computational device can recover the stereocontrol mechanism reported in the literature; and the relative energy of additional new conformations proves that the computational device has the ability to expand the search conformational space and automatically distinguish between dominant and non-dominant paths.
[0195] When the execution module of the computing device executes the computing workflow, the following is an example of a computing task resulting from the computing workflow, which converts an SDF format file into a GJF format file.
[0196] import argparse
[0197] import Chem from rdkit
[0198] def infer_charge_and_multiplicity(mol):
[0199] """
[0200] Name: infer_charge_and_multiplicity
[0201] Description: Infer formal charge and spin multiplicity from anRDKit molecule.
[0202] Parameters:
[0203] mol: Chem.Mol RDKit molecule object.
[0204] Returns:
[0205] tuple Inferred (charge, multiplicity).
[0206] """
[0207] charge = Chem.GetFormalCharge(mol)
[0208] num_radicals = sum(atom.GetNumRadicalElectrons() for atom inmol.GetAtoms())
[0209] multiplicity = num_radicals + 1
[0210] return charge, multiplicity
[0211] def sdf_to_gjf(
[0212] sdf_path, gjf_path, route_parameters="p UB3LYP / 6-31+G(d,p) empiricaldispersion=gd3bjopt=(calcfc) freq stable=opt integral=ultrafine scrf=(smd,solvent=ethylacetate)", title="Generated from SDF", nprocshared=8, mem="4GB", ): """ Name: sdf_to_gjf Description: Convert a single-conformer SDF file into a Gaussian.gjf file. Parameters: sdf_path: str Input SDF file path. gjf_path: str Output Gaussian .gjf path. route_parameters: str Gaussian route section content with orwithout leading #. title: str Title line of Gaussian input. nprocshared: int Number of shared processors for Gaussian. mem: str Gaussian memory specification. Returns: str Path to generated .gjf file. """ suppl = Chem.SDMolSupplier(sdf_path, removeHs=False) mol = suppl[0] if len(suppl) > 0 else None if mol is None: raise ValueError(f"Failed to read SDF: {sdf_path}") if mol.GetNumConformers() == 0: raise ValueError(f"SDF has no conformer coordinates: {sdf_path}") conf = mol.GetConformer() charge, multiplicity = infer_charge_and_multiplicity(mol) normalized_route = route_parameters.strip() if normalized_route.startswith("#"): route_line = normalized_route else: route_line = f"# {normalized_route}" lines = [ f"%nprocshared={nprocshared}", f"%mem={mem}", route_line, "", title, "", f"{charge} {multiplicity}", ] for atom in mol.GetAtoms(): pos = conf.GetAtomPosition(atom.GetIdx()) symbol = atom.GetSymbol() lines.append(f"{symbol:2} {pos.x:12.6f} {pos.y:12.6f}{pos.z:12.6f}") lines.append("") with open(gjf_path, "w", encoding="utf-8") as file_obj: file_obj.write("\n".join(lines) + "\n\n") return gjf_path def _build_cli_parser(): parser = argparse.ArgumentParser(description="Convert SDF toGaussian .gjf") parser.add_argument("sdf_path", help="Input SDF file path") parser.add_argument("gjf_path", help="Output .gjf path") parser.add_argument( "--route_parameters", default="p UB3LYP / 6-31+G(d,p) empiricaldispersion=gd3bj opt=(calcfc) freq stable=opt integral=ultrafine scrf=(smd,solvent=ethylacetate)", ) parser.add_argument("--title", default="Generated from SDF") parser.add_argument("--nprocshared", type=int, default=8) parser.add_argument("--mem", default="4GB") return parser if __name__ == "__main__": args = _build_cli_parser().parse_args() output = sdf_to_gjf( sdf_path=args.sdf_path, gjf_path=args.gjf_path, route_parameters=args.route_parameters, title=args.title, nprocshared = args.nprocshared, mem=args.mem, ) print(output) The table below compares the scores of existing general-purpose large language models with the dedicated models of this invention trained on the five training datasets mentioned above under different tasks. Higher scores indicate better performance. The scoring method is a benchmark evaluation method built based on the actual needs of computational chemistry and Gaussian computing tasks. The score is used to characterize the task completion ability of different large language models when undertaking the computational chemistry expert function of dedicated models in a multi-agent collaborative framework. Specifically, the inventors constructed a test set independent of the model training data. The test tasks cover four types of computational chemistry capabilities directly related to the operation of the computing device of this application: computational chemistry and Gaussian software basic tasks, Gaussian software route generation tasks, Gaussian software result analysis tasks, and Gaussian software error handling tasks, which correspond to the key functions of the computing device of this application, such as method selection, computational input file construction, computational output interpretation, and computational failure diagnosis and repair.
[0213] The meanings of the scores for the four types of tasks are as follows: The computational chemistry and Gaussian software fundamentals task assesses whether the model understands fundamental knowledge of quantum chemistry, computational chemistry methods, the meaning of Gaussian keywords, solvent models, basis sets, functionals, frequency analysis, transition state verification, etc. A higher score indicates that the model can provide a scientifically correct interpretation that conforms to Gaussian usage guidelines.
[0214] Gaussian software route generation task: This task measures whether the model can generate reasonable and executable Gaussian route segments or keyword combinations based on a given computational objective, such as structure optimization, frequency calculation, transition state search, IRC verification, and single-point energy correction. A higher score indicates that the generated route better meets the requirements of the computational purpose, software syntax, and computational chemistry methods.
[0215] Gaussian software results analysis task: This assesses whether the model can correctly interpret Gaussian outputs, such as energy, free energy, frequency, imaginary frequency, convergence information, IRC results, structural rationality, and thermodynamic information. A higher score indicates that the model is more accurate in determining whether the calculation results support the corresponding chemical conclusions.
[0216] Gaussian software error handling task: This task measures whether the model can identify the cause of failure based on Gaussian errors or abnormal outputs and propose reasonable repair solutions, such as SCF non-convergence, geometric optimization failure, inappropriate keyword settings, missing basis sets, and transition state non-convergence. A higher score indicates a stronger ability of the model to automatically diagnose and repair computational processes.
[0217] For each test question, each model being evaluated generates an answer under the same or as consistent as possible task input, chemical context, and answer requirements. During scoring, each model's answer is compared to a pre-determined reference answer, and the evaluation is based on fixed scoring rules corresponding to each task category. In the following evaluations, GPT-4o acts as the judge, comparing the model's answer with the corresponding reference answer according to the fixed scoring rules for each task category. The scoring focuses on: the chemical correctness of the answer, the appropriateness of the chosen calculation method, the feasibility of Gaussian keywords or routes, the accuracy of the interpretation of Gaussian output results, and the correctness of the diagnosis and repair suggestions for calculation errors. The average score of each model for each task within a specific task category is converted to a percentage, resulting in the percentage scores shown in the table below.
[0218] Taking the score of "70.65" of the dedicated computational model in the embodiment of the present invention in the basic tasks of computational chemistry and Gaussian software in Table 1 as an example, it means that the model has an average task completion score of 70.65 on the corresponding category test set. The higher the score, the better the model can provide answers that meet the specifications of computational chemistry, the requirements for using Gaussian software, and the task objectives.
[0219] As can be seen from the table below, the dedicated model according to the embodiments of the present invention outperforms the general large language model in each computational chemistry domain task.
[0220] Table 1: Scores of the calculation model under different tasks.
[0221]
[0222] The following is an example of how a computing device according to an embodiment of the present invention and an existing general-purpose large language model respond to the same Gaussian software route generation task.
[0223] The task problem is: "Calculate the following steps: "Use Gaussian software to calculate O=[C]c1ccc(cc1)[N+](=O)[O-],CC(=O)C, C(F)(F)(F)S(=O)(=O)O, N1CCCC1[C@H](N2CCCC2)C".
[0224] The output of the general large language model GPT-o3 is "ωB97X-D / DLPNO-CCSD(T)". However, this output is not a parameter of the Gaussian software.
[0225] The output of the computing device according to an embodiment of the present invention is: #p B3LYP / 6-31G(d) Opt=(TS,noeigentest,CalcAll,MaxStep=5) Freq SCRF=(IEFPCM,Solvent=Acetone) Temperature = 303.15 The above output demonstrates the relevance and accuracy of the question-answering capabilities of the computing device according to embodiments of the present invention in the field of computational chemistry compared to existing general-purpose large language models.
[0226] Embodiments of the invention have been described with reference to the accompanying drawings. These embodiments are illustrative and not restrictive.
Claims
1. A computational device for computational chemical question answering, characterized in that, include: Computing resources; The large language model is configured to invoke the computing resources to: Receive input from the user, including computational chemistry questions; The retrieval agent is configured to invoke the large language model to: Receive the input; Identify keywords related to computational chemistry in the input; Search for keywords to obtain search results; Suppose that an intelligent agent is generated and configured to invoke the large language model to: Receive the input, the keywords, and the search results; Based on the input, the keywords, and the search results, multiple scientific hypotheses are generated; The planning agent is configured to invoke the large language model to: Receive the multiple scientific hypotheses, the input, and the search results; Based on the multiple scientific hypotheses, the inputs, the search results, and the structured tool registry, a computational workflow is generated. The structured tool registry is provided by the computational chemistry software and indicates the functions, input formats, and output formats of the tools in the computational chemistry software. A dedicated model is configured to invoke the computing resources to: Review the professional rationale behind the aforementioned computational workflow; The large language model is configured to examine the logical consistency and input-output dependencies of the computational workflow; and The execution module is configured as follows: Receive the computation workflow; The computation workflow is converted into computation tasks based on the structured tool registry, and the computation tasks include tool name, input file and calling parameters; The computational chemistry software is used to execute the computational task to obtain the computational results, which serve as the answer to the computational chemistry question. The specialized model is obtained by training the base model on a software route generation training dataset, which is obtained through S1, and S1 includes: S11: Obtain input examples for computational tasks in different computational scenarios of computational chemistry software; S12: Extract the route segment portion from the input example, the route segment portion indicating the computational objective, theoretical method, basis set, and keywords for the computational task; S13: Based on the aforementioned route segment, construct keyword combination templates for computational chemistry software under different computational scenarios; S14: Construct software route generation training samples based on the keyword combination template, which serve as the software route generation training dataset. The software route generation training samples include computational tasks as sample inputs and route segments as sample outputs.
2. The computing device as described in claim 1, characterized in that, S1 further includes: S15: Determine whether the training samples generated by the software route conform to the verification rules; S16: In response to the software route generation training samples conforming to the verification rules, the software route generation training samples are used as the software route generation training dataset.
3. The computing device as described in claim 1, characterized in that, The specialized model is obtained by training the base model based on the software calculation results and the training dataset. The software calculation results analysis training dataset is obtained through S2, which includes: S21: Obtain the output file of the computational chemistry software's calculation task; S22: Extract the calculation information from the output file; S23: Annotate the calculated information to obtain the analysis results; S24: Based on the analysis results, construct software calculation result analysis training samples as the software calculation result analysis training dataset. The software calculation result analysis training samples include output files as sample inputs and analysis results as sample outputs.
4. The computing device as claimed in claim 1, characterized in that, The specialized model is obtained by training the base model on a software computation error handling training dataset. The software computation error handling training dataset is obtained through S3, which includes: S31: Obtain error messages from computational chemistry software for error cases; S32: Locate the error location in the error case according to the error type; S33: Generate an error solution for the error message; S34: Construct software computation error handling training samples based on the error type, the error location, and the error solution, as the software computation error handling training dataset. The software computation error handling training samples include error information as sample input, error location, and error solution as sample output.
5. The computing device as claimed in claim 1, characterized in that, The specialized model is obtained by training the base model on a basic knowledge training dataset. The basic knowledge training dataset is obtained through S4, which includes: S41: Obtain literature related to theoretical knowledge of computational chemistry and basic knowledge of computational chemistry software; S42: The literature is divided into knowledge units, which include computational chemistry theoretical knowledge units and software practical knowledge units; S43: Construct instruction-based basic knowledge training samples based on the knowledge units, as the basic knowledge training dataset.
6. The computing device as claimed in claim 1, characterized in that, The specialized model is obtained by training the base model on a general chemistry training dataset. The general chemistry training dataset is obtained through S5, which includes: S51: Obtain literature on basic chemical knowledge; S52: Perform structured analysis on the aforementioned basic chemical knowledge literature; S53: Construct general chemical knowledge training samples based on the analysis results, as the general chemical knowledge training dataset.
7. The computing device as claimed in claim 1, characterized in that, The dedicated model is configured to invoke the computing resources to determine the technical validity of the computation results, and The large language model is configured to invoke the computing resources to determine whether the computation results support the multiple scientific hypotheses.
8. The computing device as claimed in claim 7, characterized in that, The execution module is configured to: If the calculation result is invalid, the invalid result and the repair suggestion of the special model are fed back to the planning agent to adjust the calculation workflow, or the invalid result and the repair suggestion of the special model are fed back to the user to adjust the input; If the calculation result is valid but does not support the multiple scientific hypotheses, the calculation result is fed back to the hypothesis-generating agent to adjust the multiple scientific hypotheses.
9. The computing device as claimed in claim 1, characterized in that, The computation workflow includes multiple computation steps, each step including a step number, task description, tool name, input object, and output object.
10. The computing device as claimed in claim 1, characterized in that, The scientific hypothesis includes one or more of the following: reaction pathway, key intermediate, transition state configuration, selective control factor, or structure-property relationship explanation.
Citation Information
Patent Citations
Environmental chemistry research automation system and method based on large language model
CN119473265A
Multi-agent arrangement method and device based on service grid and storage medium
CN120653405A